An image recognition method and system based on an image processing algorithm

By employing image processing algorithms and template matching models in vending machines, the problems of easy wear and tear on mechanical structures and low recognition accuracy have been solved, enabling efficient and accurate product recognition in complex environments, improving equipment stability and reducing maintenance costs.

CN120451946BActive Publication Date: 2025-11-04河北盛马电子科技有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510557168.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-11-04
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

Automated vending machines suffer from problems in product recognition, such as easy wear and tear on the mechanical structure, low recognition accuracy, and susceptibility to changes in lighting and product placement.

Method used

An image processing algorithm-based approach is adopted. By preprocessing the image data when the vending machine door is open, the appearance and position information of the goods are extracted. Then, a template matching model is used to perform image recognition after the door is closed. Combined with a deep learning convolutional neural network model and feature point matching algorithm, the interference of light and position changes is overcome.

Benefits of technology

It improves the accuracy of image recognition, reduces misjudgments, lowers the probability of equipment failure, extends the service life of equipment, and reduces maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451946B_ABST
    Figure CN120451946B_ABST
Patent Text Reader

Abstract

The application provides an image recognition method and system based on an image processing algorithm, and belongs to the technical field of image processing. The method comprises the following steps: in response to the opening of a vending machine door, processing first image data to obtain target image data, wherein the first image data is image data of goods in the vending machine collected when the vending machine door is opened; the target image data comprises appearance information and position information of the goods at an initial moment; in response to the closing of the vending machine door, inputting second image data of the goods in the vending machine and the target image data into a template matching model to obtain an image recognition result, wherein the second image data is image data of the goods in the vending machine collected after the vending machine door is closed. The image recognition method and system based on the image processing algorithm provided by the application improve the accuracy of image recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of image processing technology, and more specifically, relates to an image recognition method and system based on image processing algorithms. Background Technology

[0002] Vending machines have significant shortcomings in product recognition, relying heavily on mechanical structures or weight sensors. Mechanical structures are prone to wear and tear due to frequent use, leading to frequent malfunctions; weight sensing methods are difficult to accurately distinguish between products of similar weight, resulting in low recognition accuracy.

[0003] With the development of image processing technology, image recognition has begun to be introduced into the field of vending machines; however, some image processing-based recognition methods suffer significant drops in accuracy under complex environments, such as changes in lighting and product placement. Therefore, there is an urgent need for a stable, efficient, and accurate image recognition method that can operate in various real-world scenarios. Summary of the Invention

[0004] The purpose of this application is to provide an image recognition method and system based on image processing algorithms to improve the accuracy of image recognition.

[0005] A first aspect of this application provides an image recognition method based on an image processing algorithm, comprising:

[0006] In response to the opening of the vending machine cabinet door, the first image data is processed to obtain target image data. The first image data is the image data of the goods inside the vending machine cabinet collected when the vending machine cabinet door opens. The target image data includes: the appearance information and position information of the goods at the initial moment.

[0007] In response to the closing of the vending machine cabinet door, the second image data and target image data of the goods inside the vending machine cabinet are input into the template matching model to obtain the image recognition result. The second image data is the image data of the goods inside the vending machine cabinet collected after the vending machine cabinet door is closed.

[0008] A second aspect of this application provides an image recognition system based on an image processing algorithm, comprising:

[0009] The data processing module is used to process the first image data in response to the opening of the vending machine cabinet door to obtain the target image data. The first image data is the image data of the goods inside the vending machine cabinet collected when the vending machine cabinet door opens. The target image data includes: the appearance information and position information of the goods at the initial moment.

[0010] The image recognition module is used to respond to the closing of the vending machine cabinet door by inputting the second image data of the goods inside the vending machine cabinet and the target image data into the template matching model to obtain the image recognition result. The second image data is the image data of the goods inside the vending machine cabinet collected after the vending machine cabinet door is closed.

[0011] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the image recognition method based on the image processing algorithm described above.

[0012] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the image recognition method based on the image processing algorithm described above.

[0013] The beneficial effects of the image recognition method and system based on image processing algorithms provided in this application are as follows: This application effectively overcomes the interference of complex factors such as changes in lighting and product placement by refining image data and matching image data before and after the door opens using a template matching model, thus accurately identifying products. Compared to traditional recognition methods, it improves the accuracy of image recognition, reduces misjudgments, and protects the rights of consumers and merchants. Furthermore, this application, through its image processing algorithm-based design, reduces reliance on easily worn mechanical structures, lowers the probability of equipment failure, improves the stability of vending machines during long-term operation, reduces maintenance costs, and extends the service life of the equipment. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 A flowchart illustrating an image recognition method based on an image processing algorithm, provided as an embodiment of this application;

[0016] Figure 2 A structural block diagram of an image recognition system based on an image processing algorithm provided in one embodiment of this application;

[0017] Figure 3 This is a schematic block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0018] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0019] To make the objectives, technical solutions, and advantages of this application clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.

[0020] Please refer to Figure 1 , Figure 1 A flowchart illustrating an image recognition method based on an image processing algorithm, provided in one embodiment of this application, includes:

[0021] S101: In response to the opening of the vending machine cabinet door, process the first image data to obtain target image data. The first image data is the image data of the goods inside the vending machine cabinet collected when the vending machine cabinet door is opened. The target image data includes: the appearance information and position information of the goods at the initial moment.

[0022] In this embodiment, when the vending machine door is opened, the image acquisition device is activated to acquire first image data of the goods inside the machine. The first image data consists of images of the goods inside the vending machine. After acquisition, this embodiment also requires image preprocessing to improve image quality and the accuracy of subsequent processing. After preprocessing, feature extraction is performed on the image to obtain the appearance and location information of the goods, thereby obtaining target image data. The appearance information can be used to distinguish different types of goods; the location information can be used to determine the bounding box coordinates of the goods and the specific location of each goods within the vending machine.

[0023] In this embodiment, image preprocessing operations include: grayscale conversion, filtering and noise reduction, and image enhancement. Grayscale conversion converts the color first image data into a grayscale image, reducing the data volume while preserving the image's main features. Filtering and noise reduction uses algorithms such as Gaussian filtering or median filtering to remove noise interference from the image. Noise may originate from factors such as the image acquisition device itself and ambient lighting; filtering can smooth the image and highlight the product's features. Image enhancement uses methods such as histogram equalization to enhance the contrast of the first image data, making the product's outline and details clearer.

[0024] The appearance information extraction method uses edge detection algorithms to extract the contour information of the product, and combines texture analysis methods (such as gray-level co-occurrence matrix) to obtain the texture features of the product.

[0025] The location information determination method is to use a target detection algorithm to locate the position of the product in the image and determine the bounding box coordinates of the product; in this embodiment, the coordinates can accurately determine the specific location of each product in the vending machine cabinet.

[0026] S102: In response to the closing of the vending machine cabinet door, the second image data and target image data of the goods inside the vending machine cabinet are input into the template matching model to obtain the image recognition result. The second image data is the image data of the goods inside the vending machine cabinet collected after the vending machine cabinet door is closed.

[0027] In this embodiment, after the vending machine door is closed, the image acquisition device is restarted to collect second image data of the goods inside the machine. This second image data reflects the state of the goods inside the vending machine after user operation. This embodiment also utilizes a template matching model to compare and analyze the second image data and the target image data, ultimately outputting an image recognition result. This result reflects whether the goods inside the machine have been taken away, returned, or their position has changed.

[0028] In this embodiment, the template matching model used in this application can be a scale-invariant feature transformation algorithm based on feature point matching, an accelerated robust feature algorithm, or a convolutional neural network model based on deep learning. In practical applications, a convolutional neural network model based on deep learning is preferred because it has powerful feature learning capabilities and high recognition accuracy in image recognition tasks. Before using the template matching model for recognition, it needs to be trained; the training data includes a large number of image samples of goods in the vending machine, as well as the corresponding product labels and location information. Specifically, the template matching model calculates the similarity between the features of the goods in the second image and the features of the goods in the target image (determining whether the goods have been taken, replaced, or their positions have changed based on the magnitude of the similarity); it outputs the image recognition result based on the magnitude of the similarity; the output image recognition result clearly indicates which goods have been taken out, which goods have been put back, or whether the positions of the goods have changed. The vending machine can automatically update the product inventory information and perform corresponding billing operations based on the recognition results.

[0029] As can be seen from the above, the embodiments of this application, through refined processing of image data and matching image data before and after the door opens based on a template matching model, effectively overcome the interference of complex factors such as changes in lighting and product placement, and can accurately identify products. Compared with traditional recognition methods, it improves the accuracy of image recognition, reduces misjudgments, and protects the rights and interests of consumers and merchants. This application, through its image processing algorithm-based design, reduces reliance on easily worn mechanical structures, lowers the probability of equipment failure, improves the stability of the vending machine during long-term operation, reduces maintenance costs, and extends the equipment's service life.

[0030] In one embodiment of this application, the first image data is processed to obtain target image data, including:

[0031] The backbone network parameters of the target detection algorithm are determined based on the types and quantities of goods in the vending machine;

[0032] The number of attention heads for the target detection algorithm is determined based on lighting conditions;

[0033] The target query quantity for the object detection algorithm is determined based on the quantity of goods in the vending machine.

[0034] The first image data is processed based on the backbone network parameters, the number of attention heads, and the number of target queries to obtain the target image data.

[0035] In this embodiment, the performance of the target detection algorithm is affected by various factors when processing image data in different scenarios. Specifically, in the process of processing the first image data collected when the vending machine door opens to obtain target image data, it is crucial to reasonably adjust the various parameters of the target detection algorithm. The setting of these parameters directly affects the target detection algorithm's ability to extract and recognize product features, thereby affecting the accuracy and completeness of the final target image data.

[0036] In this embodiment, the backbone network parameters of the target detection algorithm are determined based on the types and quantities of goods in the vending machine.

[0037] Specifically, the variety and quantity of goods in a vending machine is an important factor to consider. Different types of goods have different appearance characteristics, such as shape, color, and texture. When there are many types of goods, the object detection algorithm needs to have stronger feature extraction capabilities to accurately distinguish between different types of goods.

[0038] As the core component of object detection algorithms, the backbone network's parameter settings directly affect the feature extraction results. For scenarios with a large number of product types, a deeper and more complex backbone network structure is needed, along with corresponding parameter adjustments. Parameters such as kernel size, stride, and padding can be adjusted based on the number of product types. When the number of product types is small, the kernel size can be appropriately reduced to decrease computation; conversely, when the number of product types is large, the kernel size needs to be increased to capture a wider range of feature information, thereby improving the efficiency and accuracy of feature extraction.

[0039] In this embodiment, lighting conditions significantly impact image quality and feature extraction; the appearance of goods may change noticeably under different lighting conditions. To enable the object detection algorithm to accurately identify goods under various lighting conditions, the number of attention heads needs to be adjusted according to the lighting conditions. This helps the model focus more on important regions in the image, thereby improving feature extraction. When lighting conditions are poor, the features of goods in the image may be obscured, requiring an increase in the number of attention heads. This allows the model to focus more meticulously on different parts of the goods and extract features more effectively. For example, in dimly lit environments, the number of attention heads can be increased from the default 8 to 16, allowing the model to extract features from the image from multiple different angles, thus improving sensitivity to goods features. Conversely, in well-lit and uniformly lit environments, the number of attention heads can be reduced to lower computational costs and improve the algorithm's efficiency.

[0040] In this embodiment, increasing the number of target queries can improve the algorithm's ability to detect multiple targets. Specifically, the number of items in the vending machine also affects the performance of the target detection algorithm. When there are many items, the target detection algorithm needs to process more targets simultaneously, which requires increasing the number of target queries to ensure that all items can be accurately detected.

[0041] In addition, when determining the target number of queries, the arrangement and density of the products also need to be considered. If the products are arranged densely, the target number of queries may need to be increased appropriately to avoid missing any results.

[0042] For example, the first image data is processed based on the backbone network parameters, the number of attention heads, and the number of target queries to obtain target image data, including:

[0043] First, the first image data is input into the backbone network after parameter adjustment for feature extraction. The backbone network converts the input image data into a series of feature maps, which contain various feature information of the product.

[0044] Then, the extracted feature map is input into the attention module with an adjusted number of attention heads; the attention module automatically assigns attention weights based on the feature information in the image, making the model pay more attention to the important areas of the product, thereby further enhancing the feature representation.

[0045] Finally, the target detection algorithm with the target query count set is used to perform target detection on the feature map after the attention module is processed; the algorithm will search for possible product targets in the feature map and output the appearance information (such as category, color, shape, etc.) and location information (such as bounding box coordinates) of the product.

[0046] In one embodiment of this application, processing the first image data to obtain target image data further includes:

[0047] The first image data is processed into grayscale based on an adaptive grayscale algorithm to obtain the first grayscale image data.

[0048] Calculate the grayscale mean and variance of the first grayscale image data;

[0049] If the difference between the variance and the mean of the first grayscale image data is less than the first threshold, then the number of attention heads is increased by one step length.

[0050] If the difference between the variance and the mean of the first grayscale image data is greater than or equal to the first threshold, the number of attention heads is increased by a second step size; the second step size is less than the first step size.

[0051] In this embodiment, the key steps in processing the product image data in the vending machine to achieve accurate recognition are image grayscale processing and dynamically adjusting the parameters of the target detection algorithm based on its statistical features. These steps can effectively improve the algorithm's ability to capture product features under complex environments such as different lighting conditions, thereby improving the accuracy of product recognition.

[0052] In this embodiment, the number of attention heads in the object detection algorithm affects the model's attention level to different regions in the image and its feature extraction capability. By analyzing the difference between the grayscale mean and variance of the first grayscale image data, the contrast and sharpness of the image can be determined, and the number of attention heads can be dynamically adjusted to improve the algorithm's performance. Specifically, the grayscale mean and variance are important statistics describing the grayscale distribution characteristics of an image; the grayscale mean reflects the overall brightness level of the image, while the variance reflects the dispersion of the image's grayscale values.

[0053] When the difference between the variance and the mean of the first grayscale image data is less than a first threshold, it indicates that the grayscale values ​​of the image are relatively concentrated, the contrast is low, and there may be problems such as uneven lighting or blurring. This makes the product features in the image less obvious, requiring the model to pay more attention to image details. Therefore, the number of attention heads is increased by the first step length to extract features from the image, thereby capturing product feature information more comprehensively and improving the ability to identify products in low-contrast images. For example, the first step length can be set to 2; when the difference between the variance and the mean of grayscale is detected to be less than the first threshold, the number of attention heads is increased from the current 8 to 10.

[0054] When the difference between the variance and the mean of the first grayscale image data is greater than or equal to the first threshold, it indicates that the image has high contrast and the product features are relatively obvious; there is no need to significantly increase the number of attention points; therefore, the second step size is increased, and since the second step size is less than the first step size, the second step size can be set to 1; this can further optimize feature extraction in high-contrast images and avoid excessive increase in computation.

[0055] This application utilizes a method that dynamically adjusts the number of attention heads based on image grayscale statistical features, enabling the target detection algorithm to better adapt to images of different qualities and improving the accuracy and stability of product recognition in vending machines.

[0056] In one embodiment of this application, the first image data is processed to obtain target image data, including:

[0057] The first image data is denoised to obtain the target visual image data, and the depth image data is reconstructed into a 3D point cloud to obtain the target depth image data.

[0058] The target visual image data and the target depth image data are fused to obtain the target image data.

[0059] In this embodiment, in the product recognition scenario of an automated vending machine, in order to obtain more accurate and comprehensive product information, it is necessary to perform fine processing on the collected first image data. By processing and fusing the visual image data and depth image data separately, target image data containing rich product features can be obtained. Among them, the depth image data is the data collected by the depth sensor, which includes: depth camera, LiDAR, and binocular camera, etc.

[0060] In this embodiment, visual image data may be affected by various factors, such as fluctuations in ambient light and noise from the image sensor. These noises can affect the subsequent extraction and recognition of product features. Therefore, it is necessary to perform noise reduction processing on the visual image portion of the first image data to make the visual image data clearer and the outline and details of the product more obvious.

[0061] In this embodiment, the depth image data records the distance information between the goods in the vending machine and the image acquisition device. By performing three-dimensional point cloud reconstruction processing on the depth image data, the two-dimensional depth image can be converted into a three-dimensional point cloud model, thereby presenting the spatial structure and shape information of the goods more intuitively.

[0062] The 3D point cloud reconstruction process includes:

[0063] First, the original depth image data is preprocessed to remove invalid depth values ​​and fill in missing depth information.

[0064] Secondly, the pixel coordinates in the depth image data are converted into three-dimensional spatial coordinates; based on the converted three-dimensional coordinates, three-dimensional point cloud data is generated; where each point represents the spatial location of a pixel in the image.

[0065] Secondly, further processing of the generated point cloud data, such as denoising, smoothing, and sampling, can remove noise points in the point cloud, making the point cloud surface smoother, reducing the amount of point cloud data, and improving processing efficiency.

[0066] Finally, the target depth image data presents the spatial structure and shape information of the goods inside the vending machine in the form of a three-dimensional point cloud.

[0067] In one embodiment of this application, target visual image data and target depth image data are fused to obtain target image data, including:

[0068] Align the RGB channels of the target visual image data with the depth channels of the target depth image data at the pixel level;

[0069] The aligned RGB channels and depth channels are weighted and fused to obtain the fused image data;

[0070] Edge enhancement processing is performed on the fused image data to obtain the target image data.

[0071] In this embodiment, the acquisition methods and viewpoints of the target visual image data and the target depth image data may differ, and their pixel positions may not correspond. Therefore, pixel-level alignment is required. Pixel-level alignment ensures that the RGB channels of the target visual image data and the depth channel of the target depth image data are spatially consistent.

[0072] In this embodiment, pixel-level alignment includes: feature extraction, feature matching, transform estimation, and image alignment. Feature extraction involves extracting feature points from the target visual image data and the target depth image data, respectively. Feature point extraction algorithms include scale-invariant feature transformation and accelerated robust feature transformation. Feature matching involves matching the extracted feature points to find corresponding feature point pairs in the two types of image data. Matching algorithms can employ descriptor-based matching methods, such as nearest neighbor matching or random sample consensus algorithms. Transform estimation estimates the transformation relationship between the target visual image data and the target depth image data based on the matched feature point pairs, such as affine transformation or perspective transformation. Image alignment uses the estimated transformation relationship to transform the RGB channels of the target visual image data or the depth channel of the target depth image data, ensuring a one-to-one correspondence between their pixel positions.

[0073] In this embodiment, the specific weighted fusion method is as follows:

[0074] First, let the original RGB image be... Depth map is Pixel-level alignment is achieved by constructing a geometric deformation field T:

[0075]

[0076]

[0077] Where M is the bilinear sampling operator, This is the depth gradient scaling factor. The gradient at depth x; The gradient is the y-depth gradient; The coordinates of the geometric deformation field; For aligned RGB images.

[0078] The second step is to introduce deep confidence. Controlling the fusion weights to build an adaptive fusion model:

[0079]

[0080] in, To integrate data, For the Sigmoid function, For depth edge threshold, This is a non-linear mapping from the depth map to the feature space (which can be parameterized as an MLP). It is a cross-modal gating unit. For depth gradient.

[0081] The third step involves constructing a depth-guided Laplacian pyramid sharpening operator to process and fuse the data, obtaining the target image data:

[0082]

[0083] in, For target image data, For the k-th level of the Gaussian pyramid, β is the Laplacian kernel for the corresponding layer, and β is the sharpening gain coefficient. For depth-appearance related masks;

[0084]

[0085] in, is the scale factor for the k-th layer.

[0086] This application obtains fused image data through weighted fusion, which contains both visual features and depth information of the product. While the fused image data integrates visual and depth information, the edges of the product may not be clear enough in some cases, affecting subsequent product recognition performance. Therefore, edge enhancement processing is needed to highlight the edge features of the product in the fused image data; for example, methods based on convolution or Laplacian operators can be used.

[0087] In one embodiment of this application, the aligned RGB channels and depth channels are weighted and fused to obtain fused image data, and the method further includes:

[0088] The first weight adjustment step size is determined based on the distance between the product and the image acquisition device;

[0089] The first weight is obtained by adjusting the first weight reference value based on the first weight adjustment step size; the second weight is obtained by adjusting the second weight reference value based on the first weight adjustment step size.

[0090] The first weight is the weight corresponding to the RGB channel, and the second weight is the weight corresponding to the depth channel. The adjustment directions of the first weight reference value and the second weight reference value are different.

[0091] In this embodiment, the reliability and importance of visual and depth information vary depending on the distance between the product and the image acquisition device. When the product is closer to the image acquisition device, visual information (such as color and texture) may be clearer and more accurate, and the weight of the RGB channel may need to be relatively increased. However, when the product is farther away, depth information may be more critical for determining the spatial position and shape of the product, and the weight of the depth channel may need to be increased.

[0092] In this embodiment, in order to dynamically adjust the weights according to the distance between the commodity and the image acquisition device, it is first necessary to determine the first weight adjustment step size; this goal can be achieved by establishing a mapping relationship between the distance and the adjustment step size. For example, the distance can be divided into several intervals, and each interval corresponds to a specific adjustment step size. Suppose the distance is divided into three intervals: a short-distance interval d1 (0 < d ≤ D1), a medium-distance interval d2 (D1 < d ≤ D2), and a long-distance interval d3 (d > D2), where d represents the distance between the commodity and the image acquisition device, and D1 and D2 are pre-set distance thresholds; for the short-distance interval, a relatively small adjustment step size Δw1 can be set because at short distances, the visual information itself is already relatively reliable and there is no need to significantly adjust the weights; for the medium-distance interval, a moderate adjustment step size Δw2 is set; for the long-distance interval, a relatively large adjustment step size Δw3 is set to highlight the importance of depth information.

[0093] In this application, when the commodity is relatively close, enhancing the weight of the RGB channel can more clearly identify the color and texture of the commodity; when the commodity is relatively far away, enhancing the weight of the depth channel can more accurately judge the spatial position and shape of the commodity, thus providing stronger support for the accurate identification of the commodity.

[0094] In one embodiment of this application, the second image data and the target image data of the commodity in the automatic vending cabinet are input into the template matching model, and the image recognition result is obtained, including:

[0095] The target image data is input into the template matching model as the template data;

[0096] The second image data is processed to obtain the image data to be detected, and the image data to be detected is input into the template matching model;

[0097] The similarity between the image data to be detected and the target image data is calculated according to the template matching model;

[0098] When the similarity is greater than the first similarity threshold, it is determined that the image data to be detected and the target image data are the image recognition results of the same commodity.

[0099] In this embodiment, the target image data contains the appearance information and position information of the commodity at the initial moment when the automatic vending machine door is opened, and is an accurate record of the initial state of the commodity; inputting it into the template matching model as the template data provides a benchmark for the subsequent matching process; the template matching model uses the target image data as a basis to judge the similarity between the subsequent input image data and this template.

[0100] The second image data consists of product images captured after the vending machine door is closed, reflecting the product's state after user interaction. To improve template matching accuracy, the second image data needs to be preprocessed to obtain the image data to be detected.

[0101] Preprocessing steps typically include: image enhancement, noise reduction, and resizing; among which, resizing is to ensure that the size of the second image data is consistent with the size of the target image data so as to make accurate comparisons during template matching.

[0102] In this embodiment, different image processing algorithms can be used to process image data, such as traditional image processing algorithms, scale-invariant feature transformation algorithms, and deep learning-based convolutional neural network models.

[0103] In this embodiment, a first similarity threshold is set to determine whether the image data to be detected and the target image data represent the same product. This threshold is determined based on a large amount of experimental data and practical application requirements. When the similarity is greater than the first similarity threshold, it indicates that the image data to be detected and the target image data have high consistency in features, and can be considered to represent the same product. In this case, the image recognition result output by the template matching model shows that the product has not changed during the opening and closing of the cabinet door, i.e., it has not been taken or replaced. Conversely, if the similarity is less than or equal to the first similarity threshold, it indicates that the product may have changed, such as being taken out or replaced with other products. Based on this image recognition result, the vending machine can update the product inventory information and perform corresponding billing operations.

[0104] In one embodiment of this application, target visual image data and target depth image data are fused to obtain target image data, including:

[0105] Depth confidence is calculated using the calibration error parameters of the depth sensor and a real-time noise model.

[0106] The target image data is obtained by weighting the target visual image data and the target depth image data based on the depth confidence score.

[0107] In this embodiment, the depth sensor experiences calibration errors and real-time noise when acquiring depth image data, which affects the reliability of the depth information. Therefore, this embodiment proposes to calculate a depth confidence score and then perform a weighted calculation based on this score. This weighted calculation method based on depth confidence score can fully utilize the advantages of both target visual image data and target depth image data, while reducing the impact of depth sensor errors and noise, resulting in more accurate and reliable target image data.

[0108] In this embodiment, calculating the depth confidence requires determining the calibration error parameters and establishing a real-time noise model. Based on the calibration error parameters and the real-time noise model, the depth confidence of each pixel can be calculated. The depth confidence represents the reliability of the depth value of the pixel, and its value ranges from 0 to 1. The closer the value is to 1, the more reliable the depth value is.

[0109] In one embodiment of this application, the image recognition method of the image processing algorithm further includes:

[0110] The second weight adjustment step size is determined based on the deep confidence level.

[0111] The third weight is obtained by adjusting the third weight reference value based on the second weight adjustment step size; the fourth weight is obtained by adjusting the fourth weight reference value based on the second weight adjustment step size.

[0112] The third weight is the weight of the corresponding target visual image data, and the fourth weight is the weight of the corresponding target depth image data. The adjustment directions of the third weight reference value and the fourth weight reference value are different.

[0113] In this embodiment, the depth confidence reflects the reliability of the depth value of each pixel in the target depth image data. Its value range is usually between 0 and 1. The closer the value is to 1, the more reliable the depth value is.

[0114] Establish a mapping relationship between depth confidence and the second weight adjustment step size to dynamically adjust the weights of the target visual image data and the target depth image data based on the depth confidence.

[0115] For example, the depth confidence can be divided into several intervals, each interval corresponding to a specific second weight adjustment step size.

[0116] When the depth confidence is in a low range (e.g., 0≤C) d <0.3, C d When the weight is expressed as depth confidence, it indicates that the reliability of the depth image data is poor. In this case, a smaller second weight adjustment step size should be set to avoid excessive reliance on unreliable depth information to make large adjustments to the weights.

[0117] When the depth confidence level is in the medium range (e.g., 0.3 ≤ C) d When the value is less than 0.7, the depth image data has a certain degree of reliability. A suitable second weight adjustment step size can be set so as to properly consider the changes in depth information during the fusion process.

[0118] When the depth confidence is in a high range (e.g., 0.7 ≤ C) d When the value is ≤1), the depth image data is relatively reliable, and a larger second weight adjustment step size can be set, so as to make fuller use of reliable depth information in the fusion.

[0119] In this embodiment, the third weight reference value w3 and the fourth weight reference value w4 are preset initial weight values, and w3 + w4 = 1; wherein, the third weight is the weight corresponding to the target visual image data, and the fourth weight is the weight corresponding to the target depth image data, and the adjustment directions of the third weight reference value and the fourth weight reference value are different.

[0120] This application utilizes a dynamic weighting method based on depth confidence to flexibly allocate the importance of target visual image data and target depth image data during the fusion process according to the reliability of the depth image data. This results in the fused target image data more accurately reflecting the actual characteristics of the product, thereby improving the accuracy and reliability of product recognition in vending machines. In practical applications, a more reasonable weight allocation allows the system to operate stably and efficiently under different environments and conditions, providing strong support for the intelligent management of vending machines.

[0121] Corresponding to the above embodiment, an image recognition method based on an image processing algorithm, Figure 2 This is a structural block diagram of an image recognition system based on an image processing algorithm, provided as an embodiment of this application. For ease of explanation, only the parts relevant to the embodiment of this application are shown. Reference Figure 2 The image recognition system 20 based on image processing algorithms includes a data processing module 21 and an image recognition module 22.

[0122] The data processing module 21 is used to process the first image data in response to the opening of the vending machine cabinet door to obtain target image data. The first image data is the image data of the goods inside the vending machine cabinet collected when the vending machine cabinet door is opened. The target image data includes the appearance information and position information of the goods at the initial moment. The image recognition module 22 is used to input the second image data of the goods inside the vending machine cabinet and the target image data into the template matching model in response to the closing of the vending machine cabinet door to obtain the image recognition result. The second image data is the image data of the goods inside the vending machine cabinet collected after the vending machine cabinet door is closed.

[0123] In one embodiment of this application, the data processing module 21 is specifically used for:

[0124] The backbone network parameters of the target detection algorithm are determined based on the types and quantities of goods in the vending machine;

[0125] The number of attention heads for the target detection algorithm is determined based on lighting conditions;

[0126] The target query quantity for the object detection algorithm is determined based on the quantity of goods in the vending machine.

[0127] The first image data is processed based on the backbone network parameters, the number of attention heads, and the number of target queries to obtain the target image data.

[0128] In one embodiment of this application, the data processing module 21 is specifically used for:

[0129] The first image data is processed into grayscale based on an adaptive grayscale algorithm to obtain the first grayscale image data.

[0130] Calculate the grayscale mean and variance of the first grayscale image data;

[0131] If the difference between the variance and the mean of the first grayscale image data is less than the first threshold, then the number of attention heads is increased by one step length.

[0132] If the difference between the variance and the mean of the first grayscale image data is greater than or equal to the first threshold, the number of attention heads is increased by a second step size; the second step size is less than the first step size.

[0133] In one embodiment of this application, the data processing module 21 is specifically used for:

[0134] The first image data is denoised to obtain the target visual image data. The depth image data is then reconstructed into a 3D point cloud to obtain the target depth image data. The depth image data is the data collected by the depth sensor.

[0135] The target visual image data and the target depth image data are fused to obtain the target image data.

[0136] In one embodiment of this application, the data processing module 21 is specifically used for:

[0137] Align the RGB channels of the target visual image data with the depth channels of the target depth image data at the pixel level;

[0138] The aligned RGB channels and depth channels are weighted and fused to obtain the fused image data;

[0139] Edge enhancement processing is performed on the fused image data to obtain the target image data.

[0140] In one embodiment of this application, the data processing module 21 is specifically used for:

[0141] The first weight adjustment step size is determined based on the distance between the product and the image acquisition device;

[0142] The first weight is obtained by adjusting the first weight reference value based on the first weight adjustment step size; the second weight is obtained by adjusting the second weight reference value based on the first weight adjustment step size.

[0143] The first weight is the weight corresponding to the RGB channel, and the second weight is the weight corresponding to the depth channel. The adjustment directions of the first weight reference value and the second weight reference value are different.

[0144] In one embodiment of this application, the image recognition module 22 is specifically used for:

[0145] The target image data is used as template data and input into the template matching model;

[0146] The second image data is processed to obtain the image data to be detected, and the image data to be detected is input into the template matching model;

[0147] The similarity between the image data to be detected and the target image data is calculated based on the template matching model.

[0148] When the similarity is greater than the first similarity threshold, the image recognition result is determined to be the same product as the image data to be detected and the target image data.

[0149] See Figure 3 , Figure 3 This is a schematic block diagram of an electronic device provided according to an embodiment of this application. Figure 3 The electronic device 300 in this embodiment may include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memories 304 store computer programs, including program instructions. The processors 301 execute the program instructions stored in the memories 304. Specifically, the processors 301 are configured to invoke the program instructions to perform the functions of the modules in the aforementioned device embodiments, for example... Figure 2 The functions of the data processing module 21 and the image recognition module 22 are shown.

[0150] It should be understood that, in the embodiments of this application, the processor 301 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0151] Input device 302 may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint orientation information), a microphone, etc., and output device 303 may include a display (LCD, etc.), a speaker, etc.

[0152] The memory 304 may include read-only memory and random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include non-volatile random access memory. For example, the memory 304 may also store device type information.

[0153] In specific implementations, the processor 301, input device 302, and output device 303 described in the embodiments of this application can execute the implementation methods described in the first and second embodiments of the image recognition method based on image processing algorithms provided in the embodiments of this application, or they can execute the implementation methods of the electronic devices described in the embodiments of this application, which will not be repeated here.

[0154] In another embodiment of this application, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program, which includes program instructions. When executed by a processor, the program instructions implement all or part of the processes in the methods described above. Alternatively, the computer program can instruct related hardware to implement these processes. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0155] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., provided on the electronic device. Furthermore, the computer-readable storage medium can include both internal and external storage units of the electronic device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0156] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0157] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the electronic devices and units described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0158] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the couplings or direct couplings or communication connections shown or discussed may be indirect couplings or communication connections through some interfaces or units, or they may be electrical, mechanical, or other forms of connection.

[0159] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.

[0160] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0161] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An image recognition method based on image processing algorithms, characterized in that, include: In response to the opening of the vending machine cabinet door, the first image data is processed to obtain target image data. The first image data is image data of the goods inside the vending machine cabinet collected when the vending machine cabinet door opens. The target image data includes: the appearance information and position information of the goods at the initial moment. In response to the closing of the vending machine cabinet door, the second image data of the goods inside the vending machine cabinet and the target image data are input into the template matching model to obtain the image recognition result. The second image data is the image data of the goods inside the vending machine cabinet collected after the vending machine cabinet door is closed. The process of processing the first image data to obtain the target image data includes: The backbone network parameters of the target detection algorithm are determined based on the types and quantities of goods in the vending machine; The number of attention heads for the target detection algorithm is determined based on lighting conditions; The target query quantity for the object detection algorithm is determined based on the quantity of goods in the vending machine. The first image data is processed based on the backbone network parameters, the number of attention heads, and the number of target queries to obtain target image data; Also includes: The first image data is processed into grayscale based on an adaptive grayscale algorithm to obtain the first grayscale image data. Calculate the grayscale mean and variance of the first grayscale image data; When the difference between the variance of the first grayscale image data and the grayscale mean is less than the first threshold, the number of attention heads is increased by a step length. When the difference between the variance of the first grayscale image data and the grayscale mean is greater than or equal to the first threshold, the number of attention heads is increased by a second step size; the second step size is less than the first step size.

2. The image recognition method of the image processing algorithm as described in claim 1, characterized in that, The process of processing the first image data to obtain the target image data includes: The first image data is denoised to obtain target visual image data, and the depth image data is reconstructed into a three-dimensional point cloud to obtain target depth image data. The depth image data is data collected by a depth sensor. The target visual image data and the target depth image data are fused to obtain the target image data.

3. The image recognition method of the image processing algorithm as described in claim 2, characterized in that, The process of fusing the target visual image data and the target depth image data to obtain target image data includes: Align the RGB channels of the target visual image data with the depth channels of the target depth image data at the pixel level; The aligned RGB channels and depth channels are weighted and fused to obtain the fused image data; Edge enhancement processing is performed on the fused image data to obtain the target image data.

4. The image recognition method of the image processing algorithm as described in claim 3, characterized in that, Also includes: The first weight adjustment step size is determined based on the distance between the product and the image acquisition device; The first weight is obtained by adjusting the first weight reference value based on the first weight adjustment step size; The second weight is obtained by adjusting the second weight reference value based on the first weight adjustment step size; The first weight is the weight corresponding to the RGB channel, and the second weight is the weight corresponding to the depth channel. The adjustment directions of the first weight reference value and the second weight reference value are different.

5. The image recognition method of the image processing algorithm as described in claim 1, characterized in that, The step of inputting the second image data of the goods in the vending machine and the target image data into a template matching model to obtain the image recognition result includes: The target image data is used as template data and input into the template matching model; The second image data is processed to obtain the image data to be detected, and the image data to be detected is input into the template matching model; The similarity between the image data to be detected and the target image data is calculated based on the template matching model. When the similarity is greater than the first similarity threshold, the image recognition result is determined to be the same product as the image data to be detected and the target image data.

6. An image recognition system based on an image processing algorithm, characterized in that, include: The data processing module is used to process the first image data in response to the opening of the vending machine cabinet door to obtain the target image data. The first image data is the image data of the goods in the vending machine cabinet collected when the vending machine cabinet door is opened. The target image data includes: the appearance and location information of the product at the initial moment; An image recognition module is used to respond to the closing of the vending machine cabinet door by inputting the second image data of the goods inside the vending machine cabinet and the target image data into a template matching model to obtain an image recognition result. The second image data is the image data of the goods inside the vending machine cabinet collected after the vending machine cabinet door is closed. The data processing module is specifically used for: The backbone network parameters of the target detection algorithm are determined based on the types and quantities of goods in the vending machine; The number of attention heads for the target detection algorithm is determined based on lighting conditions; The target query quantity for the object detection algorithm is determined based on the quantity of goods in the vending machine. The first image data is processed based on the backbone network parameters, the number of attention heads, and the number of target queries to obtain the target image data; The data processing module is specifically used for: The first image data is processed into grayscale based on an adaptive grayscale algorithm to obtain the first grayscale image data. Calculate the grayscale mean and variance of the first grayscale image data; If the difference between the variance and the mean of the first grayscale image data is less than the first threshold, then the number of attention heads is increased by one step length. If the difference between the variance and the mean of the first grayscale image data is greater than or equal to the first threshold, the number of attention heads is increased by a second step size; the second step size is less than the first step size.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Target trajectory optimization method based on DS evidence theory

    CN110675418A

  • Image processing method and device, equipment and storage medium

    CN111126264A

  • Commodity dynamic identification method and system based on computer vision and commodity container

    CN115393366A