Goods shelf inspection algorithm based on image recognition technology and big data mining technology
By applying image recognition and big data mining technology in shelf inspections, we automatically analyze shelf images and combine sales data, the high cost, low efficiency and subjectivity problems of traditional manual inspections are solved, and efficient and reliable shelf management and sales strategy optimization are achieved.
Patent Information
- Application Number
- CN202411867772.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-27
AI Technical Summary
Traditional manual shelf inspections have high cost, low efficiency and subjectivity problems, resulting in the persistence of out-of-stock or placement errors, and the optimization of shelf placement strategies cannot be achieved.
The shelf inspection algorithm based on image recognition technology and big data mining technology is adopted to automatically collect and analyze shelf images through intelligent devices, identify product types, out of stock, placement status and label information, and combine sales data analysis to generate optimization suggestions.
Automatic inspections have been realized, manual participation and related costs have been reduced, inspection efficiency and objectivity and reliability of results have been improved, shelf management and sales strategies have been optimized, and the overall operational efficiency and sales of the store have been improved.
Smart Images

Figure FDA0005194909800000011 
Figure FDA0005194909800000012
Abstract
Description
Technical Field
[0001] The present invention relates to the field of digital inspection technology, and in particular to a shelf inspection algorithm based on image recognition technology and big data mining technology. Background Art
[0002] Shelf inspection is an important part of the daily operation of supermarkets and chain snack shops. Its purpose is to ensure the normal operation of the following aspects: (1) Inventory management: timely detection of out-of-stock situations to ensure sufficient supply of goods on the shelves. (2) Display specifications: ensuring that goods are displayed in accordance with the display methods specified by the headquarters to maintain the brand image and attract customers. (3) Service quality: improving customer experience and promoting sales through good shelf management.
[0003] Most of the current technical solutions are based on traditional manual inspections. The typical workflow is as follows: (1) The headquarters regularly dispatches inspectors to stores for inspections; (2) The inspectors check one by one whether there are any out-of-stock items on the shelves, whether the merchandise is placed in a standardized manner, and whether the price tags are correct; (3) Manually record abnormal shelf conditions (such as out-of-stock items, incorrect placement, etc.) and fill out inspection reports manually; (4) The inspection reports are uploaded to the headquarters, which decides on replenishment plans or corrective measures based on the reports.
[0004] Traditional shelf inspections rely on the headquarters to dispatch inspectors or hire third-party teams, which involves costs such as transportation and labor costs; the higher the inspection frequency, the greater the cost. High costs lead to a lower inspection frequency, and the headquarters cannot grasp the store situation in real time, which will miss the opportunity to replenish and optimize.
[0005] The manual inspection process is lengthy, including shelf inspection, manual recording, summary reporting, etc., and it takes time to upload and analyze information. This leads to low inspection efficiency, delayed feedback affects the headquarters' timely response to store issues, and may cause out-of-stock or incorrect placement issues to persist.
[0006] Inspection personnel may miss shelf anomalies or fail to report store conditions truthfully due to negligence, heavy workload or other human factors, which leads to unreliable inspection results and affects the accuracy of headquarters' decisions.
[0007] The main goal of manual inspections is to find out-of-stock or improperly placed shelves, but there is no systematic method to analyze the relationship between sales data and placement strategies. Inspections are limited to "error checking" and cannot optimize shelf placement strategies, missing opportunities to increase sales and customer satisfaction. Summary of the invention
[0008] The present invention mainly addresses the deficiencies in the prior art and provides a shelf inspection algorithm based on image recognition technology and big data mining technology, which solves the problems of high cost, low efficiency, and subjectivity existing in traditional manual inspections and existing technical solutions. It improves the shelf management level and sales efficiency of stores through data optimization.
[0009] The above technical problems of the present invention are mainly solved by the following technical solutions:
[0010] A shelf inspection algorithm based on image recognition technology and big data mining technology includes the following operating steps:
[0011] First step: Collect data of the display pictures of the shelves and the pictures of the commodities to obtain commodity data based on the pictures.
[0012] Second step: Preprocess the acquired data images, and use median filtering and mean filtering for denoising algorithms to reduce noise interference in the images.
[0013] Median filtering is a non-linear filtering method that replaces the original value of each pixel in the image with the median value of the pixels in the neighborhood of that pixel, thereby removing noise.
[0014] The specific implementation is as follows: First, select a sliding window, and the size of the window is usually odd; for each pixel in the image, move the window to the area around that pixel to form a set of pixel values; sort the pixel values in this area and select the middle value after sorting as the new value of that pixel; perform the above operations on each pixel in the image, and finally obtain the denoised image. Mean filtering is a linear filtering method that replaces the original value of each pixel in the image with the average value of all pixels in the neighborhood of that pixel. Its specific implementation is as follows: First, select a sliding window; for each pixel in the image, move the window to the area around that pixel to form a set of pixel values; calculate the average value of all pixels in this area as the new value of that pixel; perform the above operations on each pixel in the image to obtain the denoised image.
[0015] Third step: Perform object detection and identification of commodity information.
[0016] Fourth step: Detect the out-of-stock situation, placement, and labels of the commodities.
[0017] Fifth step: Conduct sales data analysis and generate optimization suggestions.
[0018] Sixth step: Generate a report based on the results of anomaly detection, the results of data analysis, and optimization suggestions, push it to the store management system and the headquarters monitoring platform, and notify relevant personnel via text message or email.
[0019] Preferably, images of each commodity are obtained through the AI weighing equipment in the store for the identification of commodity types; images of the store shelves are obtained through the mobile robots and cameras in the store, including the placement status of the shelves, the commodities placed, the placement density, and the labels attached to the shelves.
[0020] Preferably, the clarity of the image is improved through histogram equalization or contrast-limited adaptive histogram equalization algorithm to make the edges of the commodities more obvious.
[0021] Preferably, histogram equalization is an image processing method. By adjusting the gray distribution of the image, the histogram of the image is made to be as evenly distributed as possible, thereby enhancing the contrast of the image. The specific implementation is as follows: First, calculate the gray histogram of the original image, that is, count the number of pixels at each gray level in the image; Second, calculate the cumulative distribution function (CDF). The cumulative distribution function is the accumulation of the gray histogram, and calculate the cumulative probability of the pixels at each gray level. where P(i) is the pixel probability of gray level i and k is the current gray level; then calculate the new gray value. According to the cumulative distribution function, map the gray value of each pixel to the new gray value. where CDF min is the smallest non-zero cumulative probability value, N is the total number of pixels in the image, and L is the total number of gray levels; finally, apply the new gray value, replace the gray value of each pixel in the original image with the calculated new gray value, so as to obtain the equalized image.
[0022] Preferably, a convolutional neural network is used to classify products and identify the type of each product; a positioning algorithm is used to identify the position of the product on the shelf and calculate its coordinates on the shelf; image analysis is used to determine the placement state of the product; the convolutional neural network is a network architecture widely used in image processing in deep learning, and extracts and learns the features of images through multiple layers such as convolutional layers, pooling layers, and fully connected layers; the convolutional layer is the core of the CNN, responsible for extracting local features from the image, performing a convolution operation on the input image through a convolution kernel - filter, and extracting features such as edges, textures, and colors in the image; the pooling layer reduces the size of the feature map through downsampling, reduces the amount of computation, and prevents overfitting; after the convolutional and pooling layers, the CNN usually contains one or more fully connected layers, which are used to map the extracted features to classification labels, and the features and classification labels are used to associate the features with the products in combination with the KNN algorithm, so as to identify the type of each product; the positioning algorithm is used to identify the position of the product on the shelf and calculate its coordinates on the shelf, and usually involves multiple technologies and steps. First is image preprocessing, using the median filtering algorithm to remove image noise, enhancing the brightness and contrast of the image, converting the image into a black and white binary image through thresholding operations to highlight the product outline; then using YOLO image analysis to extract useful information from the image and output the bounding box of the product, that is, the upper left corner coordinates, width, and height; then a shelf model is established, obtaining the 3D model or 2D plane coordinates of the shelf through the image captured by the camera or through lidar and depth sensors. The coordinate system of the shelf is usually a two - dimensional coordinate system, where the x - axis represents the horizontal direction and the y - axis represents the vertical direction. The position and perspective of the camera are determined through a calibration method, so that the pixel points in the image are mapped into the shelf coordinate system; finally, perspective transformation is used to convert the pixels in the image into the coordinates on the actual shelf; image analysis is to combine the convolutional neural network and the positioning algorithm, and after comprehensive comparison, give the placement state of the product.
[0023] Preferably, an optical character recognition technology is used to recognize the labels of commodities, including prices, promotional information, and expiration dates. The optical character recognition technology, namely OCR technology, generally includes steps such as image preprocessing, character segmentation, feature extraction, character recognition, and post-processing. The specific implementation is as follows: First is image preprocessing, which includes grayscale conversion, that is, converting a color image into a grayscale image; binarization, that is, converting the grayscale image into a black-and-white image; denoising, that is, using methods such as median filtering for denoising; skew correction, that is, detecting the tilt angle of the text through the Hough transform method and performing correction; edge enhancement, that is, using the Laplacian operator for edge detection. Secondly is character segmentation. Line segmentation is to segment the text into lines by analyzing the blank areas in the image; character segmentation is to segment the characters in each line. Then is feature extraction, using a convolutional neural network for feature extraction. Then is character recognition, using KNN to map the characters in the image to corresponding texts through the extracted features. Finally is post-processing, performing spelling correction through a dictionary and correcting errors by comparing the words output by OCR with the words in the dictionary. Through these means, the text information on the label in the image is finally output.
[0024] Preferably, according to the output placement status of the commodities, where the placement status includes normal, empty, or half-full, the out-of-stock or vacant commodities on the shelf are given, and the out-of-stock commodities are recorded; according to the output types of the commodities and by comparing the coordinates of the commodities with the standard shelf, it is judged whether the commodities are placed incorrectly and whether they exceed the boundaries of the shelf, and the incorrectly placed commodities are recorded; the label information in the image is compared with the commodity information in the background database to detect whether the label is correct and meets the standards. The commodities with incorrect labels are recorded.
[0025] Preferably, the sales situations of different commodities are analyzed, including sales volume, turnover rate, etc., to identify which commodities are hot-selling, which commodities are slow-selling, and which commodities are associated commodities.
[0026] Preferably, based on the analysis of the sales situation of commodities, the best-selling commodities are preferentially displayed in prominent positions and related commodities are placed together; based on historical sales data and the ARIMA prediction model, reasonable replenishment time and quantity are generated. The ARIMA model models the dependencies in time series data by combining autoregressive - AR, differencing - I, and moving average - MA components for prediction and modeling; it is represented as ARIMA(p, d, q), where p is the order of the autoregressive term, d is the differencing order used to make a non - stationary time series stationary, and q is the order of the moving average term. First is autoregression, which refers to the linear relationship between the current value in the model and past values or lagged values. The core idea is that the current value can be predicted by the observed values at previous moments; the differencing part is used to eliminate the non - stationarity of the time series. Non - stationarity usually manifests as the mean, variance, or covariance of the data changing over time. The purpose of differencing is to make the time series stationary by subtracting adjacent observations; the moving average part reflects the relationship between the error term and its past values. Through the moving average model, past errors are used to explain the observed values of the current time series.
[0027] Preferably, the training step first checks whether the time series is stationary; if the data is non - stationary, the data is made stationary through differencing operations. Through one or more differencing processes, the mean and variance of the data are stabilized; determine the p, d, q parameters. The orders p and q of the AR and MA parts are selected through the autocorrelation function - ACF and the partial autocorrelation function - PACF. The parameter d is determined through differencing operations, indicating the stationarity of the sequence; fit the ARIMA model, use the selected parameters p, d, q to construct the ARIMA model, and estimate the parameters of the model through the method of least squares or maximum likelihood estimation MLE; finally, for prediction, use the trained ARIMA model to predict future values.
[0028] The present invention can achieve the following effects:
[0029] The present invention provides a shelf inspection algorithm based on image recognition technology and big data mining technology. Compared with the prior art, it solves the problems of high cost, low efficiency, and subjectivity existing in traditional manual inspections and existing technical solutions. It improves the shelf management level and sales efficiency of the store through data optimization.
[0030] (1) Replace manual inspection with intelligent inspection equipment, and automatically analyze shelf images through algorithms, greatly reducing manual participation and related costs.
[0031] (2) Realize the real - time upload and automatic analysis of shelf images, shorten the information feedback chain, and ensure that the headquarters can quickly obtain the inspection results of the store.
[0032] (3) The shelf recognition and data analysis based on algorithms avoid human omissions and subjective biases, improving the objectivity and reliability of the inspection results.
[0033] (4) Combining big data mining technology, analyze the relationship between historical sales volume and placement strategies, and put forward scientific optimization suggestions to help increase sales. Specific implementation manners
[0034] The technical solution of the invention will be further specifically described below through embodiments.
[0035] Embodiment: A shelf inspection algorithm based on image recognition technology and big data mining technology includes the following operation steps:
[0036] The first step: Collect data of the display pictures of the shelves and the pictures of the commodities to obtain commodity data based on pictures.
[0037] Through the AI weighing equipment in the store, obtain pictures of each commodity for the identification of commodity types; through the mobile robots and cameras in the store, obtain pictures of the store shelves, including the placement status of the shelves, the placed commodities, the placement density, and the labels pasted on the shelves.
[0038] The second step: Preprocess the acquired data images, and use median filtering and mean filtering for denoising algorithms to reduce noise interference in the images. Improve the clarity of the images through histogram equalization or contrast-limited adaptive histogram equalization algorithms to make the edges of the commodities more obvious.
[0039] Histogram equalization is an image processing method. By adjusting the gray distribution of the image, the histogram of the image is made to be as evenly distributed as possible, thereby enhancing the contrast of the image. Its specific implementation is as follows: First, calculate the gray histogram of the original image, that is, count the number of pixels at each gray level in the image; secondly, calculate the cumulative distribution function (CDF). The cumulative distribution function is the accumulation of the gray histogram, and calculate the cumulative probability of the pixels at each gray level where P(i) is the pixel probability of gray level i, and k is the current gray level; then calculate the new gray value. According to the cumulative distribution function, map the gray value of each pixel to the new gray value where CDF min is the smallest non-zero cumulative probability value, N is the total number of pixels in the image, and L is the total number of gray levels; finally, apply the new gray value, and replace the gray value of each pixel in the original image with the calculated new gray value, so as to obtain the equalized image.
[0040] Step 3: Conduct object detection and identification of product information. Classify products through a convolutional neural network to identify the type of each product; use a positioning algorithm to identify the position of the product on the shelf and calculate its coordinates on the shelf; use image analysis to determine the placement status of the product. Use optical character recognition technology to identify product labels, including price, promotion information, and expiration date.
[0041] A convolutional neural network is a network architecture widely used in deep learning for image processing. It extracts and learns image features through multiple layers such as convolutional layers, pooling layers, and fully connected layers. The convolutional layer is the core of the CNN and is responsible for extracting local features from the image. It performs a convolution operation on the input image through a convolution kernel - filter to extract features such as edges, textures, and colors in the image. The pooling layer downsamples to reduce the size of the feature map, reducing computational complexity and preventing overfitting. After the convolutional and pooling layers, the CNN usually contains one or more fully connected layers, which are used to map the extracted features to classification labels. Using the features and classification labels, combined with the KNN algorithm, the features are associated with the products to identify the type of each product. The positioning algorithm is used to identify the position of the product on the shelf and calculate its coordinates on the shelf, usually involving multiple technologies and steps. First is image preprocessing, using the median filtering algorithm to remove image noise, enhance the brightness and contrast of the image, and convert the image into a black - and - white binary image through thresholding operations to highlight the product contour. Then use YOLO image analysis to extract useful information from the image and output the bounding box of the product, that is, the upper - left coordinate, width, and height. Next, establish a shelf model. Obtain the 3D model or 2D planar coordinates of the shelf through images captured by a camera or through lidar and depth sensors. The coordinate system of the shelf is usually a two - dimensional coordinate system, where the x - axis represents the horizontal direction and the y - axis represents the vertical direction. Determine the position and perspective of the camera through calibration methods so that the pixel points in the image are mapped to the shelf coordinate system. Finally, use perspective transformation to convert the pixels in the image into coordinates on the actual shelf. Image analysis combines the convolutional neural network and the positioning algorithm. After comprehensive comparison, it gives the placement status of the product.
[0042] Optical character recognition technology, i.e., OCR technology, usually includes steps such as image preprocessing, character segmentation, feature extraction, character recognition, and post-processing. Its specific implementation is as follows: First is image preprocessing, which includes grayscale conversion, i.e., converting a color image into a grayscale image; binarization, i.e., converting the grayscale image into a black-and-white image; denoising, i.e., using methods such as median filtering for denoising; skew correction, i.e., detecting the skew angle of the text through the Hough transform method and performing correction; edge enhancement, i.e., using the Laplacian operator for edge detection. Secondly is character segmentation. Line segmentation is to segment the text into lines by analyzing the blank areas in the image; character segmentation is to segment the characters in each line. Then is feature extraction, using a convolutional neural network for feature extraction. Then is character recognition, using KNN to map the characters in the image to the corresponding text based on the extracted features. Finally is post-processing, performing spelling correction through a dictionary, and correcting errors by comparing the words output by OCR with the words in the dictionary. Through these means, the text information above the label in the image is finally output.
[0043] Fourth step: Detect the out-of-stock situation, placement, and labels of the goods.
[0044] Based on the output placement status of the goods, where the placement status includes normal, empty, or half-full, identify the out-of-stock or vacant goods on the shelf and record the out-of-stock goods. Compare the type of goods, the coordinates of the goods, and the standard shelf according to the output to determine whether the goods are placed incorrectly and whether they exceed the boundary of the shelf, and record the incorrectly placed goods. Compare the label information in the image with the goods information in the background database to detect whether the label is correct and meets the standards. Record the goods with incorrect labels.
[0045] Fifth step: Conduct sales data analysis and generate optimization suggestions.
[0046] Analyze the sales situation of different goods, including sales volume, turnover rate, etc., to identify which goods are hot-selling, which goods are slow-selling, and which goods are associated goods. Based on the analysis of the sales situation of the goods, prioritize the display of hot-selling goods in prominent positions and place associated goods together. Generate reasonable replenishment time and quantity according to historical sales data and the ARIMA prediction model.
[0047] The ARIMA model models the dependencies in time series data by combining autoregressive - AR, differencing - I, and moving average - MA components for prediction and modeling; it is represented as ARIMA(p, d, q), where p is the order of the autoregressive term, d is the order of differencing used to make a non - stationary time series stationary, and q is the order of the moving average term. First is autoregression, which refers to the linear relationship between the current value and past values or lagged values in the model. The core idea is that the current value can be predicted by the observed values at previous time points; the differencing part is used to eliminate the non - stationarity of the time series. Non - stationarity usually manifests as the mean, variance, or covariance of the data changing over time. The purpose of differencing is to make the time series stationary by subtracting adjacent observations; the moving average part reflects the relationship between the error term and its past values. Through the moving average model, past errors are used to explain the observed values of the current time series.
[0048] The training steps first check whether the time series is stationary; if the data is non - stationary, make the data stationary through differencing operations. Through one or more differencing processes, make the mean and variance of the data stable; determine the p, d, q parameters. Select the orders p and q of the AR and MA parts through the autocorrelation function - ACF and partial autocorrelation function - PACF. The parameter d is determined through differencing operations, indicating the stationarity of the sequence; fit the ARIMA model, use the selected parameters p, d, q to construct the ARIMA model, and estimate the parameters of the model through the method of least squares or maximum likelihood estimation MLE; finally is prediction, use the trained ARIMA model for prediction to predict future values.
[0049] Step 6: Generate a report with the results of anomaly detection, data analysis results, and optimization suggestions, push it to the store management system and the headquarters monitoring platform, and notify relevant personnel via text message or email.
[0050] Compared with manual inspections, which require checking the shelf goods one by one, are time - consuming and laborious, and have extremely low efficiency, especially when the number of stores is large. By using deep learning algorithms such as YOLO and Faster R - CNN, high - precision commodity recognition and shelf status analysis are achieved; the entire process from image acquisition, recognition to anomaly report generation is automatically completed by the system, completely getting rid of manual intervention. This makes high - frequency inspections possible, with a fast response speed, supports real - time feedback, and the efficiency is increased several times.
[0051] Compared with manual inspections, the cost is high, including the salaries of inspection personnel, transportation costs, and time losses. By automating the inspection tasks, inspection personnel only need to operate equipment to collect images or use inspection robots to perform tasks; this significantly reduces labor and transportation costs, and the shortened inspection cycle reduces the overall overhead.
[0052] Compared with manual inspections, subjective biases or omissions cannot be completely avoided. By using a deep learning model to train a standard shelf display model, anomaly detection is automatically completed by comparing with the standard model; the system can update the shelf standard model in real time to dynamically adapt to the actual layout changes of the store; significantly improving the accuracy of inspections.
[0053] Compared with manual inspections, there is a lack of data analysis capabilities, only basic inspection tasks are completed, and no support for optimizing shelf management can be provided; the analysis of the correlation between product display and sales volume relies on manual experience, and the shelf utilization rate cannot be systematically improved. By combining historical sales data with inspection data, the correlation between product placement and sales volume is mined to optimize the display strategy; personalized replenishment suggestions are provided to ensure that popular products occupy prominent positions first, increasing the sales conversion rate; effectively improving the overall operational efficiency and sales volume of the store.
[0054] Compared with manual inspections, real-time feedback cannot be achieved, and the lag in information leads to slow problem handling; there is a lack of comprehensive report support, which cannot help the headquarters or the store make quick decisions. Based on automatically generating an inspection report containing inspection results, anomaly points, and optimization suggestions, and pushing it to the management platform in real time; supporting the headquarters to monitor the shelf status of the store in real time and making quick decisions on replenishment and adjustments.
[0055] In summary, the shelf inspection algorithm based on image recognition technology and big data mining technology solves the problems of high cost, low efficiency, and subjectivity existing in traditional manual inspections and existing technical solutions. Through data optimization, the shelf management level and sales efficiency of the store are improved.
[0056] The above are only specific embodiments of the present invention, but the structural features of the present invention are not limited thereto. Any changes or modifications made by those skilled in the art within the scope of the present invention are covered by the patent scope of the present invention.
Claims
1. A shelf inspection algorithm based on image recognition technology and big data mining technology, characterized in that The steps are as follows: Step 1: Collect data on shelf display pictures and product pictures to obtain product data based on pictures; Step 2: Preprocess the acquired data image and use median filtering and mean filtering to perform denoising algorithms to reduce noise interference in the image; Step 3: Target detection and product information identification; Step 4: Check the out-of-stock, placement and labeling of goods; Step 5: Analyze sales data and generate optimization suggestions; Step 6: Generate a report with the results of anomaly detection, data analysis and optimization suggestions, push it to the store management system and headquarters monitoring platform, and notify relevant personnel via SMS or email.
2. The shelf inspection algorithm based on image recognition technology and big data mining technology according to claim 1 is characterized by: Through the store's AI weighing equipment, images of each product are obtained for product type identification; through the store's mobile robots and cameras, images of the store's shelves are obtained, including the placement of the shelves, the products on display, the density of the shelves, and the labels on the shelves.
3. The shelf inspection algorithm based on image recognition technology and big data mining technology according to claim 1 is characterized by: The histogram equalization or contrast-limited adaptive histogram equalization algorithm is used to improve the clarity of the image and make the edges of the products more obvious.
4. The shelf inspection algorithm based on image recognition technology and big data mining technology according to claim 3 is characterized by: Histogram equalization is an image processing method that adjusts the grayscale distribution of an image to make the histogram of the image as evenly distributed as possible, thereby enhancing the contrast of the image. Its specific implementation is as follows: first, the grayscale histogram of the original image is calculated, that is, the number of pixels at each grayscale level in the image is counted; secondly, the cumulative distribution function (CDF) is calculated. The cumulative distribution function is the accumulation of the grayscale histogram, and the cumulative probability of pixels at each grayscale level is calculated. Where P(i) is the pixel probability of gray level i, k is the current gray level; then calculate the new gray value, and map the gray value of each pixel to the new gray value according to the cumulative distribution function Where CDF min is the minimum non-zero cumulative probability value, N is the total number of pixels in the image, and L is the total number of gray levels; finally, the new gray value is applied to replace the gray value of each pixel in the original image with the calculated new gray value to obtain the equalized image.
5. The shelf inspection algorithm based on image recognition technology and big data mining technology according to claim 1 is characterized by: The convolutional neural network is used to classify the goods and identify the type of each product. The positioning algorithm is used to identify the position of the goods on the shelf and calculate its coordinates on the shelf. The image analysis is used to determine the placement status of the goods. The convolutional neural network is a network architecture widely used in image processing in deep learning. It extracts and learns the features of the image through multiple layers such as convolutional layers, pooling layers, and fully connected layers. The convolutional layer is the core of CNN and is responsible for extracting local features from the image. It performs convolution operations on the input image through convolution kernels-filters to extract features such as edges, textures, and colors in the image. The pooling layer reduces the size of the feature map by downsampling, reducing the amount of calculation and preventing overfitting. After the convolution and pooling layers, CNN usually contains one or more fully connected layers to map the extracted features to classification labels, and use the features and classification labels in combination with the KNN algorithm to associate the features with the goods, thereby identifying the type of each product. The positioning algorithm is used to identify the location of the goods on the shelf and calculate its coordinates on the shelf, which usually involves multiple technologies and steps. The first is image preprocessing, using the median filter algorithm to remove image noise, enhance the brightness and contrast of the image, and convert the image into a black and white binary image through a thresholding operation to highlight the outline of the product; then use YOL0 image analysis to extract useful information from the image and output the bounding box of the product, that is, the coordinates of the upper left corner, width and height; then build a shelf model, obtain the 3D model or 2D plane coordinates of the shelf through the image taken by the camera or through the laser radar and depth sensor. The coordinate system of the shelf is usually a two-dimensional coordinate system, where the x-axis represents the horizontal direction and the y-axis represents the vertical direction. The position and viewing angle of the camera are determined by the calibration method, so that the pixels in the image are mapped to the shelf coordinate system; finally, the perspective transformation is used to convert the pixels in the image into the actual coordinates on the shelf; image analysis is to combine the convolutional neural network and the positioning algorithm, and after comprehensive comparison, give the placement status of the product.
6. The shelf inspection algorithm based on image recognition technology and big data mining technology according to claim 5 is characterized by: Optical character recognition technology is used to identify product labels, including price, promotional information and expiration date. Optical character recognition technology, also known as OCR technology, usually includes steps such as image preprocessing, character segmentation, feature extraction, character recognition and post-processing. Its specific implementation is as follows: First, image preprocessing, including grayscale conversion of color images to grayscale images, binarization conversion of grayscale images to black and white images, denoising using methods such as median filtering, bevel correction using the Hough transform method to detect the tilt angle of the text and correct it, and edge enhancement using the Laplace operator for edge detection; second, character segmentation, line segmentation is to divide the text into lines by analyzing the blank areas in the image; character segmentation is to divide the characters in each line; then feature extraction, using convolutional neural networks to extract features; then character recognition, using KNN to map the characters in the image to the corresponding text through the extracted features; finally, post-processing, spelling correction through the dictionary, and error correction by comparing the words output by OCR with the words in the dictionary; through these means, the text information on the label in the image is finally output.
7. The shelf inspection algorithm based on image recognition technology and big data mining technology according to claim 1 is characterized by: According to the output product placement status, which includes normal, empty or half-full, out-of-stock or vacant products on the shelf are given, and out-of-stock products are recorded; according to the output product type and product coordinates and standard shelf comparison, it is determined whether the product is placed incorrectly or exceeds the shelf boundary, and the incorrectly placed products are recorded; the label information in the image is compared with the product information in the background database to detect whether the label is correct and meets the standard. Products with incorrect labels are recorded.
8. The shelf inspection algorithm based on image recognition technology and big data mining technology according to claim 1 is characterized by: Analyze the sales of different commodities, including sales volume, turnover rate, etc., to identify which commodities are hot-selling, which commodities are slow-selling, and which commodities are related commodities.
9. The shelf inspection algorithm based on image recognition technology and big data mining technology according to claim 8 is characterized by: According to the sales analysis of the goods, the hot-selling goods are displayed in a prominent position and related goods are placed together; according to the historical sales data and ARIMA forecasting model, reasonable replenishment time and quantity are generated; the ARIMA model combines autoregression-AR, difference-I and moving average-MA components to model the dependency in time series data for prediction and modeling; it is expressed as ARIMA (p, d, q), where p is the order of the autoregression term, d is the difference order, which is used to make the non-stationary time series become a stationary series, and q is the order of the moving average term. First is autoregression, which refers to the linear relationship between the current value and the past value or lagged value in the model. The core idea is that the current value can be predicted by the observations of the previous moments; the difference part is used to eliminate the non-stationarity of the time series; non-stationarity is usually manifested as the mean, variance or covariance of the data changing over time. The purpose of the difference is to make the time series stable by subtracting adjacent observations; the moving average part reflects the relationship between the error term and its past value. Through the moving average model, the past error is used to explain the observations of the current time series.
10. The shelf inspection algorithm based on image recognition technology and big data mining technology according to claim 9 is characterized in that: The training step first checks whether the time series is stationary; if the data is non-stationary, the data is made stationary through differential operation, and the mean and variance of the data are stabilized through one or more differential processes; the p, d, and q parameters are determined, and the orders p and q of the AR and MA parts are selected through the autocorrelation function-ACF and the partial autocorrelation function-PACF. The parameter d is determined through differential operation, which indicates the stationarity of the series; Fit the ARIMA model, use the selected parameters p, d, q to build the ARIMA model, and estimate the parameters of the model by the least squares method or maximum likelihood estimation MLE method; finally, predict, use the trained ARIMA model to predict future values.
Citation Information
Cited By
Process image recognition method based on YOLO model
CN121074870A
A process image recognition method based on a YOLO model
CN121074870B
Goods shelf inspection method and system based on machine vision and storage medium
CN121640356A