Intelligent label management method and system based on image recognition
Through the intelligent label management method based on image recognition, the information on clothes on the shelves is recognized and updated in real time, solving the problem of lagging information update under the traditional label management method, and achieving efficient and accurate product information display and management.
Patent Information
- Application Number
- CN202510219628.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
AI Technical Summary
Traditional label management methods are difficult to achieve dynamic management of product information in a fast fashion environment, resulting in lagging label information updates, affecting customers' shopping experience, and increasing store operation workload.
Using an intelligent tag management method based on image recognition, the image information of clothes on the shelf is obtained through the camera, the characteristics of clothes are identified and the brand model is determined, and the product information on the electronic display screen is updated in real time.
Real-time dynamic management of product information is realized, the accuracy and real-time nature of information display are improved, the workload of manual verification is reduced, and retail operation efficiency and customer shopping experience are improved.
Smart Images

Figure CN120147720A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of label management, and in particular, to an intelligent label management method and system based on image recognition. Background Art
[0002] With the rapid development of the fast fashion industry, the clothing retail market is characterized by a high frequency of product updates and a fast inventory turnover rate. In fast fashion brand stores, such as Uniqlo, ZARA, etc., the clothing in the store is usually displayed on the shelves in a stacked form, and customers can select their favorite products by themselves. To facilitate customers to understand product information, corresponding product labels are usually equipped in front of the shelves to display relevant information such as the brand, model, size, price, and material of the clothing.
[0003] However, in practical applications, the traditional label management method has many deficiencies and is difficult to meet the demand for dynamic management of product information in the fast fashion environment. First, the clothing is changed frequently, and the store staff needs to manually adjust the corresponding product labels. However, in actual operation, the label update often has a lag, resulting in the label information not being synchronized with the actual product in real time, which affects the shopping experience of customers. Second, during the selection process, customers often move or stack the clothing in other positions, resulting in the correspondence between the label and the clothing being disrupted. Even if the store staff replenishes the products in time, it is difficult to fully ensure that each piece of clothing matches the corresponding label accurately, resulting in incorrect display of product information.
[0004] In addition, traditional labels rely on manual placement and management, which not only increases the workload of store operation but also easily causes incorrect label placement due to operational errors. For example, when replenishing goods or changing the display, the label fails to correctly correspond to the newly stocked product, resulting in a situation where when customers make a purchase based on the label information, they find that the actual clothing does not match what they expected, thus affecting the shopping decision and even causing disputes at the checkout. Summary of the Invention
[0005] In order to achieve dynamic management of product information and significantly improve the efficiency of label management, the present application provides an intelligent label management method and system based on image recognition.
[0006] In the first aspect, the present application provides an intelligent label management method based on image recognition, adopting the following technical solution:
[0007] An intelligent label management method based on image recognition includes the following steps:
[0008] S1. Obtain image information based on a camera facing the shelf, wherein the image contains a number of stacks of clothing placed on the shelf;
[0009] S2. Obtain the characteristics of the clothing based on image recognition of the number of buttons, button styles, the ratio of button diameter to button spacing, collar styles, fabric colors, and fabric patterns of the clothing, so as to determine the brand model of the clothing;
[0010] S3. Obtain the horizontal dimension of the clothing based on image recognition and generate a horizontal cross-section of the clothing. Here, the horizontal dimension of the clothing is the dimension along the length direction of the shelf when folded on the shelf, and the horizontal cross-section is the cross-section of the clothing in the direction of facing the shelf;
[0011] S4. Identify the center of the horizontal cross-section, determine the corresponding position on the electronic display screen based on the center of the horizontal cross-section, and divide the corresponding display area on the electronic display screen based on this corresponding position and the horizontal dimension;
[0012] S5. Obtain the corresponding product information in the database based on the recognized brand model of the clothing and display it in the corresponding display area.
[0013] Optionally, S2 includes the following steps:
[0014] S21. Identify and segment each stack of clothing on the shelf;
[0015] S22. Perform instance segmentation on the top clothing of each stack of clothing and use it as the target clothing;
[0016] S23. Segment the buttons, collars, and other parts on the target clothing;
[0017] S24. Identify the vertical row of buttons below the collar, identify the style and number of the vertical row of buttons, obtain the adjacent button spacing and button diameter in the image, and calculate the ratio of the button diameter to the adjacent button spacing in the image;
[0018] S25. Identify the fabric color and fabric pattern of the other parts of the target clothing;
[0019] S26. Input the number of vertical rows of buttons, button styles, the ratio of button diameter to button spacing, collar styles, fabric colors, and fabric patterns of the clothing as recognition information into a pre-trained machine model and obtain the brand model of the clothing.
[0020] Optionally, the training steps of the pre-trained machine model include:
[0021] S261. Enter the clothing recognition information of each type of clothing. Here, the clothing information includes the style and number of vertical rows of buttons on the clothing, the ratio of adjacent button spacing to button diameter, fabric color, and fabric pattern;
[0022] S262. Use the recognition information of each type of clothing as input features and the corresponding brand model of the clothing as the output result;
[0023] S263. Normalize the input image to ensure the consistency of the mean and variance of the input data;
[0024] S264. Divide the dataset into a training set and a validation set;
[0025] S265. Input the training images of the training set into an existing neural network model and finally output the prediction results;
[0026] S266. Use a loss function to calculate the error between the prediction results and the output results;
[0027] S267. Backpropagate the error by calculating the gradient and adjust the weights and biases in the network to reduce the loss;
[0028] S268. Use an optimizer to update the network parameters to gradually reduce the value of the loss function;
[0029] S269. Evaluate the classification accuracy, loss value, and confusion matrix of the model on the validation set.
[0030] Optionally, the S3 includes the following steps:
[0031] S31. Perform geometric correction and color standardization on the shelf image;
[0032] S32. Detect and extract the edges and contours of the clothing stacks;
[0033] S33. Use the bottom edge of the shelf as a length marker and determine the pixel - actual size conversion formula;
[0034] S34. Obtain the horizontal pixel length based on the shelf image and convert it to the actual size for the horizontal dimension;
[0035] S35. Perform three - dimensional reconstruction of the horizontal cross - section to obtain the horizontal cross - section of the top clothing and the horizontal cross - section of the non - top clothing.
[0036] Optionally, the S4 includes the following steps:
[0037] S41. Perform vertex sampling of the cross - section polygon;
[0038] S42. Calculate the geometric center coordinates based on the polygon vertices;
[0039] S43. Obtain the mapping relationship between the preset spatial coordinates and the display space of the display screen and determine the center coordinates of the display area;
[0040] S44. Based on the horizontal dimension, extend from the center of the display area to both sides to obtain the boundaries of the display area, thereby obtaining the display area.
[0041] Optionally, the existing neural network model includes:
[0042] An input layer for receiving input features;
[0043] A multi-modal fusion layer for integrating visual and parametric features;
[0044] A fully connected classification layer for weighting different features to generate a probability distribution of brand models, and obtaining a predicted brand model based on the Softmax function;
[0045] An output layer for outputting brand models.
[0046] Optionally, S5 includes the following steps:
[0047] S51. Obtain the product information corresponding to the top clothing in the database and display it in the corresponding display area; wherein, the product information includes the product name, brand, price, and / or prompt displayed in separate lines;
[0048] S52. Identify whether the clothing color and pattern of the horizontal cross-section of the top clothing match the clothing color and pattern of the horizontal cross-section of the non-top clothing; if not, prompt in the display area that the upper and lower layer clothing does not match.
[0049] In a second aspect, the present application provides an intelligent label management method based on image recognition, adopting the following technical solution:
[0050] An intelligent label management method based on image recognition, including a processor, and a program of the intelligent label management method according to any one of the above is running in the processor.
[0051] In a third aspect, the present application provides a storage medium, adopting the following technical solution:
[0052] A storage medium stores a program of the intelligent label management method according to any one of the above.
[0053] In summary, the present application includes at least one of the following beneficial technical effects:
[0054] 1. Through the linkage of image recognition and the display screen, the system can capture the features of the clothing on the shelf in real time and automatically update the corresponding product information on the electronic display screen. When the clothing is replaced, moved, or restocked, the system can quickly identify the changes and synchronously adjust the brand, model, price, and promotion information on the display screen, avoiding the problems of misalignment and misguidance caused by untimely information update in traditional label management, and significantly improving the accuracy and real-time nature of product information display.
[0055] 2. By comparing the folded edges of the upper and lower layers of clothing through color histograms and texture analysis, it is possible to effectively detect whether there is a mismatch between the top layer of clothing and the lower layer. If there are differences in color, pattern, or material, the system will issue a real-time prompt on the display screen to remind customers and store clerks to verify the product information. This automated matching and verification mechanism not only improves the accuracy of product information management but also reduces the workload of manual checking, further enhancing the efficiency of retail operations and the shopping experience of customers. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 is a flowchart of an intelligent label management method based on image recognition. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] The following describes in detail the embodiments of the present application, and examples of the embodiments are shown in the drawings.
[0058] In the description of this specification, the description with reference to the terms "certain embodiments", "one embodiment", "some embodiments", "schematic embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiments or examples. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0059] The embodiments of the present application disclose an intelligent label management method based on image recognition, with reference to Figure 1 , the method includes the following steps S1 - S5.
[0060] S1. Obtain image information based on a camera facing the shelf, where the image contains several stacks of clothing placed on the shelf.
[0061] The camera periodically captures static images of the front of the shelf or transmits continuous frame images in real time through a video stream. The camera is usually installed in front of or on top of the shelf to ensure that the stacked clothing on the shelf surface can be clearly captured. Each frame of the image will go through a preliminary image processing process, including grayscale conversion, noise suppression, and edge enhancement, to improve the image quality. For example, Gaussian Blur is used to smooth the image to reduce detail interference, or the Canny edge detection algorithm is used to highlight the clothing contour. On this basis, each pixel point captured by the camera will be assigned a grayscale value or an RGB color value to form a three-dimensional matrix, facilitating subsequent feature extraction and recognition.
[0062] It should be noted that the camera is aligned with the shelf at an angle from top to bottom diagonally, and can completely capture the upper surfaces of the clothes on each layer of the shelf. The clothes on the shelf are all in a standard folded state. The two sleeves and the hem of the clothes are neatly folded to the back, and the chest position and collar area of the clothes are exposed at the front. And there are usually vertical collar buttons under the collar. If it is a shirt, there are also vertical body buttons at the same time. Each stack of clothes contains several pieces, and each piece of clothing is independently arranged in the above folding manner, so that only the folded edge part of the lower clothes is exposed at the front end. During the camera shooting process, the front appearance features of the topmost clothes can be clearly captured, including the collar, buttons and fabric patterns, and at the same time, the folding edges of each piece of clothing can be recognized, providing complete and continuous image information for subsequent feature extraction and recognition processing.
[0063] In addition, a long strip-shaped electronic display screen is provided on each layer of the shelf. The screen extends along the length direction of the shelf and is arranged parallel to the front edge of the shelf. The electronic display screen has a flexible display function and can dynamically display different contents according to actual needs, including product names, brands, prices, sizes, materials and promotional information, etc. The display area of the screen can be divided according to the actual placement of the clothes on the shelf, and the product information of the corresponding clothes is displayed in real time in the corresponding area, so as to achieve the accurate correspondence between the product information and the shelf display. When the clothes are replaced, moved or sold out, the content on the display screen can be updated in real time to ensure that the displayed information is always consistent with the actual clothes.
[0064] S2. Obtain the features of the clothes, such as the number of buttons, button styles, the ratio of button diameter to button spacing, collar styles, fabric colors and fabric patterns, based on image recognition, so as to determine the brand model of the clothes.
[0065] In this step, the feature information of the clothes is obtained based on image recognition, and these feature information are used to determine the brand and model of the clothes. Since the clothes on the shelf are usually stacked, it is difficult to apply the traditional management methods based on barcodes or RFID tags. Therefore, computer vision technology needs to be used to automatically extract the key features of the clothes and use deep learning models for classification and recognition.
[0066] Specifically, in a certain embodiment, S2 includes the following steps S21-S26.
[0067] S21. Identify and segment each stack of clothes on the shelf.
[0068] First, preprocess the original image obtained by the camera, including grayscale conversion, histogram equalization, and noise suppression. Grayscale conversion converts the color image into a single-channel grayscale image to reduce computational complexity; histogram equalization is used to enhance contrast and make the boundaries of different clothes clearer; noise suppression usually uses a Gaussian filter to reduce random noise in the image by smoothing the image. After the preprocessing is completed, edge detection needs to be performed on the image to initially identify the boundaries of each stack of clothes. The commonly used method is Canny edge detection, and its basic principle is to highlight the edge positions by calculating the gradient of the image. The formula is as follows: where G x and G y are the gradient values in the horizontal and vertical directions respectively, and G is the edge intensity.
[0069] After edge detection is completed, the system will use an instance segmentation algorithm to segment each stack of clothes. The commonly used deep learning model is Mask R-CNN. During the inference process, the input shelf image first undergoes feature extraction through a convolutional neural network (such as ResNet) to obtain feature maps of different scales. Then, the region proposal network generates candidate regions that may contain clothes. Next, the candidate regions are mapped to the feature map of a fixed size through the ROIAlign operation. Finally, the segmentation result is generated through the mask branch.
[0070] After instance segmentation is completed, the regions of each stack of clothes will be marked and assigned a unique instance ID. Taking an actual scenario as an example, assume there are three stacks of T-shirts on the shelf. After instance segmentation, the system will output three corresponding segmentation mask maps, respectively marking the boundaries and pixel regions of each stack of T-shirts. The instance ID of each stack of T-shirts will be associated with its position, size, and color information, facilitating subsequent feature extraction and brand model recognition.
[0071] S22. Perform instance segmentation on the top clothes of each stack of clothes and use them as the target clothes.
[0072] Instance segmentation is usually implemented through deep learning models such as Mask R-CNN. Based on detecting the boundaries of each stack of clothes, the model will generate a pixel-level mask (Mask) for each instance, identifying all the pixels belonging to that instance. At this time, the model will analyze the different depths within each stack of clothes region and preferentially extract the topmost instance. The judgment criterion for the top clothes is usually based on depth sorting and overlapping relationships, that is, among the same instance set, the instance located in the foreground is regarded as the top clothes.
[0073] During the instance segmentation process, the input shelf image first undergoes feature extraction through a convolutional neural network (such as ResNet or EfficientNet) to obtain multi-scale feature maps. Subsequently, the Region Proposal Network (RPN) generates candidate regions, and these regions are mapped to feature maps of a fixed size through the ROIAlign operation. Finally, pixel-level masks are generated through the Mask branch, and a confidence score is given to each instance. According to the confidence ranking, the instance with the highest score and the most forward position is selected as the top clothing. Specifically, it can be completed through the following logical steps:
[0074] 1. Extract all instances and their masks of each stack of clothing from the instance segmentation results.
[0075] 2. Calculate the Z-axis depth of each instance, which can be achieved through occlusion relationships or multi-view reconstruction.
[0076] 3. Sort the instances according to depth and confidence, and select the most foreground instance as the top clothing.
[0077] 4. Generate a segmentation mask for the top clothing, and label the coordinates, dimensions, and pixel region of this instance as "target clothing".
[0078] For example, assume there are three T-shirts stacked on a shelf, with colors red, blue, and white respectively. After instance segmentation, the system identifies three overlapping instances, and through depth-first sorting, it determines that the red T-shirt is on the top. At this time, the segmentation mask of the red T-shirt will be marked as the target clothing and enter the next feature extraction process.
[0079] S23. Segment the buttons, collars, and other parts on the target clothing.
[0080] S24. Identify the vertical row of buttons below the collar, identify the style and quantity of the vertical row of buttons, obtain the spacing between adjacent buttons and the button diameter in the image, and calculate the ratio of the button diameter to the spacing between adjacent buttons in the image.
[0081] Through a convolutional neural network, the input image is transformed into multi-scale feature maps. These feature maps contain detailed information about the clothing surface, such as edges, textures, and color distributions. Next, candidate regions are generated through the Region Proposal Network (RPN). These regions correspond to different functional parts of the clothing, such as the collar, button area, and main fabric part. The working principle of the RPN is to generate multiple candidate boxes based on the feature density in the image and screen the regions most likely to contain the target parts according to the confidence scores.
[0082] Once the candidate regions are generated, the system will further map these regions to a fixed-size feature map through the ROIAlign operation to perform more refined instance segmentation. During the segmentation process, semantic segmentation technology is used to classify each pixel into different functional categories, such as buttons, collars, fabrics, etc. In this process, the cross entropy loss function is usually used to measure the error between the segmentation result and the true label. The formula is as follows:
[0083]
[0084] Among them, y i is the actual label, is the probability value predicted by the model. Through back propagation and gradient descent, the model will continuously optimize the weights to improve the segmentation accuracy.
[0085] For example, assuming that the target clothing is a white shirt, the system first determines the collar area, button area, and fabric area of the shirt through edge detection and feature map analysis. The collar area is usually "U" or "V" shaped, and the system extracts its boundaries through morphological processing (such as opening and closing operations). The button area is identified by a circle detection algorithm (such as Hough circle transform). Suppose three buttons are found during the detection process, and their coordinates are (x 1 ,y 1 )、(x 2 ,y 2 ) and (x 3 ,y 3 ), we can verify the arrangement pattern by calculating the distance between buttons:
[0086]
[0087] If these distances are close within a certain error range, it can be inferred that the clothing item is a standard button-down shirt.
[0088] In this step, each key part of the clothing is instantiated and segmented to provide more recognizable features for brand and model recognition. By analyzing the collar, buttons, fabric and other parts independently, recognition errors caused by clothing deformation, wrinkles or lighting changes can be effectively avoided.
[0089] S25. Identify the cloth color and cloth pattern of the other parts of the target clothing.
[0090] In this step, it is first necessary to extract the color and pattern features of the target clothing area obtained in steps S23 and S24. Color features are usually analyzed through color histograms, which can count the number of pixels with different intensity values in each color channel in the image, thereby reflecting the overall color tone of the clothing.
[0091] In the process of pattern recognition, texture analysis methods are usually adopted, combined with technologies such as gray-level co-occurrence matrix, local binary pattern, and Gabor filter to extract the surface texture features of clothing.
[0092] S26. Input the number of vertical buttons on the clothing, button style, ratio of button diameter to button spacing, collar style, fabric color, and fabric pattern as recognition information into a pre-trained machine model, and obtain the brand model of the clothing.
[0093] This step constructs and trains a deep learning model, taking the structural features of the clothing (such as button style, number, spacing ratio) and appearance features (such as collar style, fabric color, pattern) as inputs, and the brand and model of the clothing as outputs, forming a mapping relationship between input features and recognition results. By continuously optimizing the model weights, fast and accurate recognition of unknown clothing can be achieved.
[0094] In the implementation process, it is first necessary to construct a training dataset, including images of a large number of clothing items of different brands and models and their corresponding feature labels. The feature information of each piece of clothing includes: the number of buttons, the ratio of button spacing to diameter, collar style, fabric color, and fabric pattern, etc. In data preprocessing, all input features will undergo normalization processing.
[0095] The training process of the model is usually based on a deep neural network (DNN) or a convolutional neural network (CNN), and through iterative training of forward propagation and backward propagation, an accurate mapping between features and brand models is achieved. In the forward propagation process, the input layer receives the standardized features, and through non-linear transformations of several hidden layers, higher-level feature representations are gradually extracted. During the training process, the model calculates the error between the predicted result and the true label through a cross-entropy loss function. To ensure the generalization ability of the model, the dataset is usually divided into a training set, a validation set, and a test set. The training set is used to train the model, the validation set is used to tune hyperparameters, and the test set is used to evaluate the performance of the model on unseen data. Commonly used optimization algorithms in the model training process are stochastic gradient descent (SGD) and its variants, such as the Adam optimizer.
[0096] Specifically, in one embodiment, the training steps of the pre-trained machine model include S261 - S269.
[0097] S261. Enter the clothing recognition information for each type of clothing, where the clothing information includes the style and number of vertical buttons on the clothing, the ratio of adjacent button spacing to button diameter, fabric color, and fabric pattern.
[0098] In this step, the recognition information of each piece of clothing is entered. This information includes the key features on the clothing, such as the style and quantity of vertical row buttons, the ratio of the distance between adjacent buttons to the button diameter, the fabric color, and the fabric pattern. These features are used to characterize the differences between clothing of different brands and models. For example, a shirt of brand A may have 4 vertical row buttons, a button diameter of 10 mm, an adjacent button spacing of 25 mm, a diameter-to-spacing ratio of 0.4, a fabric color of light blue, and a pattern of thin stripes; while the shirt of brand B has differences in button ratio, color, or pattern.
[0099] S262. Use the recognition information of each piece of clothing as input features, and the corresponding brand model of the clothing as the output result.
[0100] Use the above recognition information as input features and the corresponding brand model of the clothing as the output result, thus forming the "input-output" pairs in the training dataset.
[0101] S263. Normalize the input image to ensure that the mean and variance of the input data are consistent.
[0102] S264. Divide the dataset into a training set and a validation set.
[0103] S265. Input the training images of the training set into the existing neural network model and finally output the prediction result.
[0104] S266. Use the loss function to calculate the error between the prediction result and the output result.
[0105] S267. Through calculating the gradient, backpropagate the error and adjust the weights and biases in the network to reduce the loss.
[0106] S268. Use the optimizer to update the network parameters to gradually reduce the value of the loss function.
[0107] S269. Evaluate the classification accuracy, loss value, and confusion matrix of the model on the validation set.
[0108] Specifically, the existing neural network model mentioned above includes:
[0109] An input layer for receiving input features;
[0110] A multi-modal fusion layer for integrating visual and parametric features;
[0111] A fully connected classification layer for weighting different features to generate the probability distribution of the brand model and obtaining the predicted brand model based on the Softmax function;
[0112] An output layer for outputting the brand model.
[0113] The input layer is the starting point of the neural network and is mainly used to receive multi-dimensional features of clothing, including structural features (such as the number of buttons, the ratio of button diameter to spacing, collar style, etc.) and visual features (such as fabric color and pattern). These input features are represented in numerical form, with each feature corresponding to a dimension, forming a feature vector. For example, the feature vector of a target piece of clothing can be represented as:
[0114] X = [4, 0.4, 1, 0, 0, 0.8, 0.3]
[0115] Among them, "4" represents the number of buttons, "0.4" represents the ratio of button diameter to spacing, "1, 0, 0" represents the one-hot encoding of the collar style, and "0.8, 0.3" represents the feature weights of color and pattern. The role of the input layer is to normalize these features and pass them to the next layer for further processing.
[0116] Multi-modal fusion refers to representing different types of data sources (such as images and parametric features) in a unified space so that the neural network can learn richer patterns from them. This process is usually achieved through methods such as feature concatenation, weighted summation, or attention mechanisms.
[0117] After completing feature fusion, the model enters the fully connected classification layer. This layer further processes the fused features through several fully connected neurons and generates the probability distribution of the brand models. The core of the fully connected layer is to map the input features to a higher-dimensional feature space through the weight matrix W and the bias term b. The output of each layer can be expressed as:
[0118] h (l) = f(W (l) h (l-1) + b (l) )
[0119] Among them, h (l) is the output of the l-th layer, and f is the activation function, usually the ReLU function, to introduce non-linearity and enhance the model's expressive ability.
[0120] The output of the fully connected layer will pass through the Softmax function to generate the probability distribution of each brand model. The expression of the Softmax function is:
[0121]
[0122] Among them, z k is the score of the k-th brand model, and ∑ j exp(z j ) is the sum of the exponentials of all brand scores. The output of the Softmax function represents the predicted probability of each brand model, and the brand with the highest probability is the final prediction result of the model.
[0123] The output layer is the end point of the model, which is used to receive the output of the Softmax function and generate the final brand model prediction result. In practical applications, assume that the input clothing features are: 4 buttons, a diameter-to-spacing ratio of 0.4, a standard pointed collar style, a dark blue fabric color, and thin stripes as the pattern. After multiple layers of processing by the model, the output brand probability distribution may be as follows:
[0124] P(A) = 0.72, P(B) = 0.18, P(C) = 0.10
[0125] Since the probability of brand A is the highest, the system identifies this piece of clothing as brand A and displays the corresponding product information on the electronic screen of the shelf.
[0126] S3. Obtain the horizontal dimension of the clothing based on image recognition and generate a horizontal cross-section of the clothing, where the horizontal dimension of the clothing is the dimension along the length direction of the shelf when folded on the shelf, and the horizontal cross-section is the cross-section of the clothing in the direction of facing the shelf.
[0127] Specifically, in a certain embodiment, S3 includes the following steps S31 - S35.
[0128] S31. Perform geometric correction and color normalization on the shelf image.
[0129] First, perform geometric correction and color normalization on the shelf image. The purpose of geometric correction is to eliminate the perspective distortion caused by the camera shooting at an angle, so that the positions of the shelf and the clothing in the image are consistent with their actual physical positions. This is usually achieved by perspective transformation. Based on the coordinates of four known corner points, the transformation matrix is calculated to correct the tilted shelf in the input image to a front view. The purpose of color normalization is to eliminate the influence of ambient light changes, usually achieved through white balance algorithms or histogram equalization, to ensure the consistency of the color distribution of images taken at different times.
[0130] S32. Detect and extract the contour of the clothing stack edge.
[0131] This step detects the stack edge of the clothing and extracts the contour. Edge detection usually uses the Canny edge detection algorithm to highlight the clothing boundary by calculating the change in pixel gradients in the image.
[0132] S33. Use the bottom edge of the shelf as a length marker and determine the pixel - actual size conversion formula.
[0133] In this step, the bottom edge of the shelf is used as a length marker, and the conversion relationship between pixels and actual dimensions is determined. The key to this step lies in establishing the pixel-actual dimension conversion formula, which is usually achieved by referring to a calibration object of known length. For example, assuming a ruler with a bottom length of 100 cm on the shelf and the length of this ruler in the image being 2000 pixels, the pixel-actual dimension conversion ratio can be calculated. With this ratio, the actual lateral dimension can be deduced based on the pixel width of the clothing in the image.
[0134] S34. Obtain the lateral pixel length based on the shelf image and convert it into the actual dimension for the lateral dimension.
[0135] In this step, the lateral pixel length of each piece of clothing is obtained based on the shelf image and converted into the actual dimension. By linearly fitting the front folding edge of the clothing, its lateral dimension can be accurately measured. Assume that the folding edge range of a certain piece of clothing in the image extends from x 1 = 300 pixels to x 2 = 900 pixels, then the lateral pixel length is:
[0136] L = x 2 - x 1 = 900 - 300 = 600
[0137] Combined with the pixel-actual dimension conversion ratio, the actual lateral dimension is obtained:
[0138] L = 600 × 0.05 = 30 cm
[0139] S35. Perform three-dimensional reconstruction of the lateral cross-section to obtain the lateral cross-section of the top clothing and the lateral cross-section of the non-top clothing.
[0140] In this step, the lateral cross-section of the clothing is generated through three-dimensional reconstruction methods. Based on the edge detection results of the front view and combined with the depth information of the shelf, the three-dimensional structure of the clothing can be reconstructed through stereovision technology. In the case of no stereo camera, a similar effect can also be achieved through the perspective transformation of a monocular image and known reference dimensions. For example, through the width of the front folding edge and the pixel density in the vertical direction, the thickness of the clothing can be deduced, thereby generating the lateral cross-section.
[0141] S4. Identify the center of the lateral cross-section, determine the corresponding position on the electronic display screen based on the center of the lateral cross-section, and divide the corresponding display area on the electronic display screen based on this corresponding position and the lateral dimension.
[0142] Specifically, in a certain embodiment, S4 includes the following steps S41 - S44.
[0143] S41. Perform vertex sampling of the cross-section polygon.
[0144] First, perform polygon vertex sampling on the horizontal cross-section of the clothing to accurately describe the position and shape of the clothing on the shelf. The horizontal cross-section usually appears as a regular rectangle or an approximately rectangular shape, and vertex sampling can be completed through contour detection and polygon approximation algorithms. In the specific implementation process, the
[0145] Ramer-Douglas-Peucker algorithm is usually used to simplify the clothing edge. The core idea of this algorithm is to remove redundant points by setting a tolerance and only retain the key vertices. Suppose the four vertices of the horizontal cross-section boundary of a piece of clothing in the image coordinate system are (x 1 y 1 ), (x 2 y 2 ), (x 3 y 3 ) and (x 4 y 4 ). The system will generate a polygon based on these vertices for subsequent area positioning.
[0146] S42. Calculate the geometric center coordinates based on the polygon vertices.
[0147] After completing vertex sampling, step S42 calculates the geometric center coordinates based on the vertices of the polygon. The geometric center, that is, the centroid, is a key indicator to describe the position of the polygon and is usually calculated by the average value of all vertex coordinates. For a polygon with n vertices, its geometric center coordinates (x c y c ) calculation formula is:
[0148]
[0149] Taking a quadrilateral as an example, if the vertex coordinates are (100, 200), (300, 200), (300, 400) and (100, 400), then its geometric center coordinates are:
[0150]
[0151] This geometric center coordinate will be used as the initial reference point for the display area of the display screen.
[0152] S43. Obtain the mapping relationship between the preset space coordinates and the display space of the display screen, and determine the center coordinates of the display area.
[0153] In step S43, the system obtains the mapping relationship between the preset spatial coordinates and the display space of the display screen, and further converts the geometric center coordinates into the corresponding positions on the display screen. This mapping relationship is usually implemented through the mapping matrix M, and the matrix M is calculated based on the corresponding relationship between the actual coordinates and the screen coordinates obtained during the calibration process. Assuming that the pixel range corresponding to the length of the shelf is [0, 2000], and the length range of the display screen is [0, 1000], then the mapping relationship is:
[0154]
[0155] If the horizontal coordinate x of the geometric center c = 200 pixels, then the corresponding screen coordinates are:
[0156]
[0157] Through this mapping relationship, the system can correspond the physical position of each piece of clothing to the display position on the display screen one by one.
[0158] S44. Based on the horizontal dimension, extend from the center of the display area to both sides to obtain the boundary of the display area, thereby obtaining the display area.
[0159] Based on the horizontal dimension of the clothing, extend from the center coordinates of the display area to both sides to determine the boundary of the display area. Assuming that the actual horizontal dimension of a certain piece of clothing is 30 cm, and the position of the geometric center on the screen coordinates is x 屏幕 = 100, then the boundary of the display area can be expressed as:
[0160]
[0161] Among them, L is the length of the display area, which depends on the actual size of the clothing and the mapping ratio of the screen. For example, if the mapping ratio is 1:1, then the boundary of the display area is:
[0162] x 左 = 100 - 15 = 85, x 右 = 100 + 15 = 115
[0163] S5. Obtain the corresponding product information in the database based on the recognized brand model of the clothing, and display it in the corresponding display area.
[0164] Specifically, in one embodiment, S5 includes the following steps S51 - S52.
[0165] S51. Obtain the corresponding product information of the top clothing in the database, and display it in the corresponding display area; among them, the product information includes the product name, brand, price and / or prompt displayed in rows.
[0166] Retrieve the corresponding product information from the product database. The database usually contains the complete product description of each piece of clothing, including brand, model, product name, price, material, and promotion tips, etc. For example, assume a light blue pinstriped shirt is identified, and its product information may include: brand "Brand A", model "Spring 2024 model", price "299 yuan", material "100% cotton", tip "New product on the market, 10% discount", etc.
[0167] Once the product information is retrieved, the system will display the product information in real time at the corresponding position on the electronic display based on the display area determined in steps S41 to S44. The display format is usually line-by-line display to ensure that users can read it clearly.
[0168] S52. Identify whether the clothing color and pattern of the horizontal cross-section of the top clothing match the clothing color and pattern of the horizontal cross-section of the non-top clothing; if not, prompt in the display area that the upper and lower layer clothes do not match.
[0169] In step S52, the system further verifies the consistency of the upper and lower layer clothes through the matching of color and pattern. This verification process is based on image processing and feature extraction technologies, mainly completed through color histogram and texture analysis. The color histogram is used to count the pixel distribution of different color channels (such as RGB or HSV) in the clothing image.
[0170] Texture analysis usually uses methods such as gray-level co-occurrence matrix and local binary pattern to extract the texture features of the clothing surface. After comparing the color histograms and texture features of the upper and lower layer clothes, the system will calculate their similarity. If the similarity is lower than the set threshold (such as 0.8), it is considered that the upper and lower layer clothes do not match, and the user will be prompted on the display screen. The prompt content may be: "Attention: The upper and lower layer clothes do not match. Please verify the product information."
[0171] The embodiment of the present application also discloses an intelligent label management method based on image recognition, including a processor, and a program of the intelligent label management method based on image recognition described in any one of the above is run in the processor.
[0172] The embodiment of the present application also discloses a storage medium storing a program of the intelligent label management method based on image recognition described in any one of the above.
[0173] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. An intelligent label management method based on image recognition, characterized in that: The following steps are involved: S1. Acquire image information based on a camera facing the shelf, wherein the image includes a number of stacks of clothes placed on the shelf; S2. Obtaining the number of buttons, button style, ratio of button diameter to button spacing, collar style, fabric color and fabric pattern of the clothing based on image recognition to obtain the characteristics of the clothing, thereby determining the brand model of the clothing; S3. Acquire the transverse dimension of the clothing based on image recognition, and generate a transverse cross section of the clothing, wherein the transverse dimension of the clothing is the dimension along the length direction of the shelf when folded on the shelf, and the transverse cross section is the cross section of the clothing in the direction of looking directly at the shelf; S4. Identify the center of the transverse cross section, and determine the corresponding position on the electronic display screen based on the center of the transverse cross section, and divide the corresponding display area on the electronic display screen based on the corresponding position and the transverse size; S5. Based on the identified brand and model of the clothing, corresponding product information is obtained from the database and displayed in the corresponding display area.
2. The method for managing smart labels based on image recognition according to claim 1, characterized in that: The S2 comprises the following steps: S21. Identify and segment each stack of clothing on the shelf; S22. Segment the top clothing of each pile of clothing as an instance and use it as the target clothing; S23. Segmenting buttons, collars and other parts of the target clothing; S24. Identify the longitudinal buttons below the collar, and identify the style and number of the longitudinal buttons, obtain the spacing between adjacent buttons and the button diameter in the image, and calculate the ratio of the button diameter to the spacing between adjacent buttons in the image; S25. Identify the fabric color and fabric pattern of the other parts of the target clothing; S26. Input the number of longitudinal buttons of the clothing, button style, ratio of button diameter to button spacing, collar style, fabric color and fabric pattern as identification information into the pre-trained machine model to obtain the brand model of the clothing.
3. The method for managing smart labels based on image recognition according to claim 2, characterized in that: The training steps of the pre-trained machine model include: S261. Enter clothing identification information for each type of clothing, wherein the clothing information includes the style and number of longitudinal buttons on the clothing, the ratio of the spacing between adjacent buttons and the button diameter, the color of the fabric, and the pattern of the fabric; S262. Taking the identification information of each clothing item as input feature and the brand and model number corresponding to the clothing item as output result; S263. Normalize the input image to ensure that the mean and variance of the input data are consistent; S264. Divide the data set into a training set and a validation set; S265. Input the training images of the training set into the existing neural network model, and finally output the prediction results; S266. Calculate the error between the prediction result and the output result using the loss function; S267. By calculating the gradient, the error is back-propagated and the weights and biases in the network are adjusted to reduce the loss; S268. Use the optimizer to update the network parameters to gradually reduce the loss function value; S269. Evaluate the classification accuracy, loss value, and confusion matrix of the model on the validation set.
4. The method for managing smart labels based on image recognition according to claim 3, characterized in that: The S3 comprises the following steps: S31. Perform geometric correction and color standardization on the shelf image; S32. Edge detection and contour extraction of clothing stack; S33. Using the bottom edge of the shelf as a length marker and determining a pixel-to-actual size conversion formula; S34. Obtain the horizontal pixel length based on the shelf image to convert it into the actual size for the horizontal size; S35. Perform three-dimensional reconstruction of the transverse cross section to obtain the transverse cross section of the top clothing and the transverse cross section of the non-top clothing.
5. The method for managing smart labels based on image recognition according to claim 4, characterized in that: The S4 comprises the following steps: S41. Perform cross-section polygon vertex sampling; S42. Calculate the geometric center coordinates based on the polygon vertices; S43. Obtain the mapping relationship between the preset spatial coordinates and the display space of the display screen, and determine the center coordinates of the display area; S44. Based on the lateral size, the display area boundary is obtained by extending from the center of the display area to both sides, thereby obtaining the display area.
6. The method for managing smart labels based on image recognition according to claim 3, characterized in that: The existing neural network model includes: Input layer, used to receive input features; Multimodal fusion layer, used to integrate visual and parametric features; The fully connected classification layer is used to weight different features to generate the probability distribution of brand models and obtain the predicted brand model based on the Softmax function; The output layer is used to output the brand model.
7. The method for managing smart labels based on image recognition according to claim 5, characterized in that: The S5 comprises the following steps: S51. Obtain the corresponding product information of the top clothing in the database and display it in the corresponding display area; wherein the product information includes the product name, brand, price and / or prompt displayed in a branch; S52. Identify whether the clothing color and pattern of the transverse cross section of the top clothing match the clothing color and pattern of the transverse cross section of the non-top clothing; if not, indicate in the display area that the upper and lower clothing do not match.
8. An intelligent label management system based on image recognition, characterized in that: It comprises a processor, in which runs a program of the intelligent label management method based on image recognition as described in any one of claims 1 to 7.
9. A storage medium, characterized in that: A program for the intelligent label management method based on image recognition as described in any one of claims 1 to 7 is stored.