Automatic commodity information input method based on hang tag scanning and related device

Through image recognition and machine learning models, tag information is automatically recognized and entered, which solves the problem of identification failure caused by inconvenience in operation and barcode occlusion, and realizes efficient product information entry.

CN120147719APending Publication Date: 2025-06-13HANGZHOU YIKE INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510219591.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

There are problems such as inconvenient operation during the identification of existing tags, failure in identification resulting in barcode occlusion, and inefficient manual operation.

Method used

Automatic product information entry method based on image recognition and machine learning models is adopted, and automatic entry is realized by obtaining the image information of the tag, calling the pre-trained machine model, identifying edge radians, text areas and barcode information.

Benefits of technology

Complete tag scanning and product information entry in a few seconds, avoiding the tedious process of manually searching barcodes and improving operation flexibility and comfort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147719A_ABST
    Figure CN120147719A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of target detection, and discloses an automatic commodity information input method based on hang tag scanning and a related device, and the method comprises the following steps: obtaining the image information of a hang tag; a pre-trained machine model is called, and according to the pre-trained machine model, the edge radian of the hang tag, the proportion of the edge length of the hang tag to the font size, the color of the hang tag and the character area serve as input, and the commodity type serves as output; and inputting the image into the pre-training model to obtain a commodity type, and inputting the commodity type into a current statistical table of the system. According to the application, through image recognition and a machine learning model, the scanning of the hang tag and the entry of the commodity information can be completed within a few seconds, and the tedious process of manually searching a bar code is avoided. A user can complete hang tag recognition by operating the scanning gun with one hand, so that inconvenience of traditional double-hand operation is reduced, and the flexibility and comfort of operation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of object detection, and in particular, to an automatic commodity information entry method and related device based on tag scanning. Background Art

[0002] In the fields of clothing retail, warehousing management, and related commodity management, the identification and sorting of tags are essential links in the commodity management process, directly affecting the efficiency of commodity inventory, shelving, and sales processes. Tags usually contain basic information of commodities, such as brand, model, size, price, and barcodes, etc., which are used to quickly identify commodities and achieve inventory management. However, the existing tag identification and sorting methods have certain limitations, bringing many inconveniences to actual operations.

[0003] In the traditional commodity inventory and placement process, operators usually need to hold a barcode scanner in one hand and scan the barcodes on the tags one by one to check the inventory information and complete the classification of commodities. During this process, usually one hand holds the barcode scanner, and the other hand uses the arm to push aside other clothes on the hanger, and then holds the tag with the palm to scan the tag in the gap between the clothes. However, commodity tags usually exist in multiple forms and are fixed together by perforation to form a tag group. These tags may include various types such as product descriptions, price tags, trademark tags, and anti-theft tags. When scanning the barcode, it is often necessary to separate multiple tags first to find the one containing the barcode.

[0004] In ergonomic research, when users view tag information, the most comfortable gesture is to stagger multiple tags around the tag perforation to present a trident-like arrangement, which is convenient for viewing the text information in the lower half of the tag. However, in reality, the position of the barcode is not fixed. Sometimes it is located in the upper half, which does not match the conventional reading habit. Since the front and back of the tag are not fixed, and the shapes of different tags may vary (i.e., irregular tags), it often causes the operator to be unable to operate the tag with one hand in a narrow space to push aside other tags. In the common grasping action, the barcode position of some tags may be located in the upper half or blocked by other tags, and cannot be fully exposed. Especially for some commodity tags, there are plastic sheets with frosted materials attached, which will block the trademark of the tag and cannot directly use image recognition to identify the trademark. Summary of the Invention

[0005] In order to solve the problems of inconvenient operation, barcode occlusion resulting in recognition failure, and low efficiency of manual operation in the existing tag recognition process, this application provides an automatic commodity information entry method and related device based on tag scanning.

[0006] In a first aspect, the present application provides an automatic product information entry method based on tag scanning, adopting the following technical solution:

[0007] An automatic product information entry method based on tag scanning, comprising the following steps:

[0008] S1. Obtain the image information of the tag;

[0009] S2. Invoke a pre-trained machine model, wherein the pre-trained machine model takes the edge radian of the tag, the ratio of the tag edge length to the font size, the tag color, and the text area as inputs, and takes the product category as the output;

[0010] S3. Input the image into the pre-trained model to obtain the product category and enter it into the current statistical table of the system.

[0011] Optionally, S3 includes:

[0012] S31. Based on image recognition, determine whether there is a complete barcode in the image. If it exists, recognize the barcode information to obtain the product category and enter it into the current statistical table of the system;

[0013] S32. If not, obtain the brand corresponding to the tag in the image based on image recognition;

[0014] S33. Based on the brand, identify the preset text area in the tag and obtain the target text information in the detected preset text area, where the preset text area corresponds to the brand.

[0015] Optionally, S32 includes:

[0016] S321. Based on image recognition, obtain the trademark image in the image and determine the product brand based on the trademark image;

[0017] S322. If the recognition fails, perform image recognition on the tag in the image and perform edge recognition on the recognized tag;

[0018] S323. Calculate the radian of the obtained tag edge, where the tag edge includes the side and the corner;

[0019] S324. Recognize the font area and edge of the tag, obtain the font size in the specific text area, and calculate the ratio of the specific tag edge length to the font size;

[0020] S325. Input the tag edge radian, text area function, tag edge length, and the ratio of the font size in the specific text area into the pre-trained machine model to obtain the brand corresponding to the tag in the image.

[0021] Optionally, the pre-trained machine model includes the following steps:

[0022] Enter the hangtags of each piece of clothing, and obtain the arc of the hangtag edge, the function of the text area, the length of the hangtag edge, and the ratio of the font size in a specific text area through image recognition, and use them as marks;

[0023] Use the arc of the hangtag edge, the function of the text area, the length of the hangtag edge, and the ratio of the font size in a specific text area of the clothing as input features, and the brand of the clothing as the output result;

[0024] Normalize the input image to ensure that the mean and variance of the input data are consistent;

[0025] Divide the dataset into a training set and a validation set;

[0026] Input the training images of the training set into the existing ResNet model, pass through the convolutional layer, pooling layer, and fully connected layer, and finally output the prediction result;

[0027] Use the loss function to calculate the error between the prediction result and the output result;

[0028] By calculating the gradient, backpropagate the error, and adjust the weights and biases in the network to reduce the loss;

[0029] Use the optimizer to update the network parameters to gradually reduce the value of the loss function;

[0030] Evaluate the classification accuracy, loss value, and confusion matrix of the model on the validation set.

[0031] Optionally, the existing ResNet model includes:

[0032] Input layer, used to receive the input features obtained after image preprocessing;

[0033] Initial convolutional layer, used to extract the spatial features of the image;

[0034] Residual block, used to connect the input to the output through a skip connection;

[0035] Global average pooling layer, used to reduce the data dimension through average pooling;

[0036] Fully connected layer, used to input the features extracted by convolution and pooling into the fully connected layer for high-level feature combination and classification;

[0037] Output layer, used to output the brand prediction result and the hangtag edge recognition result.

[0038] Optionally, the S33 includes:

[0039] S331. Call the relative relationship between the preset text area and the reference features in the recognized hangtag based on the recognized brand;

[0040] S332. Use the recognized short side and the corners on both sides as reference features to determine the preset text area of the hangtag in the recognized image, where the recognized short side is the short side far from the center of the punching hole;

[0041] S333. Obtain the text information in the preset text area.

[0042] Optionally, if it is determined during image recognition that there is a complete barcode in the image, confirm that the recognition is successful and emit a first prompt sound; when obtaining the target text information in the detected preset text area, confirm that the recognition is successful and emit a second prompt sound.

[0043] In a second aspect, the present application provides an automatic commodity information entry method based on hangtag scanning, adopting the following technical solution:

[0044] An automatic commodity information entry method based on hangtag scanning, including a processor, and a program of the automatic commodity information entry method according to any one of the above is run in the processor.

[0045] In a third aspect, the present application provides a storage medium, adopting the following technical solution:

[0046] A storage medium stores a program of the automatic commodity information entry method according to any one of the above.

[0047] In summary, the present application includes at least one of the following beneficial technical effects: Through image recognition and machine learning models, the scanning of hangtags and the entry of commodity information can be completed within a few seconds, avoiding the cumbersome process of manual barcode search. The user can complete hangtag recognition by operating the scanning gun with one hand, reducing the inconvenience of traditional two-handed operations and improving the flexibility and comfort of operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a flowchart of a program of an automatic commodity information entry method based on hangtag scanning in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] The following details the embodiments of the present application, and examples of the embodiments are shown in the drawings.

[0050] In the description of this specification, the description with reference to the terms "certain embodiments", "one embodiment", "some embodiments", "schematic embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of this application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0051] An embodiment of this application discloses an automatic commodity information entry method based on tag scanning. Referring to Figure 1 , it includes the following steps S1 - S3.

[0052] S1. Obtain the image information of the tag.

[0053] In this step, the image information of the tag is obtained through an image acquisition device (such as a camera or a scanning gun). The core of this step is to ensure that the obtained image is clear and contains complete tag information, including the edges of the tag, the font area, barcodes, trademarks, etc.

[0054] S2. Invoke a pre - trained machine model, where the pre - trained machine model takes the edge curvature of the tag, the ratio of the tag edge length to the font size, the tag color, and the text area as inputs and the commodity type as the output.

[0055] Specifically, in a certain embodiment, the pre - trained machine model includes the following steps S201 - S209.

[0056] S201. Enter the tags of each type of clothing, and obtain the edge curvature of the tag, the function of the text area, the ratio of the tag edge length to the font size within a specific text area through image recognition, and use them as marks.

[0057] The tag information needs to be entered into the database in advance. First, the tag is captured by an image acquisition device such as a camera or a scanning gun. To ensure the consistency and usability of the entered tag images, the system will perform basic pre - processing on the images, including operations such as removing background noise, enhancing contrast, and edge detection, to ensure that the edge contours, text areas, and font information of the tags can be clearly extracted.

[0058] The edge curvature is calculated by detecting the curvature changes of the sides and corners of the tag, and these changes can reflect the geometric shape of the tag, such as rectangular, rounded rectangular or irregular shape, etc. The function of the text area is to detect and label the text part on the surface of the tag based on optical character recognition (OCR) technology, and to clarify which areas contain product information, price, size and brand name. At the same time, by calculating the font size within a specific text area and taking the ratio with the edge length of the tag, a stable and brand-specific feature set can be obtained.

[0059] For example, in a certain embodiment, when entering the tag of a certain brand of clothing, the system detects that the edge of the tag of this brand is a rounded rectangle, the edge length is about 10 cm, and the font size of the price label area is about 0.5 cm. Through ratio calculation, the specific "edge length: font size" ratio of the tag of this brand is obtained as 20:1.

[0060] S202. Take the edge curvature of the clothing tag, the function of the text area, the edge length of the tag, and the ratio of the font size within a specific text area as input features, and the brand of the clothing as the output result.

[0061] In this step, first, the feature data extracted from the tags of different brands of clothing need to be used as input variables, and the corresponding brand names as output labels to form a training data set. Each piece of training data contains a set of input features and a corresponding brand label. For example, for the tag of a certain brand, the edge curvature, the ratio of the edge length to the font size is within a specific numerical range, and at the same time, the arrangement of the text area is also specific. These information are recorded as input features, and the corresponding brand name is marked as the output result in the data set.

[0062] S203. Normalize the input image to ensure that the mean and variance of the input data are consistent.

[0063] S204. Divide the data set into a training set and a validation set.

[0064] A complete data set usually contains a large number of tag images and their corresponding feature annotations. The common ratios for dividing these data into a training set and a validation set are 8:2 or 7:3, that is, 80% or 70% of the data is used for model training, and 20% or 30% of the data is used for model validation. The training set is used for learning the model parameters, and the validation set is used for performance evaluation during the model training process.

[0065] Taking an actual scenario as an example, suppose there are 10,000 tag images, containing feature data of different brands, shapes, colors, and text areas. In step S204, the system first shuffles the dataset to ensure that the data distribution is random and uniform, avoiding a situation where a certain brand or specific feature occupies too high a proportion in a certain subset. Then, according to an 8:2 ratio, 8,000 images are assigned to the training set and 2,000 images are assigned to the validation set.

[0066] The training set is used to input into the neural network. The model continuously optimizes its internal parameters by learning the correspondence between the tag features and brands in the training set. The validation set is used to evaluate the performance of the model after each round of training. Usually, the performance of the model is measured by calculating metrics such as accuracy and loss value on the validation set.

[0067] S205. Input the training images of the training set into an existing ResNet model. After passing through the convolutional layer, pooling layer, and fully connected layer, the prediction result is finally output.

[0068] The input image first passes through the input layer, which converts the original image data into a tensor format that the model can process. Then, the image enters the initial convolutional layer. This layer scans the image through the convolutional kernel to extract low-level features such as edges, corners, and color blocks. Suppose the input image size is 224*224*3 (RGB three-channel image). After the initial convolutional operation, the dimension of the feature map will change according to the convolutional kernel size, stride, and padding strategy. For example, when using a 3*3 convolutional kernel, a stride of 1, and "same" padding, the dimension of the output feature map remains unchanged.

[0069] During the convolution process, the pixel values at specific positions calculate the feature values through the convolutional kernel weights. The formula is as follows:

[0070]

[0071] Among them, X represents the pixel value of the input image, W is the weight of the convolutional kernel, b is the bias term, and Y(i,j) is the value of the output feature map at position (i,j). Through the convolution operation, local features such as edges, arcs, and text areas of the tag image can be extracted.

[0072] Next, after passing through the pooling layer, the model downsamples the feature map, reducing the data dimension, improving the calculation efficiency, and at the same time retaining the most significant features. Common pooling methods include max pooling and average pooling. In this step, if a 2*2 max pooling kernel with a stride of 2 is used, the size of the feature map will be reduced to half of the original. For example, the input of 224*224 becomes 112*112 after pooling.

[0073] The key innovation of ResNet lies in the introduction of residual blocks, which directly add the input to the output through "skip connections" to alleviate the vanishing gradient problem in deep networks. Assuming the output of the convolutional layer is F(x), the output of the skip connection can be expressed as: y = F(x) + x

[0074] After multiple layers of convolution and pooling, the feature map enters the global average pooling layer, which converts the two-dimensional feature map into a one-dimensional vector by taking the average of all spatial positions for each feature channel. For example, a feature map with an input size of 7*7*512 outputs 1*1*512 after global average pooling.

[0075] Finally, the extracted feature vector is fed into the fully connected layer to map the low-dimensional features to the brand classification space, and the final brand prediction result is generated through the output layer. Assuming the model needs to identify 10 brands, the dimension of the output layer will be 10, and the output of each neuron represents the prediction probability of the corresponding brand, which is usually normalized through the Softmax function:

[0076]

[0077] where, z i is the linear output of the fully connected layer, and P(y i ) is the prediction probability of the i-th brand. Finally, the model takes the category with the highest probability as the final output of brand recognition.

[0078] S206. Calculate the error between the prediction result and the output result using the loss function.

[0079] The model generates a brand prediction result based on the input tag image, and this prediction result is usually output in the form of a probability distribution. For example, after an input image is processed by the ResNet model, the following probability distribution is output:

[0080]

[0081] Assume that the category corresponding to the true label (true brand) is the 4th category, that is, y = [0, 0, 0, 1, 0, 0, 0, 0, 0, 0]. At this time, the probability of the 4th category predicted by the model is 0.7, which deviates from the true label. S207. Calculate the gradient and backpropagate the error to adjust the weights and biases in the network to reduce the loss.

[0082] In classification tasks, the commonly used loss function is the cross-entropy loss. Cross-entropy measures the difference between two probability distributions, and its calculation formula is:

[0083]

[0084] where, yi is the probability distribution of the true label, and is the predicted probability output by the model. According to the above example, the cross-entropy loss is calculated as follows:

[0085] L = -[0×log(0.1) + 0×log(0.05) + 0×log(0.03) + 1×log(0.7) +...] = -log(0.7) ≈ 0.357

[0086] In this step, the system calculates the corresponding loss value for each training sample and takes the average of the samples in the entire batch as the overall loss for the current training epoch. As the training progresses, the model continuously optimizes the weights through gradient descent to reduce the value of the loss function, thereby improving the model's prediction ability.

[0087] S207. Calculate the gradient, backpropagate the error, and adjust the weights and biases in the network to reduce the loss.

[0088] The model first performs forward propagation based on the input data to generate prediction results and calculates the loss between the predicted value and the true label through step S206. At this time, the output of the loss function serves as the basis for measuring the model's performance. Then, the backpropagation algorithm calculates the gradients of the weights and biases in each layer of the network through the loss function and propagates these gradients layer by layer backward until the input layer.

[0089] S208. Use the optimizer to update the network parameters to gradually reduce the value of the loss function.

[0090] The update of the optimizer is based on the gradients calculated in the previous step S207. Suppose a weight parameter is w and its gradient is and the learning rate is α. Then the optimizer updates the weight according to the following formula:

[0091]

[0092] where the learning rate α controls the step size of each update. A too large value will lead to instability in the training process and even non-convergence; a too small value will result in slow training speed and increased training time. Commonly used optimizers include Stochastic Gradient Descent (SGD), Momentum, Adam, etc. Each optimizer has slightly different strategies for gradient calculation and parameter update.

[0093] S209. Evaluate the classification accuracy, loss value, and confusion matrix of the model on the validation set.

[0094] Optionally, the existing ResNet model includes an input layer, an initial convolutional layer, residual blocks, a global average pooling layer, and a fully connected layer. The input layer is used to receive the input features obtained after image preprocessing. The initial convolutional layer is used to extract the spatial features of the image. The residual blocks are used to jump-connect the input to the output. The global average pooling layer is used to reduce the data dimension through average pooling. The fully connected layer is used to input the features extracted by convolution and pooling into the fully connected layer for the combination and classification of high-level features. The output layer is used to output the brand prediction result and the tag edge recognition result.

[0095] The ResNet model is a deep convolutional neural network for image recognition and feature extraction. Its core lies in solving the common problems of gradient vanishing and gradient explosion in deep networks through "residual blocks", thereby achieving more efficient and accurate feature learning. The structure of this model includes an input layer, an initial convolutional layer, residual blocks, a global average pooling layer, a fully connected layer, and an output layer.

[0096] The input layer of the model is responsible for receiving the preprocessed image features. The input images are usually normalized to ensure that the pixel values are within a fixed range (such as [0,1] or [-1,1]). This processing not only helps to accelerate the training convergence speed but also enhances the robustness of the model under different lighting, contrast, and background conditions. For example, when inputting a tag image with three RGB channels and a size of 224*224*3, after passing through the input layer, the image is converted into a tensor form that the model can process.

[0097] The initial convolutional layer extracts features from the input image through several convolutional kernels. The role of the convolutional kernel is to scan the local area of the image and capture low-level features such as edges, corners, and textures. The mathematical formula for the convolution operation is as follows:

[0098]

[0099] Among them, X represents the input feature map, W is the weight of the convolutional kernel, b is the bias term, and Y(i,j) is the value at position (i,j) in the output feature map. Assuming that the initial convolutional layer uses a 3*3 convolutional kernel, a stride of 1, and a padding method of "same", the size of the output feature map remains unchanged.

[0100] The residual block is the core component of ResNet. It directly passes the input to the output end through a "skip connection", thus solving the problem of gradient vanishing in deep networks. Its mathematical expression is as follows:

[0101] y = F(x) + x

[0102] Among them, x is the input feature, F(x) is the output feature of the convolutional layer, and the residual block realizes more effective feature transmission by directly adding the input x to the output.

[0103] The feature map processed by the residual block enters the global average pooling layer (GLOBAL AVERAGE POOLING, GAP). This layer converts the two-dimensional feature map into a one-dimensional vector by taking the average of all spatial positions in each feature channel, thereby reducing the number of parameters and preventing overfitting.

[0104] Next, this low-dimensional feature vector is input into the fully connected layer, which is used to further combine and classify the features extracted by convolution and pooling. In this step, the model generates brand prediction results based on the input feature vector and maps the identified features to different brand categories. For example, the feature vector of a certain input tag may output the following probability distribution after being processed by the fully connected layer: [0.1, 0.05, 0.7, 0.1, 0.05]

[0105] In this example, the prediction probability of the third brand is the highest (70%), and the model takes it as the final brand recognition result.

[0106] Finally, the output layer receives the results of the fully connected layer and generates two main outputs: the brand prediction result and the tag edge recognition result. The brand prediction result is used to confirm the brand to which the commodity belongs, while the tag edge recognition result is used to further analyze information such as the shape, arc, font area, and text size of the tag, supporting subsequent commodity information entry.

[0107] S3. Input the image into the pre-trained model to obtain the commodity category and enter it into the current statistical table of the system.

[0108] Specifically, in one embodiment, S3 includes sub-steps S31 - S33.

[0109] S31. Based on image recognition, determine whether there is a complete barcode in the image. If it exists, identify the barcode information, obtain the commodity category, and enter it into the current statistical table of the system.

[0110] In this step, first, preprocess the input tag image, including steps such as grayscale conversion, binarization, noise removal, and edge detection. The purpose of these processing operations is to enhance the contrast of the barcode area, making it more prominent in the image, thus facilitating subsequent recognition. For example, use a Gaussian filter for smoothing to remove background noise, and then extract the significant boundaries in the image through the Canny edge detection algorithm to initially locate the area where the barcode may exist.

[0111] In the barcode detection stage, a method based on morphological operations and region growing algorithm is adopted. Barcodes usually appear as a series of parallel black and white stripes with certain regularities in width and spacing. Therefore, the connectivity of the barcode can be enhanced through morphological closing operation, and then the regions that conform to the barcode characteristics can be identified through connected component analysis. The closing operation, that is, the operation of first dilating and then eroding, fills the small holes in the barcode region. By calculating the aspect ratio, density and parallelism of these regions, the possible barcode candidate regions are further screened out.

[0112] Once the candidate region is determined, the OCR algorithm is used to perform character recognition on this region. At this time, the key lies in the integrity judgment of the barcode, that is, to ensure that the barcode is not blocked, not bent and the stripes are clear. If there are missing or distorted parts in the barcode, OCR may not be able to generate accurate decoding results, that is, subsequent steps are needed to identify the tag without relying on the barcode.

[0113] S32. If not, the brand corresponding to the tag in the image is obtained based on image recognition.

[0114] In this step, the method of template matching or feature matching is adopted, and the brand is identified by comparing with the trademark templates in the brand database. Taking the ORB (Oriented FAST and Rotated BRIEF) feature extraction algorithm as an example, this algorithm can quickly identify the trademark pattern on the tag through key point detection and descriptor generation, and compare it with the trademark templates in the database. If the matching score exceeds the set threshold, the system can confirm the brand.

[0115] Specifically, in a certain embodiment, S32 includes sub-steps S321 - S325.

[0116] S321. The trademark image in the image is obtained based on image recognition, and the product brand is judged based on the trademark image.

[0117] In the process of extracting the trademark image, the region growing algorithm or the method based on color segmentation is usually adopted to locate the trademark region. These methods separate the region containing the trademark from the background by analyzing the continuity of color, texture and shape in the image.

[0118] Then the system extracts the feature points of the trademark through a key point detection algorithm (such as SIFT, SURF or ORB) and generates the corresponding descriptors. Taking the ORB (Oriented FAST and Rotated BRIEF) algorithm as an example, this algorithm identifies the significant feature points in the trademark through the FAST corner detector and uses the BRIEF descriptor to encode these features. Each feature point can be regarded as a multi-dimensional vector representing the local characteristics of the trademark in space.

[0119] Suppose 100 feature points are extracted from a certain trademark image, and each feature point is represented by a 128-dimensional vector. Then the feature description of this trademark can be expressed in the following matrix form:

[0120]

[0121] In the feature matching stage, the system compares the trademark features extracted from the hangtag with the features of the known trademark templates in the brand database. Common methods include calculating the Euclidean distance and the Hamming distance. Taking the Euclidean distance as an example, the distance d between two feature vectors p and q can be expressed as:

[0122]

[0123] If the number of matching feature points exceeds a set threshold (such as more than 30), and the average distance of the matches is less than a certain preset value (such as 10 pixel units), it is considered that the trademark recognition is successful, and the corresponding brand can be inferred.

[0124] S322. If the recognition fails, perform image recognition on the hangtag in the image, and perform edge recognition on the recognized hangtag.

[0125] It is usually implemented using the Canny edge detection algorithm. This algorithm identifies the boundaries with significant gray-scale changes by calculating the gradients of the pixel gray-scale values in the image. The implementation of the Canny algorithm includes the following steps: The Gaussian filter is used to smooth the image and reduce noise interference; the gradient magnitude and direction of the image are calculated through the Sobel operator; double-threshold processing is applied to divide the edge pixels into strong edges, weak edges, and non-edge regions; finally, the strong edges and eligible weak edges are connected through edge tracking to generate a complete edge contour.

[0126] S323. Calculate the radian of the obtained hangtag edge, where the hangtag edge includes the side and the corner.

[0127] The radian is a geometric quantity reflecting the degree of curve bending, usually achieved by fitting an arc or a quadratic curve. Suppose a section of curve on the hangtag edge is represented by a set of coordinate points (x i , y i ), and the radian can be calculated by fitting an arc using the least squares method. The standard formula for arc fitting is as follows:

[0128] (x - h) 2 +(y - k) 2 =R 2

[0129] Among them, (h, k) are the coordinates of the center of the circle, and R is the radius. By fitting the edge points, the center of the circle and the radius can be solved, thereby calculating the radian of the edge. The radian K can be calculated by the following formula:

[0130]

[0131] If the radius R is smaller, the radian K is larger, indicating a higher degree of curvature of the edge; conversely, if R is larger, the radian is smaller, indicating that the edge is more tending to be straight.

[0132] Taking practical applications as an example, assume that the edges of a certain brand of hangtags usually present a rounded rectangle design, and the radius of its rounded corners is about 10 millimeters, and the corresponding radian is K = 1 / 10 = 0.1. When the system recognizes that the edge radian of a certain hangtag is close to 0.1, it can be inferred that the hangtag is very likely to belong to this brand. And if the radian significantly deviates from this range, it may belong to other brands.

[0133] S324. Identify the font area and edge of the hangtag, obtain the font size within a specific text area, and calculate the ratio of the specific edge length to the font size of the hangtag.

[0134] This step adopts a method based on connected component analysis. By analyzing the pixel regions connected to each other in the binary image, the regions that may contain text are identified. Each connected component will be marked as an independent region and further screened through features such as its area, aspect ratio, and density. For example, the text regions usually present relatively regular rectangles

[0135] Once the candidate text regions are identified, the font sizes within these regions need to be measured next. The font size is usually represented by calculating the height of the text region. In actual operation, the size of the text region can be determined by the way of the bounding box. Assume that the bounding box coordinates of a certain text region are (x 1 , y 1 ) and (x 2 , y 2 ), then the height h of this region can be calculated by the following formula: h = y 2 - y 1 .

[0136] In this step, the system usually analyzes multiple text regions and takes the average height as the representative value of the font size. For example, if the average font height of the price label, size label, and product description regions of a certain hangtag is 8 millimeters, then this value can be used as the typical font size feature of this hangtag.

[0137] The system then measures the edge length of the tag. The edge length refers to the perimeter of the physical contour of the tag and is usually achieved through edge detection and contour extraction algorithms. Canny edge detection is one of the commonly used methods. It identifies boundaries with significant gray-scale changes by calculating the pixel gray-scale gradient. After edge detection, the edge length can be calculated through contour analysis, and the formula is as follows:

[0138]

[0139] where (x i , y i ) are consecutive pixel points on the edge, and L is the total length of the edge.

[0140] After obtaining the font size and the edge length, the system generates key features for brand recognition by calculating their ratio. The ratio calculation formula is:

[0141] Suppose the edge length of a certain tag is 200 millimeters and the font size is 8 millimeters, then its ratio is:

[0142] Tags of different brands often have specific styles in design. These styles are not only reflected in the shape and color but also in the ratio relationship between the text and the edge. For example, a certain brand may prefer a larger font, with a ratio of edge length to font size of 20:1, while another brand may prefer a more compact design with a ratio of 30:1. Therefore, by analyzing this ratio, the brand corresponding to the tag can be further inferred.

[0143] S325. Input the ratio of the tag edge radian, the function of the text area, the tag edge length, and the font size within a specific text area into a pre-trained machine model to obtain the brand corresponding to the tag in the image.

[0144] In this step, the key features extracted in S321 to S324 are standardized and feature-fused. The purpose of standardization is to convert feature values of different scales into the same dimension so that the model can better understand and process the input data. Standardization usually adopts the Z-Score method, and its formula is as follows:

[0145]

[0146] where x is the original feature value, μ is the mean of the feature, σ is the standard deviation, and x′ is the standardized feature value. For example, the ratios of the tag edge length, radian, and font size may vary significantly among different brands. After standardization, these features are mapped to the same numerical range, thus avoiding the deviation caused by inconsistent feature value ranges in model training.

[0147] In the feature fusion stage, the system combines features such as the edge curvature, the position of the text area, the edge length, and the ratio of the font size into a multi-dimensional vector. For example, assuming the extracted features include the edge curvature K, the edge length L, the font size h, the ratio R of the edge to the font, and the text area coordinates (x, y, w, h), the fused input feature vector can be expressed as: X = [K, L, h, R, x, y, w, h].

[0148] Then, these features are input into a pre-trained machine learning model for brand prediction. In this application, ResNet is used as the base model. ResNet solves the problem of gradient vanishing in deep neural networks through residual blocks, thus ensuring the stability of training and the accuracy of recognition while maintaining the depth of the network. The structure of the model includes an input layer, an initial convolutional layer, residual blocks, a global average pooling layer, a fully connected layer, and an output layer, and finally outputs the probability corresponding to each brand.

[0149] During the forward propagation process of the model, the input features first pass through the initial convolutional layer to extract low-level features such as edges and corners. Then, the residual blocks further learn more complex feature expressions through skip connections. In the global average pooling layer, the system takes the average value of all spatial positions of each feature channel to generate a feature vector of a fixed length. Finally, through the fully connected layer, the model generates the probability distribution of each brand, and the output layer selects the brand with the highest probability as the final prediction result.

[0150] S33. Based on the preset text area within the brand recognition label, and obtain the target text information in the detected preset text area, where the preset text area corresponds to the brand.

[0151] The label design of each brand usually has certain standardized features. For example, information such as sizes, prices, and product descriptions of different brands are often located in fixed positions. Therefore, by constructing the correspondence between the brand and the text area, the system can quickly locate the specific text area on the label of the brand after identifying the brand. For example, the size information of a certain brand is usually located at the lower left of the label, and the price information is located at the lower right. Through the correspondence between the brand and these area positions, the system can directly perform text recognition on specific areas at the preset positions.

[0152] Then, further accurately locate the preset text area through edge detection and morphological processing. In this process, the Canny edge detection algorithm is usually used to highlight the edge features of the label, and then the closing operation is used to connect the discontinuous edges to form a complete candidate box for the text area. For the determined candidate area, the system filters out the areas that conform to the brand characteristics according to the brand recognition result in the previous step, avoiding misrecognition of irrelevant areas.

[0153] During the positioning process of the preset text area, the system also uses the short side and corners as reference features to determine the position of the text area. Generally, the short side of the tag refers to the side away from the center of the perforation, and the corners are the endpoints of the tag's edges. By analyzing the geometric relationship between the short side and the corners, the preset text area can be determined more accurately. For example, assume that the text area of a certain brand's tag is located 2 cm below the short side and 1.5 cm to the right of the left corner. The system will generate a region of interest based on this relative relationship and perform OCR recognition within this region.

[0154] Common OCR algorithms include Tesseract, EAST, CRAFT, etc. They can perform high-precision text recognition under different fonts, sizes, and background conditions. Taking Tesseract as an example, the process of recognizing text includes the following steps: binarization, character segmentation, feature extraction, and pattern matching. Assume that the input text area contains the information "XL 199 yuan". The OCR output result is the string "XL 199 yuan". The system, according to the preset field mapping, takes "XL" as the value of the size field and "199 yuan" as the value of the price field and enters them into the product database.

[0155] Specifically, in one embodiment, S33 includes S331 - S333.

[0156] S331. Based on the recognized brand, call the relative relationship between the preset text area and the reference features in the recognized tag.

[0157] In this step, first, based on the brand recognized in the previous step, the system searches in the brand-specific tag template database for the corresponding preset text area configuration of this brand. The tag templates of each brand are usually stored in the form of coordinates, which define the relative position of the specific text area in the tag image. These coordinates are usually determined by using the short side and corners of the tag as reference features. For example, for a certain brand, the size information is located at the lower left of the tag, and the price information is located at the lower right corner. The system can obtain these preset positions through the "brand - position mapping table".

[0158] Assume that in the tag template of a certain brand, the coordinates of the size area are (x 1 , y 1 , w 1 , h 1 ), and the coordinates of the price area are (x 2 , y 2 , w 2 , h 2 ), where (x,y) represents the upper left corner coordinates of the area, and (w,h) represents the width and height of the area. According to these coordinates, the system can crop the corresponding text area in the tag image for subsequent text recognition.

[0159] Mathematically, assume the size of the input tag image is W×H, and the relative coordinates of the preset text area are (r x , r y , r w , r h ), where r x , r y represent the horizontal and vertical offsets of the area relative to the upper left corner of the tag, and r w , r h represent the proportional relationship of the area relative to the width and height of the tag. Then the absolute coordinates in the actual image can be calculated by the following formula:

[0160] x = r x × W, y = r y × H,, w = r w × W,, h = r h × H

[0161] Through these calculations, the system can accurately locate the preset text area in tag images of different sizes.

[0162] S332. Use the recognized short side and the corners of both sides as reference features to determine the preset text area of the tag in the recognized image, where the recognized short side is the short side far from the center of the hole.

[0163] After obtaining the edge image, the system analyzes the contour of the tag to determine the positions of the short side and the corners. The short side refers to the side of the tag far from the center of the perforation, which usually appears as the boundary line with a shorter length in a standard rectangular tag.

[0164] After determining the positions of the short side and the corners, the system infers the absolute coordinates of the text area based on the relative relationship of the preset text area corresponding to the brand. Assume that the size information of a certain brand tag is located 2 cm below the short side and 1.5 cm to the right of the left corner, then the upper left corner coordinates (x, y) of the text area can be calculated by the following formula:

[0165] x = x corner + Δx, y = y short_edge + Δy

[0166] where Δx and Δy represent the offsets relative to the corner point and the short side respectively.

[0167] S333. Obtain the text information in the preset text area.

[0168] By accurately identifying the key text information in the hangtag, the information entry and management process of the commodity is further improved. Since the preset text area is defined based on the brand-specific template, in the hangtags of the same brand, the layout of relevant information usually remains consistent. This template-based text recognition method greatly improves the accuracy of recognition. Even when the hangtag is slightly rotated, partially blocked, or has uneven illumination, the system can still accurately obtain the target text information.

[0169] In addition, the original text results obtained by OCR recognition can be cleaned and normalized. Since OCR may introduce noise characters when dealing with complex backgrounds, distorted fonts, or low-quality images, string cleaning is required to remove invalid characters, whitespace, and non-standard symbols. For example, the recognized text may contain extra spaces, line breaks, or unnecessary punctuation marks, and the cleaning process can be quickly completed through regular expressions.

[0170] The parsed commodity attributes can be further compared and verified with the standardized database of the system to ensure the accuracy and consistency of the data. For example, if the recognized size does not conform to the predefined size standards (such as XS, S, M, L, XL, etc.), the system can automatically mark the anomaly and prompt for manual review.

[0171] Optionally, if it is determined that there is a complete barcode in the image during image recognition, it is confirmed that the recognition is successful and a first prompt sound is emitted; when the target text information is obtained in the detected preset text area, it is confirmed that the recognition is successful and a second prompt sound is emitted.

[0172] The embodiment of the present application also discloses an automatic commodity information entry method based on hangtag scanning, including a processor, and a program of the automatic commodity information entry method based on hangtag scanning described in any one of the above is run in the processor.

[0173] The embodiment of the present application also discloses a storage medium storing a program of the automatic commodity information entry method based on hangtag scanning described in any one of the above.

[0174] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A method for automatically entering product information based on tag scanning, characterized in that: The following steps are involved: S1. Get the image information of the tag; S2. Calling a pre-trained machine model, wherein the pre-trained machine model takes the edge curvature of the tag, the ratio of the edge length of the tag to the font size, the color of the tag and the text area as input, and takes the product type as output; S3. Input the image into the pre-trained model, obtain the product category, and enter it into the current statistical table of the system.

2. The automatic commodity information entry method based on tag scanning according to claim 1 is characterized in that: The S3 includes: S31. Based on image recognition, determine whether there is a complete barcode in the image. If so, identify the barcode information, obtain the product type, and enter it into the current statistics table of the system; S32. If it does not exist, obtain the brand corresponding to the tag in the image based on image recognition; S33. Identify a preset text area in the hangtag based on the brand, and obtain target text information in the detected preset text area, wherein the preset text area corresponds to the brand.

3. The automatic commodity information entry method based on tag scanning according to claim 1 is characterized in that: The S32 includes: S321. Acquire the trademark image in the image based on image recognition, and determine the brand of the product based on the trademark image; S322. If the recognition fails, image recognition is performed on the hangtag in the image, and edge recognition is performed on the recognized hangtag; S323. Calculate the arc of the acquired tag edge, wherein the tag edge includes the side and the corner; S324. Identify the font area and edge of the tag, obtain the font size within the specific text area, and calculate the ratio of the specific edge length and font size of the tag; S325. Input the ratio of the curvature of the tag edge, the function of the text area, the length of the tag edge and the font size in the specific text area into the pre-trained machine model to obtain the brand corresponding to the tag in the image.

4. The automatic commodity information entry method based on tag scanning according to claim 1 is characterized in that: The pre-trained machine model includes the following steps: S201. Enter the hangtag of each clothing item, and obtain the ratio of the hangtag edge curvature, text area function, hangtag edge length and font size in a specific text area through image recognition, and use it as a mark; S202. The clothing, the curvature of the tag edge, the text area function, the length of the tag edge and the ratio of the font size in the specific text area are used as input features, and the brand of the clothing is used as the output result; S203. Normalize the input image to ensure that the mean and variance of the input data are consistent; S204. Divide the data set into a training set and a validation set; S205. Input the training images of the training set into the existing ResNet model, pass through the convolution layer, pooling layer, and fully connected layer, and finally output the prediction results; S206. Calculate the error between the prediction result and the output result using the loss function; S207. By calculating the gradient, the error is back-propagated and the weights and biases in the network are adjusted to reduce the loss; S208. Use the optimizer to update the network parameters to gradually reduce the loss function value; S209. Evaluate the classification accuracy, loss value, and confusion matrix of the model on the validation set.

5. The automatic commodity information entry method based on tag scanning according to claim 1 is characterized in that: The existing ResNet model includes: The input layer is used to receive input features obtained after image preprocessing; The initial convolutional layer is used to extract the spatial features of the image; Residual block, used to jump-connect the input to the output; Global average pooling layer, used to reduce data dimension through average pooling; The fully connected layer is used to input the features extracted by convolution and pooling into the fully connected layer to combine and classify high-level features; The output layer is used to output brand prediction results and tag edge recognition results.

6. The automatic commodity information entry method based on tag scanning according to claim 1 is characterized in that: The S33 includes: S331. Based on the identified brand, the relative relationship between the preset text area in the identified tag and the reference feature is called; S332. Using the identified short side and the corners of both sides as reference features, determining the preset text area of ​​the tag in the recognition image, wherein the identified short side is the short side away from the center of the rotation hole; S333. Obtain text information in a preset text area.

7. The automatic commodity information entry method based on tag scanning according to claim 1 is characterized in that: If a complete barcode is determined to exist in the image during image recognition, the recognition is confirmed to be successful and a first prompt sound is emitted; when target text information is obtained in the detected preset text area, the recognition is confirmed to be successful and a second prompt sound is emitted.

8. An automatic commodity information entry system based on tag scanning, characterized in that: It comprises a processor, in which runs a program of the automatic commodity information entry method based on tag scanning as described in any one of claims 1 to 7.

9. A storage medium, characterized in that: A program for the automatic commodity information entry method based on tag scanning as described in any one of claims 1 to 7 is stored.