Label character recognition method and system, electronic device and readable storage medium
Patent Information
- Application Number
- CN202311385118.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-24
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2043-10-24
AI Technical Summary
如果标签信息被错误识别,可能会导致一系列问题,例如产品批号的错误归类、产品生产时间的误判等
[0011] By fitting the target text image with a Bézier curve, the target text image is converted into a planar text image based on the fitting result. Then, a pre-defined character angle detection model is used to detect the characters in the image, and the image is rotated based on the detection results. The target text information is then obtained from the rotated character image. This method not only converts curved label text into planar text using Bézier curves but also rotates angle-shifted character images to their normal orientation using a character angle detection model before performing text recognition. This reduces the difficulty of text recognition when product labels are applied to product surfaces with complex shapes, thereby improving recognition accuracy.
Smart Images

Figure CN117409418B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of label recognition technology, and in particular to a label character recognition method, system, electronic device, and readable storage medium. Background Technology
[0002] Accurate identification of product labels is crucial for managing the product lifecycle during the transportation of products such as steel. Incorrect label information can lead to a series of problems, such as misclassification of product batch numbers and misjudgment of product production dates. These issues not only affect the accuracy of product lifecycle management but can also trigger subsequent problems, such as chaotic inventory management and failures in product quality control.
[0003] Existing identification devices are typically used to identify flat product labels. When product labels are applied to product surfaces with complex shapes, they often deform due to the curvature of the surface, affecting the accuracy of the identification device. Furthermore, due to the variability and complexity of product shapes, fixed label designs and spraying techniques are often unable to adapt to various shapes and surfaces, resulting in poor label recognition adaptability and failure to meet production needs. Summary of the Invention
[0004] To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary is not intended as a general commentary, nor is it intended to identify key / important components or describe the scope of protection of these embodiments, but rather as a prelude to the detailed description that follows.
[0005] In view of the shortcomings of the prior art described above, the present invention discloses a label character recognition method, system, electronic device and readable storage medium to improve the accuracy of label character recognition.
[0006] This invention provides a label character recognition method, comprising: acquiring a target text image of a product label; fitting the target text image using a Bézier curve to obtain a fitting result, and converting the target text image into a planar text image based on the fitting result, wherein the planar text image includes one or more character images; detecting the character images according to a preset character angle detection model, and rotating the character images according to the detection result, wherein the character angle detection model is obtained by training a model using character sample images with tilt angle labels; and performing text recognition on the rotated character image to obtain target text information corresponding to the target text image.
[0007] This invention provides a label character recognition system, comprising: an acquisition module for acquiring a target text image of a product label; a conversion module for fitting the target text image using a Bézier curve to obtain a fitting result, and converting the target text image into a planar text image based on the fitting result, wherein the planar text image includes one or more character images; a rotation module for detecting the character images according to a preset character angle detection model, and rotating the character images based on the detection result, wherein the character angle detection model is obtained through model training using character sample images with tilt angle labels; and a recognition module for performing text recognition on the rotated character image to obtain target text information corresponding to the target text image.
[0008] The present invention provides an electronic device, comprising: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to cause the electronic device to perform the above-described method.
[0009] The present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-described method.
[0010] The beneficial effects of this invention are:
[0011] By fitting the target text image with a Bézier curve, the target text image is converted into a planar text image based on the fitting result. Then, a pre-defined character angle detection model is used to detect the characters in the image, and the image is rotated based on the detection results. The target text information is then obtained from the rotated character image. This method not only converts curved label text into planar text using Bézier curves but also rotates angle-shifted character images to their normal orientation using a character angle detection model before performing text recognition. This reduces the difficulty of text recognition when product labels are applied to product surfaces with complex shapes, thereby improving recognition accuracy. Attached Figure Description
[0012] Figure 1 This is a schematic diagram of the structure of an application environment for implementing a label character recognition method in an embodiment of the present invention;
[0013] Figure 2 This is a flowchart illustrating a label character recognition method according to an embodiment of the present invention;
[0014] Figure 3 This is a schematic diagram of the structure of a text image sample in an embodiment of the present invention;
[0015] Figure 4This is a flowchart illustrating another label character recognition method in an embodiment of the present invention;
[0016] Figure 5 This is a schematic diagram of the structure of a label character recognition system according to an embodiment of the present invention;
[0017] Figure 6 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0018] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, unless otherwise specified, the following embodiments and sub-samples in the embodiments can be combined with each other.
[0019] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0020] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention.
[0021] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this disclosure described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion.
[0022] Unless otherwise stated, the term "multiple" means two or more.
[0023] In this embodiment of the disclosure, the character " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B means: A or B.
[0024] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.
[0025] Combination Figure 1 As shown in the embodiments of this disclosure, an application environment is provided for implementing a label character recognition method, including a server and a user terminal, wherein the server communicates with the user terminal through a network.
[0026] The user terminal is used to send user commands to the server.
[0027] The server is used to execute user instructions, which include at least one of the following: acquiring the target text image of the product label; fitting the target text image with a Bézier curve to obtain the fitting result, and converting the target text image into a planar text image based on the fitting result; detecting the character image according to a preset character angle detection model, and rotating the character image based on the detection result; performing text recognition on the rotated character image to obtain the target text information corresponding to the target text image.
[0028] Combination Figure 2 As shown, this disclosure provides a label character recognition method, including:
[0029] Step S201: Obtain the target text image of the product label;
[0030] Among them, the product label is used to carry one or more of the following: product production date, product batch, product ID (Identity document);
[0031] Step S202: Fit the target text image using a Bézier curve to obtain the fitting result, and convert the target text image into a planar text image based on the fitting result.
[0032] Among them, planar text images include one or more character images;
[0033] Step S203: Detect the character image according to the preset character angle detection model, and rotate the character image according to the detection result;
[0034] The character angle detection model is trained using character sample images labeled with tilt angles.
[0035] Step S204: Perform text recognition on the rotated character image to obtain the target text information corresponding to the target text image.
[0036] The label character recognition method provided in this disclosure uses a Bézier curve to fit the target text image, obtains a fitting result, converts the target text image into a planar text image based on the fitting result, and then detects the character image according to a preset character angle detection model. The character image is then rotated based on the detection result, and the target text information is obtained from the rotated character image. In this way, not only is the curved label text converted into planar text using a Bézier curve, but the character angle detection model also rotates the angularly offset character image to its normal orientation before performing text recognition. This reduces the difficulty of text recognition when product labels are applied to product surfaces with complex shapes, thereby improving recognition accuracy.
[0037] Optionally, acquiring the target text image of the product label includes: pre-setting the product label in text form in a preset product label area; setting the angle between the product label area and the image acquisition device according to a preset angle, and acquiring an image of the product label area through the image acquisition device to obtain a target label image; performing text detection on the target label image, and extracting the label text image from the target label image based on the text detection results; and performing image processing on the label text image to obtain the target text image, wherein the image processing includes one or more of perspective transformation, image smoothing, noise reduction, image sharpening, boundary enhancement, and contrast enhancement.
[0038] In some embodiments, product labels are set in a preset product label area in the form of text by means of inkjet printing, laser printing, etc., wherein the product label area has a planar area and a curved area.
[0039] In some embodiments, the preset angle includes 45° to 90°.
[0040] In some embodiments, the text detection results include Where x1 is the x-coordinate of the top left corner of the label text image, y1 is the y-coordinate of the top left corner of the label text image, x2 is the x-coordinate of the bottom left corner of the label text image, y2 is the y-coordinate of the bottom left corner of the label text image, x3 is the x-coordinate of the top right corner of the label text image, y3 is the y-coordinate of the top right corner of the label text image, x4 is the x-coordinate of the bottom right corner of the label text image, and y4 is the y-coordinate of the bottom right corner of the label text image.
[0041] In some embodiments, the label text image is subjected to perspective transformation using the following formula:
[0042]
[0043] Where (x=x' / w,y=y' / h) are the image coordinates of the target text image, w is the length of the target text image, h is the height of the target text image, (u,v) are the image coordinates of the label text image, k is a preset scaling factor, and perspective transformation matrix. Let T2 be the first matrix used for linear transformation of the image. 13 a 23 ] T T3 is the second matrix used for image perspective transformation, where T3 = [a 31 a 32 ] is the third matrix used for image translation.
[0044] Optionally, before performing text detection on the target label image, the method further includes: acquiring multiple label image samples; annotating the text regions in each label image sample to obtain text coordinate information corresponding to each label image sample, and generating label image attribute information corresponding to each label image sample, wherein the label image attribute information includes at least one of label image name (Filename), label image width (Width), label image height (Height), and label image depth (Depth); using the text coordinate information and label image attribute information as training information for the label image samples, and establishing a detection sample set based on each label image sample with training information; training a preset initial detection model based on the detection sample set to obtain a text detection model, wherein the text detection model is used to perform text detection on the target label image.
[0045] In some embodiments, image depth refers to the number of bits used to store each pixel, and it is also used to measure the color resolution of an image. In a color image, the number of colors that each pixel can have is determined by the image depth; for a grayscale image, the number of gray levels that each pixel can have is also determined by the image depth.
[0046] In some embodiments, the detection sample set includes a detection training set, a detection test set, and a detection validation set; the detection training set is used to train the initial detection model to obtain an intermediate detection model; the detection test set is used to test the intermediate detection model and adjust the model parameters of the intermediate detection model based on the model test results; the detection validation set is used to validate the intermediate detection model after adjusting the model parameters, and the intermediate detection model that passes the validation is determined as the text detection model.
[0047] In some embodiments, an initial detection model is established based on the DB (Differentiable Binarization) algorithm, the LK-PAN (Large Kernel PAN) module, the RSE-FPN (Residual Segmentation Embed FPN) structure, and the DML (Distillation with Maximum Logit Balance) distillation strategy.
[0048] In some embodiments, the DB algorithm performs binarization in the segmentation network, adaptively setting the binarization threshold to obtain more accurate text shape information and simplifying post-processing.
[0049] In some embodiments, the LK-PAN module is a lightweight PAN (Path Aggregation Network) module with a large receptive field. The main idea is to increase the size of the convolutional kernel in the path enhancement of the PAN structure, which can increase the receptive field of each pixel in the feature map, making it easier to detect large fonts and text with extreme aspect ratios.
[0050] In some embodiments, the RSE-FPN structure introduces a residual attention mechanism by replacing the convolutional layers in the FPN (Feature Pyramid Network) with RSEConv (Residual Segmentation Embedded Convolutional layers) to improve the representational power of the feature maps. RSEConv consists of a residual structure and a Squeeze-and-Excitation (SE) block, and its introduction can effectively improve text detection performance.
[0051] In some embodiments, the DML distillation strategy can effectively improve the accuracy of a text detection model by having two structurally identical models learn from each other.
[0052] Optionally, the target text image is fitted using a Bézier curve to obtain a fitting result, including: obtaining the curvature of the product label area; if the curvature of the product label area is less than a preset curvature threshold, the target text image is determined as a planar text image; if the curvature of the product label area is greater than or equal to the preset curvature threshold, the target text image is fitted using a Bézier curve to obtain a fitting result.
[0053] In this way, the smaller the curvature, the flatter the product label area, and the easier it is to recognize the text image. When the curvature is small, the target text image can be directly treated as a planar text image, reducing computing power consumption and thus improving recognition efficiency.
[0054] Optionally, before fitting the target text image using a Bézier curve to obtain the fitting result, the method further includes: acquiring multiple text image samples; using a Bézier curve to delineate the text region in each text image sample to obtain the text region boundary, and marking text feature points on the text region boundary; establishing a fitting sample set based on the text image samples with text feature points; and training a preset initial fitting model based on the fitting sample set to obtain a text fitting model, wherein the text fitting model is used to fit the target text image using a Bézier curve to obtain target feature points, and the target feature points are determined as the fitting result.
[0055] In some embodiments, the curvature and preset angle of the target text image and the product label area are used as input data and input into the text fitting model, so that the text fitting model fits the target text image according to the input data to obtain target feature points, and the target feature points are determined as the fitting result.
[0056] In some embodiments, a Bézier curve is a mathematical curve that can be defined by a series of control points. Each control point corresponds to a Bézier basis function, which determines how the curve behaves around the control point. By adjusting the position and number of control points, the shape of the curve can be changed.
[0057] like Figure 3 As shown, the text image sample is XXXXXX; the text region in the text image sample is delineated by the upper Bézier curve A and the lower Bézier curve B, thus obtaining the text region boundary; starting from the upper left corner, n text feature points are marked clockwise on the text region boundary, and the coordinates of the text feature points are denoted as [(x1,y1),(x2,y2),(x3,y3),...,(x... n ,y n )).
[0058] In some embodiments, the initial fitting model includes an ABCNet (Adaptive Bezier Curve Network) network.
[0059] Optionally, a fitting sample set is established based on text image samples with text feature points, including: obtaining the angle parameters and curvature parameters of the text image samples, wherein the text image samples are obtained by image acquisition devices of the text sample regions, the angle parameters are used to characterize the angle between the image acquisition devices and the text sample regions, and the curvature parameters are used to characterize the curvature of the text sample regions; using the text image samples, angle parameters, and curvature parameters as training samples, and using the text feature points corresponding to the text image samples as training labels for the training samples; establishing a fitting sample set based on multiple training samples with training labels, wherein the fitting sample set includes a fitting training set, a fitting test set, and a fitting validation set; the fitting training set is used to train the initial fitting model to obtain an intermediate fitting model; the fitting test set is used to test the intermediate fitting model and adjust the model parameters of the intermediate fitting model based on the model test results; the fitting validation set is used to validate the intermediate fitting model after adjusting the model parameters, and the intermediate fitting model that passes the validation is determined as the text fitting model.
[0060] In this way, since both angle and curvature parameters affect text image acquisition, the angle, curvature parameters, and text image samples are used together as training samples to train the model, resulting in a text fitting model. When the input data of the text fitting model includes the target text image, the curvature of the product label area, and the preset angle, compared to training the model on a single image, the text fitting model can refer to more input data, thus improving the accuracy of the output results.
[0061] In some embodiments, the training samples may also include text image name, text image width, text image height, text image depth, etc.
[0062] Optionally, converting the target text image into a planar text image based on the fitting result includes: using a Bézier curve alignment method to convert the target text image into a planar text image based on the fitting result, wherein the shape of the sampling grid for Bézier alignment is not rectangular, each column of the arbitrary shape grid is orthogonal to the Bézier curve boundary of the text, the sampling points are equally spaced in width and height, and bilinear interpolation is performed on the coordinates.
[0063] In some embodiments, the pixel size of the planar text image is set to H. out ×W out The i-th pixel in a planar text image is p i , pixel p i The coordinates in the planar text image are Through formula Calculate the correspondence coefficient between the coordinate positions of the planar text image and the target text image; calculate the target feature points based on the definition formula of t and Bézier curves to obtain the upper Bézier curve boundary A' and lower Bézier curve boundary B' of the target text image; and use the formula... The sampling point C' is linearly indexed; bilinear interpolation is used to calculate the coordinates of the sampling point C' in the target text image, thus obtaining the coordinates of the sampling point C' in the planar text image.
[0064] Optionally, the character angle detection model can be obtained by performing self-supervised training on the RotNet neural network for predicting image rotation using multiple character sample images with tilt angle labels.
[0065] Optionally, before performing text recognition on the rotated character image to obtain the target text information corresponding to the target text image, the method further includes: acquiring multiple initial text images; labeling the sample text information corresponding to each initial text image; establishing a recognition sample set based on the initial text images with sample text information; establishing an initial recognition model based on the SVTR_Tiny network structure, the CTC (Connectionist Temporal Classification) module, the TextConAug (Text Conditioned Data Augmentation) data augmentation strategy, and the TextRotNet (Text Rotation and Language Modeling) pre-trained model; and training the initial recognition model based on the recognition sample set to obtain a text recognition model, wherein the text recognition model is used to perform text recognition on the rotated character image to obtain the target text information corresponding to the target text image.
[0066] In some embodiments, the SVTR_Tiny network structure is obtained by fusing the SVTR (Survivable Virtual Topology Routing) network and the lightweight CNN network PP-LCNet (Lightweight CPU-Lightweight Convolutional Neural Network). It uses a visual model similar to the Swing Transformer, a fully connected layer, and a CTC decoder for text sequence prediction, and has good performance in terms of model performance and inference speed.
[0067] In some embodiments, the expression of multiple text features is fused through the CTC module.
[0068] In some embodiments, the TextConAug data augmentation strategy is used to enhance the method to enrich the contextual information of the training data.
[0069] In some embodiments, the initial weights of SVTR_LCNet are initialized using a TextRotNet pre-trained model.
[0070] In some embodiments, the feature maps of PP-LCNet, the output of the SVTR module, and the output of the Attention module between two different SVTR_LCNet and Attention structures are simultaneously trained under supervision.
[0071] Combination Figure 4 As shown, this disclosure provides a label character recognition method, including:
[0072] Step S401: Set the product label in text form in the preset product label area;
[0073] Among them, the product label is used to carry one or more of the following: product production date, product batch, product ID, etc.
[0074] Step S402: Image acquisition is performed on the product label area to obtain the target label image;
[0075] Step S403: Perform text detection on the target label image, and extract the label text image from the target label image based on the text detection results.
[0076] Step S404: Perform image processing on the label text image to obtain the target text image;
[0077] Image processing includes one or more of the following: perspective transformation, image smoothing, noise reduction, image sharpening, boundary enhancement, and contrast enhancement.
[0078] Step S405: Obtain the curvature of the product label area;
[0079] Step S406: Determine whether the curvature of the product label area is less than the preset curvature threshold. If yes, proceed to step S407; otherwise, proceed to step S408.
[0080] Step S407: Determine the target text image as a planar text image, then proceed to step S410;
[0081] Step S408: Fit the target text image using a Bézier curve to obtain the fitting result;
[0082] Step S409: Convert the target text image into a planar text image based on the fitting result, then proceed to step S410;
[0083] Step S410: Detect the character image according to the preset character angle detection model, and rotate the character image according to the detection result;
[0084] Among them, planar text images include one or more character images;
[0085] The character angle detection model is trained using character sample images labeled with tilt angles.
[0086] Step S411: Perform text recognition on the rotated character image to obtain the target text information corresponding to the target text image.
[0087] The label character recognition method provided in this disclosure uses a Bézier curve to fit the target text image, obtains the fitting result, converts the target text image into a planar text image based on the fitting result, and detects the character image according to a preset character angle detection model. The character image is then rotated based on the detection result, thereby obtaining the target text information from the rotated character image. This method has the following advantages:
[0088] First, it not only uses Bézier curves to convert the label text on curved surfaces into planar text, but also uses a character angle detection model to rotate the character image with an angle offset to the normal direction, and then performs text recognition on the character image, reducing the difficulty of text recognition when product labels are applied to product surfaces with complex shapes, thereby improving the recognition accuracy.
[0089] Secondly, since the smaller the curvature, the flatter the product label area, and the easier it is to recognize the text image, when the curvature is small, the target text image can be directly treated as a planar text image, reducing computing power consumption and thus improving recognition efficiency.
[0090] Third, since both angle and curvature parameters affect text image acquisition, angle, curvature parameters, and text image samples are used together as training samples to train the model and obtain a text fitting model. When the input data of the text fitting model includes the target text image, the curvature of the product label area, and the preset angle, compared with training the model on a single image, the text fitting model can refer to more input data and improve the accuracy of the output results.
[0091] Combination Figure 5 As shown, this embodiment of the present disclosure provides a label character recognition system, including an acquisition module 501, a conversion module 502, a rotation module 503, and a recognition module 504.
[0092] The acquisition module 501 is used to acquire the target text image of the product label.
[0093] The conversion module 502 is used to fit the target text image with a Bézier curve, obtain the fitting result, and convert the target text image into a planar text image based on the fitting result, wherein the planar text image includes one or more character images.
[0094] The rotation module 503 is used to detect character images according to a preset character angle detection model and rotate the character images according to the detection results. The character angle detection model is obtained by training the model on character sample images with tilt angle labels.
[0095] The recognition module 504 is used to perform text recognition on the rotated character image to obtain the target text information corresponding to the target text image.
[0096] The label character recognition system provided in this disclosure uses Bézier curves to fit the target text image, obtains the fitting result, converts the target text image into a planar text image based on the fitting result, and then detects the character image according to a preset character angle detection model. Based on the detection result, the character image is rotated, and the target text information is obtained from the rotated character image. In this way, not only is the curved label text converted into planar text using Bézier curves, but the character angle detection model also rotates the angularly offset character image to its normal orientation before performing text recognition. This reduces the difficulty of text recognition when product labels are applied to product surfaces with complex shapes, thereby improving recognition accuracy.
[0097] Figure 6 A schematic diagram of a computer system suitable for implementing the embodiments of this application is shown. It should be noted that... Figure 6 The computer system 600 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0098] like Figure 6 As shown, the computer system 600 includes a Central Processing Unit (CPU) 601, which can perform various appropriate actions and processes, such as executing the methods described in the above embodiments, based on a program stored in Read-Only Memory (ROM) 602 or a program loaded from Storage Section 608 into Random Access Memory (RAM) 603. The RAM 603 also stores various programs and data required for system operation. The CPU 601, ROM 602, and RAM 603 are interconnected via a bus 604. An Input / Output (I / O) interface 605 is also connected to the bus 604.
[0099] The following components are connected to I / O interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 609 performs communication processing via a network such as the Internet. Drive 160 is also connected to I / O interface 605 as needed. Removable media 611, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 160 as needed so that computer programs read from them can be installed into storage section 608 as needed.
[0100] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 609, and / or installed from removable medium 611. When the computer program is executed by central processing unit (CPU) 601, it performs various functions defined in the system of this application.
[0101] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.
[0102] This disclosure also provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements any of the methods in this embodiment.
[0103] The computer-readable storage medium in the embodiments of this disclosure will be understood by those skilled in the art: all or part of the steps of the above method embodiments can be implemented by hardware related to computer programs. The aforementioned computer program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disk, or optical disk.
[0104] The electronic device disclosed in this embodiment includes a processor, a memory, a transceiver, and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication between them. The memory is used to store computer programs, the communication interface is used to perform communication, and the processor and the transceiver are used to run the computer programs, so that the electronic device performs the various steps of the above method.
[0105] In this embodiment, the memory may include random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device.
[0106] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), graphics processing units (GPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0107] The foregoing description and accompanying drawings fully illustrate embodiments of this disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, procedural, and other changes. The embodiments represent only possible variations. Individual components and functions are optional unless explicitly required, and the order of operation may vary. Parts and subsamples of some embodiments may be included in or replace parts and subsamples of other embodiments. Moreover, the terminology used in this application is for describing embodiments only and is not intended to limit the claims. As used in the description of embodiments and claims, the singular forms “a,” “an,” and “the” are intended to equally include the plural forms unless the context clearly indicates otherwise. Similarly, the term “and / or” as used herein means including one or more of the associated listed items and all possible combinations thereof. Additionally, when used in this application, the term "comprise" and its variations "comprises" and / or "comprising" refer to the presence of stated subsamples, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other subsamples, wholes, steps, operations, elements, components, and / or groups thereof. Without further limitations, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes the element. In this document, each embodiment may focus on the differences from other embodiments, and similar or identical parts between embodiments can be referred to mutually. For methods, products, etc., disclosed in the embodiments, if they correspond to the method section disclosed in the embodiments, the relevant parts can be referred to the description of the method section.
[0108] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this disclosure. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0109] The methods and products (including but not limited to devices and equipment) disclosed in the embodiments herein can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units may be merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some sub-samples may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms. Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected to implement this embodiment according to actual needs. Furthermore, the functional units in the embodiments of this disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0110] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
Claims
1. A method for recognizing label characters, characterized in that, include: Obtain the target text image of the product label; The target text image is fitted using a Bézier curve to obtain a fitting result, and the target text image is converted into a planar text image based on the fitting result, wherein the planar text image includes one or more character images; Before fitting the target text image using Bézier curves to obtain the fitting result, multiple text image samples are acquired. The text regions in each of the text image samples are delineated using Bézier curves to obtain the text region boundaries, and text feature points are marked on the text region boundaries. A fitting sample set is established based on the text image samples with text feature points. A preset initial fitting model is trained using the fitting sample set to obtain a text fitting model. The text fitting model is used to fit the target text image using Bézier curves to obtain target feature points, and these target feature points are determined as the fitting result. Establishing a fitting sample set based on text image samples with text feature points includes: acquiring the angle parameters and curvature parameters of the text image samples, wherein the text image samples are obtained by image acquisition devices capturing images of text sample regions; the angle parameters are used to characterize the angle between the image acquisition devices and the text sample regions; and the curvature parameters are used to characterize the curvature of the text sample regions. The text image samples, the angle parameters, and the curvature parameters are used as training samples, and the text feature points corresponding to the text image samples are used as training labels for the training samples. A fitting sample set is established based on multiple training samples with training labels, wherein the fitting sample set includes a fitting training set, a fitting test set, and a fitting validation set. The fitting training set is used to train the initial fitting model to obtain an intermediate fitting model. The fitting test set is used to test the intermediate fitting model and adjust the model parameters of the intermediate fitting model based on the model test results. The fitting validation set is used to validate the intermediate fitting model after adjusting the model parameters, and the validated intermediate fitting model is determined as the text fitting model. The character image is detected according to a preset character angle detection model, and the character image is rotated according to the detection result. The character angle detection model is obtained by training the model on character sample images with tilt angle labels. Text recognition is performed on the rotated character image to obtain the target text information corresponding to the target text image.
2. The method according to claim 1, characterized in that, Obtain the target text image of the product label, including: The product label is pre-set in text form in a preset product label area; The angle between the product label area and the image acquisition device is set according to a preset angle, and the image acquisition device is used to acquire an image of the product label area to obtain a target label image; Text detection is performed on the target label image, and the label text image is extracted from the target label image based on the text detection results; The label text image is processed to obtain the target text image, wherein the image processing includes one or more of perspective transformation, image smoothing, noise reduction, image sharpening, boundary enhancement, and contrast enhancement.
3. The method according to claim 2, characterized in that, Before performing text detection on the target label image, the method further includes: Obtain multiple labeled image samples; The text regions in each of the label image samples are labeled to obtain the text coordinate information corresponding to each label image sample, and the label image attribute information corresponding to each label image sample is generated respectively. The label image attribute information includes at least one of the following: label image name, label image width, label image height, and label image depth. The text coordinate information and the label image attribute information are used as training information for the label image samples, and a detection sample set is established based on each label image sample with training information; The preset initial detection model is trained based on the detection sample set to obtain a text detection model, wherein the text detection model is used to perform text detection on the target label image.
4. The method according to claim 1, characterized in that, The target text image is fitted using a Bézier curve to obtain the fitting result, including: Obtain the curvature of the product label area; If the curvature of the product label area is less than a preset curvature threshold, the target text image is determined to be a planar text image. If the curvature of the product label area is greater than or equal to the preset curvature threshold, then the target text image is fitted using a Bézier curve to obtain the fitting result.
5. The method according to any one of claims 1 to 4, characterized in that, Before performing text recognition on the rotated character image to obtain the target text information corresponding to the target text image, the method further includes: Get multiple initial text images; Label the sample text information corresponding to each of the initial text images; A recognition sample set is established based on the initial text image containing sample text information; An initial recognition model is established based on the SVTR_Tiny network structure, CTC module, TextConAug data augmentation strategy, and TextRotNet pre-trained model; The initial recognition model is trained based on the recognition sample set to obtain a text recognition model, wherein the text recognition model is used to perform text recognition on the rotated character image to obtain the target text information corresponding to the target text image.
6. A label character recognition system, characterized in that, include: The acquisition module is used to acquire the target text image of the product label; The conversion module is used to fit the target text image with a Bézier curve to obtain a fitting result, and convert the target text image into a planar text image based on the fitting result, wherein the planar text image includes one or more character images; The conversion module is further configured to fit the target text image using a Bézier curve, and before obtaining the fitting result, acquire multiple text image samples; use the Bézier curve to delineate the text region in each of the text image samples to obtain the text region boundary, and mark text feature points on the text region boundary; establish a fitting sample set based on the text image samples with text feature points; and train a preset initial fitting model based on the fitting sample set to obtain a text fitting model, wherein the text fitting model is used to fit the target text image using a Bézier curve to obtain target feature points, and determine the target feature points as the fitting result; The conversion module establishes a fitting sample set based on text image samples with text feature points in the following manner: It obtains the angle and curvature parameters of the text image samples, wherein the text image samples are obtained by image acquisition devices capturing images of text sample regions; the angle parameters characterize the angle between the image acquisition devices and the text sample regions; and the curvature parameters characterize the curvature of the text sample regions. The text image samples, the angle parameters, and the curvature parameters are used as training samples, and the text feature points corresponding to the text image samples are used as training labels for the training samples. A fitting sample set is established based on multiple training samples with training labels, including a fitting training set, a fitting test set, and a fitting validation set. The fitting training set is used to train the initial fitting model to obtain an intermediate fitting model. The fitting test set is used to test the intermediate fitting model and adjust its model parameters based on the test results. The fitting validation set is used to validate the intermediate fitting model after parameter adjustment, and the validated intermediate fitting model is determined as the text fitting model. A rotation module is used to detect the character image according to a preset character angle detection model and rotate the character image according to the detection result. The character angle detection model is obtained by training the model on character sample images with tilt angle labels. The recognition module is used to perform text recognition on the rotated character image to obtain the target text information corresponding to the target text image.
7. An electronic device, characterized in that, include: Processor and memory; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to cause the electronic device to perform the method as described in any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Character angle recognition method based on deep learning
CN113963356A