A tobacco retail license identification method and system

Through the combination of convolutional neural networks and bidirectional long short-term memory modules, the feature extraction difficulties caused by diverse acquisition conditions and interference factors in tobacco retail license image recognition were solved, and high-precision license information extraction was achieved.

CN114913516BActive Publication Date: 2025-09-12CHINA TOBACCO ZHEJIANG IND CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210383762.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-12
Publication Date
2025-09-12
Estimated Expiration
2042-04-12

AI Technical Summary

Technical Problem

The existing tobacco retail license feature extraction method is affected by diverse collection conditions and interference factors, and the feature extraction effect is poor.

Method used

A combination of convolutional neural networks and bidirectional long short-term memory modules is adopted to achieve non-segmentation recognition and information extraction of tobacco retail licenses through preprocessing, convolutional feature extraction and encoding and decoding technologies, reducing the sensitivity to collection conditions and interference factors.

Benefits of technology

The accuracy of tobacco retail license image recognition is improved, the impact of diverse collection conditions and interference factors is reduced, and efficient feature extraction is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114913516B_ABST
    Figure CN114913516B_ABST
Patent Text Reader

Abstract

The present application discloses a method and system for identifying tobacco retail licenses. The identification method includes: preprocessing an image of a tobacco retail license to obtain a preprocessed image; inputting the preprocessed image into a convolutional neural network to obtain a first visual feature sequence; wherein the convolutional neural network includes at least one convolution submodule, each convolution submodule includes a first densely connected network module, a convolution layer, and a first pooling layer connected in sequence, and the input data of each convolution layer is densely connected data of the output data of all convolution submodules before the convolution layer; encoding and decoding the first visual feature sequence to obtain the license number, company name, and validity period on the tobacco retail license. The present application realizes non-segmentation recognition of tobacco retail license images and extraction of tobacco sales information, reduces sensitivity to diverse acquisition conditions and other interference factors, and improves the accuracy of feature extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and more specifically, to a method and system for identifying tobacco retail licenses. Background Art

[0002] Tobacco marketing Robotic Process Automation (RPA) involves intelligently identifying multiple objects in tobacco operations. For example, to verify the effectiveness of tobacco marketing campaigns on product marketing, it is necessary to conduct in-depth analysis and mining of the various text and image information involved in the retail process. Robotic process automation (sometimes referred to as business process automation) involves using "robots" as digital labor to perform mundane tasks, such as processing invoices or transferring data from one database or spreadsheet to another. Much of the data that must be processed is unstructured, such as text in emails, or even videos and images, which requires complex algorithms. AI robots can use technologies such as computer vision to identify different types of documents or natural language processing to understand the context of email messages. However, intelligent tobacco operations management requires identifying and extracting information from tobacco retail licenses from different regions. Tobacco retail license images have complex, unstructured features, which increases the difficulty of extracting information. Existing feature extraction methods for tobacco retail licenses are affected by diverse acquisition conditions and other interfering factors, resulting in poor feature extraction results. Summary of the Invention

[0003] The present application provides a tobacco retail license recognition method and system, which realizes non-segmentation recognition of tobacco retail license images and extraction of tobacco sales information, reduces sensitivity to diverse acquisition conditions and other interference factors, and improves the accuracy of feature extraction.

[0004] This application provides a method for identifying a tobacco retail license, including:

[0005] Preprocessing the tobacco retail license image to obtain a preprocessed image;

[0006] Inputting the preprocessed image into a convolutional neural network to obtain a first visual feature sequence; wherein the convolutional neural network includes at least one convolution submodule, each convolution submodule includes a first densely connected network module, a convolution layer, and a first pooling layer connected in sequence, and the input data of each convolution layer is densely connected data of the output data of all convolution submodules before the convolution layer;

[0007] The first visual feature sequence is encoded and decoded to obtain the license number, company name and validity period on the tobacco retail license.

[0008] Preferably, encoding and decoding the first visual feature sequence includes:

[0009] Inputting the first visual feature sequence into the first bidirectional long short-term memory module for serialization association processing to obtain a second visual feature sequence of the first bidirectional long short-term memory module;

[0010] fusing the second visual feature sequence obtained by the first bidirectional long short-term memory module with its input data to obtain an original feature map corresponding to the first bidirectional long short-term memory module, performing a one-dimensional self-attention operation on the original feature map to obtain an operation result, and inputting the operation result into a fully connected layer to obtain a first attention feature map corresponding to the first bidirectional long short-term memory module;

[0011] For each second bidirectional long short-term memory module after the first bidirectional long short-term memory module, performing serialized association processing on the output data of the first bidirectional long short-term memory module or the second bidirectional long short-term memory module before the second bidirectional long short-term memory module as input, and obtaining a first attention feature map and an original feature map corresponding to the second bidirectional long short-term memory module;

[0012] Calculate the pixel products between the first attention feature map and the original feature map corresponding to the first bidirectional long short-term memory module and each second bidirectional long short-term memory module respectively, add all pixel products of the same pixel points to generate a second attention feature map, identify the license number, company name and validity period based on the second attention feature map, and output the recognition result.

[0013] Preferably, the preprocessed image is input into a convolutional neural network to obtain a first visual feature sequence, which specifically includes:

[0014] Obtain the output data of all convolution submodules in sequence;

[0015] The densely connected data of the output data of all convolutional sub-modules are input into the second pooling layer, and the obtained pooling result is used as the first visual feature sequence.

[0016] Preferably, the densely connected data input to the convolutional layer is a tensor obtained by mapping and connecting the output data of all convolutional sub-modules before the convolutional layer.

[0017] Preferably, preprocessing the image of the tobacco retail license to obtain a preprocessed image specifically includes:

[0018] Filter and enhance images of tobacco retail licenses;

[0019] Content-based alignment is performed on the results of the filtering and enhancement processes.

[0020] Preferably, the content-based alignment process specifically includes:

[0021] Column scanning algorithm is used to extract the edge of the image;

[0022] Calculate the average slope of the extracted edge clusters;

[0023] The image is rotated based on the average slope so that horizontal lines in the processed image are parallel to the horizontal edges of the image.

[0024] The present application also provides a tobacco retail license recognition system, comprising a pre-processing module, a first visual feature sequence acquisition module, and an encoding and decoding module;

[0025] The preprocessing module is used to preprocess the image of the tobacco retail license to obtain a preprocessed image;

[0026] The first visual feature sequence acquisition module is used to input the preprocessed image into a convolutional neural network to obtain a first visual feature sequence; the convolutional neural network includes at least one convolution submodule, the convolution submodule includes a first densely connected network module, a convolution layer, and a first pooling layer connected in sequence, and the first densely connected network module is connected to the first pooling layers of all convolution submodules before the convolution submodule;

[0027] The encoding and decoding module is used to encode and decode the first visual feature sequence to obtain the license number, company name and validity period on the tobacco retail license.

[0028] Preferably, the convolutional neural network further includes a second densely connected network module and a second pooling layer, the second densely connected network module is connected to the first pooling layers of all convolution submodules, and the second pooling layer is connected to the second densely connected network module;

[0029] The second densely connected network module is used to connect the output data of all convolution submodules into densely connected data;

[0030] The pooling result of the second pooling layer forms a first visual feature sequence.

[0031] Preferably, the codec module includes a first codec submodule, at least one second codec submodule and a first decoder;

[0032] The first encoding and decoding submodule includes a first bidirectional long short-term memory module and a second decoder. The first bidirectional long short-term memory module is used to perform serialized association processing on the first visual feature sequence to obtain a second visual feature sequence of the first bidirectional long short-term memory module; the second decoder is used to fuse the second visual feature sequence obtained by the first bidirectional long short-term memory module with its input data to obtain an original feature map corresponding to the first bidirectional long short-term memory module, perform a one-dimensional self-attention operation on the original feature map to obtain an operation result, and input the operation result into a fully connected layer to obtain a first attention feature map corresponding to the first bidirectional long short-term memory module;

[0033] The second encoding and decoding submodule includes a second bidirectional long short-term memory module and a third decoder, wherein the second bidirectional long short-term memory module is used to perform serialized association processing on the output data of the first bidirectional long short-term memory module or the second bidirectional long short-term memory module before the second bidirectional long short-term memory module as input to obtain a second visual feature sequence of the second bidirectional long short-term memory module; the third decoder is used to obtain a first attention feature map and an original feature map corresponding to the second bidirectional long short-term memory module;

[0034] The first decoder is used to calculate the pixel products between the first attention feature map and the original feature map corresponding to the first bidirectional long short-term memory module and each second bidirectional long short-term memory module, and add all pixel products of the same pixel points to generate a second attention feature map, identify the license number, company name and validity period based on the second attention feature map, and output the recognition result.

[0035] Preferably, the pre-processing module includes a filtering enhancement module and an alignment module;

[0036] The filtering and enhancement module is used to filter and enhance the image of the tobacco retail license;

[0037] The alignment module is used to perform content-based alignment processing on the processing results of filtering and enhancement processing.

[0038] Other features and advantages of the present application will become apparent from the following detailed description of exemplary embodiments of the present application with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments of the application and, together with the description, serve to explain the principles of the application.

[0040] Figure 1 A flowchart of the method for identifying a tobacco retail license provided for this application;

[0041] Figure 2 A structural diagram of the tobacco retail license identification system provided for this application. DETAILED DESCRIPTION

[0042] Various exemplary embodiments of the present application will now be described in detail with reference to the accompanying drawings. It should be noted that unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present application.

[0043] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.

[0044] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.

[0045] In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not limiting. Therefore, other examples of the exemplary embodiments may have different values.

[0046] The present application provides a tobacco retail license recognition method and system, which realizes non-segmentation recognition of tobacco retail license images and extraction of tobacco sales information, reduces sensitivity to diverse acquisition conditions and other interference factors, and improves the accuracy of feature extraction.

[0047] like Figure 1 As shown, the identification methods of tobacco retail licenses include:

[0048] S110: Preprocess the image of the tobacco retail license to obtain a preprocessed image.

[0049] Specifically, as an embodiment, preprocessing the image of the tobacco retail license to obtain the preprocessed image includes the following steps:

[0050] S1101: Filter and enhance the image of the tobacco retail license.

[0051] As an embodiment, a Gaussian filtering method is used for filtering, and a dilation-erosion method is used for enhancement.

[0052] It is understandable that other filtering methods (such as mean filtering) can be used to filter the image. Other image enhancement processing methods (such as neighborhood enhancement) can be used to perform enhancement processing.

[0053] S1102: Perform content-based alignment processing on the processing results of the filtering and enhancement processing.

[0054] Specifically, content-based alignment processing includes:

[0055] S11021: Use a column scanning algorithm to extract edges from the image and obtain edge clusters.

[0056] S11022: Calculate the average slope k of the extracted edge clusters.

[0057] S11023: Rotate the image based on the average slope k so that horizontal lines in the processed image are parallel to the horizontal edges of the image.

[0058] S120: Input the preprocessed image into a convolutional neural network to obtain a first visual feature sequence I1.

[0059] Among them, the convolutional neural network includes at least one convolution submodule (see Figure 2 ), each convolution submodule includes a first densely connected network module, a convolution layer and a first pooling layer connected in sequence, and the input data of each convolution layer is the densely connected data of the output data of all convolution submodules before the convolution layer.

[0060] Since features do not change after they are generated in a densely connected structure, some shallow features may have great potential in the deep layer and can play a big role after fine-tuning. However, they are still regarded as redundant in the deep layer of the network and are often pruned. Based on the above considerations, preferably, the first densely connected network module also includes an activation function, the above densely connected data is input into the activation function, and the output data of the activation function is used as the input data of the convolution layer. The activation function can extract potential redundant features and reactivate them, so that they can better adapt to the feature learning of the deep network, thereby maximizing the feature reuse efficiency of the network.

[0061] Thus, each convolution layer is connected in a feedforward mode with residual connection, so that the lth convolution layer takes the feature maps of all previous layers as input, and calculates the visual feature sequence I of the lth convolution layer through the nonlinear change function. l :

[0062] I l =H l ([I0,I1,…,I l-1 ])

[0063] Among them H l is the nonlinear transformation function of each convolutional layer, I0,I1,…,I l-1 is the sequence of visual features output by all convolutional submodules before the lth layer.

[0064] Specifically, as an embodiment, the densely connected data of the input convolutional layer is a tensor obtained by mapping and connecting the output data of all convolutional sub-modules before the convolutional layer.

[0065] Preferably, the convolutional neural network further includes a second densely connected network module and a second pooling layer, wherein the second densely connected network module is arranged between the last convolution submodule and the second pooling layer. Specifically, the second pooling layer is a maximum pooling layer.

[0066] On the basis of the above preferred embodiment, the preprocessed image is input into the convolutional neural network to obtain the first visual feature sequence I1, which specifically includes:

[0067] S1201: Obtain output data of all convolution submodules in sequence.

[0068] S1202: Input the densely connected data of the output data of all convolution sub-modules into the second pooling layer, and use the obtained pooling result as the first visual feature sequence I1.

[0069] Because densely connected network modules link features from all layers together, each layer receives gradient signals from all previous layers during backpropagation, alleviating the gradient vanishing problem during training. Furthermore, because a large number of features are reused, a large number of features can be generated using a small number of convolution kernels, resulting in a relatively small convolutional neural network model size.

[0070] In the above convolutional neural network structure, since each convolutional layer accepts the features of all previous layers as input, in order to prevent the feature dimension from growing too fast as the number of network layers increases, when downsampling, the feature dimension is first compressed to half of the current input through the convolutional layer, and then the pooling operation is performed.

[0071] S130: Encode and decode the first visual feature sequence I1 to obtain the license number, company name, and validity period on the tobacco retail license.

[0072] Specifically, the first visual feature sequence is encoded and decoded using multiple encoding and decoding sub-modules and a first decoder, each encoding and decoding sub-module includes a bidirectional long short-term memory module and a decoder, wherein the bidirectional long short-term memory module connected to the convolutional neural network is recorded as a first bidirectional long short-term memory module, and the subsequent bidirectional long short-term memory module is recorded as a second bidirectional long short-term memory module.

[0073] Based on the above encoding and decoding structure, encoding and decoding the first visual feature sequence includes:

[0074] S1301: Input the first visual feature sequence I1 into the first bidirectional long short-term memory module for serialization association processing to obtain the second visual feature sequence I2 of the first bidirectional long short-term memory module.

[0075] S1302: The second visual feature sequence I2 obtained by the first bidirectional long short-term memory module is fused with its input data (i.e., the first visual feature sequence I1) to obtain an original feature map D0 corresponding to the first bidirectional long short-term memory module, a one-dimensional self-attention operation is performed on the original feature map D0 to obtain an operation result, and the operation result is input into the fully connected layer to obtain a first attention feature map D1 corresponding to the first bidirectional long short-term memory module.

[0076] S1303: For each second bidirectional long short-term memory module after the first bidirectional long short-term memory module, the second visual feature sequence I2 output by the first bidirectional long short-term memory module before the second bidirectional long short-term memory module or the second visual feature sequence I2 output by the second bidirectional long short-term memory module is used as input for serialized association processing, and the first attention feature map D1 and the original feature map D0 corresponding to the second bidirectional long short-term memory module are obtained.

[0077] S1304: Calculate the pixel products between the first attention feature map D1 corresponding to the first bidirectional long short-term memory module and each second bidirectional long short-term memory module and the original feature map D0 respectively (therefore, for each pixel point, there is a pixel product corresponding to the first bidirectional long short-term memory module and each second bidirectional long short-term memory module), and add all pixel products of the same pixel point to generate a second attention feature map D2, identify the license number, company name and validity period based on the second attention feature map D2 and output the recognition result.

[0078] Based on the above recognition method, the present application provides a tobacco retail license recognition system, including a preprocessing module 210 , a first visual feature sequence acquisition module 220 and a coding and decoding module 230 .

[0079] The preprocessing module 210 is used to preprocess the image of the tobacco retail license to obtain a preprocessed image.

[0080] The pre-processing module 210 includes a filtering and enhancement module and an alignment module. The filtering and enhancement module is used to filter and enhance the image of the tobacco retail license. The alignment module is used to perform content-based alignment on the results of the filtering and enhancement processes.

[0081] The first visual feature sequence obtaining module 220 is used to input the preprocessed image into a convolutional neural network to obtain a first visual feature sequence.

[0082] like Figure 2As shown, the convolutional neural network includes at least one convolution sub-module 2201, and the convolution sub-module includes a first densely connected network module 22011, a convolution layer 22012 and a first pooling layer 22013 connected in sequence. The first densely connected network module 22011 is connected to the first pooling layers of all convolution sub-modules before the convolution sub-module, and the pooling result of the first pooling layer of the last convolution sub-module is used as the first visual feature sequence.

[0083] Preferably, the convolutional neural network also includes a second densely connected network module 2202 and a second pooling layer 2203, the second densely connected network module 2202 is connected to the first pooling layers of all convolution sub-modules, and the second pooling layer 2203 is connected to the second densely connected network module 2202.

[0084] The second densely connected network module 2202 is used to connect the output data of all convolution submodules into densely connected data.

[0085] The pooling result of the second pooling layer 2203 forms a first visual feature sequence.

[0086] The encoding and decoding module 230 is used to encode and decode the first visual feature sequence to obtain the license number, company name and validity period on the tobacco retail license.

[0087] The codec module includes a first codec submodule 2301 , at least one second codec submodule 2302 and a first decoder 2303 .

[0088] The first encoding / decoding submodule 2301 includes a first bidirectional long short-term memory module 23011 and a second decoder 23012. The first bidirectional long short-term memory module 23011 is configured to perform serialized association processing on the first visual feature sequence to obtain a second visual feature sequence for the first bidirectional long short-term memory module. The second decoder 23012 is configured to fuse the second visual feature sequence obtained by the first bidirectional long short-term memory module with its input data to obtain an original feature map corresponding to the first bidirectional long short-term memory module, perform a one-dimensional self-attention operation on the original feature map to obtain an operation result, and input the operation result into a fully connected layer to obtain a first attention feature map corresponding to the first bidirectional long short-term memory module.

[0089] The second encoding / decoding submodule 2302 includes a second bidirectional long short-term memory module 23021 and a third decoder 23022. The second bidirectional long short-term memory module 23021 is configured to perform serialized association processing on the output data of the first bidirectional long short-term memory module or the second bidirectional long short-term memory module preceding the second bidirectional long short-term memory module, thereby obtaining a second visual feature sequence for the second bidirectional long short-term memory module. The third decoder 23022 is configured to obtain a first attention feature map and an original feature map corresponding to the second bidirectional long short-term memory module.

[0090] The first decoder 2303 is used to calculate the pixel products between the first attention feature map and the original feature map corresponding to the first bidirectional long short-term memory module and each second bidirectional long short-term memory module, and add all pixel products of the same pixel points to generate a second attention feature map, identify the license number, company name and validity period based on the second attention feature map, and output the recognition result.

[0091] In this application, the pooling layer in the convolutional neural network can be viewed as a special average-weighted attention mechanism. As is well known, the bidirectional long short-term memory model is also a type of attention mechanism. Therefore, this application uses a dual attention mechanism to achieve non-segmented recognition of tobacco retail license images and extract tobacco sales information, reducing sensitivity to diverse acquisition conditions and other interfering factors, and achieving good results in most practical application scenarios.

[0092] Although some specific embodiments of the present application have been described in detail by way of examples, it should be understood by those skilled in the art that the above examples are for illustration only and are not intended to limit the scope of the present application. It should be understood by those skilled in the art that the above embodiments may be modified without departing from the scope and spirit of the present application. The scope of the present application is defined by the appended claims.

Claims

1. A method for identifying a tobacco retail license, characterized in that: include: Preprocessing the tobacco retail license image to obtain a preprocessed image; Inputting the preprocessed image into a convolutional neural network to obtain a first visual feature sequence; wherein the convolutional neural network includes at least one convolution submodule, each convolution submodule includes a first densely connected network module, a convolution layer, and a first pooling layer connected in sequence, and the input data of each convolution layer is densely connected data of the output data of all convolution submodules before the convolution layer; Encoding and decoding the first visual feature sequence to obtain the license number, company name, and validity period on the tobacco retail license, specifically including: Inputting the first visual feature sequence into a first bidirectional long short-term memory module for serialization association processing to obtain a second visual feature sequence of the first bidirectional long short-term memory module; fusing the second visual feature sequence obtained by the first bidirectional long short-term memory module with its input data to obtain an original feature map corresponding to the first bidirectional long short-term memory module, performing a one-dimensional self-attention operation on the original feature map to obtain an operation result, and inputting the operation result into a fully connected layer to obtain a first attention feature map corresponding to the first bidirectional long short-term memory module; For each second bidirectional long short-term memory module after the first bidirectional long short-term memory module, the output data of the first bidirectional long short-term memory module or the second bidirectional long short-term memory module before the second bidirectional long short-term memory module is used as input for serialized association processing, and a first attention feature map and an original feature map corresponding to the second bidirectional long short-term memory module are obtained: Calculate the pixel products between the first attention feature map and the original feature map corresponding to the first bidirectional long short-term memory module and each second bidirectional long short-term memory module respectively, add all pixel products of the same pixel points to generate a second attention feature map, identify the license number, company name and validity period based on the second attention feature map, and output the recognition result.

2. The tobacco retail license identification method according to claim 1, characterized in that: Inputting the preprocessed image into a convolutional neural network to obtain a first visual feature sequence specifically includes: Obtain the output data of all convolution submodules in sequence; The densely connected data of the output data of all convolutional sub-modules are input into the second pooling layer, and the obtained pooling result is used as the first visual feature sequence.

3. The tobacco retail license identification method according to claim 1, characterized in that: The densely connected data input to the convolutional layer is a tensor obtained by mapping and connecting the output data of all convolutional sub-modules before the convolutional layer.

4. The tobacco retail license identification method according to claim 1, characterized in that: Preprocess the tobacco retail license image to obtain a preprocessed image, specifically including: performing filtering and enhancement processing on the image of the tobacco retail license; Content-based alignment is performed on the results of the filtering and enhancement processes.

5. The tobacco retail license identification method according to claim 4, characterized in that: The content-based alignment process specifically includes: Column scanning algorithm is used to extract the edge of the image; Calculate the average slope of the extracted edge clusters; The image is rotated based on the average slope so that horizontal lines in the processed image are parallel to the horizontal edges of the image.

6. A tobacco retail license identification system, characterized in that: It includes a pre-processing module, a first visual feature sequence acquisition module and a coding and decoding module; The preprocessing module is used to preprocess the image of the tobacco retail license to obtain a preprocessed image; The first visual feature sequence acquisition module is used to input the preprocessed image into a convolutional neural network to obtain a first visual feature sequence; the convolutional neural network includes at least one convolution submodule, the convolution submodule includes a first densely connected network module, a convolution layer and a first pooling layer connected in sequence, and the first densely connected network module is connected to the first pooling layers of all convolution submodules before the convolution submodule; The encoding and decoding module is used to encode and decode the first visual feature sequence to obtain the license number, company name and validity period on the tobacco retail license; The encoding and decoding module specifically includes: a first encoding and decoding submodule, at least one second encoding and decoding submodule and a first decoder; The first encoding and decoding submodule includes a first bidirectional long short-term memory module and a second decoder, wherein the first bidirectional long short-term memory module is used to perform serialized association processing on the first visual feature sequence to obtain a second visual feature sequence of the first bidirectional long short-term memory module; the second decoder is used to fuse the second visual feature sequence obtained by the first bidirectional long short-term memory module with its input data to obtain an original feature map corresponding to the first bidirectional long short-term memory module, perform a one-dimensional self-attention operation on the original feature map to obtain an operation result, and input the operation result into a fully connected layer to obtain a first attention feature map corresponding to the first bidirectional long short-term memory module; The second encoding and decoding submodule includes a second bidirectional long short-term memory module and a third decoder, wherein the second bidirectional long short-term memory module is used to perform serialized association processing on the output data of the first bidirectional long short-term memory module or the second bidirectional long short-term memory module before the second bidirectional long short-term memory module as input to obtain a second visual feature sequence of the second bidirectional long short-term memory module; the third decoder is used to obtain a first attention feature map and an original feature map corresponding to the second bidirectional long short-term memory module; The first decoder is used to calculate the pixel products between the first attention feature map and the original feature map corresponding to the first bidirectional long short-term memory module and each second bidirectional long short-term memory module, and add all pixel products of the same pixel points to generate a second attention feature map, identify the license number, company name and validity period based on the second attention feature map and output the recognition result.

7. The tobacco retail license identification system according to claim 6, characterized in that: The convolutional neural network also includes a second densely connected network module and a second pooling layer, the second densely connected network module is connected to the first pooling layers of all convolution submodules, and the second pooling layer is connected to the second densely connected network module; The second densely connected network module is used to connect the output data of all convolution submodules into densely connected data; The pooling result of the second pooling layer forms the first visual feature sequence.

8. The tobacco retail license identification system according to claim 6, characterized in that: The pre-processing module includes a filtering enhancement module and an alignment module; The filtering and enhancing module is used to perform filtering and enhancing processing on the image of the tobacco retail license; The alignment module is used to perform content-based alignment processing on the processing results of the filtering and enhancement processing.

Citation Information

Patent Citations

  • Image classification method based on ultra-dense connection neural network

    CN113902026A