A method, system, device and medium for text recognition in electronic books

By combining convolutional neural networks and Transformer networks and optimizing the attention mechanism, the problem of OCR technology's difficulty in recognizing illustrated text in complex backgrounds has been solved, achieving higher recognition accuracy and efficiency.

CN120599633BActive Publication Date: 2025-10-31UNICOM WOYUEDU TECH CULTURE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511113111.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-10-31
Estimated Expiration
2045-08-11

AI Technical Summary

Technical Problem

Existing OCR technology struggles to accurately identify text boxes when recognizing text in illustrations, especially artistic text against complex backgrounds, resulting in a high rate of false recognition.

Method used

This paper employs a combination of convolutional neural networks and Transformer networks. By extracting local texture features from visible light images and global contextual information from gradient images, it optimizes the attention mechanism of the Transformer network using LU decomposition of column principal components, combines a multi-layer interaction mechanism for shape contour segmentation, and uses OCR tools to recognize text.

Benefits of technology

It improves the accuracy and efficiency of text recognition and reduces the misrecognition rate caused by OCR tools' errors in locating text boxes in complex backgrounds, especially when illustrations contain complex shapes and artistic fonts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599633B_ABST
    Figure CN120599633B_ABST
Patent Text Reader

Abstract

This application relates to a method, system, device, and medium for text recognition in e-books. Through the collaborative work of Transformer and CNN, it accurately segments text regions within complex shape contours while preserving overall structure and detailed features, eliminating interference from background patterns. By defining a precise recognition range through shape segmentation, it avoids missed or incorrect text recognition caused by text box positioning deviations in traditional methods, thus improving the accuracy of text extraction from e-book illustrations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular to a method, system, device and medium for text recognition in electronic books. Background Technology

[0002] E-books are publications that digitize text, images, and audio information and read them on electronic devices. Compared to paper books, e-books have advantages such as portability and applicability to various scenarios. However, prolonged monotonous reading of text can cause visual fatigue. Therefore, adding illustrations to e-books is a common way to re-engage users' reading attention.

[0003] OCR (Optical Character Recognition) is a technology that converts text in images into editable and searchable text data. Currently, OCR is commonly used to recognize text in illustrations. However, some illustrations contain text that is often located within shapes (such as artistic text inserted into a cloud shape). In such complex contexts, OCR is prone to misrecognition because it cannot accurately determine the text box. Summary of the Invention

[0004] The following is an overview of the subject matter described in detail herein. This overview is not intended to limit the scope of the claims.

[0005] The main objective of this disclosure is to propose a text recognition method for electronic books, which can segment shape contours and accurately locate text regions by fusing different image modal features. This solves the problem of difficulty in recognition caused by blurred boundaries between text and background in illustrations, and has the advantages of improving the accuracy and efficiency of text recognition.

[0006] A first aspect of this application proposes a text recognition method for electronic books, the method comprising:

[0007] In response to a text recognition instruction in a first illustration of an e-book, the first illustration is converted into a second illustration; wherein the first illustration is a visible light image and the second illustration is a gradient image;

[0008] The method involves extracting a first feature from the first illustration based on a convolutional neural network and extracting a second feature from the second illustration based on a Transformer network. The extraction of the second feature from the second illustration using the Transformer network includes: performing feature embedding and positional encoding on the second illustration to obtain an input feature matrix; linearly transforming the input feature matrix to generate a query matrix, a value matrix, and a key matrix; performing LU decomposition of the query matrix with column pivots to obtain a unit lower triangular matrix and an upper triangular matrix; calculating an attention coefficient matrix based on the unit lower triangular matrix and the upper triangular matrix; calculating an output feature matrix based on the attention coefficient matrix and the value matrix; and inputting the key matrix and the output feature matrix into a feedforward neural network to obtain the second feature output by the feedforward neural network.

[0009] The shape outline in the first illustration is segmented based on the fusion result of the first feature and the second feature;

[0010] A text box is set on the shape outline, and text is recognized from the text box.

[0011] In some implementations, calculating the attention coefficient matrix based on the unit lower triangular matrix and the upper triangular matrix includes:

[0012] Multiply the lower triangular matrix and the upper triangular matrix to obtain the similarity matrix;

[0013] The attention coefficient matrix is ​​obtained by activating the similarity matrix using an activation function;

[0014] The calculation of the output feature matrix based on the attention coefficient matrix and the value matrix includes:

[0015] Multiply the attention coefficient matrix and the value matrix by matrix multiplication to obtain the weighted value matrix;

[0016] The output feature matrix is ​​extracted from the weighted value matrix by the fully connected layer.

[0017] In some implementations, the Transformer network includes N Transformer sub-networks; and each Transformer sub-network extracts a corresponding second feature, with the corresponding second feature extracted by the previous Transformer sub-network serving as the input feature of the next Transformer sub-network.

[0018] The convolutional neural network includes N convolutional blocks, and each convolutional block extracts a corresponding first feature. The element-wise multiplication result between the corresponding first feature extracted by the previous convolutional block and the corresponding second feature extracted by a Transformer sub-network is used as the input feature of the next convolutional block.

[0019] N is a positive integer greater than 1;

[0020] The segmentation of the shape outline in the first illustration based on the fusion result of the first feature and the second feature includes:

[0021] The second feature output by the last Transformer subnetwork among the N Transformer subnetworks is fused with the first feature output by the last convolutional block among the N convolutional blocks to obtain the third feature;

[0022] The shape outline in the first illustration is segmented based on the third feature.

[0023] In some implementations, the process of extracting the corresponding first feature from any convolutional block includes:

[0024] ;

[0025] in, For the mapping function of the multilayer perceptron, This is the mapping function for channel average pooling. For channel max pooling mapping function, The mapping function for the first feature extracted from the convolutional block. It is the sigmoid activation function;

[0026] ;

[0027] in, A mapping function for element rearrangement operations;

[0028] ;

[0029] ;

[0030] ;

[0031] in, The mapping function for a 7x7 convolution;

[0032] ;

[0033] ;

[0034] ;

[0035] in, The SiLU activation function is used. The mapping function for a 1x3 convolution. The mapping function for a 3x1 convolution. The mapping function is a 1x5 convolution. The mapping function for a 5x1 convolution. The mapping function is a 1x7 convolution. The mapping function is a 7x1 convolution. These are the input features for the convolutional block.

[0036] In some implementations, before fusing the second feature output by the last Transformer subnetwork among the N Transformer subnetworks with the first feature output by the last convolutional block among the N convolutional blocks to obtain the third feature, the method further includes:

[0037] The difference image between the first inset and the second inset is calculated based on the LR operator;

[0038] Extract the difference features from the difference images;

[0039] The process of fusing the first feature and each of the second features to obtain the third feature includes:

[0040] The third feature is obtained by fusing the second feature output by the last Transformer subnetwork among the N Transformer subnetworks, the first feature output by the last convolutional block among the N convolutional blocks, and the third feature.

[0041] In some embodiments, setting a text box on the shape outline and recognizing text from the text box includes:

[0042] The rectangular text box is obtained by cropping the area based on the coordinates of the shape.

[0043] Use an OCR tool to detect the text in the rectangular text box.

[0044] In some implementations, prior to using an OCR tool to detect the text in the rectangular text box, the method further includes:

[0045] Detect the text state within the rectangular text box;

[0046] When the text is in a tilted and perspective state, the affine transformation and perspective transformation of the text are corrected to be horizontal.

[0047] A second aspect of this application provides a text recognition system for electronic books, the system comprising:

[0048] An image conversion module is used to convert the first illustration into a second illustration in response to a text recognition instruction in the first illustration of an e-book; wherein the first illustration is a visible light image and the second illustration is a gradient image;

[0049] The feature calculation module is used to extract a first feature of the first illustration based on a convolutional neural network, and to extract a second feature of the second illustration based on a Transformer network. The extraction of the second feature of the second illustration based on the Transformer network includes: performing feature embedding and position encoding on the second illustration to obtain an input feature matrix; linearly transforming the input feature matrix to generate a query matrix, a value matrix, and a key matrix; performing LU decomposition of the query matrix with column pivots to obtain a unit lower triangular matrix and an upper triangular matrix; calculating an attention coefficient matrix based on the unit lower triangular matrix and the upper triangular matrix; calculating an output feature matrix based on the attention coefficient matrix and the value matrix; and inputting the key matrix and the output feature matrix into a feedforward neural network to obtain the second feature output by the feedforward neural network.

[0050] The contour segmentation module is used to segment the shape contour in the first illustration based on the fusion result of the first feature and the second feature;

[0051] A text recognition module is used to set a text box on the shape outline and recognize text from the text box.

[0052] A third aspect of this application provides an electronic device including at least one controller and a memory for communicatively connecting to the controller; the memory stores instructions executable by the at least one controller, the instructions being executed by the at least one controller to cause the at least one controller to perform a text recognition method in an electronic book as described above.

[0053] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed, implements the text recognition method in an electronic book as described above.

[0054] The text recognition method for electronic books provided in this embodiment has the following advantages:

[0055] Current methods only utilize a single image modality and cannot collaboratively leverage color and edge features. This method extracts gradient images from visible light images, which carry color and texture information. A Transformer network is used to capture long-range dependencies and identify semantic features of text. The gradient image highlights shape edges, and a convolutional neural network extracts local structural features. Through the collaborative work of Transformer and CNN, both overall structure and detailed features are preserved. This method can accurately segment text regions within complex shape contours and eliminate interference from background patterns. Furthermore, during the feature extraction process of the Transformer network, the query matrix is ​​decomposed using LU decomposition with column pivots, providing sequence information for the entire query matrix. This reduces the impact of individual sequences on the attention coefficient, allowing for the extraction of features from the entire sequence. The calculated attention coefficient improves the interpretability and accuracy of the second feature in the second illustration by the Transformer network, thereby improving the accuracy of subsequent shape contour segmentation and reducing the text misrecognition rate caused by OCR tools' errors in text box positioning.

[0056] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0057] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0058] Figure 1 This is a flowchart illustrating a text recognition method in an e-book provided in an embodiment of this application;

[0059] Figure 2 This is a schematic diagram illustrating feature extraction using Transformer and convolutional neural networks provided in the embodiments of this application;

[0060] Figure 3 This is a schematic diagram of the structure of a text recognition system in an electronic book provided in an embodiment of this application;

[0061] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0062] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0063] In the description of this application, the use of terms such as "first," "second," etc., is for the purpose of distinguishing technical features only and should not be construed as indicating or implying relative importance or implicitly indicating the number of technical features indicated or the order of the technical features indicated.

[0064] In the description of this application, it should be understood that the orientation descriptions, such as up, down, etc., are based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application.

[0065] like Figure 1 One embodiment of this application provides a text recognition method for electronic books, the method comprising the following steps S100 to S400:

[0066] Step S100: In response to the text recognition instruction in the first illustration of the e-book, the first illustration is converted into a second illustration; wherein the first illustration is a visible light image and the second illustration is a gradient image.

[0067] E-books contain many illustrations, and many of the irregular shapes in the illustrations contain text descriptions. In order to extract this text, the first step is to respond to the text recognition instruction in the first illustration of the e-book. Here, the first illustration is the illustration used to extract the text. In the editing process, the first illustration is often taken by the author and uploaded to enrich the reader's reading experience. Here, a visible light image is used as an example.

[0068] Since it's necessary to first segment the shape's outline to facilitate text extraction, a gradient image is used here; that is, it's converted to a gradient image to utilize gradient information to help generate more accurate edge contours. It should be noted that gradient image conversion is common knowledge in this field and is not specifically limited here.

[0069] Step S200: Extract the first feature of the first illustration based on a convolutional neural network, and extract the second feature of the second illustration based on a Transformer network.

[0070] Because convolutional neural networks excel at extracting local texture features, while Transformer networks excel at extracting global contextual information, visible light images carry color and texture information, and gradient images highlight shape edges, convolutional neural networks are used to extract local texture features from the visible light image of the first inset and identify semantic features therein, while Transformer networks are used to extract global contextual information from the gradient image of the second inset and capture shape edges.

[0071] In some embodiments, the Transformer network consists of a multi-head attention mechanism and a feedforward neural network, which will not be described in detail here. However, since the multi-head attention mechanism calculates the corresponding attention coefficients through dot product, the relative importance of different tokens is inconsistent, resulting in poor consistency of the extracted global context information. To solve this technical problem, a technique of column pivoting LU decomposition is introduced. This technique is commonly used in text feature processing. This step replaces the dot product operation between the query matrix and the value matrix in the multi-head attention mechanism. The decomposition of the query matrix is ​​achieved by column pivoting LU decomposition, providing the sequence information of the entire query matrix. This can reduce the influence of individual sequences on the attention coefficients.

[0072] Step S200, which involves extracting the second feature of the second illustration based on the Transformer network, includes the following steps S210 to S220:

[0073] Step S210: Perform feature embedding and position encoding on the second illustration to obtain the input feature matrix;

[0074] Step S220: Linearly transform the input feature matrix to generate a query matrix, a value matrix, and a key matrix;

[0075] The steps described above are the same as those of the multi-head attention mechanism, and will not be elaborated here. The key lies in the matrix decomposition of the query matrix.

[0076] Step S230: Perform LU decomposition with column pivoting on the query matrix to obtain a unit lower triangular matrix and an upper triangular matrix; the LU decomposition with column pivoting divides the query matrix into three matrices:

[0077] 1) Permutation matrix; not involved in subsequent calculations;

[0078] 2) Unit lower triangular matrix; A unit lower triangular matrix is ​​a matrix in which all elements on the main diagonal are 1 and the lower triangular part is non-zero.

[0079] 3) Upper triangular matrix. An upper triangular matrix is ​​a matrix in which all elements below the main diagonal are zero.

[0080] The specific operation method of LU decomposition of column pivoting is not described in detail here.

[0081] Step S240: Calculate the attention coefficient matrix based on the unit lower triangular matrix and the upper triangular matrix; in some embodiments, step S240 includes the following steps:

[0082] Step S2410: Multiply the unit lower triangular matrix and the upper triangular matrix to obtain the similarity matrix; the similarity matrix is ​​used to characterize the degree of correlation between features at different locations.

[0083] Step S2420: Activate the similarity matrix based on the activation function to obtain the attention coefficient matrix; the activation function refers to the operation of performing non-linear transformation on the similarity matrix. For example, the Softmax function can be used to normalize the similarity matrix so that the attention coefficient matrix can highlight the weights of important feature regions.

[0084] Step S250: Calculate the output feature matrix based on the attention coefficient matrix and the value matrix; in some embodiments, this includes:

[0085] Step S2510: Multiply the attention coefficient matrix and the value matrix to obtain a weighted value matrix. The weighted value matrix is ​​the result of weighted fusion of the attention coefficients and the value matrix, which can be achieved through matrix multiplication and is used to enhance the feature representation of key regions.

[0086] Step S2520: Extract the output feature matrix from the weighted value matrix using the fully connected layer. A fully connected layer refers to a neural network layer with a fully connected structure, such as a structure combining linear transformations and nonlinear activation functions, used to extract high-dimensional semantic features from the weighted value matrix.

[0087] Step S260: Input the key matrix and output feature matrix into the feedforward neural network to obtain the second feature output by the feedforward neural network. The feedforward neural network is another structure of the Transformer network, and its specific operation will not be demonstrated here.

[0088] Specifically, when calculating the attention coefficient matrix, the query matrix is ​​first decomposed into a unit lower triangular matrix and an upper triangular matrix using column-pivoting LU decomposition. The product of these two matrices generates a similarity matrix. This step reduces computational complexity through matrix decomposition while preserving key information of the query features. Subsequently, an activation function is used to perform a non-linear transformation on the similarity matrix, generating an attention coefficient matrix with normalized weights. When generating the output feature matrix, the attention coefficient matrix is ​​multiplied by the value matrix to achieve feature weighting, allowing the model to focus on features related to text regions. Finally, a fully connected layer performs dimensionality transformation and feature enhancement on the weighted features, providing highly discriminative feature representations for subsequent shape contour segmentation.

[0089] This method effectively solves the problems of low computational efficiency of the attention mechanism and unreasonable feature weight allocation in text recognition of illustrations in complex backgrounds. This method decomposes the query matrix by using LU decomposition with column pivots, providing the sequence information of the entire query matrix. This reduces the influence of individual sequences on the attention coefficient, thereby enabling the mining of features of the entire sequence. The calculated attention coefficient can improve the interpretability and accuracy of the second feature of the second illustration in the Transformer network, thereby improving the accuracy of subsequent shape contour segmentation and reducing the text misrecognition rate caused by text box positioning errors in OCR tools.

[0090] Step S300: Segment the shape outline in the first illustration based on the fusion result of the first feature and the second feature.

[0091] In some embodiments, the Transformer network includes N Transformer subnetworks, and any two Transformer subnetworks are connected by downsampling blocks; the associated convolutional neural network includes N convolutional blocks, and any two convolutional blocks are connected by downsampling blocks; N is a positive integer greater than 1;

[0092] Each Transformer subnetwork extracts a corresponding second feature, and the corresponding second feature extracted by the previous Transformer subnetwork among the N Transformer subnetworks is used as the input feature of the next Transformer subnetwork.

[0093] Each convolutional block extracts a corresponding first feature, and the element-wise multiplication result between the corresponding first feature extracted by the previous convolutional block and the corresponding second feature extracted by a Transformer sub-network is used as the input feature of the next convolutional block.

[0094] Step S300, which segments the shape contour in the first illustration based on the fusion result of the first and second features, includes steps S310-320:

[0095] Step S310: Fuse the second feature output by the last Transformer subnetwork among the N Transformer subnetworks with the first feature output by the last convolutional block among the N convolutional blocks to obtain the third feature;

[0096] Step S320: Segment the shape outline in the first illustration based on the third feature.

[0097] like Figure 2Specifically, by stacking Transformer subnetworks and convolutional blocks hierarchically and establishing a cross-modal feature interaction mechanism between each level, global semantic information in visible light images and local edge information in gradient images can be progressively fused. For example, when N is 3, the low-level features extracted by the first-level Transformer subnetwork (belonging to the first feature extracted by the first-level Transformer subnetwork) and the edge features output by the corresponding convolutional block (belonging to the second feature extracted by the corresponding convolutional block) are multiplied by a feature matrix, which can enhance the sensitivity of subsequent convolutional blocks to text contours; the second-level Transformer subnetwork further fuses intermediate semantic features (belonging to the first feature extracted by the second-level Transformer subnetwork) and refined edge features (belonging to the second feature extracted by the corresponding convolutional block); the third-level Transformer subnetwork integrates high-level semantic information (belonging to the first feature extracted by the third-level Transformer subnetwork) and accurate gradient responses, and finally achieves accurate segmentation of shape contours by fusing multi-level features.

[0098] Compared with existing technologies, this method effectively combines the semantic information of visible light images with the edge information of gradient images through a hierarchical network structure and a cross-modal layer-by-layer interaction mechanism. By enhancing feature complementarity through layer-by-layer interaction, it improves the segmentation accuracy of complex shape contours in illustrations, thereby providing an accurate text box localization basis for subsequent text recognition and overcoming the problem of inaccurate contour segmentation under complex background interference.

[0099] In some embodiments, the process of extracting the corresponding first feature from any convolutional block includes:

[0100] ;

[0101] in, For the mapping function of the multilayer perceptron, This is the mapping function for channel average pooling. For channel max pooling mapping function, The mapping function for the first feature extracted from the convolutional block. It is the sigmoid activation function;

[0102] ;

[0103] in, A mapping function for element rearrangement operations;

[0104] ;

[0105] ;

[0106] ;

[0107] in, The mapping function for a 7x7 convolution;

[0108] ;

[0109] ;

[0110] ;

[0111] in, The SiLU activation function is used. The mapping function for a 1x3 convolution. The mapping function for a 3x1 convolution. The mapping function is a 1x5 convolution. The mapping function for a 5x1 convolution. The mapping function is a 1x7 convolution. The mapping function is a 7x1 convolution. These are the input features for the convolutional block.

[0112] Among these, the mapping function of a multilayer perceptron refers to performing a non-linear transformation on the input data through fully connected layers to enhance feature representation capabilities. The channel average pooling mapping function performs a global average calculation on the feature map along the channel dimension to extract global contextual information. The channel max pooling mapping function extracts the maximum value of the feature map along the channel dimension to preserve significant local features. The element rearrangement mapping function rearranges the channel order of the feature map, achieved by adjusting channel positions, to enhance feature diversity. The mapping function for convolutions of different sizes uses combinations of 1x3, 3x1, 1x5, 5x1, 1x7, and 7x1 convolutional kernels to extract multi-scale features from the input. This can be achieved by stacking convolutional blocks with different aspect ratios in parallel, to simultaneously capture fine-grained texture features both horizontally and vertically.

[0113] Specifically, when processing input features in the convolutional block, global statistical features are first extracted using channel-wise average pooling and max pooling, respectively. These two types of features are then input into a multilayer perceptron to generate channel attention weights, which are then normalized using the sigmoid function. The input features are then multiplied by the attention weights to enhance the channel-level features. Next, the enhanced features undergo element rearrangement to break the fixed correlation between channels, followed by spatial feature fusion using a 7x7 convolution. Finally, the SiLU activation function is used to perform a non-linear transformation on the outputs of multiple convolutional kernels with different aspect ratios, and the features from all branches are concatenated to form the final second feature. This multi-scale convolutional kernel combination effectively captures the edge and contour information of text regions in different directions.

[0114] Compared to existing technologies, this method, by combining multiple sets of horizontal and vertical asymmetric convolution kernels, can simultaneously extract horizontal, vertical, and diagonal edge features of text regions, significantly improving the ability to perceive complex shape contours. Furthermore, the combination of channel attention mechanism and element rearrangement operation enhances the response intensity of key regions in the feature map and suppresses redundant information. This method can more accurately extract text region features from inset gradient images, and is particularly suitable for recognizing artistic characters inside irregular shapes such as clouds and bubbles.

[0115] Furthermore, before fusing the first and second features to obtain the third feature, the method also includes:

[0116] The difference image between the first and second insets is calculated based on the LR operator. The LR operator is an image difference calculation algorithm based on local region contrast. Specifically, it can be implemented by traversing the image pixels with a sliding window and calculating the sum of squares of the gray value differences between adjacent regions. This is used to capture subtle changes between the visible light image and the gradient image.

[0117] Extracting difference features from difference images: Difference images are grayscale images generated through pixel-level comparison that reflect local differences between two images. Specifically, they can be generated by normalizing the results of the LR operator calculation and used to locate the boundary between text regions and the background in illustrations. Difference features refer to the deep semantic information extracted from difference images. Specifically, convolutional neural networks can be used to extract features from difference images at multiple scales to enhance the focus on text regions during shape contour segmentation. Feature fusion refers to jointly modeling features from different sources. Specifically, this can be achieved by concatenating channels and then reducing dimensionality through fully connected layers, to integrate the complementary information of visible light features, gradient features, and difference features. Based on the fusion of the first feature and each second feature, a third feature is obtained, including:

[0118] The third feature is obtained by fusing the first feature, the second feature, and the difference feature.

[0119] Specifically, after converting the visible light image into a gradient image, a pixel-by-pixel difference analysis is performed on the two images using the LR operator to generate a difference image that highlights the contrast difference between the text region and the background. Subsequently, a convolutional network is used to extract multi-level features from the difference image, forming difference features. In the feature fusion stage, the first feature from the visible light image, the second feature from the gradient image, and the difference feature are concatenated along the channel dimension, and the number of channels is compressed through a fully connected layer, ultimately generating a third feature containing multimodal information. This third feature combines the texture information, gradient edge information, and contrast information of the difference region from the original image, providing more comprehensive feature support for subsequent shape contour segmentation.

[0120] Compared with existing technologies, this method can enhance the sensitivity of the shape contour segmentation stage to the text region, especially in scenarios where illustrations contain complex textures or artistic fonts. By supplementing the deficiencies of gradient features and visible light features through difference features, it reduces the misrecognition rate of OCR tools caused by text box positioning deviations and improves the recognition accuracy of illustration text in e-books.

[0121] Step S400: Set a text box on the shape outline and identify the text from the text box.

[0122] Furthermore, a text box is set on the shape outline, and text is recognized from the text box, including:

[0123] The rectangular text box is obtained by cropping the area based on the coordinates of the shape.

[0124] Use an OCR tool to detect text within a rectangular text box.

[0125] Here, a rectangular text box refers to the smallest bounding rectangle region determined by the coordinate range of its shape outline; for example, it can be implemented using the `minAreaRect` function in the OpenCV library. OCR tools are optical character recognition engines capable of recognizing text in images, such as Tesseract or PaddleOCR. Text state detection determines whether text is tilted or has perspective distortion; this can be achieved using orientation classification models or geometric feature analysis, such as detecting line angles through Hough transform. Affine transformation and perspective transformation involve performing linear geometric transformations on the image to eliminate tilt and perspective distortion; this can be achieved using homography matrix calculations, for example, through the `warpPerspective` function in OpenCV.

[0126] This method generates precise text box regions by combining shape contour segmentation results, effectively isolating background noise and significantly improving recognition robustness in complex layout scenarios.

[0127] Furthermore, before using an OCR tool to detect text within a rectangular text box, the following steps are also included:

[0128] Detect the text status within a rectangular text box;

[0129] When the text is in a slanted and perspective state, correct the affine transformation and perspective transformation of the text to a horizontal direction.

[0130] The text state refers to the geometric shape of the text in the image. This can be achieved through edge detection algorithms or Hough transform for angle analysis, used to determine if the text is tilted or has perspective distortion. Tilt refers to a non-zero angle between the text axis and the horizontal direction; if this angle exceeds a preset threshold, it is considered tilted. Perspective refers to trapezoidal distortion caused by projection transformation; when the vertex coordinates do not conform to the rectangular rules, it is considered perspective distortion. Affine transformation refers to linear transformations of the image through translation, rotation, and scaling operations. This can be implemented using OpenCV's `warpAffine` function, which can correct tilted text to a horizontal orientation. Perspective transformation adjusts the spatial relationships of the image through a projection matrix. This can be implemented using OpenCV's `warpPerspective` function, which can eliminate trapezoidal distortion and restore the planar shape of the text.

[0131] This method improves text recognition accuracy in complex layout scenarios, especially in e-book illustrations with tilted artistic fonts or 3D layouts. It restores the standard shape of the text through geometric correction, avoiding character misrecognition caused by deformation. For example, after rotation correction, the character spacing and stroke features of tilted artistic fonts in a cloud shape are accurately preserved, and the OCR engine can correctly recognize cursive and decorative text.

[0132] The character recognition method for electronic books provided in this application has at least the following beneficial effects:

[0133] Current methods only utilize a single image modality and cannot collaboratively leverage color and edge features. This method extracts gradient images from visible light images, which carry color and texture information. A Transformer network is used to capture long-range dependencies and identify semantic features of text. The gradient image highlights shape edges, and a convolutional neural network extracts local structural features. Through the collaborative work of Transformer and CNN, both overall structure and detailed features are preserved. This method can accurately segment text regions within complex shape contours and eliminate interference from background patterns. Furthermore, during the feature extraction process of the Transformer network, the query matrix is ​​decomposed using LU decomposition with column pivots, providing sequence information for the entire query matrix. This reduces the impact of individual sequences on the attention coefficient, allowing for the extraction of features from the entire sequence. The calculated attention coefficient improves the interpretability and accuracy of the second feature in the second illustration by the Transformer network, thereby improving the accuracy of subsequent shape contour segmentation and reducing the text misrecognition rate caused by OCR tools' errors in text box positioning.

[0134] like Figure 3 One embodiment of this application provides a text recognition system for electronic books, the system comprising:

[0135] The image conversion module 1100 is used to convert the first illustration into a second illustration in response to a text recognition instruction in the first illustration of the e-book; wherein the first illustration is a visible light image and the second illustration is a gradient image;

[0136] The feature calculation module 1200 is used to extract the first feature of the first illustration based on a convolutional neural network and the second feature of the second illustration based on a Transformer network. The extraction of the second feature of the second illustration based on the Transformer network includes: performing feature embedding and position encoding on the second illustration to obtain an input feature matrix; linearly transforming the input feature matrix to generate a query matrix, a value matrix, and a key matrix; performing LU decomposition of the query matrix with column pivots to obtain a unit lower triangular matrix and an upper triangular matrix; calculating the attention coefficient matrix based on the unit lower triangular matrix and the upper triangular matrix; calculating the output feature matrix based on the attention coefficient matrix and the value matrix; and inputting the key matrix and the output feature matrix into a feedforward neural network to obtain the second feature output by the feedforward neural network.

[0137] The contour segmentation module 1300 is used to segment the shape contour in the first illustration based on the fusion result of the first feature and the second feature;

[0138] The text recognition module 1400 is used to set a text box on a shape outline and to recognize the text from the text box.

[0139] It should be noted that the text recognition method in an e-book provided in this embodiment is based on the same inventive concept as the text recognition method in an e-book described above. Therefore, the relevant content of the text recognition method in an e-book described above is also applicable to the text recognition system in an e-book. Therefore, it will not be repeated here.

[0140] Reference Figure 4 This application also provides an electronic device, which includes:

[0141] At least one memory;

[0142] At least one processor;

[0143] At least one program;

[0144] The program is stored in memory, and the processor executes at least one program to implement the character recognition method in the electronic book described above in this disclosure.

[0145] This electronic device can be any smart terminal, including mobile phones, tablets, personal digital assistants (PDAs), and in-vehicle computers.

[0146] The electronic devices according to embodiments of this application will now be described in detail.

[0147] The processor 1600 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0148] The memory 1700 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 1700 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the program code is stored in the memory 1700 and is called and executed by the processor 1600 to execute the text recognition method in the electronic book of this application.

[0149] The input / output interface 1800 is used to implement information input and output.

[0150] The communication interface 1900 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0151] Bus 2000 transmits information between various components of the device (e.g., processor 1600, memory 1700, input / output interface 1800, and communication interface 1900);

[0152] The processor 1600, memory 1700, input / output interface 1800 and communication interface 1900 are connected to each other within the device via bus 2000.

[0153] This application also provides a storage medium, which is a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the text recognition method in the above-described electronic book.

[0154] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0155] The embodiments described in this application are intended to more clearly illustrate the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0156] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0157] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0158] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0159] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0160] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0161] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0162] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0163] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0164] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0165] The above is a detailed description of the preferred embodiments of this application. However, the embodiments of this application are not limited to the above-described implementation methods. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the embodiments of this application. All such equivalent modifications or substitutions are included within the scope defined by the claims of the embodiments of this application.

Claims

1. A text recognition method for electronic books, characterized in that, The method includes: In response to a text recognition instruction in a first illustration of an e-book, the first illustration is converted into a second illustration; wherein the first illustration is a visible light image and the second illustration is a gradient image; The method involves extracting a first feature from the first illustration based on a convolutional neural network and extracting a second feature from the second illustration based on a Transformer network. The extraction of the second feature from the second illustration using the Transformer network includes: embedding and encoding features in the second illustration to obtain an input feature matrix; linearly transforming the input feature matrix to generate a query matrix, a value matrix, and a key matrix; performing LU decomposition of the query matrix with column pivots to obtain a unit lower triangular matrix and an upper triangular matrix; calculating an attention coefficient matrix based on the unit lower triangular matrix and the upper triangular matrix; calculating an output feature matrix based on the attention coefficient matrix and the value matrix; and inputting the key matrix and the output feature matrix into a feedforward neural network to obtain the second feature output by the feedforward neural network. The Transformer network includes N Transformer sub-networks, and each Transformer sub-network extracts a corresponding second feature, with the corresponding second feature extracted by the previous Transformer sub-network serving as the input feature of the next Transformer sub-network. The convolutional neural network includes N convolutional blocks, and each convolutional block extracts a corresponding first feature. The element-wise multiplication result between the corresponding first feature extracted by the previous convolutional block and the corresponding second feature extracted by a Transformer sub-network is used as the input feature of the next convolutional block. N is a positive integer greater than 1; The shape outline in the first illustration is segmented based on the fusion result of the first feature and the second feature; A text box is set on the shape outline, and text is recognized from the text box; The segmentation of the shape outline in the first illustration based on the fusion result of the first feature and the second feature includes: The second feature output by the last Transformer subnetwork among the N Transformer subnetworks is fused with the first feature output by the last convolutional block among the N convolutional blocks to obtain the third feature; The shape outline in the first illustration is segmented based on the third feature.

2. The text recognition method in electronic books according to claim 1, characterized in that, The calculation of the attention coefficient matrix based on the unit lower triangular matrix and the upper triangular matrix includes: Multiply the lower triangular matrix and the upper triangular matrix to obtain the similarity matrix; The attention coefficient matrix is ​​obtained by activating the similarity matrix using an activation function; The calculation of the output feature matrix based on the attention coefficient matrix and the value matrix includes: Multiply the attention coefficient matrix and the value matrix by matrix multiplication to obtain the weighted value matrix; The output feature matrix is ​​extracted from the weighted value matrix by the fully connected layer.

3. The text recognition method in electronic books according to claim 1, characterized in that, The process of extracting the corresponding first feature from any convolutional block includes: CAM(F)=σ(MLP(AvaPl(F))+MLP(MaxPl(F))); Where MLP is the mapping function of the multilayer perceptron, AvaPl is the mapping function of channel average pooling, MaxPl is the mapping function of channel max pooling, CAM(F) is the mapping function of the first feature extracted from the convolutional block, and σ is the sigmoid activation function. F=Shuffle(SAM(F1), SAM(F2), SAM(F3)); Where Shuffle is the mapping function for element rearrangement operations; SAM(F1)=σ(Conv 7×7 ([AvaPl(F1);MaxPl(F1)]))); SAM(F2)=σ(Conv 7×7 ([AvaPl(F2);MaxPl(F2)])); SAM(F3)=σ(Conv 7×7 ([AvaPl(F3);MaxPl(F3)])); Among them, Conv 7×7 The mapping function for a 7x7 convolution; F1=Conv 3×1 (SiLU(Conv 1×3 (F0))); F2=Conv 5×1 (SiLU(Conv 1×5 (F0))); F3=Conv 7×1 (SiLU(Conv 1×7 (F0))); Where SiLU is the SiLU activation function, Conv 1×3 The mapping function for a 1x3 convolution, Conv 3×1 The mapping function for a 3x1 convolution, Conv 1×5 The mapping function for a 1x5 convolution, Conv 5×1 The mapping function for a 5x1 convolution, Conv 1×7 The mapping function for a 1x7 convolution, Conv 7×1 F0 is the mapping function for a 7x1 convolution, and F0 is the input feature of the convolution block.

4. The text recognition method in electronic books according to claim 1, characterized in that, Before fusing the second feature output by the last Transformer subnetwork among the N Transformer subnetworks with the first feature output by the last convolutional block among the N convolutional blocks to obtain the third feature, the method further includes: The difference image between the first inset and the second inset is calculated based on the LR operator; Extract the difference features from the difference images; The process of fusing the first feature and each of the second features to obtain the third feature includes: The third feature is obtained by fusing the second feature output by the last Transformer subnetwork among the N Transformer subnetworks, the first feature output by the last convolutional block among the N convolutional blocks, and the third feature.

5. The text recognition method in electronic books according to claim 1, characterized in that, The step of setting a text box on the shape outline and recognizing text from the text box includes: The rectangular text box is obtained by cropping the area based on the coordinates of the shape. Use an OCR tool to detect the text in the rectangular text box.

6. The text recognition method in electronic books according to claim 5, characterized in that, Before using the OCR tool to detect the text in the rectangular text box, the method further includes: Detect the text state within the rectangular text box; When the text is in a tilted and perspective state, the affine transformation and perspective transformation of the text are corrected to be horizontal.

7. A text recognition system for electronic books, characterized in that, The system includes: An image conversion module is used to convert the first illustration into a second illustration in response to a text recognition instruction in the first illustration of an e-book; wherein the first illustration is a visible light image and the second illustration is a gradient image; A feature calculation module is used to extract a first feature of the first illustration based on a convolutional neural network, and to extract a second feature of the second illustration based on a Transformer network. The extraction of the second feature of the second illustration based on the Transformer network includes: performing feature embedding and position encoding on the second illustration to obtain an input feature matrix; linearly transforming the input feature matrix to generate a query matrix, a value matrix, and a key matrix; performing LU decomposition of the query matrix with column pivots to obtain a unit lower triangular matrix and an upper triangular matrix; calculating an attention coefficient matrix based on the unit lower triangular matrix and the upper triangular matrix; calculating an output feature matrix based on the attention coefficient matrix and the value matrix; and inputting the key matrix and the output feature matrix into a feedforward neural network to obtain the second feature output by the feedforward neural network. The Transformer network includes N Transformer sub-networks; and each Transformer sub-network extracts a corresponding second feature, with the corresponding second feature extracted by the previous Transformer sub-network serving as the input feature of the next Transformer sub-network. The convolutional neural network includes N convolutional blocks, and each convolutional block extracts a corresponding first feature. The element-wise multiplication result between the corresponding first feature extracted by the previous convolutional block and the corresponding second feature extracted by a Transformer sub-network is used as the input feature of the next convolutional block. N is a positive integer greater than 1; The contour segmentation module is used to segment the shape contour in the first illustration based on the fusion result of the first feature and the second feature; The segmentation of the shape outline in the first illustration based on the fusion result of the first feature and the second feature includes: The second feature output by the last Transformer subnetwork among the N Transformer subnetworks is fused with the first feature output by the last convolutional block among the N convolutional blocks to obtain the third feature; The shape outline in the first illustration is segmented based on the third feature; A text recognition module is used to set a text box on the shape outline and recognize text from the text box.

8. An electronic device, characterized in that, It includes at least one controller and a memory for communicatively connecting with the controller; the memory stores instructions executable by the at least one controller, the instructions being executed by the at least one controller to cause the at least one controller to perform the text recognition method in an electronic book as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions for causing a computer to perform the text recognition method in an electronic book as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Transform-based scene image character modification method and device, electronic equipment and storage medium

    CN115908639A

  • Double-branch adaptive coding and decoding colorectal polyp segmentation method

    CN117765012A