Picture text extraction method and device, equipment and storage medium

By optimizing the visual language model through dynamic input pruning and sparse attention matrix, the problems of resource consumption and inference time consumption of large models in image text extraction are solved, and efficient image text extraction is achieved.

CN120612702APending Publication Date: 2025-09-09CHINA MERCHANTS BANK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510753891.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing methods for extracting text from images using large models have the problem that resource consumption and inference time increase geometrically with the image size.

Method used

A dynamic input pruning strategy and sparse attention matrix are used to optimize the visual language model. The images are processed by dynamically cropping redundant areas, compressing and reducing the dimensionality, and a sparse attention matrix is ​​introduced into the model to reduce the amount of computation.

Benefits of technology

While retaining the effective features of the original data, the input data scale is reduced, the model inference speed is improved, resource consumption is reduced, and recognition accuracy and generalization ability are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612702A_ABST
    Figure CN120612702A_ABST
Patent Text Reader

Abstract

The invention discloses a picture text extraction method and device, equipment and a storage medium, and relates to the technical field of computer vision, and the method comprises the steps: obtaining a to-be-processed picture; the to-be-processed picture is input into a pre-constructed picture text extraction model, a picture text is output, the picture text extraction model is based on a visual language model and further comprises a dynamic input pruning strategy and a sparse attention matrix, and resource consumption and reasoning time consumption can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular to a method, apparatus, device, and storage medium for extracting text from images. Background Art

[0002] With the explosive growth of image data, traditional manual transcription methods are unable to meet the needs of quickly acquiring and analyzing text information in massive images. Therefore, technologies that automatically recognize and extract text from images through computers have emerged.

[0003] The current method of extracting text from images using large models has the problem that resource consumption and inference time increase geometrically with the image size.

[0004] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention

[0005] The main purpose of this application is to provide a method, device, equipment and storage medium for extracting text from images, aiming to solve the technical problem that the current method of extracting text from images through large models has resource consumption and inference time that grows geometrically with the size of the image.

[0006] To achieve the above objectives, this application proposes a method for extracting text from an image, which includes:

[0007] Get the image to be processed;

[0008] The image to be processed is input into a pre-built image-text extraction model to output the image text. The image-text extraction model is based on a visual language model and also includes a dynamic input pruning strategy and a sparse attention matrix.

[0009] In one embodiment, the step of inputting the image to be processed into a pre-built image text extraction model and outputting the image text includes:

[0010] Performing image processing on the image to be processed by using the dynamic input pruning strategy;

[0011] The processed image is input into the improved Tongyi Qianwen visual language model, and the image text is output. The improved Tongyi Qianwen visual language model includes the sparse attention matrix.

[0012] In one embodiment, the step of performing image processing on the image to be processed using the dynamic input pruning strategy includes:

[0013] Performing text area recognition on the image to be processed using a pre-built optical character recognition model to obtain the text area;

[0014] Based on the text area, the invalid part of the image to be processed is cropped to retain the information area;

[0015] The information region is compressed and / or dimensionally reduced to obtain the processed image.

[0016] In one embodiment, the step of compressing and / or reducing the dimension of the information region to obtain the processed image includes:

[0017] Calculating information density based on the text area and the information area;

[0018] Calculating a scaling factor of the image to be processed based on the information area, the information density, and a pre-established dynamic scaling factor formula;

[0019] Performing geometric scaling on the information area based on the scaling factor;

[0020] Grayscale processing is performed on the scaled information area to obtain the processed image.

[0021] In one embodiment, the step of inputting the processed image into the improved Tongyi Qianwen visual language model and outputting the image text includes:

[0022] Obtain the Tongyi Qianwen visual language model;

[0023] The BigBird sparse matrix is ​​introduced into the Tongyi Qianwen visual language model and the Tongyi Qianwen visual language model is improved to obtain the improved Tongyi Qianwen visual language model.

[0024] In one embodiment, the step of introducing the BigBird sparse matrix into the Tongyi Qianwen visual language model and improving the Tongyi Qianwen visual language model to obtain the improved Tongyi Qianwen visual language model includes:

[0025] Replacing the random attention in the BigBird sparse matrix with the information region attention to obtain the sparse attention matrix;

[0026] Performing a logical AND operation on the sparse attention matrix and the attention mask to obtain a sparse attention mask;

[0027] Performing a weighted operation on the sparse attention mask based on a preset attention weight to obtain a sparse attention weight;

[0028] The Tongyi Qianwen visual language model is improved based on the sparse attention weight to obtain the improved Tongyi Qianwen visual language model.

[0029] In one embodiment, the step of calculating the scaling factor of the image to be processed based on the information area, the information density, and a pre-established dynamic scaling factor formula includes:

[0030] Get the height and width of the information area;

[0031] Calculating a minimum zoom ratio based on the height, the width, a preset minimum pixel threshold, and a preset minimum zoom threshold;

[0032] The scaling factor of the image to be processed is calculated based on the minimum scaling ratio, the information density, a preset threshold for triggering dynamic adjustment of the scaling factor, and the dynamic scaling factor formula.

[0033] In addition, to achieve the above-mentioned purpose, the present application also proposes a picture text extraction device, which includes:

[0034] Image acquisition module, used to obtain images to be processed;

[0035] The text extraction module is used to input the image to be processed into a pre-built image-text extraction model and output the image text. The image-text extraction model is based on a visual language model and also includes a dynamic input pruning strategy and a sparse attention matrix.

[0036] In addition, to achieve the above-mentioned purpose, the present application also proposes a picture text extraction device, which includes: a memory, a processor, and a computer program stored on the memory and runnable on the processor, and the computer program is configured to implement the steps of the picture text extraction method described above.

[0037] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by the processor, the steps of the image text extraction method described above are implemented.

[0038] One or more technical solutions proposed in this application have at least the following technical effects:

[0039] This application uses a dynamic input pruning strategy to reduce the size of input data as much as possible while retaining the effective features of the original data, thereby reducing resource consumption and improving the reasoning speed of the model. By introducing a sparse attention matrix, the amount of calculation can be reduced without changing the model weights. This application is based on a visual language model and combines two optimization strategies, dynamic input pruning strategy and sparse attention matrix. Starting from the two aspects of input processing and model calculation, it optimizes the different performance bottlenecks of large models, and ultimately achieves high recognition accuracy in image text extraction scenarios, strong generalization ability, low resource consumption, and short reasoning time. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0041] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0042] Figure 1 A flowchart of the first embodiment of the image text extraction method of this application is provided;

[0043] Figure 2 This is a flow chart of the image text extraction method of this application for processing an image using a dynamic input pruning strategy;

[0044] Figure 3 A schematic diagram of improving the BigBird sparse matrix;

[0045] Figure 4 Schematic diagram for calculating sparse attention weights;

[0046] Figure 5 This is the information density-scaling factor curve for this application;

[0047] Figure 6 This is a schematic diagram of the module structure of the image text extraction device according to an embodiment of the present application;

[0048] Figure 7 This is a schematic diagram of the device structure of the hardware operating environment involved in the image text extraction method in the embodiment of the present application.

[0049] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0050] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.

[0051] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0052] The main solution of the embodiment of the present application is: obtaining the image to be processed; inputting the image to be processed into a pre-built image-text extraction model, and outputting the image text. The image-text extraction model is based on a visual language model and also includes a dynamic input pruning strategy and a sparse attention matrix.

[0053] In this embodiment, for ease of description, the following description is made with the image text extraction system as the execution subject.

[0054] The current method of extracting text from images using large models has the problem that resource consumption and inference time increase geometrically with the image size.

[0055] The present application provides a solution. Through the dynamic input pruning strategy, the present application can reduce the size of the input data as much as possible while retaining the effective features of the original data, thereby reducing resource consumption while improving the reasoning speed of the model. By introducing a sparse attention matrix, the amount of calculation can be reduced without changing the model weights. The present application is based on a visual language model, combining two optimization strategies, dynamic input pruning strategy and sparse attention matrix, and starting from the two aspects of input processing and model calculation, respectively, to optimize the different performance bottlenecks of large models, and ultimately achieve the effect of high recognition accuracy in image text extraction scenarios, strong generalization ability, low resource consumption, and short reasoning time.

[0056] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or an electronic device, image text extraction device, etc. that can implement the above functions, or an electronic system, image text extraction system, etc. that can implement the above functions. The following uses the image text extraction system as an example to illustrate this embodiment and the following embodiments.

[0057] Based on this, the present application embodiment provides a method for extracting text from an image. Figure 1 , Figure 1 A flowchart of the first embodiment of the image text extraction method of this application is provided.

[0058] In this embodiment, the image text extraction method includes steps S10 to S20:

[0059] Step S10, obtaining the image to be processed;

[0060] In step S20, the image to be processed is input into a pre-built image-text extraction model to output the image text. The image-text extraction model is based on a visual language model and also includes a dynamic input pruning strategy and a sparse attention matrix.

[0061] It should be noted that the visual language model is a multimodal AI model that can process both visual (image / video) and textual information. Visual language models include many models and frameworks, covering different types such as open source, closed source, general, and specialized. In this embodiment of the application, the Tongyi Qianwen visual language model can be used.

[0062] It's important to note that dynamic input pruning is a technique for optimizing computational efficiency, ultimately aiming to reduce computational effort and increase processing speed. This strategy can include dynamically trimming redundant areas (such as solid backgrounds) based on the complexity of the input image content, retaining only key areas likely to contain text for subsequent processing. It can also include image compression and dimensionality reduction.

[0063] It's important to note that sparse attention is an optimized variant of traditional dense attention. By forcing most elements of the attention weight matrix to zero, only connections at a few key locations are retained, significantly reducing computational and memory overhead. The core idea is that not all input elements need to interact with each other; only the most relevant parts need to be focused. In text extraction tasks, sparse attention allows the model to focus more closely on text regions, reducing redundant computation.

[0064] This embodiment provides a method for extracting text from an image. Through a dynamic input pruning strategy, this application can reduce the size of the input data as much as possible while retaining the effective features of the original data, thereby reducing resource consumption while improving the reasoning speed of the model. By introducing a sparse attention matrix, the amount of calculation can be reduced without changing the model weights. This application is based on a visual language model and combines two optimization strategies, a dynamic input pruning strategy and a sparse attention matrix. Starting from the two aspects of input processing and model calculation, the application optimizes the different performance bottlenecks of large models, and ultimately achieves high recognition accuracy, strong generalization ability, low resource consumption, and short reasoning time in the image text extraction scenario.

[0065] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the first embodiment can be referred to the above introduction, and no further details will be given later. On this basis, step S20, the image to be processed is input into the pre-built image text extraction model, and the output of the image text also includes steps S21 to S24:

[0066] Step S21, performing image processing on the image to be processed using the dynamic input pruning strategy;

[0067] It should be noted that in the reasoning process of large models, the scale of input data directly affects the computational complexity and reasoning speed of the model.

[0068] Large multimodal models not only need to process plain text, but also need to process high-data-volume inputs such as images, which leads to longer inference time and greater resource consumption.

[0069] For example, an image with a resolution of 1024*768 and a bit depth of 24 has a data size of 2359296 bytes, which is approximately equivalent to 1179648 Chinese characters. This shows that large multimodal models take a long time to process images and consume a lot of resources, so dynamic input pruning strategies are needed to process images.

[0070] The dynamic input pruning strategy may include cropping the image, compressing the image, and / or reducing its dimensionality.

[0071] Step S24: input the processed image into the improved Tongyi Qianwen visual language model and output the image text, wherein the improved Tongyi Qianwen visual language model includes the sparse attention matrix.

[0072] It should be noted that the Tongyi Qianwen visual language model is a multimodal large model and a visually enhanced version of the Tongyi Qianwen (Qwen) large model series. Its core goal is to achieve joint understanding of images and text, supporting cross-modal tasks such as visual question answering, image-text generation, and text extraction.

[0073] Among them, the Tongyi Qianwen Visual Language Model (Qwen-VL) can be improved to adapt to text extraction tasks, especially by introducing a sparse attention matrix to optimize computational efficiency and accuracy.

[0074] This application processes the image to be processed using a dynamic input pruning strategy, which can minimize the size of the input data while preserving the valid features of the original data, thereby reducing resource consumption and improving the model's inference speed. By introducing a sparse attention matrix, the amount of computation can be reduced without changing the model weights.

[0075] Currently, there are two main technologies for image text extraction: one uses traditional machine learning small model recognition, and the other uses multimodal large model recognition.

[0076] Small models have fast recognition speed but low accuracy, high requirements for image quality, poor generalization ability, and weak versatility. They require specialized models to be trained for different scenarios.

[0077] Large models have high recognition accuracy and strong generalization capabilities, but they consume a lot of resources and take a long time to reason. Resource consumption and reasoning time increase geometrically with image size.

[0078] Based on the above problems and the above embodiments of this application, the third embodiment of this application is proposed. In the third embodiment of this application, the same or similar contents as those of the above embodiments can be referred to the above introduction and will not be repeated hereafter. On this basis, step S21, performing image processing on the image to be processed by the dynamic input pruning strategy, further includes steps S211 to S213:

[0079] Step S211, performing text area recognition on the image to be processed using a pre-built optical character recognition model to obtain the text area;

[0080] Among them, text region recognition is the first step of input pruning. Figure 2 , Figure 2 This is a flow chart of the image text extraction method of this application that processes the image to be processed through a dynamic input pruning strategy. This application uses an optical character recognition model to identify text areas in the input image and obtain the area containing text.

[0081] As an implementation method, the optical character recognition model can be a lightweight pre-trained small model PaddlePaddle optical character recognition fourth generation PP-OCRv4, or EasyOCR, TrOCR, etc.

[0082] The above optical character recognition model has low computational complexity and can quickly locate the text area in the image, providing a basis for subsequent pruning operations. The coordinates of the text area T can be marked as

[0083] [x left-top ,y left - top ],[x right-top ,y right-top ],[x left-bottom ,y left-bottom ],[x right-bottom ,y right-bottom ]]

[0084] Step S212: cropping the invalid portion of the image to be processed based on the text area to retain the information area;

[0085] Among them, reference Figure 2 , Figure 2 This is a flowchart of the image text extraction method of this application using a dynamic input pruning strategy to process the image to be processed. After the text area recognition is completed, the invalid part outside the area can be cropped to retain the area containing the text.

[0086] It’s important to note that the purpose of cropping is to reduce the size of the input data, thereby reducing the computational complexity of subsequent processing. For example, for an image containing text and background, cropping the background and leaving only the text area can significantly reduce the size of the input data.

[0087] In order to preserve the original typesetting information and reduce the computational complexity, the present invention defines the coordinates of the information area I as

[0088] [[x min ,y min ],[x max ,y min ],[x max ,y max ],[x min ,y max ]]

[0089] Step S213: compress and / or reduce the dimension of the information region to obtain the processed image.

[0090] As an implementation method, the information area may be compressed through lossless compression, such as PNG optimization or binary compression, to obtain a processed image.

[0091] As another implementation, the information area can be compressed by at least one of resolution adjustment (such as adaptive downsampling), color space dimensionality reduction (such as grayscale, color quantization, etc.), frequency domain compression (such as discrete cosine transform, wavelet transform, etc.), and CNN feature extraction to obtain a processed image.

[0092] Optionally, adaptive downsampling can dynamically adjust the resolution according to the text height (eg, when the text height is less than 20 pixels, scale it to 50% of the original size).

[0093] Optionally, grayscale conversion may be performed by converting RGB into a single-channel grayscale image (3 channels → 1 channel), thereby reducing the amount of calculation by 2 / 3.

[0094] Optionally, color quantization may be to reduce the number of colors (eg, 256 colors→16 colors), which is suitable for color text background separation.

[0095] Optionally, discrete cosine transform is performed on image blocks (eg, 8×8), which can retain low-frequency coefficients and discard high-frequency noise, thereby reducing the amount of calculation.

[0096] Optionally, multi-resolution analysis and data compression can be achieved through Haar / Dobeshi wavelet transform.

[0097] Alternatively, a lightweight CNN can be used to extract local feature maps instead of the original high-resolution input.

[0098] This invention is based on a large model (Tongyi Qianwen visual language model), drives the large model through a small model (optical character recognition model), and solves the problems of low recognition accuracy and poor generalization ability of the small model in image text extraction, high resource consumption of the large model, and long inference time through algorithm optimization.

[0099] Based on the above embodiments of the present application, in the fourth embodiment of the present application, the same or similar contents as those in the above embodiments can be referred to the above introduction and will not be described in detail. On this basis, step S213 compresses and / or reduces the dimension of the information area to obtain the processed image, which also includes steps S2131 to S2134:

[0100] Step S2131, calculating information density based on the text area and the information area;

[0101] Information density is an important metric for measuring the information content of input data. The calculated information density can be used to dynamically adjust the scaling of the input data. A higher information density means the input data contains more useful information, so a higher scaling ratio may be required to preserve more detail. A lower information density, on the other hand, means the input data contains more redundant information, so a lower scaling ratio can be used to reduce the computational effort.

[0102] In this scheme, the information density d is defined as the text area S T Information area S I The ratio of

[0103]

[0104] Step S2132: Calculate the scaling factor of the image to be processed based on the information area, the information density, and a pre-established dynamic scaling factor formula;

[0105] The dynamic scaling factor formula is the core of input pruning. By substituting the information density d into the dynamic scaling factor formula, we can calculate the scaling factor that is appropriate for the current input data.

[0106] Step S2133, scaling the information area proportionally based on the scaling factor;

[0107] The scaling factor is defined as f, and the information area I is scaled proportionally according to the calculated scaling factor f to reduce the input amount and retain the clarity of the image text.

[0108] Step S2134: grayscale the scaled information area to obtain the processed image.

[0109] Among them, the scaled image is grayscaled, which can reduce the subsequent calculation complexity by reducing the number of color channels and the dimension of input data.

[0110] The present application calculates the information density based on the text area and the information area, calculates the scaling factor of the image to be processed based on the information area, the information density and a pre-constructed dynamic scaling factor formula, scales the information area proportionally based on the scaling factor, and grayscales the scaled information area to obtain the processed image. This can reduce the size of the input data as much as possible while retaining the effective features of the original data, thereby reducing resource consumption while improving the inference speed of the model.

[0111] Based on the above embodiments of the present application, in the fifth embodiment of the present application, the same or similar contents as those in the above embodiments can be referred to the above introduction, and no further description will be given later. On this basis, step S24, the processed image is input into the improved Tongyi Qianwen visual language model, and steps S22 to S23 are also included before outputting the image text:

[0112] Step S22, obtaining the Tongyi Qianwen visual language model;

[0113] Among them, the Tongyi Qianwen visual language model can be obtained from official channels.

[0114] Step S23 , introducing the BigBird sparse matrix into the Tongyi Qianwen visual language model and improving the Tongyi Qianwen visual language model to obtain the improved Tongyi Qianwen visual language model.

[0115] It's important to note that attention mechanisms are computationally intensive during inference in large models. The visual module of the Tongyi Qianwen visual language model is based on SDPA (Scaled Dot-Product Attention), which requires computing features at all positions. The computational complexity increases with the square of the sequence length, resulting in high computational resource consumption and inference latency.

[0116] In the image text extraction scenario, close areas in the image have strong local correlation in physical location, semantics, and vision, while areas farther away have weaker correlation.

[0117] Based on this, this application converts the attention matrix of the visual module into a sparse matrix without changing the model weights and structure, reducing the amount of calculation, improving the reasoning speed, and maintaining a high recognition accuracy.

[0118] Among them, BigBird sparse matrix is ​​a sparse attention mechanism that reduces redundant attention connections in the Transformer converter, thereby reducing computational complexity. BigBird sparse attention has been shown to have good results in long sequence tasks.

[0119] As an implementation method, the fully connected matrix of the original Transformer can be replaced with the sparse matrix of BigBird in the Tongyi Qianwen visual language model to obtain the improved Tongyi Qianwen visual language model.

[0120] As another implementation, the Tongyi Qianwen visual language model can also be modified using a local window sparse matrix so that each image block or text word only interacts with elements within a neighboring fixed window (such as a 3×3 area) rather than global calculations.

[0121] As another implementation, the Tongyi Qianwen visual language model can also be modified using a block sparse matrix to divide the input image and text sequence into blocks of fixed size (such as 8×8 pixel blocks or text blocks consisting of 16 words). Each block maintains full connection within, while only some key connections are retained between blocks.

[0122] This application introduces an improved sparse matrix to reduce the amount of calculation without changing the model weights.

[0123] Based on the above embodiments of the present application, in the sixth embodiment of the present application, the same or similar contents as those in the above embodiments can be referred to the above introduction and will not be repeated hereafter. On this basis, step S23 introduces the BigBird sparse matrix into the Tongyi Qianwen visual language model and improves the Tongyi Qianwen visual language model to obtain the improved Tongyi Qianwen visual language model, further comprising steps S231 to S234:

[0124] Step S231, replacing the random attention in the BigBird sparse matrix with the information region attention to obtain the sparse attention matrix;

[0125] It should be noted that the present invention improves the BigBird algorithm by replacing the random attention in BigBird with information region attention to achieve full coverage of effective features.

[0126] The improved sparse attention Figure 3 As shown, Figure 3 A schematic diagram of improving the BigBird sparse matrix.

[0127] Among them, the blue part represents global attention, and the blue part focuses on global features and is used to integrate long-distance dependencies.

[0128] The green part represents the local window attention, which focuses on local features and limits the attention range to a fixed-size neighborhood, reducing the complexity from O(n 2 ) is reduced to O(n·k).

[0129] The orange part represents the information region attention, which includes the information region I calculated in the input pruning and is used to cover the valid features.

[0130] By merging the above three regions, a sparse mask can be constructed.

[0131] Step S232, performing a logical AND operation on the sparse attention matrix and the attention mask to obtain a sparse attention mask;

[0132] Step S233, performing a weighted operation on the sparse attention mask based on a preset attention weight to obtain a sparse attention weight;

[0133] Among them, you can refer to Figure 4 , Figure 4 This is a diagram showing how to calculate sparse attention weights. The VisionSdpaAttention source code for the multimodal model uses the Attention Mask to add the Attention Weight to mask the padding weights.

[0134] Based on this, the Sparse Mask sparse mask is logically ANDed with the original Attention Mask attention mask to obtain a new Sparse Attention Mask sparse attention mask, which replaces the original Attention Mask attention mask. In this way, when the scaled_dot_product_attentionScaled scaling dot product attention operation is performed, the original weight matrix will be converted into a sparse matrix by the Sparse Attention Mask sparse attention mask, thereby reducing the amount of calculation.

[0135] Step S234: improving the Tongyi Qianwen visual language model based on the sparse attention weights to obtain the improved Tongyi Qianwen visual language model.

[0136] This application introduces an improved sparse matrix to reduce the amount of calculation without changing the model weights.

[0137] Based on the above embodiments of the present application, in the seventh embodiment of the present application, the same or similar contents as those in the above embodiments can be referred to above and will not be described in detail. On this basis, step S2132, based on the information area, the information density and the pre-constructed dynamic scaling factor formula, calculates the scaling factor of the image to be processed, including steps S21321 to S21323:

[0138] Step S21321, obtaining the height and width of the information area;

[0139] The formulas for obtaining the height h and width w are as follows:

[0140] w=x max -x min

[0141] h=y max -y min

[0142] Among them, x can be obtained based on the coordinates of information area I max 、x min 、y max 、y min .

[0143] Step S21322: Calculate a minimum zoom ratio based on the height, the width, a preset minimum pixel threshold, and a preset minimum zoom threshold;

[0144] The minimum scaling ratio s can be calculated as follows:

[0145] s=max(min(MIN_PIEXL / min(w,h),1.0),MIN_SCALE)

[0146] h is the height of the information area I, and w is the width of the information area I.

[0147] Among them, MIN_PIEXL is the preset minimum pixel threshold, which is used to ensure the resolution of low-pixel images.

[0148] MIN_SCALE is the preset minimum scaling threshold, which is used to ensure the scaling ratio of low-pixel images.

[0149] Step S21323: Calculate the scaling factor of the image to be processed based on the minimum scaling ratio, the information density, a preset threshold for triggering dynamic adjustment of the scaling factor, and the dynamic scaling factor formula.

[0150] The scaling factor f can be calculated as:

[0151]

[0152] Where MIN_DENSITY is the preset threshold that triggers dynamic adjustment of the scaling factor, which is used to control the sensitivity of the impact of information density on the scaling ratio. s is the minimum scaling ratio, and d is the information density.

[0153] Reference Figure 5 , Figure 5 This is the information density-scaling factor curve for this application. When the information density d>MIN_DENSITY, the exponential term will quickly approach 0, causing the scaling factor f to approach 1.0;

[0154] When the information density d is less than MIN_DENSITY, the exponential term increases and the denominator becomes larger overall, so that the scaling factor f approaches the minimum scaling ratio s.

[0155] The final scaling factor f is taken in a continuous range between MIN_SCALE and 1.0 under the influence of multiple factors such as threshold and information density d.

[0156] By dynamically adjusting the scaling factor, the scaling ratio of the input data can be made more flexible, thereby reducing the size of the input data as much as possible and improving the inference speed while ensuring information integrity.

[0157] It should be noted that the above examples are only used to understand this application and do not constitute a limitation on the image text extraction method of this application. More simple transformations based on this technical concept are all within the scope of protection of this application.

[0158] This application also provides a device for extracting text from images. Figure 6 , the image text extraction device includes:

[0159] The image acquisition module 10 is used to acquire the image to be processed;

[0160] The text extraction module 20 is used to input the image to be processed into a pre-built image-text extraction model and output the image text. The image-text extraction model is based on a visual language model and also includes a dynamic input pruning strategy and a sparse attention matrix.

[0161] The image-text extraction device provided in this application, which employs the image-text extraction method of the aforementioned embodiment, can address the technical issues with current methods for extracting image text using large models, where resource consumption and inference time increase geometrically with image size. Compared to the prior art, the image-text extraction device provided in this application has the same beneficial effects as the image-text extraction method provided in the aforementioned embodiment, and the other technical features of the image-text extraction device are the same as those disclosed in the aforementioned embodiment, and are not further elaborated here.

[0162] The present application provides a picture text extraction device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the picture text extraction method in the above-mentioned embodiment one.

[0163] Reference below Figure 7 , which shows a schematic structural diagram of a device for extracting text from an image suitable for implementing an embodiment of the present application. The device for extracting text from an image in an embodiment of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The picture text extraction device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0164] like Figure 7 As shown, the image text extraction device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory 1002 or programs loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the image text extraction device. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the image text extraction device to communicate with other devices wirelessly or wired to exchange data. Although the figure shows an image text extraction device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or provided instead.

[0165] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.

[0166] The image text extraction device provided in this application, which employs the image text extraction method of the aforementioned embodiment, can address the technical issues with current methods for extracting image text using large models, where resource consumption and inference time increase geometrically with image size. Compared to the prior art, the image text extraction device provided in this application has the same beneficial effects as the image text extraction method provided in the aforementioned embodiment, and the other technical features of the image text extraction device are the same as those disclosed in the method of the previous embodiment, and are not further elaborated here.

[0167] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0168] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0169] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, a computer program) stored thereon, and the computer-readable program instructions are used to execute the image text extraction method in the above embodiment.

[0170] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0171] The computer-readable storage medium may be included in the image-text extraction device; or it may exist independently without being assembled into the image-text extraction device.

[0172] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the image-text extraction device, the image-text extraction device: obtains the image to be processed; inputs the image to be processed into a pre-built image-text extraction model, and outputs the image text. The image-text extraction model is based on a visual language model and also includes a dynamic input pruning strategy and a sparse attention matrix.

[0173] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0174] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0175] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.

[0176] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., a computer program) for executing the above-mentioned image text extraction method. This computer-readable storage medium can solve the technical problem that the current method of extracting text from images using large models has resource consumption and inference time that increases geometrically with the size of the image. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the image text extraction method provided in the above-mentioned embodiment, and will not be repeated here.

[0177] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A method for extracting text from an image, characterized in that: The method comprises: Get the image to be processed; The image to be processed is input into a pre-built image-text extraction model to output the image text. The image-text extraction model is based on a visual language model and also includes a dynamic input pruning strategy and a sparse attention matrix.

2. The method according to claim 1, wherein The step of inputting the image to be processed into a pre-built image text extraction model and outputting the image text comprises: Performing image processing on the image to be processed by using the dynamic input pruning strategy; The processed image is input into the improved Tongyi Qianwen visual language model, and the image text is output. The improved Tongyi Qianwen visual language model includes the sparse attention matrix.

3. The method according to claim 2, wherein The step of performing image processing on the image to be processed by using the dynamic input pruning strategy includes: Performing text area recognition on the image to be processed using a pre-built optical character recognition model to obtain the text area; Based on the text area, the invalid part of the image to be processed is cropped to retain the information area; The information region is compressed and / or dimensionally reduced to obtain the processed image.

4. The method according to claim 3, wherein The step of compressing and / or reducing the dimension of the information region to obtain the processed image comprises: Calculating information density based on the text area and the information area; Calculating a scaling factor of the image to be processed based on the information area, the information density, and a pre-established dynamic scaling factor formula; Performing geometric scaling on the information area based on the scaling factor; Grayscale processing is performed on the scaled information area to obtain the processed image.

5. The method according to claim 2, wherein The step of inputting the processed image into the improved Tongyi Qianwen visual language model and outputting the image text includes: Obtain the Tongyi Qianwen visual language model; The BigBird sparse matrix is ​​introduced into the Tongyi Qianwen visual language model and the Tongyi Qianwen visual language model is improved to obtain the improved Tongyi Qianwen visual language model.

6. The method according to claim 5, wherein The step of introducing the BigBird sparse matrix into the Tongyi Qianwen visual language model and improving the Tongyi Qianwen visual language model to obtain the improved Tongyi Qianwen visual language model comprises: Replacing the random attention in the BigBird sparse matrix with the information region attention to obtain the sparse attention matrix; Performing a logical AND operation on the sparse attention matrix and the attention mask to obtain a sparse attention mask; Performing a weighted operation on the sparse attention mask based on a preset attention weight to obtain a sparse attention weight; The Tongyi Qianwen visual language model is improved based on the sparse attention weight to obtain the improved Tongyi Qianwen visual language model.

7. The method according to claim 4, wherein The step of calculating the scaling factor of the image to be processed based on the information area, the information density and a pre-constructed dynamic scaling factor formula includes: Get the height and width of the information area; Calculating a minimum zoom ratio based on the height, the width, a preset minimum pixel threshold, and a preset minimum zoom threshold; The scaling factor of the image to be processed is calculated based on the minimum scaling ratio, the information density, a preset threshold for triggering dynamic adjustment of the scaling factor, and the dynamic scaling factor formula.

8. A device for extracting text from an image, characterized in that: The device comprises: Image acquisition module, used to obtain images to be processed; The text extraction module is used to input the image to be processed into a pre-built image-text extraction model and output the image text. The image-text extraction model is based on a visual language model and also includes a dynamic input pruning strategy and a sparse attention matrix.

9. A device for extracting text from an image, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the image text extraction method according to any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the image text extraction method according to any one of claims 1 to 7 are implemented.