Text detection method and device
By using the Harris corner detection algorithm and convolutional neural network model to detect and feature extraction of natural scene text images in natural scene text detection, the problems of poor adaptability and high detection complexity of text images in the prior art are solved, and higher detection accuracy and process simplification are achieved.
Patent Information
- Application Number
- CN202510272554.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-10
AI Technical Summary
The prior art has poor adaptability to the detection and recognition process of text images in natural scenes, and has a high complexity.
The original image is detected by the Harris corner point detection algorithm, and weighted fusion is performed with preset transparent parameters. Then, the convolutional neural network model and the Vision Transformer network model are used for feature extraction, and mapped to character probabilities through linear classifiers and softmax functions to generate extracted text data.
It improves the adaptability to text data detection with rich and diverse shapes, reduces the complexity of the text detection process, and improves the accuracy of the final extracted text data.
Smart Images

Figure CN119785362B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to a text detection method and device. Background Art
[0002] The integration of multiple forms of expression such as pictures, videos, text, and voice has led to a myriad of ways for people to capture information. Recently, the demand for text detection applications in natural scenes has gradually increased, covering areas such as information notices, smart transportation, guide systems, image retrieval, and content prompts. These applications have put forward higher-level requirements on the ability of algorithms to extract information from the environment, requiring algorithms to accurately detect and recognize text in complex natural environments to provide useful information and improve user experience. In particular, in the context of data informatization and network security, the exchange of information on online networks requires content security filtering of text on pictures or videos, that is, identifying and detecting threatening, offensive, fraudulent, and sensitive information content.
[0003] Traditional OCR algorithms mainly recognize and detect text for standard documents with clean backgrounds, concise text, and a single style. However, when faced with natural scene text content with different and changing styles, traditional OCR algorithms perform poorly. For example, text images in natural scenes have problems such as different regional text sizes and font aspect ratio differences due to perspective deformation, such as larger near and smaller far away due to the angle of the camera; image texts in natural scenes have different directions and forms, such as tilt, bend, twist, etc.; compared with the uniform and standardized background of common data documents, text images in natural scenes have problems such as strong light reflection, incomplete occlusion, and texture similarity caused by light exposure, field of view angle, and background material. The above-mentioned problems of text images in natural scenes are all interference factors for OCR algorithms to detect and recognize text, which undoubtedly greatly reduces the effect of OCR algorithms on text image recognition in natural scenes. In addition, due to the different display purposes and requirements of text images in natural scenes, the printed form or font form of text is often not displayed in a unified way, but is either neat and standardized with clear strokes, or handwritten with varied strokes and artistic connection. The resulting unclear boundaries between letters and difficulty in segmenting strokes reduce the recognition accuracy of the OCR algorithm.
[0004] In order to improve the accuracy of text image detection and recognition in natural scenes, the following two algorithms have been proposed in recent years to improve the traditional OCR algorithm, including: one is a sliding window-based method: this method uses a sliding window of fixed or variable size on the image to gradually scan the entire image to detect possible text areas. Whenever the sliding window covers an area in the image, the algorithm will analyze the area to determine whether the area contains text. The other is a semantic segmentation-based method: this method realizes pixel-level recognition of text in the image, accurately identifies the text area in the image by distinguishing text pixels from non-text pixels, and accurately locates the position of the text. However, the sliding window-based method has the problem that the preset text box cannot cover all special-shaped texts. The semantic segmentation-based method requires complex post-processing after the pixel features are aggregated into the result text line. Therefore, the existing technology still has the problem of poor adaptability to the detection of text images in natural scenes and high complexity of the detection and recognition process of text images. Summary of the invention
[0005] The present disclosure provides a text detection method and device, the main purpose of which is to solve the problem that the existing technology has poor adaptability to the detection of text images in natural scenes and the detection and recognition process of text images is highly complex.
[0006] According to a first aspect of the present disclosure, a text detection method is provided, comprising:
[0007] Perform corner detection on the text data of the original image based on the Harris corner detection algorithm to obtain a corner detection map;
[0008] Performing weighted fusion on the corner point detection image and the original image based on a preset transparency parameter to obtain a fused image;
[0009] Extracting features from the fused image based on a preset convolutional neural network model and a Vision Transformer network model, and mapping the extracted features to character probabilities based on a linear classifier and a softmax function;
[0010] Extracted text data is generated based on the character probability of each extracted feature.
[0011] Optionally, performing corner point detection on the text data of the original image based on the Harris corner point detection algorithm to obtain a corner point detection map includes:
[0012] Defining a local window on the original image based on the Harris corner detection algorithm, and obtaining a first window image grayscale corresponding to the local window;
[0013] By moving the local window in the horizontal direction and the vertical direction of the original image, a second window image grayscale corresponding to the moved local window is obtained;
[0014] Constructing a local window gradient change function based on the grayscale of the first window image, the grayscale of the second window image and a preset window function;
[0015] Defining a second-order covariance matrix based on a Taylor expansion of the local window gradient variation function;
[0016] Calculating a first eigenvalue and a second eigenvalue of the second-order covariance matrix;
[0017] Calculate a corner point evaluation coefficient value based on a corner point evaluation function combined with the first eigenvalue and the second eigenvalue;
[0018] Based on the corner point evaluation coefficient value, determine whether the local window moving area is a target corner point or a target endpoint. Similarly, traverse the original image based on the local window to obtain the corner point detection map composed of each target corner point and each target endpoint in the original image.
[0019] Optionally, determining whether the local window moving area is a target corner point or a target endpoint based on the corner point evaluation coefficient value includes:
[0020] When the corner point evaluation coefficient value is less than a first preset threshold, determining the local window moving area as a target endpoint; or
[0021] When the corner point evaluation coefficient value is greater than a second preset threshold, the local window moving area is determined to be a target corner point.
[0022] Optionally, the formula of the corner point evaluation function is expressed as:
[0023]
[0024] in, is the first eigenvalue of the second-order covariance matrix, is the second-order covariance moment The second eigenvalue of , k is a constant weight coefficient, is the corner point evaluation coefficient value.
[0025] Optionally, the formula of the local window gradient change function is expressed as:
[0026]
[0027] in, is the center coordinate of the local window on the original image, is the grayscale of the first window image, x is the displacement of the local window in the x-axis direction, y is the displacement of the local window in the y-axis direction, is the grayscale of the second window image, is the preset window function, used to represent the coordinates of the local window on the original image The weight of each pixel in the local window.
[0028] Optionally, performing weighted fusion on the corner detection image and the original image based on a preset transparency parameter to obtain a fused image includes:
[0029] Calculating a first product of the preset transparency parameter and the corner detection image and calculating a first difference value and a second product of the original image, wherein the first difference value is a difference value of a value one minus the preset transparency parameter;
[0030] The sum of the first product and the second product is calculated to obtain the fused image.
[0031] According to a second aspect of the present disclosure, a text detection device is provided, comprising:
[0032] A detection unit, used for performing corner point detection on the text data of the original image based on the Harris corner point detection algorithm to obtain a corner point detection map;
[0033] A fusion unit, configured to perform weighted fusion on the corner detection image and the original image based on a preset transparency parameter to obtain a fused image;
[0034] An extraction unit, used for performing feature extraction on the fused image based on a preset convolutional neural network model and a Vision Transformer network model, and mapping the extracted features to character probabilities based on a linear classifier and a softmax function;
[0035] The generating unit is used to generate the extracted text data based on the character probability of each extracted feature.
[0036] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0037] at least one processor; and
[0038] a memory communicatively connected to the at least one processor; wherein,
[0039] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect.
[0040] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described in the first aspect.
[0041] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method as described in the first aspect above.
[0042] The text detection method and device provided by the present invention perform corner detection on the text data of the original image based on the Harris corner detection algorithm to obtain a corner detection map; perform weighted fusion on the corner detection map and the original image based on a preset transparency parameter to obtain a fused image; perform feature extraction on the fused image based on a preset convolutional neural network model and a Vision Transformer network model, and map the extracted features to character probabilities based on a linear classifier and a softmax function; and generate extracted text data based on the character probability of each extracted feature. Compared with the related art, by performing corner point detection on the text data of the original image based on the Harris corner point detection algorithm, a corner point detection map is obtained. Even if the shapes of the text data are rich and diverse, the Harris corner point detection algorithm can still locate and extract the key corner points in the original image more accurately, thereby achieving good adaptability to the detection of text data with rich and diverse shapes. All key corner points extracted from the original image are composed of the corner point detection map, and the corner point detection map and the original image are fused, so that the key corner points form constraint supervision on the extracted features in the subsequent feature extraction process, thereby improving the accuracy of the extracted text data finally obtained. The present invention does not form text lines in the process of performing text detection on the original image, thereby reducing the complexity of the text detection process. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the accompanying drawings:
[0044] Figure 1 A flowchart of a text detection method provided by an embodiment of the present disclosure;
[0045] Figure 2 A schematic diagram of the principle of a Harris corner detection algorithm provided in an embodiment of the present disclosure;
[0046] Figure 3 A schematic diagram of the distribution of key points in a text image provided by an embodiment of the present disclosure;
[0047] Figure 4 A flowchart of another text detection method provided by an embodiment of the present disclosure;
[0048] Figure 5 A structural schematic diagram of a text detection device provided by an embodiment of the present disclosure;
[0049] Figure 6 A schematic block diagram of an exemplary electronic device 300 provided for an embodiment of the present disclosure. DETAILED DESCRIPTION
[0050] The present invention will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that the embodiments and features in the embodiments of the present invention can be combined with each other without conflict.
[0051] The following detailed description is an exemplary description, which is intended to provide further detailed description of the present invention. Unless otherwise specified, all technical terms used in the present invention have the same meaning as those generally understood by those skilled in the art to which the present invention belongs. The terms used in the present invention are only for describing specific embodiments, and are not intended to limit the exemplary embodiments according to the present invention.
[0052] In addition, the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein.
[0053] The text detection method and device of the embodiments of the present disclosure are described below with reference to the accompanying drawings.
[0054] In order to at least solve the problem that the existing technology has poor adaptability to the detection of text images in natural scenes and the detection and recognition process of text images is highly complex, this embodiment provides a text detection method.
[0055] Figure 1 The following is a flow chart of a text detection method provided by an embodiment of the present disclosure. Figure 1 As shown, the method comprises the following steps:
[0056] Step 101, performing corner point detection on the text data of the original image based on the Harris corner point detection algorithm to obtain a corner point detection map;
[0057] Step 102, performing weighted fusion on the corner point detection image and the original image based on a preset transparency parameter to obtain a fused image;
[0058] Step 103, extracting features from the fused image based on a preset convolutional neural network model and a Vision Transformer network model, and mapping the extracted features to character probabilities based on a linear classifier and a softmax function;
[0059] As a refinement of the above step 103, the preset convolutional neural network model includes but is not limited to VGG or ResNet. Based on the preset convolutional neural network model as the backbone network, feature extraction is performed on the fused image to obtain low-dimensional deep-level feature data. The low-dimensional deep-level feature data is used as the input of the Vision Transformer network model to perform further feature extraction. The finally extracted features are mapped to character probabilities based on a linear classifier and a softmax function to complete preliminary recognition of character classification.
[0060] Step 104: Generate extracted text data based on the character probability of each extracted feature.
[0061] The text detection method provided by the present invention performs corner detection on the text data of the original image based on the Harris corner detection algorithm to obtain a corner detection map; performs weighted fusion on the corner detection map and the original image based on a preset transparency parameter to obtain a fused image; performs feature extraction on the fused image based on a preset convolutional neural network model and a Vision Transformer network model, and maps the extracted features to character probabilities based on a linear classifier and a softmax function; and generates extracted text data based on the character probability of each extracted feature. Compared with the related art, by performing corner point detection on the text data of the original image based on the Harris corner point detection algorithm, a corner point detection map is obtained. Even if the shapes of the text data are rich and diverse, the Harris corner point detection algorithm can still locate and extract the key corner points in the original image more accurately, thereby achieving good adaptability to the detection of text data with rich and diverse shapes. All key corner points extracted from the original image are composed of the corner point detection map, and the corner point detection map and the original image are fused, so that the key corner points form constraint supervision on the extracted features in the subsequent feature extraction process, thereby improving the accuracy of the extracted text data finally obtained. The present invention does not form text lines in the process of performing text detection on the original image, thereby reducing the complexity of the text detection process.
[0062] Figure 2 A schematic diagram of the Harris corner detection algorithm principle provided in the embodiment of the present disclosure is provided. In order to facilitate the understanding of the above embodiment, this embodiment is combined with Figure 2The principle of Harris corner detection algorithm is briefly explained, including: Harris corner detection algorithm defines a local window on the image, and considers three cases through the average change of image intensity during the movement of the local window on the image: In the first case, Figure 2 The first window moves in all directions of the hexagonal content area. It can be observed that the pixel values of the local window in all directions do not change significantly, so there is no corner point in the image area. In the second case, Figure 2 The second window moves in one direction (horizontally) without obvious pixel changes, but moves in another direction (vertically) with large changes. This means that the image area meets the edge features, but there are no corner points. In the third case, Figure 2 If the movement of the third window in all directions in the image area has a relatively obvious pixel gradient change, then the area can be judged as a corner point. The types of the local window include but are not limited to the first window, the second window, and the third window. In short, the core principle of the Harris corner detection algorithm is to use the local window to determine the corner points of the image target based on the intensity difference of the image pixels in all directions. That is, if the intensity change value changes greatly in all directions, the pixel point area is considered to be a corner point.
[0063] As a refinement of the embodiment of the present disclosure, when performing corner point detection on the text data of the original image based on the Harris corner point detection algorithm to obtain a corner point detection map in step 101, the following implementation methods may also be adopted but are not limited to, for example: defining a local window on the original image based on the Harris corner point detection algorithm, and obtaining the grayscale of a first window image corresponding to the local window; obtaining the grayscale of a second window image corresponding to the moved local window by moving the local window in the horizontal direction and the vertical direction of the original image; constructing a local window gradient change function based on the grayscale of the first window image, the grayscale of the second window image and a preset window function; defining a second-order covariance matrix based on the Taylor expansion of the local window gradient change function; calculating a first eigenvalue and a second eigenvalue of the second-order covariance matrix; calculating a corner point evaluation coefficient value based on a corner point evaluation function combined with the first eigenvalue and the second eigenvalue; determining whether the local window moving area is a target corner point or a target endpoint based on the corner point evaluation coefficient value, and so on, traversing the original image based on the local window to obtain the corner point detection map composed of each target corner point and each target endpoint in the original image.
[0064] In order to facilitate the understanding of the above embodiment, this embodiment expands the above embodiment in combination with formulas, including: In combination with the above description of the principle of Harris corner detection algorithm, the formula for defining the local window gradient change function is expressed as:
[0065]
[0066] in, is the center coordinate of the local window on the original image, is the grayscale of the first window image, that is, the grayscale of the image in the local window, x is the displacement of the local window in the x-axis direction, y is the displacement of the local window in the y-axis direction, is the grayscale of the second window image, that is, the grayscale of the image in the local window after translation. The change of the grayscale of the image inside the local window caused by the movement of the local window back and forth is expressed as , is the preset window function, used to represent the coordinates of the local window on the original image According to the characteristics that the rectangular edge of the local window has a greater impact on the image grayscale during the movement process, while the rectangular center has a smaller impact, the pixel weights in the local window can be Set to The Gaussian weight distribution as the center is, that is, the weight of the edge pixels of the local window is reduced, and the weight of the central pixels of the local window is increased. The local window is a rectangular window.
[0067] In order to facilitate understanding of the above embodiment, this embodiment defines the preset window function The formula is expressed as:
[0068]
[0069] in, is the center position of the window, is the standard deviation of the Gaussian function, and Relative to the center of the window The horizontal and vertical offsets of , G represents the Gaussian function, and e is the exponential logarithm.
[0070] Furthermore, the Taylor expansion formula of the local window gradient change function is expressed as:
[0071]
[0072] in, represents the second-order covariance matrix, and Relative to the center of the window The specific formula of the second-order covariance matrix is expressed as:
[0073]
[0074] in, Represents the horizontal gradient of the image grayscale in the local window, Represents the vertical gradient of the image grayscale within the local window.
[0075] To achieve efficient calculation, we define and is the gradient covariance matrix The characteristic value of , then the formula for further defining the corner point evaluation function is expressed as:
[0076]
[0077] in, is the first eigenvalue of the second-order covariance matrix, is the second-order covariance moment The second eigenvalue of , k is a constant weight coefficient, is the corner point evaluation coefficient value.
[0078] Figure 3 A schematic diagram of the distribution of key points in a text image provided by an embodiment of the present disclosure, Figure 3 The corner points and endpoints shown in the figure can be the detection results of the text image after being detected by the Harris corner point detection algorithm. It should be understood that the text image is one of the original images.
[0079] As a refinement of the above embodiment, when executing the method of determining whether the local window moving area is a target corner point or a target endpoint based on the corner point evaluation coefficient value, the following implementation methods may also be adopted but are not limited to, for example: when the corner point evaluation coefficient value is less than a first preset threshold value, the local window moving area is determined to be a target endpoint; or when the corner point evaluation coefficient value is greater than a second preset threshold value, the local window moving area is determined to be a target corner point.
[0080] In order to facilitate understanding of the above embodiment, the present embodiment provides further exemplary explanations, including: in the process of the local window moving on the text image, if the local window meets the characteristic of having no obvious gradient change in only one direction, then the area corresponding to the movement of the local window is defined as the target endpoint of the text data in the text image, that is, when the eigenvalue θ determined by the second-order covariance matrix M is a large negative value, then the area is determined as the target endpoint without obvious image intensity change in one direction, that is, less than the first preset threshold. It is defined as the target endpoint, which is the edge feature, and the formula is expressed as:
[0081]
[0082] in, is the corner point evaluation coefficient value, and 0 is the first preset threshold.
[0083] When the eigenvalue θ determined by the matrix M is a relatively large positive value, the region is determined to be a corner point with obvious image intensity changes in almost all directions, that is, when the extracted θ is greater than the second preset threshold:
[0084]
[0085] If it satisfies the situation that the local area has gradient changes in all directions, then the area is determined to be the target corner point of the text data, that is, the turning point of the text data stroke, and 0 is the second preset threshold. When |θ|≈0, this situation represents that there is no intensity change of pixel grayscale in the local window, that is, it is a flat area feature. It should be understood that in this embodiment, the value range of the first preset threshold and the second preset threshold are both set to 0 for example only. In some embodiments, the value of the first preset threshold is less than 0, the value of the second preset threshold is greater than zero, and the interval range composed of the first preset threshold and the second preset threshold is defined as a setting method approximately equal to 0, which is also an implementation method of this embodiment. This embodiment does not limit the specific value range of the first preset threshold and the second preset threshold.
[0086] Figure 4 A flowchart of another text detection method provided by an embodiment of the present disclosure is shown in FIG. Figure 4 The visual model (Vision Model) provided by the present disclosure is shown in the figure, the original image (Input Image), the local window (Adaptive Local Windows Sizes), the corner detection map (Corners Detection Map), the weighted fusion operation can also be expressed as Differential Fusion Transparency a, the fused image (FusionImage), the feature map (Features Map), the Vision Transformer network model can also be expressed as Transformer Sequence Modeling, the position information of the image block (Position Attention), the features extracted by the Vision Transformer network model (Visual Features), and the extracted text data (VisionPrediction). This embodiment is combined with Figure 4An exemplary description is given, including: first, the original image is processed by the corner detection module, and a corner detection map is generated using a predefined local window size to guide the visual model to reduce the noise interference of complex background in text detection. The visual model is a network model constructed based on the Harris corner detection algorithm in this disclosure. However, during the detection process, some corners originating from the background may be mistakenly identified as text strokes, thereby introducing additional noise. In order to solve this problem, the transparency parameter is introduced. That is, the preset transparent parameter controls the corner point detection map I corners With the original image The degree of mixing between allows to preserve the key point information, that is, the information of the corner points, while reducing the interference of the background on the original image. Such a design plays a key role in improving the accuracy and robustness of text detection and effectively copes with the challenges brought by complex backgrounds:
[0087]
[0088] in, I fusion is a function representation for weighted fusion of the corner point detection map and the original image.
[0089] Subsequently, the Vision Transformer (ViT) model using the Transformer architecture extended the spatial and sequence processing capabilities from the field of natural language processing (NLP) to the field of scene text detection. In this framework, the corner fusion map, i.e. the fused image, is passed as input to the backbone network (such as VGG or ResNet) for basic feature extraction. These extracted features are input into the Transformer network together with the position embedding of the image block. In this way, the model can make full use of the sequence modeling capabilities of the Transformer to effectively capture the text information and contextual relationships in the image.
[0090] The Transformer network performs high-level feature extraction and sequence modeling, enabling the model to more fully understand the text content in the image. By converting visual features into character probabilities, the model is able to process different features in parallel and map these features to character probabilities through a combination of linear classifiers and softmax operators, thereby achieving accurate extraction and recognition of text.
[0091] Among them, the loss function of the text detection result is as follows:
[0092]
[0093] in, is the prediction result of the visual model, i.e., the extracted text data outputted in step 104, is the label (true value), weight parameter Used to balance the importance of different samples, regularization parameters Used to control the importance of the regularization term, i represents the subscript variables of different training samples, N is the total number of samples, is the loss value of the loss function. The loss function is used to minimize the difference between the prediction result and the label to complete the training of the visual model.
[0094] As a refinement of the above embodiment, performing weighted fusion of the corner detection map and the original image based on the preset transparent parameter in step 102 to obtain the fused image includes: calculating a first product of the preset transparent parameter and the corner detection map and calculating a first difference and a second product of the original image, wherein the first difference is a difference of a value one minus the preset transparent parameter; and calculating the sum of the first product and the second product to obtain the fused image. For parts not explained in this embodiment, reference may be made to the above exemplary description.
[0095] Table 1 is a model comparison table
[0096]
[0097] In order to evaluate the effectiveness of the stroke keypoint-based text detection (SKP) method in detecting text instances, the study conducted a detailed comparison of its performance with the recent leading arbitrary shape text detector. This comparison aims to reveal the advantages and potential of the SKP method in handling text detection tasks, so as to gain a more comprehensive understanding of its utility in the field of text detection. By comparing with the state-of-the-art, the performance of the SKP method in handling various text shapes and complex scenes can be evaluated, providing important insights for further improving text detection technology. For the ICDAR 2015 dataset, following the conventional practice, the following three dictionaries are used for comparison: "Strong" (S), "Weak" (W), and "Generic" (G). The Strong dictionary contains 100 candidate words specific to each image, the Weak dictionary contains all the words that appear in the test set, and the Generic dictionary contains an extensive vocabulary of 90k words. In the TotalText dataset, "None" means no dictionary is used, and "Full" means all the words that appear in the test set are used as the dictionary. The results are shown in Table 1 above.
[0098] In summary, the embodiments of the present disclosure can achieve the following effects:
[0099] 1. By performing corner point detection on the text data of the original image based on the Harris corner point detection algorithm, a corner point detection map is obtained. Even if the text data has rich and diverse shapes, the Harris corner point detection algorithm can still accurately locate and extract the key corner points in the original image, thereby achieving good adaptability to the detection of text data with rich and diverse shapes. All key corner points extracted from the original image are composed of the corner point detection map, and the corner point detection map and the original image are fused to achieve the key corner points to form constraint supervision on the extracted features in the subsequent feature extraction process, thereby improving the accuracy of the extracted text data finally obtained. The present disclosure does not form text lines in the process of performing text detection on the original image, thereby reducing the complexity of the text detection process.
[0100] 2. The strokes of the text are regarded as components of written text characters. The structural stability of the strokes makes the strokes have key points, which include the target corner points and the target endpoints. In addition, the key points may also include but are not limited to starting points, ending points, inflection points and breakpoints, and then a text detection method based on stroke key points is proposed. For texts in different languages, the method for determining stroke key points can be universal, because these key points are based on the morphological and structural features of the strokes rather than features specific to a certain language. By extracting these key points, important features such as the starting point and end point of the stroke corresponding to the text instance can be determined. Then, by adjusting the transparency parameter, the model can balance the information fusion between the original image and the corner point detection image during training, further improving the detection effect.
[0101] 3. The clustered area of the key points is used as an auxiliary positioning for natural scene text detection. Compared with the previously proposed regression-based method, it can adapt to the positioning of text with complex shapes, reduce the noise interference caused by the redundant background framed by the preset text box, and improve the accuracy of the model.
[0102] Corresponding to the above-mentioned text detection method, the present invention also provides a text detection device. Since the device embodiment of the present invention corresponds to the above-mentioned method embodiment, details not disclosed in the device embodiment can be referred to the above-mentioned method embodiment, and will not be repeated in the present invention.
[0103] Figure 5 A structural diagram of a text detection device provided by an embodiment of the present disclosure is shown in FIG. Figure 4 As shown, including:
[0104] A detection unit 21 is used to perform corner point detection on the text data of the original image based on the Harris corner point detection algorithm to obtain a corner point detection map;
[0105] A fusion unit 22, configured to perform weighted fusion on the corner point detection image and the original image based on a preset transparency parameter to obtain a fused image;
[0106] An extraction unit 23, used for performing feature extraction on the fused image based on a preset convolutional neural network model and a Vision Transformer network model, and mapping the extracted features to character probabilities based on a linear classifier and a softmax function;
[0107] The generating unit 24 is used to generate extracted text data based on the character probability of each extracted feature.
[0108] The text detection device provided by the present disclosure performs corner detection on the text data of the original image based on the Harris corner detection algorithm to obtain a corner detection map; performs weighted fusion on the corner detection map and the original image based on a preset transparency parameter to obtain a fused image; performs feature extraction on the fused image based on a preset convolutional neural network model and a Vision Transformer network model, and maps the extracted features to character probabilities based on a linear classifier and a softmax function; and generates extracted text data based on the character probability of each extracted feature. Compared with the related art, by performing corner point detection on the text data of the original image based on the Harris corner point detection algorithm, a corner point detection map is obtained. Even if the shapes of the text data are rich and diverse, the Harris corner point detection algorithm can still locate and extract the key corner points in the original image more accurately, thereby achieving good adaptability to the detection of text data with rich and diverse shapes. All key corner points extracted from the original image are composed of the corner point detection map, and the corner point detection map and the original image are fused, so that the key corner points form constraint supervision on the extracted features in the subsequent feature extraction process, thereby improving the accuracy of the extracted text data finally obtained. The present invention does not form text lines in the process of performing text detection on the original image, thereby reducing the complexity of the text detection process.
[0109] It should be noted that the above explanation of the method embodiment is also applicable to the device of this embodiment, and the principle is the same, which is not limited in this embodiment.
[0110] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0111] Figure 6A schematic block diagram of an example electronic device 300 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0112] like Figure 6 As shown, the device 300 includes a computing unit 301, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 302 or a computer program loaded from a storage unit 308 to a RAM (Random Access Memory) 303. In the RAM 303, various programs and data required for the operation of the device 300 can also be stored. The computing unit 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An I / O (Input / Output) interface 305 is also connected to the bus 304.
[0113] A number of components in the device 300 are connected to the I / O interface 305, including: an input unit 306, such as a keyboard, a mouse, etc.; an output unit 307, such as various types of displays, speakers, etc.; a storage unit 308, such as a disk, an optical disk, etc.; and a communication unit 309, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 309 allows the device 300 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0114] The computing unit 301 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, etc. The computing unit 301 performs the various methods and processes described above, such as a text detection method. For example, in some embodiments, the text detection method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 308. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 300 via the ROM 302 and / or the communication unit 309. When the computer program is loaded into the RAM 303 and executed by the computing unit 301, one or more steps of the method described above may be performed. Alternatively, in other embodiments, the computing unit 301 may be configured to execute the aforementioned text detection method in any other appropriate manner (eg, by means of firmware).
[0115] Various embodiments of the systems and techniques described above herein may be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System On Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor that may be a dedicated or general-purpose programmable processor that may receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0116] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0117] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), optical storage device, magnetic storage device, or any suitable combination of the foregoing.
[0118] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0119] The systems and techniques described herein may be implemented in a computing system that includes a backend component (e.g., as a data server), or a computing system that includes a middleware component (e.g., an application server), or a computing system that includes a frontend component (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.
[0120] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server may also be a server of a distributed system, or a server combined with a blockchain.
[0121] It should be noted that artificial intelligence is a discipline that studies how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.), and includes both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include computer vision technology, speech recognition technology, natural language processing technology, as well as machine learning / deep learning, big data processing technology, knowledge graph technology, and other major directions.
Claims
1. A text detection method, characterized in that: include: Perform corner detection on the text data of the original image based on the Harris corner detection algorithm to obtain a corner detection map; The corner point detection image and the original image are weightedly fused based on a preset transparency parameter to obtain a fused image: Calculating a first product of the preset transparency parameter and the corner detection image and calculating a first difference value and a second product of the original image, wherein the first difference value is a difference value of a value one minus the preset transparency parameter; Calculating the sum of the first product and the second product to obtain the fused image; Performing feature extraction on the fused image based on a preset convolutional neural network model and a Vision Transformer network model; The preset convolutional neural network model includes but is not limited to VGG or ResNet, and the fused image is subjected to feature extraction based on the preset convolutional neural network model as a backbone network to obtain low-dimensional deep-level feature data, and the low-dimensional deep-level feature data is used as the input of the Vision Transformer network model to perform further feature extraction, and the extracted features are mapped to character probabilities based on a linear classifier and a softmax function; Extracted text data is generated based on the character probability of each extracted feature.
2. The method according to claim 1, characterized in that The Harris corner detection algorithm is used to perform corner detection on the text data of the original image to obtain a corner detection map including: Defining a local window on the original image based on the Harris corner detection algorithm, and obtaining a first window image grayscale corresponding to the local window; By moving the local window in the horizontal direction and the vertical direction of the original image, a second window image grayscale corresponding to the moved local window is obtained; Constructing a local window gradient change function based on the grayscale of the first window image, the grayscale of the second window image and a preset window function; Defining a second-order covariance matrix based on a Taylor expansion of the local window gradient variation function; Calculating a first eigenvalue and a second eigenvalue of the second-order covariance matrix; Calculate a corner point evaluation coefficient value based on a corner point evaluation function combined with the first eigenvalue and the second eigenvalue; Based on the corner point evaluation coefficient value, determine whether the local window moving area is a target corner point or a target endpoint. Similarly, traverse the original image based on the local window to obtain the corner point detection map composed of each target corner point and each target endpoint in the original image.
3. The method according to claim 2, characterized in that The determining whether the local window moving area is a target corner point or a target endpoint based on the corner point evaluation coefficient value comprises: When the corner point evaluation coefficient value is less than a first preset threshold, determining the local window moving area as a target endpoint; or When the corner point evaluation coefficient value is greater than a second preset threshold, the local window moving area is determined to be a target corner point.
4. The method according to any one of claims 2 or 3, characterized in that: The formula of the corner point evaluation function is expressed as: in, is the first eigenvalue of the second-order covariance matrix, is the second-order covariance moment The second eigenvalue of , k is a constant weight coefficient, is the corner point evaluation coefficient value.
5. The method according to claim 4, characterized in that The formula of the local window gradient change function is expressed as: in, is the center coordinate of the local window on the original image, is the grayscale of the first window image, x is the displacement of the local window in the x-axis direction, y is the displacement of the local window in the y-axis direction, is the grayscale of the second window image, is the preset window function, used to represent the coordinates of the local window on the original image The weight of each pixel in the local window.
6. A text detection device, characterized in that: include: A detection unit, used for performing corner point detection on the text data of the original image based on the Harris corner point detection algorithm to obtain a corner point detection map; A fusion unit, configured to perform weighted fusion on the corner detection image and the original image based on a preset transparency parameter to obtain a fused image; Calculating a first product of the preset transparency parameter and the corner detection image and calculating a first difference value and a second product of the original image, wherein the first difference value is a difference value of a value one minus the preset transparency parameter; Calculating the sum of the first product and the second product to obtain the fused image; An extraction unit, used for performing feature extraction on the fused image based on a preset convolutional neural network model and a Vision Transformer network model; The preset convolutional neural network model includes but is not limited to VGG or ResNet, and the fused image is subjected to feature extraction based on the preset convolutional neural network model as a backbone network to obtain low-dimensional deep-level feature data, and the low-dimensional deep-level feature data is used as the input of the Vision Transformer network model to perform further feature extraction, and the extracted features are mapped to character probabilities based on a linear classifier and a softmax function; The generating unit is used to generate the extracted text data based on the character probability of each extracted feature.
7. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.
8. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-6.
9. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 6.