Text detection method and device, electronic equipment and storage medium

By adjusting the neural network structure, combining image entropy value and texture complexity, and optimizing the convolution layer and text box layer, the problem of low text detection efficiency in natural scenes is solved, and efficient and accurate text detection is achieved.

CN120198903APending Publication Date: 2025-06-24SHENZHEN POWER SUPPLY BUREAU
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510287409.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art has low text detection efficiency in natural scenes, making it difficult to effectively improve the detection speed and accuracy of scene text.

Method used

By obtaining the grayscale distribution and grayscale symbiosis matrix of the target scene image, the image entropy value and texture complexity are determined, combined with the target recognition accuracy and efficiency, the convolution layer and text box layer in the neural network are adjusted, convolution and feature extraction are performed, the initial box set and text background are determined, and the detection efficiency is improved.

Benefits of technology

It realizes efficient detection of scene text in complex scenarios, improves detection speed and accuracy, adapts to various scene changes, reduces calculation costs, and enhances adaptability to complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198903A_ABST
    Figure CN120198903A_ABST
Patent Text Reader

Abstract

Embodiments of the invention disclose a text detection method and apparatus, an electronic device and a storage medium. The method comprises the steps of obtaining a target scene image; determining an image entropy value and image texture complexity of the target scene image; obtaining target recognition precision and target recognition efficiency of the target scene image; determining n convolution layers and m textbox layers in a target neural network according to the image entropy, the image texture complexity, the target recognition precision and the target recognition efficiency; performing convolution on the target scene image to obtain n feature maps; acquiring m feature maps from the n feature maps; determining m initial frame sets of the m feature maps according to the target recognition precision, the target recognition efficiency and the scale corresponding to each feature map in the m feature maps; determining m first texts and m first backgrounds of the m feature maps according to the m initial frame sets; and determining a target text according to the m first texts, the m first backgrounds and the scale corresponding to each feature map in the m feature maps.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a text detection method, apparatus, electronic device, and storage medium. Background Art

[0002] Natural scene text is the core carrier for transmitting semantic information in images and plays an irreplaceable role in fields such as intelligent transportation, image retrieval, and mobile device assistance. Especially in complex scenarios such as image geolocation and video understanding, text information can provide key semantic anchors, significantly improving the intelligence level of the system.

[0003] However, due to the diversity of foreground text and background objects, scene text detection is extremely challenging, and using traditional text detection methods will result in low detection efficiency.

[0004] Therefore, there is an urgent need for a text detection method to improve the detection efficiency of scene text. Summary of the Invention

[0005] To solve the above problems, embodiments of the present invention provide a text detection method, apparatus, electronic device, and storage medium, which can improve the detection efficiency of scene text when detecting a target scene image.

[0006] In a first aspect, embodiments of the present invention provide a text detection method, including:

[0007] Obtain a target scene image; the target scene image includes target text;

[0008] Determine the image entropy value of the target scene image according to the gray-scale distribution of the target scene image;

[0009] Determine the image texture complexity of the target scene image according to the gray-level co-occurrence matrix of the target scene image;

[0010] Obtain the target recognition accuracy and target recognition efficiency of the target scene image;

[0011] Determine n convolutional layers and m text box layers in a target neural network according to the image entropy value, the image texture complexity, the target recognition accuracy, and the target recognition efficiency; the m text box layers are connected to m corresponding convolutional layers in the n convolutional layers one by one, where m and n are positive integers, and m is less than or equal to n;

[0012] Perform convolution on the target scene image through the n convolutional layers to obtain n feature maps;

[0013] Obtain m feature maps from the n feature maps through the m text box layers;

[0014] Determine the initial bounding box set for each of the m feature maps according to the target recognition accuracy, the target recognition efficiency, and the scale corresponding to each of the m feature maps, to obtain m initial bounding box sets;

[0015] According to the m initial bounding box sets, determine the first text for each of the m feature maps, and the first background corresponding to each first text, to obtain m first texts and m first backgrounds;

[0016] Determine the target text according to the m first texts, the m first backgrounds, and the scale corresponding to each of the m feature maps.

[0017] In a second aspect, an embodiment of the present invention provides a text detection device, and the device includes an acquisition unit and a processing unit;

[0018] The acquisition unit is configured to acquire a target scene image; the target scene image includes target text;

[0019] The processing unit is configured to determine the image entropy value of the target scene image according to the gray-scale distribution of the target scene image;

[0020] Determine the image texture complexity of the target scene image according to the gray-level co-occurrence matrix of the target scene image;

[0021] Acquire the target recognition accuracy and the target recognition efficiency of the target scene image;

[0022] According to the image entropy value, the image texture complexity, the target recognition accuracy, and the target recognition efficiency, determine n convolutional layers and m text box layers in a target neural network; the m text box layers are connected to m corresponding convolutional layers in the n convolutional layers one by one, m and n are positive integers, and m is less than or equal to n;

[0023] Perform convolution on the target scene image through the n convolutional layers to obtain n feature maps;

[0024] Obtain m feature maps from the n feature maps through the m text box layers;

[0025] Determine the initial bounding box set for each of the m feature maps according to the target recognition accuracy, the target recognition efficiency, and the scale corresponding to each of the m feature maps, to obtain m initial bounding box sets;

[0026] According to the m initial bounding box sets, determine the first text for each of the m feature maps, and the first background corresponding to each first text, to obtain m first texts and m first backgrounds;

[0027] Determine the target text according to the m first texts, the m first backgrounds, and the scale corresponding to each feature map among the m feature maps.

[0028] In a third aspect, an embodiment of the present invention provides an electronic device, which includes a processor and a memory. The processor is connected to the memory. The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device executes the method described in the first aspect.

[0029] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method described in the first aspect.

[0030] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute the method described in the first aspect.

[0031] Implementing the embodiments of the present application has the following beneficial effects:

[0032] In the embodiment of the present application, first, a target scene image is obtained, where the target scene image includes target text. Then, according to the gray distribution of the target scene image, the image entropy value of the target scene image is determined, and according to the gray-level co-occurrence matrix of the target scene image, the image texture complexity of the target scene image is determined. Then, the target recognition accuracy and target recognition efficiency of the target scene image are obtained, and according to the image entropy value, the image texture complexity, the target recognition accuracy, and the target recognition efficiency, n convolutional layers and m text box layers in the target neural network are determined, where the m text box layers are connected to m corresponding convolutional layers in the n convolutional layers one by one, and m and n are positive integers, and m is less than or equal to n. Next, the target scene image is convolved through the n convolutional layers to obtain n feature maps, and m feature maps are obtained from the n feature maps through the m text box layers. Then, according to the target recognition accuracy, the target recognition efficiency, and the scale corresponding to each feature map in the m feature maps, an initial box set of each feature map in the m feature maps is determined to obtain m initial box sets, and according to the m initial box sets, the first text of each feature map in the m feature maps and the first background corresponding to each first text are determined to obtain m first texts and m first backgrounds. Finally, according to the m first texts, the m first backgrounds, and the scale corresponding to each feature map in the m feature maps, the target text is determined. Thus, by determining the n convolutional layers and m text box layers in the target neural network according to the image entropy value, the image texture complexity, the target recognition accuracy, and the target recognition efficiency, and by determining the first text of each feature map in the m feature maps and the first background corresponding to each first text according to the m initial box sets to obtain m first texts and m first backgrounds, and determining the target text according to the m first texts, the m first backgrounds, and the scale corresponding to each feature map in the m feature maps, the detection efficiency of scene text can be improved when detecting the target scene image. Description of the Drawings

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the background art, the following will describe the drawings required to be used in the embodiments of the present invention or the background art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0034] Figure 1 It is a schematic diagram of the architecture of a text detection system provided by an embodiment of the present application;

[0035] Figure 2 It is a flowchart of a text detection method provided by an embodiment of the present application;

[0036] Figure 3It is an application scenario diagram of a text detection method based on the minimum horizontal rectangle provided by an embodiment of the present application;

[0037] Figure 4 It is a flowchart of a text density determination method provided by an embodiment of the present application;

[0038] Figure 5 It is a flowchart of a display method of target text provided by an embodiment of the present application;

[0039] Figure 6 It is a schematic diagram of the neural network architecture of a text detection method provided by an embodiment of the present application;

[0040] Figure 7 It is a schematic structural diagram of a text detection device provided by an embodiment of the present application;

[0041] Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0042] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present application belong to the scope protected by the present application.

[0043] The terms "first", "second", "third", "fourth", etc. in the specification and claims of the present application and the accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not limited to the listed steps or modules, but optionally further includes steps or modules not listed, or optionally further includes other steps or modules inherent to these processes, methods, products, or devices.

[0044] Referring to

[0045] Refer to Figure 1 , Figure 1It is a schematic diagram of the architecture of a text detection system provided by an embodiment of the present application. The text detection system is based on a deep neural network architecture and realizes the accurate extraction of target text from complex scene images through multi-stage processing. After inputting the target scene image into the neural network, the neural network can output the target text corresponding to the target scene image.

[0046] Refer to Figure 2 , Figure 2 It is a flowchart of a text detection method provided by an embodiment of the present application. As Figure 2 shown, a text detection method provided by an embodiment of the present application includes but is not limited to the following steps:

[0047] Step S101: Obtain a target scene image;

[0048] Among them, the target scene image includes target text;

[0049] Step S102: Determine the image entropy value of the target scene image according to the gray-scale distribution of the target scene image;

[0050] Step S103: Determine the image texture complexity of the target scene image according to the gray-level co-occurrence matrix of the target scene image;

[0051] Step S104: Obtain the target recognition accuracy and target recognition efficiency of the target scene image;

[0052] Step S105: Determine n convolutional layers and m text box layers in the target neural network according to the image entropy value, image texture complexity, target recognition accuracy, and target recognition efficiency;

[0053] Among them, the m text box layers are connected to the corresponding m convolutional layers in the n convolutional layers one by one. m and n are positive integers, and m is less than or equal to n;

[0054] Step S106: Convolve the target scene image through the n convolutional layers to obtain n feature maps;

[0055] Step S107: Obtain m feature maps from the n feature maps through the m text box layers;

[0056] Step S108: Determine the initial box set of each feature map in the m feature maps according to the target recognition accuracy, target recognition efficiency, and the scale corresponding to each feature map in the m feature maps, and obtain m initial box sets;

[0057] Step S109: Determine the first text of each feature map in the m feature maps and the first background corresponding to each first text according to the m initial box sets, and obtain m first texts and m first backgrounds;

[0058] Step S110: Determine the target text according to the scales corresponding to each feature map among the m first texts, the m first backgrounds, and the m feature maps.

[0059] In a possible embodiment, the image entropy is expressed as the average number of bits of the set of image gray levels, describes the average information amount of the image information source, and reflects the complexity of the gray level distribution in the image. For an image with a high entropy value, there are many details and variations in brightness, indicating that the image contains a large amount of information and rich content. For an image with a low entropy value, there are not many details and variations, and the amount of information contained is relatively small. The larger the image entropy value, the richer the information carried by the image and the better the image effect.

[0060] In a possible embodiment, if the image entropy value is high, it means that the image contains rich information and high uncertainty, and a neural network with strong feature extraction and representation capabilities is required. A neural network with a deeper and more complex structure can be selected, such as the residual network series, which solves the problems of gradient disappearance and degradation that occur as the network deepens through residual connections, and can better learn the complex features in high-entropy images. For images with low entropy values, the information is relatively simple and certain, and a network with a relatively simple structure, such as a small convolutional neural network, can be used, which can not only reduce the amount of calculation but also avoid overfitting, achieving a better recognition effect. When the image texture complexity is high, the network needs to be able to capture subtle texture changes and detailed information. A network structure with multi-scale convolutional kernels or dilated convolutions can be adopted, such as the densely connected network, whose densely connected method can fuse features at different levels and better process complex textures. If the image texture is simple, a network structure combining ordinary convolutional layers and pooling layers, such as a network structure with a relatively regular structure, can be adopted, which can effectively extract the features of simple texture images while having a relatively small amount of calculation, facilitating the improvement of recognition efficiency. If extremely high requirements are placed on the target recognition accuracy, a powerful network architecture can be selected, such as optimizing dimensions such as the depth, width, and resolution of the network to improve the accuracy. In addition, the method of ensemble learning can also be adopted to fuse multiple different neural networks to further improve the recognition accuracy. When high requirements are placed on the target recognition efficiency, such as in scenarios such as real-time monitoring, a lightweight neural network architecture needs to be selected. By adopting techniques such as depthwise separable convolutions, the amount of calculation and model parameters are greatly reduced, and the inference speed is improved. If the requirements for efficiency are not high and more attention is paid to recognition accuracy, some networks with a large amount of calculation but powerful performance can be selected, which can extract rich features but have a relatively high calculation cost.

[0061] In a possible embodiment, n convolutional layers and m text box layers in the target neural network are determined according to the image entropy value, the image texture complexity, the target recognition accuracy, and the target recognition efficiency. First, the image entropy value of the image gray-scale distribution is calculated using the information entropy formula to reflect the image information complexity of the target scene image. If the image entropy value is greater than the preset entropy value, the target scene image is determined to be a high-entropy value image, and the feature extraction ability needs to be enhanced. If the image entropy value is less than or equal to the preset entropy value, it is determined to be a low-entropy value image, and the network structure of the target neural network can be simplified. Also, the gray-level co-occurrence matrix is used to calculate texture features such as contrast, correlation, and energy to quantify the texture complexity of the target scene image. If the texture complexity is high, the depth of the convolutional layer needs to be increased or a multi-scale convolutional kernel is used. If the texture is simple, the number of convolutional layers can be reduced. The number of convolutional layers can be determined using the following formula: the number of convolutional layers = the basic number of layers + the entropy adjustment term + the texture adjustment term. Among them, the basic number of layers can be set to 3 to 8 layers according to the task complexity. The entropy adjustment term can add 1 to 2 layers for high-entropy value images and subtract 1 layer for low-entropy value images. The texture adjustment term can add 1 layer according to the texture complexity and subtract 1 layer for simplicity. The number of text box layers can be determined according to the following formula: the number of text box layers = min(convolutional layer number, target scale range). If the target scale span is large, the number of text box layers can be equal to the number of convolutional layers. If the target scale is concentrated, the number of text box layers can be half of the number of convolutional layers. The corresponding rule between the text box layer and the convolutional layer is that the bottom convolutional layer is connected to the small-scale text box layer to capture details, the middle convolutional layer is connected to the medium-scale text box layer to balance semantics and details, and the high convolutional layer is connected to the large-scale text box layer to extract global features. In addition, a loss function is set, and after model training, L1 regularization is performed on the convolutional layer, redundant channels are cropped, the text detection specific metric is used as the accuracy in the evaluation metrics, and the recognition time and memory occupancy are used as the efficiency in the evaluation metrics. If the accuracy is insufficient, 1 to 2 convolutional layers are added, and the number of anchor boxes of the text box layer is increased. If the efficiency is too low, 1 convolutional layer is reduced, and adjacent text box layers are merged.

[0062] In a possible embodiment, according to the target recognition accuracy, target recognition efficiency, and the scale corresponding to each of the m feature maps, determine the initial box set for each of the m feature maps, obtaining m initial box sets. Each initial box set includes multiple initial boxes, which can also be referred to as multiple default boxes. For each of the m feature maps, the aspect ratios of the default boxes in the same feature map are inconsistent. In the same feature map, the aspect ratios of the default boxes are diversely designed according to the actual aspect ratio distribution of the text in the training dataset. If the common aspect ratios in the dataset are 1:5, which is a long text line, and 1:2, which is a shorter text, then multiple default boxes with different aspect ratios are preset at each spatial position of the same feature map. This design enables the model to cover text instances of different shapes and improves the detection ability for texts with extreme aspect ratios, such as long slogans and vertically arranged text. Multiple default boxes are generated at each spatial position of each feature map. The specific number is determined by the following parameters: the number of aspect ratio types, such as 3 types: 1:5, 1:2, 1:1; the number of vertical offset types, such as 2 offset directions or ratios; the scale coverage range, which is related to the resolution of the feature map. Then, if a certain feature map presets 3 aspect ratios and 2 vertical offsets, 6 default boxes are generated at each spatial position. Based on the balance of text distribution effects, computational efficiency, and coverage rate, and based on the statistical features such as the aspect ratio, orientation, and density of the text in the training dataset, dynamically adjust the aspect ratios and offsets of the default boxes. Through experimental verification, select a parameter combination that ensures coverage while avoiding computational redundancy. Default boxes are also generated in the text-free area. The default boxes are densely distributed on the feature map, covering all regions of the image. During the training phase, by calculating the intersection over union (IoU) between the default boxes and the ground truth boxes, only the default boxes with a high overlap with the ground truth boxes participate in the loss calculation, and the unmatched default boxes are marked as the background. During the application phase, the model filters the default boxes through the detection score, that is, the classification confidence, and only retains the candidate boxes with high confidence. The default boxes in the text-free area are filtered due to their low scores.

[0063] In the embodiments of the present application, by determining n convolutional layers and m text box layers in the target neural network according to the image entropy value, the image texture complexity, the target recognition accuracy, and the target recognition efficiency, different network architectures can be determined according to different target scene images, enabling accurate feature extraction and adaptation to task requirements, reducing the computational cost and improving the operation efficiency when the target scene image is relatively simple, and enabling text recognition adaptable to various scene changes. Moreover, by determining the first text of each of the m feature maps and the first background corresponding to each first text according to the m initial box sets, obtaining m first texts and m first backgrounds, and determining the target text according to the m first texts, the m first backgrounds, and the scale corresponding to each of the m feature maps, the background object can provide additional context information for text recognition by combining the background of the target scene image. The distinction between similar texts or polysemous texts can be achieved by constructing a complete semantic scene. The m first texts are convenient for capturing the details and global information of the target scene image, and the target text is obtained by fusing multi-dimensional semantics and can be processed in parallel, improving the recognition efficiency and accuracy.

[0064] Optionally, in step S109, determining the first text of each of the m feature maps and the first background corresponding to each first text according to the m initial box sets, obtaining m first texts and m first backgrounds may include the following steps:

[0065] Step S201: Determine the target initial box set corresponding to the target feature map in the m initial box sets; the target feature map is any one of the m feature maps; the target initial box set includes k initial boxes, and k is a positive integer;

[0066] Step S202: Determine the text density of each of the k initial boxes to obtain k text densities; the k text densities are used to reflect the density of the text in the k initial boxes;

[0067] Step S203: Obtain the height and aspect ratio of each of the k initial boxes to obtain k heights and k aspect ratios;

[0068] Step S204: Determine the vertical offset and horizontal offset of each of the k initial boxes according to the k text densities, the k heights, and the k aspect ratios to obtain k vertical offsets and k horizontal offsets;

[0069] Step S205: Determine k scale scaling factors and k rotation angles corresponding to the k initial boxes;

[0070] Step S206: Adjust the k initial boxes according to the k vertical offsets, the k horizontal offsets, the k scale scaling factors, and the k rotation angles to obtain k oriented text boxes;

[0071] Step S207: Determine the minimum horizontal rectangle of each oriented text box among the k oriented text boxes to obtain k prediction boxes;

[0072] Step S208: Determine the first confidence level and the second confidence level of each prediction box among the k prediction boxes to obtain k first confidence levels and k second confidence levels; the k first confidence levels are used to reflect the probability that the k prediction boxes cover the target text; the k second confidence levels are used to reflect the probability that the text within the k prediction boxes is recognizable;

[0073] Step S209: Screen the k prediction boxes according to the k first confidence levels and the k second confidence levels to obtain j candidate boxes; j is a positive integer less than or equal to k;

[0074] Step S210: Perform text extraction and background recognition on the j candidate boxes to obtain the first text and the first background of the target feature map;

[0075] Step S211: Determine the first text of each feature map among the m feature maps according to the first text and the first background of the target feature map, and the first background corresponding to each first text, to obtain m first texts and m first backgrounds.

[0076] In a possible embodiment, in the default box vertical offset adjustment mechanism, the determination of the text density is realized through the feature map response intensity analysis and the dynamic rule design. Specifically, for the feature map local response intensity analysis, first input the feature map source. Before generating the default box, extract the multi-scale feature map through the convolutional layer, and then use the activation value at each position in the feature map to represent the response intensity of the region to the text feature. For example, for high response regions such as edges and corners, then determine that the response intensity of the dense region is high and the distribution fluctuation is large, while the response of the background region is low and the distribution is uniform. In addition, the sliding window statistics and region classification method can also be used, including window setting, statistical indicators, and threshold determination. First, use a fixed-size sliding window, such as 5×5 pixels, and traverse the feature map at a preset step size, such as 2 pixels. Then calculate the average response intensity of all pixels within the window to measure the overall activity of the region, calculate the fluctuation degree of the response intensity within the window, distinguish the dense text, that is, high fluctuation, from the smooth background, that is, low fluctuation, and then based on the statistical distribution of the training data set, preset the average response intensity threshold and the fluctuation degree threshold. If both the average response intensity and the fluctuation degree within the window exceed the threshold, it is determined as a text-dense region, otherwise it is a sparse region.

[0077] In a possible embodiment, the dynamic vertical offset adjustment rule includes basic offset setting, dense area enhancement, sparse area reduction, and smooth transition constraint. First, a basic vertical offset is preset according to the default box height, for example, 30% of the default box height. If the current window is determined to be a dense area, the vertical offset is increased proportionally, for example, increased to 1.5 times the basic offset to cover adjacent text lines. If it is a non-dense area, the vertical offset is decreased proportionally, for example, decreased to 80% of the basic offset to avoid redundant detection boxes. The adjustment range of the offset between adjacent windows does not exceed a preset ratio, such as 20%, to prevent sudden changes in the position of the detection box.

[0078] In a possible embodiment, the implementation process of determining the vertical offset can be as follows: The input image generates multi-scale feature maps through a convolutional network. A sliding window statistic is applied to each layer of the feature maps to classify dense and non-dense areas. According to the area classification results, the vertical offset of the default boxes in this area is dynamically adjusted, and the adjusted default boxes are input into the subsequent classification and regression branches to complete end-to-end training and inference.

[0079] In a possible embodiment, during the training phase of the model, the intersection over union (IoU) is used to screen a large number of default boxes to obtain candidate boxes containing text. The calculation of the IoU involves the ground truth box, that is, the IoU is the ratio of the overlapping part of the default box and the ground truth box to their total coverage area. Among them, the confirmation and generation process of the ground truth box is as follows: During the model training process, the ground truth box, that is, the labeled text bounding box, is provided by an artificially labeled dataset. For example, the rendered text is embedded into a natural image through an algorithm, and a corresponding annotation file is automatically generated, such as the vertex coordinates of a quadrilateral or the parameters of a rotated rectangle. In addition, the artificial annotator annotates the text area in the natural scene image, and the annotation format includes text content, bounding box position, such as the center point, angle, height, etc. of the vertices of the quadrilateral or the rotated rectangle. During training, the model calculates the IoU between the default boxes and these pre-labeled ground truth boxes to generate a matching indication matrix, thereby determining which default boxes need to participate in the loss calculation. However, when the model is actually applied, there is no need for the ground truth box. At this time, the working process of the model is: directly input the image to be detected, predict candidate boxes through the trained network, including classification confidence and bounding box offset, use non-maximum suppression to screen the final detection results, and output the predicted text bounding box. The ground truth box is only used for supervised learning and model performance evaluation in the training phase, and in actual applications, it completely depends on the prediction ability of the model itself.

[0080] In a possible embodiment, the first confidence level can be represented by a detection score, and the second confidence level can be represented by an identification score. The detection score focuses on spatial localization and is used to reflect the probability that a candidate box covers a text region. The identification score focuses on semantic verification and is used to reflect the recognizability of the text content within the candidate box. The fusion of the two can effectively eliminate false detections where the localization is correct but the content is invalid, such as blurred text or non-text objects, and further improve the reliability of the detection results. Among them, the detection score is not only the basis for preliminary screening but also a necessary parameter for optimizing the detection results by fusing the identification score, solving the limitations of pure visual localization. The default box, as a spatial reference and training anchor, can ensure that the model can efficiently learn the geometric characteristics of the text and support the stability and scalability of the end-to-end process. This method can improve the accuracy and robustness of text detection in complex scenarios through the joint optimization of detection and identification. The detection score is used for preliminary screening of candidate boxes. The detection score is generated by the classification branch of the model and represents the probability that the default box contains text. In the detection stage, the network first generates a large number of default boxes and performs preliminary screening according to the detection score, such as non-maximum suppression. The purpose of this step is to quickly filter out a large number of default boxes that obviously do not contain text, retain high-confidence candidate boxes, and reduce redundant calculations in the subsequent identification stage. The detection score is also used to optimize the joint confidence level of detection and identification. Even if a candidate box passes the preliminary screening, there may still be false detections, such as misjudging background textures as text. By fusing the detection score and the identification score, such as the harmonic mean, the "text existence" and "text content rationality" of the candidate box can be comprehensively evaluated.

[0081] In a possible embodiment, the default box is a spatial anchor predefined by the model during the training stage, covering potential text regions with different positions, scales, and aspect ratios. By learning the offsets from the default box to the ground truth box, such as the center point, width, height, rotation angle, etc., the position of the candidate box is gradually corrected, and finally an accurate oriented bounding box is generated. Based on the default box, the position and shape of the candidate box are dynamically adjusted by predicting the offsets to fit the real text region. This mechanism ensures the model's adaptability to text in any direction, size, and shape.

[0082] In a possible embodiment, refer to Figure 3 , Figure 3 FIG. is an application scenario diagram of a text detection method based on the minimum horizontal rectangle provided by an embodiment of the present application. As Figure 3 shown, the black dashed box is the initial box of the target text, that is, the default box, and the red solid box is the oriented text box of the target text. The minimum horizontal rectangle of the oriented text box is the black solid box, that is, the predicted box of the target text.

[0083] Optionally, refer to Figure 4 , Figure 4It is a flowchart of a text density determination method provided by an embodiment of the present application. As Figure 4 shown, in step S202, determining the text density of each of the k initial boxes to obtain k text densities may include the following steps:

[0084] Step S301: Obtain the edge features and corner features of each of the k initial boxes to obtain k edge features and k corner features;

[0085] Step S302: Determine the activation value of each of the k initial boxes according to the k edge features and k corner features to obtain k activation values;

[0086] Among them, the k activation values are used to reflect the response intensity of the k initial boxes to the target text;

[0087] Step S303: Determine the text density of each of the k initial boxes according to the k activation values to obtain k text densities.

[0088] In a possible embodiment, to obtain the edge features and corner features of each of the k initial boxes, first, the image containing the target text is grayscale processed to convert the color image into a grayscale image, reducing the data dimension. Then, Gaussian blur is performed on the grayscale image to remove the noise in the image. An edge detection algorithm is used to perform edge detection on the preprocessed image. This algorithm includes four steps: calculating the gradient magnitude and direction of the image, non-maximum suppression, double-threshold processing, and hysteresis boundary tracking. Through these steps, a clear edge image can be obtained. For each initial box, the edge pixel information within the box, such as the number of edge pixels, the length of the edge, the angular distribution of the edge, etc., is extracted from the edge image as the edge feature of the initial box. The corner points are extracted by using a corner detection algorithm. This algorithm calculates the autocorrelation matrix of the image and determines whether it is a corner point according to the eigenvalues of the matrix. For each initial box, information such as the number of corner points within the box and the position distribution of the corner points is counted as the corner feature of the initial box.

[0089] In a possible embodiment, to determine the activation value of each of the k initial boxes according to the k edge features and k corner features, weights are assigned to the edge features and corner features respectively. For each initial box, its activation value = the edge feature score of the initial box * the edge feature weight + the corner feature score of the initial box * the corner feature weight. The edge feature score and the corner feature score can be normalized to map the feature values to the interval from 0 to 1.

[0090] In a possible embodiment, the text density of each of the k initial boxes is determined based on the k activation values. The k activation values are normalized to convert the activation values into a probability distribution. The text density of each initial box can be directly represented by the normalized activation value because the activation value reflects the response intensity of the initial box to the target text. The normalized activation value can be regarded as the probability that the initial box contains text, that is, the text density.

[0091] In the embodiments of the present application, edge features and corner features are important visual cues for text. Text usually has obvious edges and corners. By extracting these features, the text region can be more accurately located. The calculation of activation values and text density can quantify the correlation between each initial box and the target text, which helps to screen out the initial boxes that truly contain text and reduce the cases of false detection and missed detection. Combining edge features and corner features can make the model have better adaptability to texts of different fonts, sizes, and orientations. For example, the edge and corner features of texts with different fonts may be different, but by comprehensively considering these features, the model can more accurately detect the text. At the same time, the calculation of text density can, to a certain extent, resist noise and interference and improve the robustness of the model.

[0092] Optionally, step S205 of determining the k scale scaling factors and the k rotation angles corresponding to the k initial boxes may include the following steps:

[0093] Step S401: Obtain the target image resolution and the target scale of the target feature map;

[0094] Step S402: Determine the reference scaling factors of the k initial boxes according to the target image resolution and the target scale;

[0095] Step S403: Determine the k scale scaling factors according to the reference scaling factors, the k heights, and the k width-to-height ratios;

[0096] Step S404: Determine the horizontal direction gradient and the vertical direction gradient of each of the k initial boxes to obtain the k horizontal direction gradients and the k vertical direction gradients;

[0097] Step S405: Determine the k rotation angles according to the k horizontal direction gradients and the k vertical direction gradients.

[0098] In a possible embodiment, the target image resolution and target scale of the target feature map are obtained. For the target feature map, its corresponding original image resolution is usually known in the data input or preprocessing stage. If the image is a feature map obtained through network processing, the resolution information of the original image can be traced back. The target scale can be determined according to the position of the feature map in the network and the design of the network. For example, in some structures based on the feature pyramid network, feature maps at different levels correspond to different scales, with higher-level feature maps corresponding to larger scales and lower-level feature maps corresponding to smaller scales. The target scale of the target feature map can be determined through a preset scale rule or network parameters.

[0099] In a possible embodiment, according to the target image resolution and target scale, the reference scaling factors of k initial boxes are determined. According to the target scale, the scaling ratio relative to the original image is calculated. Considering the requirements of the initial boxes, the basic scaling ratio is used as a reference and adjusted in combination with the actual situation. Based on experience or analysis of the data, a reference scaling factor near the basic scaling ratio can be set for each initial box.

[0100] In a possible embodiment, according to the reference scaling factors, k heights, and k width-to-height ratios, k scale scaling factors are determined. For each initial box, according to the reference scaling factor, the given height, and the width-to-height ratio, the actual width and height of the initial box are calculated. The calculated initial box size is compared with the standard size or expected size, and the final scale scaling factor is obtained through proportional adjustment.

[0101] In a possible embodiment, the horizontal direction gradient and vertical direction gradient of each initial box among the k initial boxes are determined, obtaining k horizontal direction gradients and k vertical direction gradients. For each initial box, the image region within the box is extracted from the target feature map, and using a gradient calculation operator, the horizontal direction gradient and vertical direction gradient of the extracted image region are calculated respectively.

[0102] In a possible embodiment, according to the k horizontal direction gradients and k vertical direction gradients, k rotation angles are determined. For each initial box, the gradient direction is calculated according to its horizontal direction gradient, vertical direction gradient, and the gradient calculation formula. The calculated gradient direction is converted into a rotation angle to obtain the rotation angle of each initial box.

[0103] In the embodiments of the present application, by determining a reference scaling factor according to the target image resolution and scale, and then determining a scale scaling factor in combination with the height and aspect ratio, the initial box can better adapt to the size and shape of the target, improving the positioning accuracy of the target. At the same time, by determining the rotation angle according to the horizontal and vertical gradients, the direction information of the target can be accurately captured, further improving the detection accuracy. This method can dynamically adjust the parameters of the initial box according to different feature map scales and image resolutions, enabling the model to adapt to targets of different sizes and directions, enhancing the generalization ability of the model and its adaptability to complex scenarios. Calculating the horizontal and vertical gradients of the initial box makes full use of the local feature information of the image. These gradient information can reflect features such as the edges and textures of the target, helping to more accurately judge the position and direction of the target and improving the target recognition ability of the model. The accurately determined scale scaling factor and rotation angle provide a more accurate basis for subsequent operations such as target classification and box regression, contributing to improving the performance and efficiency of the entire target detection system.

[0104] Optionally, step S110 of determining the target text according to the m first texts, the m first backgrounds, and the scale corresponding to each feature map in the m feature maps may include the following steps:

[0105] Step S501: Determine the semantic relationship between each first text in the m first texts and the corresponding first background in the m first backgrounds to obtain m semantic relationships;

[0106] Step S502: Determine the third confidence level of each first text in the m first texts according to the m semantic relationships to obtain m third confidence levels;

[0107] Step S503: Obtain the proportion of the preset character types of each character in the m first texts to obtain m proportions; the preset character types include missing characters and / or blurred characters;

[0108] Step S504: Determine the fourth confidence level of each first text in the m first texts according to the m proportions to obtain m fourth confidence levels;

[0109] Step S505: Obtain the initial character size of each character in the m first texts in the target scene image;

[0110] Step S506: Determine the fifth confidence level of each first text in the m first texts according to the scale corresponding to each feature map in the m feature maps and the initial character size of each character in the m first texts in the target scene image to obtain m fifth confidence levels;

[0111] Step S507: Determine the target text based on a preset confidence weight group, m third confidences, m fourth confidences, m fifth confidences, and m first texts; the confidence weight group includes a first weight, a second weight, and a third weight.

[0112] In a possible embodiment, determine the semantic relationship between each first text in the m first texts and the corresponding first background in the m first backgrounds to obtain m semantic relationships. For each first text, use a word vector model in natural language processing to convert the text into a vector representation and extract the semantic features of the text. For the corresponding first background, if the background is an image, a convolutional neural network can be used to extract the feature vector of the image. For example, input the background image into a pre-trained model to obtain its high-level features, and use methods such as cosine similarity to calculate the similarity between the text feature vector and the background image feature vector to measure the semantic relationship between them.

[0113] In a possible embodiment, determine the third confidence of each first text in the m first texts according to the m semantic relationships to obtain m third confidences. According to the actual situation and experience, set the mapping relationship between the semantic relationship value and the confidence, and determine the third confidence for each first text.

[0114] In a possible embodiment, obtain the proportion of the preset character types of each first text in the m first texts to obtain m proportions. The preset character types include missing characters and / or blurred characters. Analyze each character in each first text one by one, and judge whether it is a missing character or a blurred character according to the characteristics such as the clarity and integrity of the character. For example, set a pixel threshold to judge whether the character is blurred. If the pixel value of the character is lower than a certain threshold, it is considered a blurred character. If some strokes of the character are missing, it is identified as a missing character. Then, count the number of preset character types, that is, missing characters and / or blurred characters, in each first text, and calculate the proportion of the total number of characters in the first text.

[0115] In a possible embodiment, determine the fourth confidence of each first text in the m first texts according to the m proportions to obtain m fourth confidences, establish the mapping relationship between the proportion and the confidence, and determine the fourth confidence for each first text.

[0116] In a possible embodiment, obtain the initial character size of each character in the m first texts in the target scene image. In the target scene image, locate the position of each character in the m first texts through a text detection algorithm and measure its size. For example, for the size of the character, the width and height of its circumscribed rectangle can be measured, and record the initial character size information of each character.

[0117] In a possible embodiment, according to the scale corresponding to each of the m feature maps and the initial character size of each character in the m first texts in the target scene image, determine the fifth confidence level of each first text in the m first texts, obtaining m fifth confidence levels. Analyze the relationship between the scale of each feature map and the character size, and determine the fifth confidence level according to the comparison result, and determine its fifth confidence level for each first text.

[0118] In a possible embodiment, according to a preset confidence weight group, m third confidence levels, m fourth confidence levels, m fifth confidence levels, and m first texts, determine the target text. The confidence weight group includes a first weight, a second weight, and a third weight. For each first text, calculate its comprehensive confidence level according to the preset confidence weight group. Compare the comprehensive confidence levels of the m first texts, and select the first text with the highest comprehensive confidence level as the target text. If there are multiple first texts with the same and highest comprehensive confidence level, further screening can be performed according to other conditions, such as text length, integrity of text content, etc.

[0119] In the embodiments of the present application, by considering the semantic relationship between the text and the background, it is possible to determine whether the text conforms to the background scene, avoid misjudging text irrelevant to the background as the target text, improve the accuracy of the text. At the same time, the analysis of the preset character type ratio can identify texts with lower quality and further screen out more reliable target texts. Combining the feature map scale and the character size to determine the confidence level enables the model to adapt to images of different resolutions and scales, enhancing the adaptability to complex scenes. Comprehensively considering multiple dimensions such as semantic relationships, character quality, and size information, and performing comprehensive evaluation by setting weights can more comprehensively measure the reliability of the text, avoid the limitations of single factors, and improve the scientificity and rationality of determining the target text.

[0120] Optionally, in step S105, according to the image entropy value, the image texture complexity, the target recognition accuracy, and the target recognition efficiency, determine the n convolutional layers and the m text box layers in the target neural network, which may include the following steps:

[0121] Step S601: According to the preset parameter value, the image entropy value, and the image texture complexity, determine the first parameter value; n is less than or equal to the first parameter value; the preset parameter value is used to reflect the basic number of convolutional layers;

[0122] Step S602: According to the first parameter value, the target recognition accuracy, and the target recognition efficiency, determine the n convolutional layers and the n scales; the n scales correspond one-to-one to the n convolutional layers;

[0123] Step S603: According to the number of the n convolutional layers and the target recognition efficiency, determine the second parameter value; m is less than or equal to the second parameter value;

[0124] Step S604: Determine m text box layers according to each of the n scales, the adjacent scales corresponding to the scales, and the second parameter value; the adjacent scale is smaller than the scale.

[0125] In a possible embodiment, determine the first parameter value according to a preset parameter value, an image entropy value, and an image texture complexity. First, calculate the entropy value of the image through the probabilities of pixels with different gray values in the image. The entropy value reflects the richness of information in the image. For the image texture complexity, relevant features can be calculated using a gray-level co-occurrence matrix, such as indicators like contrast, entropy, and energy to measure. According to the preset parameter value, which represents the basic number of convolutional layers, adjust this parameter in combination with the image entropy value and texture complexity. If the image entropy value is high and the texture complexity is large, it indicates that the image information is rich and complex, and more convolutional layers are needed to extract features. At this time, a certain value can be added to the preset parameter value. Conversely, if the image entropy value is low and the texture is simple, the preset parameter value can be reduced.

[0126] In a possible embodiment, determine n convolutional layers and n scales according to the first parameter value, the target recognition accuracy, and the target recognition efficiency. Within the range of the first parameter value, consider the balance between the target recognition accuracy and efficiency. If the requirement for target recognition accuracy is high and the computing resources permit, the number of convolutional layers can be appropriately increased, approaching or equal to the first parameter value. If more emphasis is placed on the target recognition efficiency, the number of convolutional layers can be reduced, but it is necessary to ensure that a certain accuracy requirement is met. For each convolutional layer, determine the corresponding scale according to its position and role in the network. Shallow convolutional layers process features at a smaller scale to capture detailed information, and deeper convolutional layers process features at a larger scale to extract more advanced semantic information. The scales can be determined according to a certain proportional relationship.

[0127] In a possible embodiment, determine the second parameter value according to the number of n convolutional layers and the target recognition efficiency. According to the determined number of convolutional layers, analyze its computational amount. The computational amount of the convolutional layer is related to factors such as the size and number of convolutional kernels and the size of the feature map. Computational amount estimation formulas, such as the multiplication and addition operation count formulas for convolutional operations, can be used to estimate the total computational amount of this convolutional layer. In combination with the target recognition efficiency requirement, determine the second parameter value on the premise of ensuring a certain efficiency. If the computational amount is large, in order to improve efficiency, the number of subsequent text box layers needs to be reduced, and at this time, the second parameter value can be smaller. If the computational amount is relatively small, the second parameter value can be appropriately increased.

[0128] In a possible embodiment, m text box layers are determined according to each of the n scales, the adjacent scales corresponding to the scales, and the second parameter value. For each scale, the difference between it and the adjacent smaller scale is analyzed, and the number of text box layers is determined according to the relationship between the scales and the second parameter value. Layers with larger scale differences can be preferentially selected to set the text box layers to cover targets of different sizes. Each text box layer corresponds to a scale, and these scales can better meet the detection requirements of targets of different sizes.

[0129] In the embodiments of the present application, the basic parameters for adjusting the number of convolutional layers according to the image entropy value and texture complexity can enable the network architecture to better adapt to the characteristics of the input image. For images with rich information and complex textures, increasing the number of convolutional layers can extract features more fully. For simple images, reducing the number of convolutional layers can avoid wasting computing resources and improve the running efficiency of the model. In the process of determining the number of convolutional layers and the number of text box layers, the requirements for target recognition accuracy and efficiency are fully considered. By reasonably adjusting these parameters, the running speed of the model can be improved on the premise of ensuring a certain accuracy, meeting the requirements of different application scenarios. Determining the text box layers according to the scale relationship can enable the network to better detect targets of different sizes. Selecting appropriate scales to set the text box layers can cover a wider range of target sizes, improving the accuracy and recall rate of target detection. For example, both small targets and large targets can be effectively detected and located through the text box layers of corresponding scales.

[0130] Optionally, refer to Figure 5 , Figure 5 which is a flowchart of a method for displaying target text provided by the embodiments of the present application. As Figure 5 shown, a text detection method provided by the embodiments of the present application further includes:

[0131] Step S701: Obtain the coordinate mapping relationship between each pixel point in the target scene image and the target text;

[0132] Step S702: Determine the display parameters of the target text;

[0133] wherein the display parameters include at least one of the following: font format, font size, font transparency, display direction;

[0134] Step S703: Display the target text in the target scene image according to the coordinate mapping relationship and the display parameters.

[0135] In a possible embodiment, the coordinate mapping relationship between each pixel point in the target scene image and the target text is obtained. The target scene image is processed using a text detection algorithm to accurately detect the position of the target text in the image, obtaining the bounding rectangle of the target text or a more refined contour of the text region. Taking the upper left corner of the target scene image as the origin, a plane rectangular coordinate system is established to determine the coordinates of each pixel point in the image. For each character or text block in the target text, its specific position coordinates in the image are determined. Then, by traversing the pixel points within the text region, the relative position relationship between each pixel point in the target scene image and the target text is calculated to establish the coordinate mapping relationship. For example, information such as the distance from each pixel point to the boundary of the target text and whether the pixel point is within the target text region can be recorded.

[0136] In a possible embodiment, the display parameters of the target text are determined; the display parameters may include at least one of the following: font format, font size, font transparency, display direction. According to the style and requirements of the target scene image, as well as the font of the target text itself in the target scene image, a suitable font format is selected from the system font library or a custom font library, such as Song typeface, Boldface, Regular script, etc. Based on the importance and the space occupied by the target text in the image, as well as the visual effect on the human eye, a suitable font size is determined. The resolution of the image and the length of the target text can be referred to, and the font size can be calculated through experiments or empirical formulas. For example, for a larger-sized image and important title text, a larger font size can be selected; for secondary explanatory text, a smaller font size can be selected. According to the complexity of the image background and the contrast between the text and the background, the transparency of the font is set. If the background is relatively complex, in order to make the text display more clearly, the font transparency can be reduced. If the background is relatively simple, appropriately increasing the font transparency can make the text better integrated with the background. The transparency can usually be represented by a value between 0 and 1. According to the layout and design requirements of the target scene image, the display direction of the target text is determined. The display direction has a horizontal direction and a vertical direction. For some special texts, the rotation angle of the text, such as 45 degrees, 90 degrees, etc., can also be set.

[0137] In a possible embodiment, the target text is displayed in the target scene image according to the coordinate mapping relationship and the display parameters. The target scene image is loaded using an image processing library and converted into an editable image object. According to the coordinate mapping relationship, the specific position of the target text in the image is determined. Then, according to the determined display parameters, the text drawing function provided by the image processing library is used to draw the target text at the corresponding position in the image.

[0138] In the embodiments of the present application, by obtaining the coordinate mapping relationship between pixel points and the target text, the accurate positioning of the target text in the image can be achieved, enabling the better integration of the text with other elements of the image and avoiding problems such as visual disharmony or inaccurate information transmission caused by improper text display positions. Determining rich display parameters such as font format, size, transparency, display direction, etc., the target text can be personalized for display settings according to different scenario requirements and design styles, which helps to highlight the key points of the text, enhance the visual effect of the image, and improve the efficiency of information transmission. Reasonably setting parameters such as font size, transparency, and display direction can improve the readability of the target text. For example, selecting an appropriate font size, a font color with high contrast, and a clear display direction can make it easier for readers to read and understand the text content, especially in complex background images.

[0139] In a possible embodiment, referring to Figure 6 , Figure 6 is a schematic diagram of the neural network architecture of a text detection method provided by the embodiments of the present application. As Figure 6 shown, after the target scene image passes through the convolutional layer, it is processed by the text box layer, and candidate boxes are obtained after non-maximum suppression. This network architecture inherits the VGG-16 architecture, retains some convolutional layers and converts the fully connected layers into convolutional layers, and at the same time adds additional convolutional layers and multiple text box layers. The text box layer is used to predict text existence and bounding boxes, and the output includes oriented bounding boxes (quadrilaterals or rotated rectangles) and the smallest horizontal bounding rectangle containing the oriented bounding boxes. The network structure is a fully convolutional structure, which can adapt to images of any size and has flexibility in the training and testing phases. The text box layer predicts text existence and bounding boxes based on the input feature map, which is achieved by predicting the offset from the pre-designed horizontal default box to the oriented text box. Multiple aspect ratios are set for the default boxes to adapt to different texts, and to cope with the case of dense text, a vertical offset is set for each default box to make the default boxes denser in the vertical direction. During training, the true word boxes are matched with the default boxes according to box overlap, and the efficiency is improved by using the smallest horizontal rectangle matching, enabling the network to learn specific regression and classification weights. For oriented text, a 3×5 convolutional filter is used, and its rectangular receptive field is more suitable for the long shape of text, avoiding the noise interference caused by the square receptive field and being superior to the 1×5 irregular convolutional filter previously used for horizontal text detection.

[0140] Among them, the number of additional convolutional layers is designed based on the depth and accuracy requirements of input image feature extraction. In the present invention, through experimental verification, it is found that adding 3 additional convolutional layers can achieve a good balance between feature extraction efficiency and computational complexity. The number of channels of these additional convolutional layers is 512, 256, and 128 respectively, which are used to capture text features of different scales, so as to enhance the network's detection ability for multi-scale text. The convolutional kernel size of each additional convolutional layer is 3×3, the stride is 1, and the padding is 1 to maintain the spatial resolution of the feature map. The number of text box layers is determined according to the scale distribution of the target text and the detection accuracy requirements. In the present invention, 6 text box layers are designed, corresponding to feature maps of different scales respectively. These text box layers are distributed on the feature maps from shallower to deeper, and can cover text instances from smaller to larger. Each text box layer is responsible for predicting the text bounding boxes within a specific scale range, so as to achieve a comprehensive detection of multi-scale text. Specifically, the default box aspect ratio of each text box layer is adjusted according to the actual aspect ratio distribution of the text in the training dataset to better adapt to text instances of different shapes.

[0141] Among them, the vertical offset of each default box can be determined specifically by the following methods. Based on the statistical analysis of the text line spacing, in the training dataset, a large number of text images are analyzed to statistically calculate the average spacing between text lines. According to the statistical results, a standard line spacing value is determined. The vertical offset of the default box is set to a certain proportion of the standard line spacing to ensure that the default box can cover adjacent text lines in the vertical direction. This proportion can be adjusted according to the specific application scenario and the characteristics of the dataset. Dynamically adjust according to the height of the default box. When designing the default box, the vertical offset is dynamically adjusted according to its height. For a higher default box, increase the vertical offset so that it can cover more possible text lines; for a shorter default box, reduce the vertical offset to avoid unnecessary computational overhead. The specific adjustment strategy is as follows: for a default box with a larger height, the vertical offset is set to a larger proportion of its height, for example, 40% to 50%; for a default box with a smaller height, the vertical offset is set to a smaller proportion of its height, for example, 20% to 30%. This dynamic adjustment method can flexibly set the vertical offset according to the actual height of the default box, so as to better adapt to text of different scales. Adaptive adjustment based on text density. In practical applications, the text density in different regions of the input image may be different. By analyzing the local features of the input image, the text density is judged, and the vertical offset of the default box is adjusted accordingly: in the text-dense region, appropriately increase the vertical offset of the default box to improve the coverage ability for dense text; in the text-sparse region, maintain the default vertical offset or appropriately reduce it to avoid over-detection. This adaptive adjustment mechanism can dynamically optimize the coverage range of the default box according to the local features of the input image, thereby improving the accuracy and efficiency of detection.

[0142] Moreover, ground truth representation, loss function, online hard example mining, and data augmentation are adopted to train the model to adapt to text in any direction. Among them, the ground truth representation includes quadrilateral representation and rotated rectangle representation. For the ground truth box of oriented text in the quadrilateral representation, first obtain its minimum horizontal rectangle bounding box, and then represent the oriented text with the vertices of the quadrilateral. Determine the vertex order through specific rules to minimize the sum of the Euclidean distances between the corresponding points of the quadrilateral and the horizontal rectangle. Then, the rotated rectangle representation adopts a special representation to avoid the model's dependence on the dataset due to uneven angle distribution, and uses two vertices of the horizontal rectangle and the height of the rotated rectangle to represent. The loss function consists of confidence loss and location loss. Determine the default boxes participating in the calculation through the matching indicator matrix, use smooth L1 loss to calculate the location loss, binary cross-entropy loss to calculate the confidence loss, and set the α parameter to balance the importance of the two. Online hard example mining divides the training into two stages according to the difficult-to-distinguish negative samples in the training dataset, such as textures and signs similar to the text, gradually adjusts the positive and negative sample ratios, suppresses hard examples, and improves the network's ability to distinguish between text and non-text. The data augmentation proposes a random cropping strategy based on object coverage constraint. Compared with the traditional strategy based on Jaccard overlap, it is more suitable for small text objects. By considering two constraints simultaneously, randomly scale the cropped area to increase the diversity of training data and improve the network's robustness to texts of different sizes.

[0143] Among them, the matching indicator matrix is generated by calculating the overlap degree between the default box and the ground truth box and matching according to the rules, and determines which default boxes participate in the loss calculation. Among them, calculating the overlap degree is to calculate the overlap degree between each default box and the ground truth annotation box. Simply put, it is to see the proportion of the overlapping part of the default box and the ground truth box in their total coverage area. If the overlap degree of a certain default box and a certain ground truth box exceeds the set threshold, it is considered that this default box matches the ground truth box. For each ground truth box, find the default box with the highest overlap degree with it. Even if this overlap degree does not reach the threshold, it also needs to be matched. If a default box does not match any ground truth box, it is marked as the background, that is, the area without text. The matching indicator matrix is a table, with rows representing default boxes and columns representing ground truth boxes. If a certain default box matches a certain ground truth box, mark 1 at the corresponding table position, otherwise mark 0. The α parameter is determined through experimental adjustment or cross-validation, and is used to balance the confidence loss and the location loss to optimize the overall performance of the model.

[0144] During testing, the prediction results of the network at six different scales are processed. First, the multi-scale prediction results are scaled to the original image size and fused into a dense confidence map, and then non-maximum suppression is performed in two steps. First, non-maximum suppression with a higher Intersection over Union (IoU) threshold is applied to the smallest horizontal rectangle containing the predicted quadrilateral or rotated rectangle to remove a large number of candidate boxes, and then non-maximum suppression with a lower IOU threshold is applied to the remaining small number of candidate boxes on the quadrilateral or rotated rectangle to obtain the final detection results. This cascaded method significantly improves the speed. Finally, in combination with text recognition for optimizing detection, a text recognizer is used to achieve word localization and end-to-end recognition tasks. The probability of the character sequence corresponding to the input image is estimated through the output layer of the text recognizer, supporting both dictionary-free and dictionary-based modes. The matching degree between the image and the word is measured by defining the recognition score. A method for optimizing the detection results by fusing the recognition score and the detection score is proposed. Since the numerical ranges of the two are different, an exponential function is first used to make them comparable, and then the harmonic mean is used to obtain the final combined score, avoiding the complexity of grid search for parameter tuning and effectively improving the detection accuracy.

[0145] Among them, the six scales cover text sizes from extremely small to extremely large, including the original image scale (1.0): Prediction is made using the size of the original input image, which is the most basic scale and is used to detect text that matches the size of the original image. Including the downscaled scale (0.5): The input image is scaled down to 50% of its original size. This scale is mainly used to detect text of larger sizes. By downscaling the image, the performance of these texts on the feature map will be closer to the scale during model training. Including the upscaled scale (2.0): The input image is scaled up to 200% of its original size. This scale is used to detect text of smaller sizes. By upscaling the image, the performance of these small texts on the feature map will be clearer, facilitating model detection. Including the intermediate scales (0.75 and 1.25): 0.75 means scaling the input image down to 75% of its original size, and 1.25 means scaling the input image up to 125% of its original size. These two scales are used to detect text of medium sizes, while filling the scale gaps between 0.5 to 1.0 and 1.0 to 2.0 to ensure that the model can effectively detect text of different sizes. Including the extreme scale (0.25): The input image is scaled down to 25% of its original size. This scale is used to detect very large text. By extreme downscaling, the performance of these texts on the feature map will be closer to the scale during model training. Exemplarily, assuming the size of the original input image is 1024×1024 pixels, the sizes of the input images at the six different scales are: original scale (1.0): 1024×1024 pixels; downscaled scale (0.5): 512×512 pixels; upscaled scale (2.0): 2048×2048 pixels; intermediate scale (0.75): 768×768 pixels; intermediate scale (1.25): 1280×1280 pixels; extreme scale (0.25): 256×256 pixels.

[0146] Among them, the recognition score is calculated by the output layer to calculate a probability value for each possible character sequence, indicating the likelihood that the sequence is the text content in the input image. After mapping the feature sequence of the input image to the character sequence, a confidence score, that is, the recognition score, is output. This score is calculated based on the probability distribution of the output layer and is used to measure the reliability of the recognition result. The detection score is calculated by the classification branch to predict whether each default box contains text. The classification branch outputs a binary classification confidence score, indicating the probability that the default box contains text. For each predicted text box, the detection score is the confidence score calculated by the classification branch, which reflects the reliability of the model's judgment on whether the text box contains text.

[0147] In the embodiments of the present application, a fully convolutional neural network that is end-to-end trainable is used to directly predict word bounding boxes in any direction. Without complex post-processing except for efficient non-maximum suppression, high-precision and high-efficiency text detection is achieved in the forward propagation of a single network. A quadrilateral or oriented rectangle is used to represent text in any direction, and a special network structure and training method are designed to adapt to this representation, including adjusting the aspect ratio of the default boxes, increasing the vertical offset to better cover dense text regions, using specific convolutional kernels to process long text lines, etc., to adapt to the representation and prediction of text in any direction. Moreover, the ground truth representation is improved to enable the network to better learn the regression and classification weights of oriented text. During the model training process, multi-scale inputs are adopted. Smaller-size inputs are used in the early stage to improve the training speed, and larger-size inputs are used in the later stage to enhance the detection ability for multi-scale text. During testing, the multi-scale prediction results are fused, and the final detection results are obtained through efficient cascaded non-maximum suppression, improving the detection accuracy and speed. In combination with a text recognizer, the semantic information of the recognition results is used to optimize the detection results. By defining a new score, the detection and recognition scores are fused to improve the detection accuracy. In the word localization and end-to-end recognition tasks, the detection results are further screened by the text recognizer to remove false positives and improve the overall performance. A data augmentation strategy of random cropping for text is proposed. By simultaneously considering the Jaccard overlap and object coverage constraints, the problem that the traditional random cropping strategy is not applicable to small text objects is solved. Online hard example mining, a suitable loss function, and an adaptive learning rate strategy are adopted to improve the network training effect.

[0148] In summary, in the embodiment of the present application, first, a target scene image is obtained, where the target scene image includes target text. Then, according to the gray-scale distribution of the target scene image, the image entropy value of the target scene image is determined. According to the gray-level co-occurrence matrix of the target scene image, the image texture complexity of the target scene image is determined. Then, the target recognition accuracy and target recognition efficiency of the target scene image are obtained, and according to the image entropy value, the image texture complexity, the target recognition accuracy, and the target recognition efficiency, n convolutional layers and m text box layers in the target neural network are determined, where the m text box layers are connected to m corresponding convolutional layers in the n convolutional layers one by one, and m and n are positive integers, and m is less than or equal to n. Next, the target scene image is convolved through the n convolutional layers to obtain n feature maps, and m feature maps are obtained from the n feature maps through the m text box layers. Then, according to the target recognition accuracy, the target recognition efficiency, and the scale corresponding to each feature map in the m feature maps, an initial box set for each feature map in the m feature maps is determined to obtain m initial box sets, and according to the m initial box sets, the first text of each feature map in the m feature maps and the first background corresponding to each first text are determined to obtain m first texts and m first backgrounds. Finally, according to the m first texts, the m first backgrounds, and the scale corresponding to each feature map in the m feature maps, the target text is determined. Thus, by determining the n convolutional layers and the m text box layers in the target neural network according to the image entropy value, the image texture complexity, the target recognition accuracy, and the target recognition efficiency, and by determining the first text of each feature map in the m feature maps and the first background corresponding to each first text according to the m initial box sets to obtain m first texts and m first backgrounds, and determining the target text according to the m first texts, the m first backgrounds, and the scale corresponding to each feature map in the m feature maps, the detection efficiency of scene text can be improved when detecting the target scene image.

[0149] The method of the embodiment of the present invention is described in detail above, and the device of the embodiment of the present invention is provided below.

[0150] Refer to Figure 7 , Figure 7 which is a schematic structural diagram of a text detection device provided by an embodiment of the present application. As Figure 7 shown, the text detection device 800 includes an acquisition unit 801 and a processing unit 802;

[0151] The acquisition unit 801 is used to acquire a target scene image; the target scene image includes target text;

[0152] The processing unit 802 is used to determine the image entropy value of the target scene image according to the gray-scale distribution of the target scene image;

[0153] Determine the image texture complexity of the target scene image according to the gray-level co-occurrence matrix of the target scene image;

[0154] Obtain the target recognition accuracy and target recognition efficiency of the target scene image;

[0155] Determine n convolutional layers and m text box layers in the target neural network according to the image entropy value, image texture complexity, target recognition accuracy, and target recognition efficiency; the m text box layers are connected to the corresponding m convolutional layers in the n convolutional layers one by one, where m and n are positive integers and m is less than or equal to n;

[0156] Perform convolution on the target scene image through n convolutional layers to obtain n feature maps;

[0157] Obtain m feature maps from the n feature maps through m text box layers;

[0158] Determine the initial box set of each feature map in the m feature maps according to the target recognition accuracy, target recognition efficiency, and the scale corresponding to each feature map in the m feature maps, and obtain m initial box sets;

[0159] Determine the first text of each feature map in the m feature maps and the first background corresponding to each first text according to the m initial box sets, and obtain m first texts and m first backgrounds;

[0160] Determine the target text according to the m first texts, m first backgrounds, and the scale corresponding to each feature map in the m feature maps.

[0161] In a possible embodiment, in terms of determining the first text of each feature map in the m feature maps and the first background corresponding to each first text according to the m initial box sets, and obtaining m first texts and m first backgrounds, the processing unit 802 is specifically configured to:

[0162] Determine the target initial box set corresponding to the target feature map in the m initial box sets; the target feature map is any one of the m feature maps; the target initial box set includes k initial boxes, where k is a positive integer;

[0163] Determine the text density of each initial box in the k initial boxes to obtain k text densities; the k text densities are used to reflect the density of the text in the k initial boxes;

[0164] Obtain the height and aspect ratio of each initial box in the k initial boxes to obtain k heights and k aspect ratios;

[0165] Determine the vertical offset and horizontal offset of each initial box in the k initial boxes according to the k text densities, k heights, and k aspect ratios, and obtain k vertical offsets and k horizontal offsets;

[0166] Determine k scaling factors and k rotation angles corresponding to the k initial boxes;

[0167] Adjust the k initial boxes according to the k vertical offsets, k horizontal offsets, k scaling factors and k rotation angles to obtain k oriented text boxes;

[0168] Determine the minimum horizontal rectangle of each oriented text box among the k oriented text boxes to obtain k prediction boxes;

[0169] Determine the first confidence level and the second confidence level of each prediction box among the k prediction boxes to obtain k first confidence levels and k second confidence levels; the k first confidence levels are used to reflect the probability that the k prediction boxes cover the target text; the k second confidence levels are used to reflect the probability that the text within the k prediction boxes is recognizable;

[0170] Filter the k prediction boxes according to the k first confidence levels and k second confidence levels to obtain j candidate boxes; j is a positive integer less than or equal to k;

[0171] Perform text extraction and background recognition on the j candidate boxes to obtain the first text and the first background of the target feature map;

[0172] According to the first text and the first background of the target feature map, determine the first text of each feature map among the m feature maps, and the first background corresponding to each first text, to obtain m first texts and m first backgrounds.

[0173] In a possible embodiment, in terms of determining the text density of each initial box among the k initial boxes to obtain k text densities, the processing unit 802 is specifically configured to:

[0174] Obtain the edge features and corner features of each initial box among the k initial boxes to obtain k edge features and k corner features;

[0175] According to the k edge features and k corner features, determine the activation value of each initial box among the k initial boxes to obtain k activation values; the k activation values are used to reflect the response intensity of the k initial boxes to the target text;

[0176] Determine the text density of each initial box among the k initial boxes according to the k activation values to obtain k text densities.

[0177] In a possible embodiment, in terms of determining the k scaling factors and k rotation angles corresponding to the k initial boxes, the processing unit 802 is specifically configured to:

[0178] Obtain the target image resolution and the target scale of the target feature map;

[0179] Determine the reference scaling factors of k initial boxes according to the target image resolution and the target scale;

[0180] Determine k scale scaling factors according to the reference scaling factors, k heights, and k width-to-height ratios;

[0181] Determine the horizontal gradient and vertical gradient of each initial box among the k initial boxes, obtaining k horizontal gradients and k vertical gradients;

[0182] Determine k rotation angles according to the k horizontal gradients and the k vertical gradients.

[0183] In a possible embodiment, when determining the target text according to m first texts, m first backgrounds, and the scale corresponding to each feature map among the m feature maps, the processing unit 802 is specifically configured to:

[0184] Determine the semantic relationship between each first text among the m first texts and the corresponding first background among the m first backgrounds, obtaining m semantic relationships;

[0185] Determine the third confidence level of each first text among the m first texts according to the m semantic relationships, obtaining m third confidence levels;

[0186] Obtain the proportion of the preset character type of each first text among the m first texts, obtaining m proportions; the preset character type includes missing characters and / or blurred characters;

[0187] Determine the fourth confidence level of each first text among the m first texts according to the m proportions, obtaining m fourth confidence levels;

[0188] Obtain the initial character size of each character in the m first texts in the target scene image;

[0189] Determine the fifth confidence level of each first text among the m first texts according to the scale corresponding to each feature map among the m feature maps and the initial character size of each character in the m first texts in the target scene image, obtaining m fifth confidence levels;

[0190] Determine the target text according to the preset confidence weight combination group, m third confidence levels, m fourth confidence levels, m fifth confidence levels, and the m first texts; the confidence weight combination group includes a first weight, a second weight, and a third weight.

[0191] In a possible embodiment, when determining n convolutional layers and m text box layers in the target neural network according to the image entropy value, the image texture complexity, the target recognition accuracy, and the target recognition efficiency, the processing unit 802 is specifically configured to:

[0192] Determine a first parameter value according to a preset parameter value, an image entropy value, and an image texture complexity; n is less than or equal to the first parameter value; the preset parameter value is used to reflect the basic number of layers of a convolutional layer;

[0193] Determine n convolutional layers and n scales according to the first parameter value, a target recognition accuracy, and a target recognition efficiency; the n scales correspond to the n convolutional layers one by one;

[0194] Determine a second parameter value according to the number of the n convolutional layers and the target recognition efficiency; m is less than or equal to the second parameter value;

[0195] Determine m text box layers according to each of the n scales, the adjacent scale corresponding to the scale, and the second parameter value; the adjacent scale is less than the scale.

[0196] In a possible embodiment, the processing unit 802 is further configured to:

[0197] Obtain the coordinate mapping relationship between each pixel point in the target scene image and the target text;

[0198] Determine the display parameters of the target text; the display parameters include at least one of the following: font format, font size, font transparency, display direction;

[0199] Display the target text in the target scene image according to the coordinate mapping relationship and the display parameters.

[0200] Refer to Figure 8 , Figure 8 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 8 shown, the electronic device 900 includes a transceiver 901, a processor 902, and a memory 903, which are connected by a bus 904. The memory 903 is used to store computer programs and data, and can transmit the data stored in the memory 903 to the processor 902. Among them, the electronic device 900 may be the above-mentioned text detection device 800, and the processor 902 may be the above-mentioned acquisition unit 801 and processing unit 802.

[0201] The processor 902 is configured to read the computer program in the memory 903 and perform the following operations:

[0202] Obtain a target scene image; the target scene image includes target text;

[0203] Determine the image entropy value of the target scene image according to the gray distribution of the target scene image;

[0204] Determine the image texture complexity of the target scene image according to the gray-level co-occurrence matrix of the target scene image;

[0205] Obtain the target recognition accuracy and target recognition efficiency of the target scene image;

[0206] Determine n convolutional layers and m text box layers in the target neural network according to the image entropy value, image texture complexity, target recognition accuracy, and target recognition efficiency; the m text box layers are connected to the corresponding m convolutional layers in the n convolutional layers one by one, where m and n are positive integers, and m is less than or equal to n;

[0207] Perform convolution on the target scene image through n convolutional layers to obtain n feature maps;

[0208] Obtain m feature maps from the n feature maps through m text box layers;

[0209] Determine the initial box set of each feature map in the m feature maps according to the target recognition accuracy, target recognition efficiency, and the scale corresponding to each feature map in the m feature maps, and obtain m initial box sets;

[0210] Determine the first text of each feature map in the m feature maps according to the m initial box sets, and the first background corresponding to each first text, and obtain m first texts and m first backgrounds;

[0211] Determine the target text according to the m first texts, m first backgrounds, and the scale corresponding to each feature map in the m feature maps.

[0212] In a possible embodiment, in terms of determining the first text of each feature map in the m feature maps according to the m initial box sets, and the first background corresponding to each first text, and obtaining m first texts and m first backgrounds, the processor 902 is specifically configured to perform the following operations:

[0213] Determine the target initial box set corresponding to the target feature map in the m initial box sets; the target feature map is any one of the m feature maps; the target initial box set includes k initial boxes, where k is a positive integer;

[0214] Determine the text density of each initial box in the k initial boxes to obtain k text densities; the k text densities are used to reflect the density of the text in the k initial boxes;

[0215] Obtain the height and aspect ratio of each initial box in the k initial boxes to obtain k heights and k aspect ratios;

[0216] Determine the vertical offset and horizontal offset of each initial box in the k initial boxes according to the k text densities, k heights, and k aspect ratios, and obtain k vertical offsets and k horizontal offsets;

[0217] Determine the k scale scaling factors and k rotation angles corresponding to the k initial boxes;

[0218] Adjust the k initial bounding boxes according to k vertical offsets, k horizontal offsets, k scale factors, and k rotation angles to obtain k oriented text bounding boxes;

[0219] Determine the minimum horizontal rectangle of each oriented text bounding box among the k oriented text bounding boxes to obtain k predicted bounding boxes;

[0220] Determine the first confidence level and the second confidence level of each predicted bounding box among the k predicted bounding boxes to obtain k first confidence levels and k second confidence levels; the k first confidence levels are used to reflect the probability that the k predicted bounding boxes cover the target text; the k second confidence levels are used to reflect the probability that the text within the k predicted bounding boxes is recognizable;

[0221] Filter the k predicted bounding boxes according to the k first confidence levels and the k second confidence levels to obtain j candidate bounding boxes; j is a positive integer less than or equal to k;

[0222] Perform text extraction and background recognition on the j candidate bounding boxes to obtain the first text and the first background of the target feature map;

[0223] According to the first text and the first background of the target feature map, determine the first text of each feature map among the m feature maps, and the first background corresponding to each first text, to obtain m first texts and m first backgrounds.

[0224] In a possible embodiment, in terms of determining the text density of each initial bounding box among the k initial bounding boxes to obtain k text densities, the processor 902 is specifically configured to perform the following operations:

[0225] Obtain the edge features and corner features of each initial bounding box among the k initial bounding boxes to obtain k edge features and k corner features;

[0226] According to the k edge features and the k corner features, determine the activation value of each initial bounding box among the k initial bounding boxes to obtain k activation values; the k activation values are used to reflect the response intensity of the k initial bounding boxes to the target text;

[0227] Determine the text density of each initial bounding box among the k initial bounding boxes according to the k activation values to obtain k text densities.

[0228] In a possible embodiment, in terms of determining the k scale factors and the k rotation angles corresponding to the k initial bounding boxes, the processor 902 is specifically configured to perform the following operations:

[0229] Obtain the target image resolution and the target scale of the target feature map;

[0230] According to the target image resolution and the target scale, determine the reference scale factor of the k initial bounding boxes;

[0231] Determine k scale scaling factors according to the reference scaling factor, k heights, and k aspect ratios.

[0232] Determine the horizontal gradient and vertical gradient of each initial box among the k initial boxes, obtaining k horizontal gradients and k vertical gradients.

[0233] Determine k rotation angles according to the k horizontal gradients and k vertical gradients.

[0234] In a possible embodiment, in determining the target text according to m first texts, m first backgrounds, and the scale corresponding to each feature map among the m feature maps, the processor 902 is specifically configured to perform the following operations:

[0235] Determine the semantic relationship between each first text among the m first texts and the corresponding first background among the m first backgrounds, obtaining m semantic relationships.

[0236] Determine the third confidence level of each first text among the m first texts according to the m semantic relationships, obtaining m third confidence levels.

[0237] Obtain the proportion of the preset character type of each first text among the m first texts, obtaining m proportions; the preset character type includes missing characters and / or blurred characters.

[0238] Determine the fourth confidence level of each first text among the m first texts according to the m proportions, obtaining m fourth confidence levels.

[0239] Obtain the initial character size of each character in the m first texts in the target scene image.

[0240] Determine the fifth confidence level of each first text among the m first texts according to the scale corresponding to each feature map among the m feature maps and the initial character size of each character in the m first texts in the target scene image, obtaining m fifth confidence levels.

[0241] Determine the target text according to the preset confidence weight combination, m third confidence levels, m fourth confidence levels, m fifth confidence levels, and the m first texts; the confidence weight combination includes a first weight, a second weight, and a third weight.

[0242] In a possible embodiment, in determining n convolutional layers and m text box layers in the target neural network according to the image entropy value, the image texture complexity, the target recognition accuracy, and the target recognition efficiency, the processor 902 is specifically configured to perform the following operations:

[0243] Determine a first parameter value according to the preset parameter value, the image entropy value, and the image texture complexity; n is less than or equal to the first parameter value; the preset parameter value is used to reflect the basic number of layers of the convolutional layer.

[0244] Determine n convolutional layers and n scales according to the first parameter value, the target recognition accuracy, and the target recognition efficiency; the n scales correspond to the n convolutional layers one by one;

[0245] Determine the second parameter value according to the number of the n convolutional layers and the target recognition efficiency; m is less than or equal to the second parameter value;

[0246] Determine m text box layers according to each of the n scales, the adjacent scale corresponding to the scale, and the second parameter value; the adjacent scale is less than the scale.

[0247] In a possible embodiment, the processor 902 is further configured to perform the following operations:

[0248] Obtain the coordinate mapping relationship between each pixel point in the target scene image and the target text;

[0249] Determine the display parameters of the target text; the display parameters include at least one of the following: font format, font size, font transparency, display direction;

[0250] Display the target text in the target scene image according to the coordinate mapping relationship and the display parameters.

[0251] The embodiment of the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement some or all of the steps of any one of the text detection methods described in the foregoing method embodiments.

[0252] The embodiment of the present application further provides a computer program product, where the computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute some or all of the steps of any one of the text detection methods described in the foregoing method embodiments.

[0253] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0254] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0255] In several embodiments provided in this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or module can be in electrical or other forms.

[0256] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0257] In addition, in each embodiment of this application, the functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software program modules.

[0258] If the above-mentioned integrated module is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. And the aforementioned memory includes: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0259] The above has introduced the embodiments of this application in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A text detection method, characterized in that: include: Acquire a target scene image; The target scene image includes target text; Determining an image entropy value of the target scene image according to the grayscale distribution of the target scene image; Determining the image texture complexity of the target scene image according to the gray level co-occurrence matrix of the target scene image; Acquire target recognition accuracy and target recognition efficiency of the target scene image; According to the image entropy value, the image texture complexity, the target recognition accuracy and the target recognition efficiency, determine n convolutional layers and m text box layers in the target neural network; the m text box layers are connected to the corresponding m convolutional layers in the n convolutional layers in a one-to-one correspondence, m and n are positive integers, and m is less than or equal to n; Convolving the target scene image through the n convolutional layers to obtain n feature maps; Obtaining m feature maps from the n feature maps through the m text box layers; Determine an initial frame set of each feature map in the m feature maps according to the target recognition accuracy, the target recognition efficiency, and a scale corresponding to each feature map in the m feature maps, to obtain m initial frame sets; Determine, according to the m initial frame sets, a first text of each feature map in the m feature maps and a first background corresponding to each first text, to obtain m first texts and m first backgrounds; The target text is determined according to the m first texts, the m first backgrounds, and a scale corresponding to each feature map in the m feature maps.

2. The method according to claim 1, characterized in that The step of determining, according to the m initial frame sets, a first text of each feature map in the m feature maps and a first background corresponding to each first text to obtain m first texts and m first backgrounds comprises: Determine a target initial frame set corresponding to a target feature map in the m initial frame sets; the target feature map is any feature map in the m feature maps; the target initial frame set includes k initial frames, where k is a positive integer; Determine the text density of each of the k initial frames to obtain k text densities; the k text densities are used to reflect the density of the texts in the k initial frames; Obtaining the height and aspect ratio of each of the k initial frames to obtain k heights and k aspect ratios; Determine a vertical offset and a horizontal offset of each of the k initial frames according to the k text densities, the k heights and the k aspect ratios, to obtain k vertical offsets and k horizontal offsets; Determine k scale scaling factors and k rotation angles corresponding to the k initial frames; The k initial frames are adjusted according to the k vertical offsets, the k horizontal offsets, the k scale scaling factors and the k rotation angles to obtain k directional text frames; Determine the minimum horizontal rectangle of each of the k oriented text boxes to obtain k prediction boxes; Determine a first confidence and a second confidence of each of the k prediction boxes to obtain k first confidences and k second confidences; the k first confidences are used to reflect the probability that the k prediction boxes cover the target text; the k second confidences are used to reflect the probability that the text in the k prediction boxes is recognizable; The k prediction boxes are screened according to the k first confidences and the k second confidences to obtain j candidate boxes, where j is a positive integer less than or equal to k; Performing text extraction and background recognition on the j candidate frames to obtain a first text and a first background of the target feature map; According to the first text and the first background of the target feature map, the first text of each feature map in the m feature maps and the first background corresponding to each first text are determined to obtain the m first texts and the m first backgrounds.

3. The method according to claim 2, characterized in that The determining of the text density of each of the k initial frames to obtain k text densities includes: Obtain edge features and corner point features of each of the k initial frames to obtain k edge features and k corner point features; Determine the activation value of each of the k initial boxes according to the k edge features and the k corner point features to obtain k activation values; the k activation values ​​are used to reflect the response strength of the k initial boxes to the target text; The text density of each of the k initial frames is determined according to the k activation values ​​to obtain the k text densities.

4. The method according to claim 2, characterized in that The determining k scale scaling factors and k rotation angles corresponding to the k initial frames includes: Obtaining a target image resolution and a target scale of the target feature map; Determining reference scaling factors of the k initial frames according to the target image resolution and the target scale; Determining the k scale scaling factors according to the reference scaling factor, the k heights and the k aspect ratios; Determine the horizontal gradient and the vertical gradient of each of the k initial frames to obtain k horizontal gradients and k vertical gradients; The k rotation angles are determined according to the k horizontal direction gradients and the k vertical direction gradients.

5. The method according to claim 1, characterized in that The determining the target text according to the m first texts, the m first backgrounds, and the scale corresponding to each feature map in the m feature maps includes: Determine a semantic relationship between each first text in the m first texts and a corresponding first background in the m first backgrounds to obtain m semantic relationships; Determining a third confidence level of each of the m first texts according to the m semantic relationships to obtain m third confidence levels; Obtaining a ratio of a preset character type of each of the m first texts to obtain m ratios; the preset character type includes missing characters and / or ambiguous characters; Determining a fourth confidence level of each of the m first texts according to the m ratios to obtain m fourth confidence levels; Obtaining an initial character size of each character in the m first texts in the target scene image; Determine a fifth confidence level of each of the m first texts according to a scale corresponding to each of the m feature maps and an initial character size of each character in the m first texts in the target scene image, to obtain m fifth confidence levels; The target text is determined according to a preset confidence weight group, the m third confidences, the m fourth confidences, the m fifth confidences and the m first texts; the confidence weight group includes a first weight, a second weight and a third weight.

6. The method according to claim 1, characterized in that The determining, according to the image entropy value, the image texture complexity, the target recognition accuracy and the target recognition efficiency, n convolutional layers and m text box layers in the target neural network comprises: Determine a first parameter value according to the preset parameter value, the image entropy value and the image texture complexity; n is less than or equal to the first parameter value; the preset parameter value is used to reflect the number of base layers of the convolution layer; Determining the n convolutional layers and the n scales according to the first parameter value, the target recognition accuracy, and the target recognition efficiency; the n scales correspond one-to-one to the n convolutional layers; Determine a second parameter value according to the number of the n convolutional layers and the target recognition efficiency; m is less than or equal to the second parameter value; The m text box layers are determined according to each scale of the n scales, an adjacent scale corresponding to the scale, and the second parameter value; the adjacent scale is smaller than the scale.

7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: Obtaining a coordinate mapping relationship between each pixel point in the target scene image and the target text; Determining display parameters of the target text; the display parameters include at least one of the following: font format, font size, font transparency, and display direction; The target text is displayed in the target scene image according to the coordinate mapping relationship and the display parameters.

8. A text detection device, characterized in that: The device comprises an acquisition unit and a processing unit; The acquisition unit is used to acquire a target scene image; the target scene image includes a target text; The processing unit is used to determine the image entropy value of the target scene image according to the grayscale distribution of the target scene image; Determining the image texture complexity of the target scene image according to the gray level co-occurrence matrix of the target scene image; Acquire target recognition accuracy and target recognition efficiency of the target scene image; According to the image entropy value, the image texture complexity, the target recognition accuracy and the target recognition efficiency, determine n convolutional layers and m text box layers in the target neural network; the m text box layers are connected to the corresponding m convolutional layers in the n convolutional layers in a one-to-one correspondence, m and n are positive integers, and m is less than or equal to n; Convolving the target scene image through the n convolutional layers to obtain n feature maps; Obtaining m feature maps from the n feature maps through the m text box layers; Determine an initial frame set of each feature map in the m feature maps according to the target recognition accuracy, the target recognition efficiency, and a scale corresponding to each feature map in the m feature maps, to obtain m initial frame sets; Determine, according to the m initial frame sets, a first text of each feature map in the m feature maps and a first background corresponding to each first text, to obtain m first texts and m first backgrounds; The target text is determined according to the m first texts, the m first backgrounds, and a scale corresponding to each feature map in the m feature maps.

9. An electronic device, characterized in that: include: A processor and a memory, the processor is connected to the memory, the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Natural scene text detection method based on semantic feature step-by-step recombination

    CN121789197A

  • A natural scene text detection method based on semantic feature step-by-step reorganization

    CN121789197B