Character detection method and device, electronic equipment and computer readable storage medium
Through the backbone network and detection head network of the image detection model combined with the spatial attention mechanism, the accuracy and stability of road sign text detection are solved, and accurate text recognition is achieved under complex backgrounds and unsatisfactory imaging conditions.
Patent Information
- Application Number
- CN202510259223.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-24
AI Technical Summary
The existing models have problems of inaccurate detection and poor stability when detecting the images of road sign text, mainly due to the diverse shapes of road sign text, complex background and unsatisfactory imaging conditions.
The backbone network of the image detection model is used for feature extraction, including convolutional layers of different convolution kernel sizes connected sequentially and feature extraction layers based on spatial attention mechanisms. The text area detection and feature segmentation are combined with the detection head network, and text recognition is performed using the codec model.
It improves the accuracy and stability of road sign text detection, and can accurately identify road sign text content under complex backgrounds and unsatisfactory imaging conditions.
Smart Images

Figure CN120198918A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology and is applied to the field of fintech, and particularly relates to a method and device for text detection, an electronic device, and a computer-readable storage medium. Background Art
[0002] Text detection technology refers to the technology of detecting text from images. There are more and more text information in the traffic environment, including road signs, traffic signs, and vehicle license plates, etc., collectively referred to as road sign text. In the auto insurance claims business in the financial scenario, an insurance company can perform text detection on relevant images of the driving road uploaded by a driver to facilitate the extraction of various road sign texts on the driving road and provide a reference for verifying matters such as auto insurance claim amounts.
[0003] Existing models often have problems with inaccurate detection when detecting images containing road sign text. The main reasons are as follows: First, the road sign text itself has diverse forms and attributes; second, the background environment is complex, there are objects similar to the text, and occlusion may occur; third, the imaging conditions are not ideal, such as low resolution, image distortion or blur, insufficient or excessive light, etc. Due to at least one of the above reasons or other reasons, the accuracy and stability of road sign text detection are poor. Summary of the Invention
[0004] The main purpose of the embodiments of this application is to propose a method and device for text detection, an electronic device, and a computer-readable storage medium, which can accurately detect images containing road sign text and improve the detection stability.
[0005] To achieve the above purpose, the first aspect of the embodiments of this application proposes a method for text detection, and the method includes:
[0006] Obtain an image containing a text area to obtain an initial text image; wherein, the text area contains road sign text;
[0007] Extract features from the initial text image through the backbone network of a preset image detection model to obtain text image features; wherein, the backbone network includes at least two sequentially connected convolutional layers and a feature extraction layer based on a spatial attention mechanism, and the convolutional kernels of any two of the convolutional layers are different;
[0008] Perform text area detection on the text image features through the detection head network of the image detection model to obtain text area position information;
[0009] Perform feature segmentation on the text image features according to the text area position information to obtain text area features;
[0010] Perform character recognition on the character region features to obtain the content of the road sign text corresponding to the road sign text;
[0011] Determine the road sign text position information and the road sign text content as road sign text detection data.
[0012] Optionally, the feature extraction of the initial text image through the backbone network of the preset image detection model to obtain text image features includes:
[0013] Perform feature extraction on the initial text image through the first convolutional layer to obtain the reference text image features output by the first convolutional layer; wherein, the first convolutional layer is the first convolutional layer among at least two convolutional layers;
[0014] Perform feature extraction on the reference text image features output by the previous convolutional layer through the second convolutional layer to obtain the reference text image features output by the second convolutional layer; wherein, the second convolutional layer is each convolutional layer after the first convolutional layer among at least two convolutional layers;
[0015] Concatenate the reference text image features output by the first convolutional layer and the reference text image features output by the second convolutional layer to obtain the target reference text image features;
[0016] Perform spatial attention calculation on the target reference text image features through the feature extraction layer to obtain the text image features.
[0017] Optionally, the performing spatial attention calculation on the target reference text image features through the feature extraction layer to obtain the text image features includes:
[0018] Perform channel average pooling on the target reference text image features through the feature extraction layer to obtain the first spatial description features;
[0019] Perform max pooling on the target reference text image features through the feature extraction layer to obtain the second spatial description features;
[0020] Perform channel concatenation on the first spatial description features and the second spatial description features through the feature extraction layer to obtain the comprehensive spatial description features;
[0021] Perform convolutional calculation on the comprehensive spatial description features through the feature extraction layer to obtain the spatial attention features;
[0022] Perform non-linear activation on the spatial attention features through the feature extraction layer to obtain the text image features.
[0023] Optionally, the convolution calculation of the comprehensive spatial description feature by the feature extraction layer to obtain a spatial attention feature includes:
[0024] Normalize the comprehensive spatial description feature through the feature extraction layer to obtain a standard comprehensive spatial description feature;
[0025] Perform non-linear activation on the standard comprehensive spatial description feature through the feature extraction layer to obtain a first spatial attention feature;
[0026] Perform 1×1 convolution on the first spatial attention feature through the feature extraction layer to obtain a second spatial attention feature;
[0027] Normalize the second spatial attention feature through the feature extraction layer to obtain a standard second spatial attention feature;
[0028] Perform non-linear activation on the standard second spatial attention feature through the feature extraction layer to obtain a spatial attention feature.
[0029] Optionally, the backbone network further includes a feature optimization layer, and the feature optimization layer is connected to the feature extraction layer;
[0030] After calculating the spatial attention of the target reference text image feature through the feature extraction layer to obtain the text image feature, the method further includes:
[0031] Perform convolution calculation on the text image feature output by the feature extraction layer through the feature optimization layer to obtain an initial text image optimization feature;
[0032] Split the initial text image optimization feature through the feature optimization layer to obtain a first split text image optimization feature and a second split text image optimization feature;
[0033] Perform spatial attention calculation on the second split text image optimization feature through the feature optimization layer to obtain a text image depth attention feature;
[0034] Concatenate the first split text image optimization feature and the text image depth attention feature through the feature optimization layer to obtain a text image concatenated attention feature;
[0035] Perform convolution calculation on the text image concatenated attention feature through the feature optimization layer to obtain a target text image optimization feature;
[0036] Determine the target text image optimization feature as the text image feature.
[0037] Optionally, the spatial attention calculation is performed on the optimized features of the second split text image through the feature optimization layer to obtain the depth attention features of the text image, including:
[0038] Performing first spatial attention calculation on the optimized features of the second split text image through the first spatial attention sub-layer of the feature optimization layer to obtain first depth attention features of the text image;
[0039] Performing second spatial attention calculation on the first depth attention features of the text image through the second spatial attention sub-layer of the feature optimization layer to obtain the depth attention features of the text image.
[0040] Optionally, the text recognition of the text region features to obtain the corresponding road sign text content of the road sign text includes:
[0041] Obtaining the feature format supported by the encoding and decoding model to obtain the target feature format;
[0042] Performing feature conversion on the text region features to obtain the text region features in the target feature format;
[0043] Performing text encoding and decoding on the text region features in the target feature format through the encoding and decoding model to obtain the road sign text content.
[0044] To achieve the above object, a second aspect of the embodiments of the present application proposes a text detection device, the device includes:
[0045] An image acquisition module, configured to acquire an image including a text region to obtain an initial text image; wherein, the text region includes road sign text;
[0046] A feature extraction module, configured to extract features from the initial text image through the backbone network of a preset image detection model to obtain text image features; wherein, the backbone network includes at least two sequentially connected convolutional layers and a feature extraction layer based on a spatial attention mechanism, and the convolutional kernels of any two of the convolutional layers are different;
[0047] A region detection module, configured to perform text region detection on the text image features through the detection head network of the image detection model to obtain text region position information;
[0048] A feature segmentation module, configured to perform feature segmentation on the text image features according to the text region position information to obtain text region features;
[0049] A text recognition module, configured to perform text recognition on the text region features to obtain the corresponding road sign text content of the road sign text;
[0050] A data determination module, configured to determine the text area position information and the road sign text content as road sign text detection data.
[0051] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the text detection method described in the first aspect above is implemented.
[0052] To achieve the above object, a fourth aspect of the embodiments of the present application provides a storage medium, which is a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, the text detection method described in the first aspect above is implemented.
[0053] The text detection method, text detection device, electronic device, and computer-readable storage medium proposed in the present application do not directly identify the text content from the image using a trained model. Instead, it first extracts features through an image detection model, and then performs text area detection. The backbone network of the image detection model includes at least two convolutional layers with different convolutional kernel sizes connected in sequence and a feature extraction layer based on a spatial attention mechanism. This can well handle the image information in the image that may affect the recognition of road sign text, and can not only obtain comprehensive and prominent text image features, but also accurately detect the text area position information. Further, according to the text area position information, the text image features are segmented into text area features, and then the text area features are recognized to obtain the text content. In this way, the text area features are relatively accurate, and the recognized road sign text content is relatively accurate. Finally, the text area position information and the road sign text content are merged together to obtain road sign text detection data. In summary, the present application can accurately detect images containing road sign text and improve the detection stability. Description of the Drawings
[0054] Figure 1 is a flowchart of the text detection method provided by the embodiments of the present application;
[0055] Figure 2 is Figure 1 a flowchart of step 102 in
[0056] Figure 3 is Figure 2 a flowchart of step 204 in
[0057] Figure 4 is Figure 3 a flowchart of step 304 in
[0058] Figure 5 is Figure 1Another flowchart of step 102 in
[0059] Figure 6 is Figure 5 the flowchart of step 503 in
[0060] Figure 7 the module structure block diagram of the text detection device provided by the embodiments of the present application;
[0061] Figure 8 the schematic hardware structure diagram of the electronic device provided by the embodiments of the present application. Detailed implementation manners
[0062] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, but not to limit the present application.
[0063] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the description and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.
[0064] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0065] First, several nouns involved in the present application are analyzed:
[0066] Artificial intelligence (AI): It is a new technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence; artificial intelligence is a branch of computer science. Artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence also uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results of theories, methods, technologies and application systems.
[0067] Natural Language Processing (NLP): NLP uses computers to process, understand, and apply human languages (such as Chinese, English, etc.). NLP is a branch of artificial intelligence and an interdisciplinary field of computer science and linguistics, often referred to as computational linguistics. Natural language processing includes syntactic analysis, semantic analysis, discourse understanding, etc. Natural language processing is commonly used in technical fields such as machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information image processing, information extraction and filtering, text classification and clustering, public opinion analysis, and opinion mining. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research related to language processing, and linguistic research related to language computing, etc.
[0068] With the rapid development of urban transportation, more and more text information has emerged in the traffic environment, including road signs, traffic signs, and vehicle license plates. These road sign texts are crucial for ensuring traffic safety and improving traffic efficiency. There are mainly three difficulties in the detection and recognition of road sign texts. First, the text itself has diverse forms and attributes. Second, the background environment is complex, with objects similar to the text and possible occlusions. Third, the imaging conditions are not ideal, such as low resolution, image distortion or blurring, insufficient or excessive light, etc.
[0069] With the rapid development of artificial intelligence and deep learning, the accuracy of text detection in natural scenes has been greatly improved. The natural scene text recognition algorithm includes a natural scene text recognition algorithm without segmentation. This algorithm usually adopts an end-to-end approach to directly recognize the text content from the image without explicitly segmenting the text area. Its model is relatively simple and has strong robustness. However, when recognizing road sign texts, problems such as reduced accuracy and poor accuracy stability are likely to occur.
[0070] Based on this, the embodiments of this application propose a text detection method, a text detection device, an electronic device, and a computer-readable storage medium, which can accurately detect images containing road sign texts and improve the detection stability.
[0071] The text detection method provided by the embodiments of this application can be applied to terminals and server sides, and can also be software running on the server side. The server side can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the text detection method, etc., but is not limited to the above forms.
[0072] This application can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: server computers, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0073] Embodiments of this application provide a text detection method, a text detection device, an electronic device, and a computer-readable storage medium, which will be specifically described through the following embodiments. First, the text detection method in the embodiments of this application will be described.
[0074] It should be noted that in each specific implementation manner of this application, when it comes to performing relevant processing based on data related to the user's identity or characteristics, such as the user's voice data, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards.
[0075] Referring to Figure 1 , Figure 1 is an optional flowchart of the text detection method provided by the embodiments of this application, which may include but is not limited to steps 101 to 106.
[0076] Step 101, obtain an image containing a text area to obtain an initial text image; wherein, the text area contains road sign text;
[0077] Step 102, perform feature extraction on the initial text image through the backbone network of a preset image detection model to obtain text image features; wherein, the backbone network includes at least two sequentially connected convolutional layers and a feature extraction layer based on a spatial attention mechanism, and the convolutional kernels of any two convolutional layers are different;
[0078] Step 103, perform text area detection on the text image features through the detection head network of the image detection model to obtain text area position information;
[0079] Step 104, perform feature segmentation on the text image features according to the text area position information to obtain text area features;
[0080] In step 105, perform character recognition on the text area features to obtain the corresponding road sign text content of the road sign text;
[0081] In step 106, determine the road sign text position information and the road sign text content as road sign text detection data.
[0082] Steps 101 to 106 illustrated in the embodiments of the present application do not directly identify the text content from the image using a trained model. Instead, feature extraction is first performed through an image detection model, and then text area detection is carried out. The backbone network of the image detection model includes at least two convolutional layers with different convolutional kernel sizes connected in sequence and a feature extraction layer based on a spatial attention mechanism. This can well handle the image information in the image that may affect the recognition of road sign text, and can not only obtain comprehensive and prominent text image features, but also accurately detect the road sign text position information. Further, according to the road sign text position information, the text area features are segmented from the text image features, and then character recognition is performed on the text area features to obtain the text content. In this way, the text area features are relatively accurate, and thus the recognized road sign text content is relatively accurate. Finally, the road sign text position information and the road sign text content are combined to obtain the road sign text detection data. In summary, the embodiments of the present application can accurately detect images containing road sign text and improve the detection stability.
[0083] When applying the embodiments of the present application to the quality inspection service of offline escort for health insurance in the financial scenario, the check-in pictures of the escort personnel can be obtained first, and text detection is performed on the check-in pictures to obtain road sign text detection data. Then, the road sign text detection data is compared with the opened longitude and latitude locations to confirm the arrival of the escort personnel, without manual detection, reducing the labor cost.
[0084] When applying the embodiments of the present application to the vehicle insurance claim settlement service in the financial scenario, the insurance company can perform text detection on the relevant images of the driving road uploaded by the driver, which is convenient for extracting various road sign texts on the driving road and providing a reference for verifying matters such as vehicle insurance claim amounts. For example, the road sign text "speed limit" is detected from the image uploaded by the driver applying for vehicle insurance claim settlement, and combined with the driver's vehicle speed record, it is automatically determined that the driver may have exceeded the speed limit. Then, it is transferred to the manual for re-verifying whether there is speeding behavior, and further determining whether to settle the vehicle insurance claim for the driver.
[0085] In step 101 of some embodiments, an image containing a text area is obtained to obtain an initial text image. Among them, the text area contains road sign text.
[0086] In one embodiment, step 101 may include: (1) using an image search engine: by searching for keywords such as "road sign" or "traffic sign" through an image search engine, many images containing road sign text may be found, and then the images are used as initial text images. (2) taking real-life photos: when taking photos using a smartphone or camera, different angles and light conditions may be selected to take photos of the road sign entity to obtain an initial text image.
[0087] In one example, the initial text image includes sign information and text information. The sign information indicates the type of sign, such as warning signs, instruction signs, and guide signs. Warning signs such as "Watch out for the intersection" and "Curve" usually use triangles. Instruction signs such as "Stop ahead" and "Speed limit" are usually circular. Guide signs such as direction signs contain destination information. Text information includes reminder information and nearby location prompt information. Reminder information includes specific warning or instruction statements, such as "No entry", "Turn right", "Drive less", etc. Nearby location prompt information, for example, some road signs will point to specific cities, attractions or facilities (such as "XX City, 20 kilometers"). Sometimes the road signs in the initial text image may contain additional information, such as updated traffic information, construction notices, precautions, etc.
[0088] In the quality inspection business of offline accompanying medical consultation for health insurance in the financial scenario, the road sign text contained in the text area of this embodiment may specifically include "XX Hospital", "XX City", "XX Road" and the like.
[0089] In step 102 of some embodiments, feature extraction is performed on the initial text image through a backbone network of a preset image detection model to obtain text image features.
[0090] The image detection model is a neural network model used to detect the text area in the initial text image. The backbone network is a part of the image detection model, which is used to convert the initial text image into text image features.
[0091] The backbone network includes at least two sequentially connected convolutional layers and a feature extraction layer based on the spatial attention mechanism, and the convolution kernels of any two convolutional layers are different. On the one hand, the convolutional layers with different convolution kernel sizes can deeply mine the image features of different scales, and on the other hand, the spatial attention of the feature extraction layer can further mine the potential associations between image features of different scales, significantly improving the feature information richness of text image features, and can highlight text information in text image features, thereby improving the accuracy of text detection.
[0092] In one example, the backbone network includes a first convolutional layer, a second convolutional layer, and a feature extraction layer based on a spatial attention mechanism that are sequentially connected. The kernel size of the first convolutional layer is smaller than that of the second convolutional layer. For example, the kernel size of the first convolutional layer is 3×3, and the kernel size of the second convolutional layer is 5×5.
[0093] The spatial attention mechanism is an attention mechanism widely used in computer vision tasks. Its core idea is to enable the model to focus on key spatial regions in the input image or feature map, and by assigning different weights to different regions, enhance the model's ability to process important information. In the embodiments of the present application, the spatial attention mechanism is used to assign higher weights to text regions, thereby enhancing the processing ability of the backbone network for text regions, and further improving the importance of the road sign text contained in the text regions in the text image features, and effectively highlighting the features corresponding to the road sign text.
[0094] In one example, the Backbone network of YOLOv11 can be selected as the backbone network of the image detection model, and the C3K2 module in the Backbone network is replaced with a spatial attention module. The spatial attention module includes at least two convolutional layers with different kernel sizes and a feature extraction layer based on a spatial attention mechanism that are sequentially connected. In this way, the accuracy of the text image features can be further improved through this backbone network.
[0095] In one embodiment, referring to Figure 2 , step 102 may include:
[0096] Step 201, perform feature extraction on the initial text image through the first convolutional layer to obtain the reference text image features output by the first convolutional layer; wherein, the first convolutional layer is the first convolutional layer among at least two convolutional layers;
[0097] Step 202, perform feature extraction on the reference text image features output by the previous convolutional layer through the second convolutional layer to obtain the reference text image features output by the second convolutional layer; wherein, the second convolutional layer is each convolutional layer after the first convolutional layer among at least two convolutional layers;
[0098] Step 203, splice the reference text image features output by the first convolutional layer and the reference text image features output by the second convolutional layer to obtain the target reference text image features;
[0099] Step 204, perform spatial attention calculation on the target reference text image features through the feature extraction layer to obtain the text image features.
[0100] In one example, the first convolutional layer includes at least one 3×3 convolutional kernel, and the second convolutional layer includes at least one 5×5 second convolutional kernel. Feature extraction is performed on the initial text image through at least one 3×3 convolutional kernel to obtain the reference text image feature p output by the first convolutional layer 1out . Feature extraction is performed on the reference text image feature p through the 5×5 convolutional kernel in the first second convolutional layer 1out to obtain the reference text image feature p output by the first second convolutional layer 2out . Feature extraction is performed on the reference text image feature p through the 5×5 convolutional kernel in the second second convolutional layer 2out to obtain the reference text image feature p output by the second second convolutional layer 3out . The reference text image features p 1out , the reference text image features p 2out , and the reference text image features p 3out are concatenated to obtain the target reference text image feature p out . Finally, spatial attention calculation is performed on the target reference text image feature p through the feature extraction layer out to obtain the text image feature.
[0101] It should be noted that only 2 second convolutional layers are shown in the above example, but the specific number can be set according to requirements.
[0102] The benefits of the embodiments of the above steps 201 to 204 are that by jointly using convolutional layers with different convolutional kernel sizes and using the spatial attention mechanism, the processing ability of the backbone network for text regions can be enhanced, and further the importance of the road sign texts included in the text regions in the text image features can be improved, and the features corresponding to the road sign texts can be effectively highlighted.
[0103] In one embodiment, referring to Figure 3 , step 204 may include:
[0104] Step 301, performing channel average pooling on the target reference text image feature through the feature extraction layer to obtain the first spatial description feature;
[0105] Step 302, performing max pooling on the target reference text image feature through the feature extraction layer to obtain the second spatial description feature;
[0106] Step 303, performing channel concatenation on the first spatial description feature and the second spatial description feature through the feature extraction layer to obtain the comprehensive spatial description feature;
[0107] Step 304, performing convolutional calculation on the comprehensive spatial description feature through the feature extraction layer to obtain the spatial attention feature;
[0108] Step 305, nonlinearly activate the spatial attention features through the feature extraction layer to obtain text image features.
[0109] Specifically, the feature extraction layer may include a channel average pooling sublayer, a maximum pooling sublayer, a splicing sublayer, a convolution sublayer, and a non-linear activation layer.
[0110] The channel-wise average pooling sublayer is used to perform channel-wise pooling on feature maps. Channel-wise Average Pooling is a pooling operation commonly used in deep learning, especially in convolutional neural networks (CNNs). It reduces the spatial dimensions (width and height) of a specific feature map by averaging each channel, but keeps the number of channels unchanged. This allows the feature map to retain important information for the model while reducing the amount of computation.
[0111] The max pooling sublayer is used to perform maximum pooling on the feature map. Max pooling is a commonly used downsampling technique that is widely used in convolutional neural networks (CNNs). Its main purpose is to reduce the size of the feature map while retaining the most important feature information, helping to improve the performance and computational efficiency of the model. The basic idea of max pooling is to select the maximum value as the output in a certain area of the feature map.
[0112] The concatenation sublayer is used to concatenate multiple features. Its main function is to concatenate multiple input layers along a specific dimension to form a new output layer. This operation helps to combine features from different sources, allowing the model to learn with more information.
[0113] The convolutional sublayer uses a set of learnable convolution kernels (or filters), the dimensions of each convolution kernel are usually small (such as 3×3, 5×5), and their number determines the number of channels of the output feature map. The size and number of convolution kernels are hyperparameters that need to be selected during network design.
[0114] The non-linear activation sublayer is used for non-linear activation, such as using activation functions such as sigmoid.
[0115] The benefit of the above-mentioned embodiment of step 301 to step 305 is that it can accurately and deeply mine the potential feature information of the target reference text image features related to the road sign text, thereby improving the detection accuracy of the road sign text.
[0116] In one embodiment, referring to Figure 4 , step 304 may include:
[0117] Step 401, standardizing the comprehensive spatial description features through the feature extraction layer to obtain standard comprehensive spatial description features;
[0118] Step 402: Non-linearly activate the standard comprehensive space description features through the feature extraction layer to obtain the first spatial attention feature;
[0119] Step 403: Perform 1×1 convolution on the first spatial attention feature through the feature extraction layer to obtain the second spatial attention feature;
[0120] Step 404: Standardize the second spatial attention feature through the feature extraction layer to obtain the standard second spatial attention feature;
[0121] Step 405: Non-linearly activate the standard second spatial attention feature through the feature extraction layer to obtain the spatial attention feature.
[0122] Specifically, the feature extraction layer mentioned in this embodiment is specifically the convolutional sub-layer in the feature extraction layer. The convolutional sub-layer may include a first normalization unit, a first non-linear activation unit, a 1×1 convolution unit, a second normalization unit, and a second non-linear activation unit, corresponding to steps 401 to 405 respectively.
[0123] Both the first normalization unit and the second normalization unit are used to normalize the features, also known as normalization. The specific explanations of the first non-linear activation unit, the 1×1 convolution unit, and the second non-linear activation unit are similar to those recorded above, and will not be elaborated here.
[0124] The benefits of the above embodiments of steps 401 to 405 are that they can improve the non-linear characteristics of the spatial attention feature, thereby improving the accuracy of text detection.
[0125] In one embodiment, the backbone network further includes a feature optimization layer, and the feature optimization layer is connected to the feature extraction layer.
[0126] Refer to Figure 5 , after step 204, step 102 may further include:
[0127] Step 501: Perform convolution calculation on the text image features output by the feature extraction layer through the feature optimization layer to obtain the initial text image optimization features;
[0128] Step 502: Split the initial text image optimization features through the feature optimization layer to obtain the first split text image optimization features and the second split text image optimization features;
[0129] Step 503: Perform spatial attention calculation on the second split text image optimization features through the feature optimization layer to obtain the text image depth attention features;
[0130] Step 504, the first split text image optimized feature and the text image depth attention feature are concatenated through the feature optimization layer to obtain the text image concatenated attention feature;
[0131] Step 505, the text image concatenated attention feature is subjected to convolution calculation through the feature optimization layer to obtain the target text image optimized feature;
[0132] Step 506, the target text image optimized feature is determined as the text image feature.
[0133] Specifically, the feature optimization layer includes a first feature convolution sub-layer, a feature splitting sub-layer, a feature attention calculation sub-layer, a feature concatenation sub-layer, a second feature convolution sub-layer, and a feature output sub-layer, corresponding to steps 501 to 506 respectively.
[0134] The feature splitting sub-layer (also known as the split layer) refers to a layer that splits the input feature into multiple parts. This layer is very useful in processing complex data and allows the model to process different information in parallel in different parts. The relevant introductions of other sub-layers in the feature optimization layer can be found above and will not be elaborated here.
[0135] The benefits of the embodiments of the above steps 501 to 506 are that they can optimize the text image features output by the feature extraction layer, thereby improving the accuracy of the text image features, and further improving the accuracy of text detection.
[0136] In one embodiment, referring to Figure 6 , step 503 may include:
[0137] Step 601, the first split text image optimized feature is subjected to the first spatial attention calculation through the first spatial attention sub-layer of the feature optimization layer to obtain the first text image depth attention feature;
[0138] Step 602, the first text image depth attention feature is subjected to the second spatial attention calculation through the second spatial attention sub-layer of the feature optimization layer to obtain the text image depth attention feature.
[0139] Specifically, the above first spatial attention sub-layer and second spatial attention sub-layer are collectively referred to as the feature attention calculation sub-layer, which is used to perform spatial attention calculation on the input feature.
[0140] The benefits of the embodiments of the above steps 601 to 602 are that they can further improve the expression ability of the text image depth attention feature for the features corresponding to the road sign text, and further improve the accuracy of text detection.
[0141] In step 103 of some embodiments, the text region detection is performed on the text image features through the detection head network of the image detection model to obtain the text region position information. The detection head network of the image detection model is used to detect the position and category of the target. For example, the detection head network is the Detection Head network in the real-time object detection machine learning algorithm model (You Only Look Once, YOLO).
[0142] In an example, the text image features include text image features of multiple different scales. Step 103 may include:
[0143] (1) Sampling the text image features of multiple different scales to the same scale by using the Feature Pyramid Network (FPN), and then performing feature fusion on the sampled features in the Detection Head by using multiple spatial attention modules. The spatial attention module includes at least two convolutional layers with different convolutional kernel sizes connected in sequence and a feature extraction layer based on the spatial attention mechanism, which can perform feature fusion faster and more efficiently.
[0144] (2) Next, perform bounding box transformation and non-maximum suppression on the fused text image features to obtain the text region position information, which can be specifically represented as the bounding box of the text region.
[0145] In step 104 of some embodiments, feature segmentation is performed on the text image features according to the text region position information to obtain the text region features.
[0146] Combined with the above example, after (2) in the above example, it further includes: (3) According to the bounding box of the detected text region, crop the corresponding text region features from the fused text image features. Coordinate transformation can also be performed to map the bounding box coordinates from the current feature scale to the feature scale of the backbone network. If the sizes of the cropped features are inconsistent, size adjustment operations are required.
[0147] In step 105 of some embodiments, text recognition is performed on the text region features to obtain the content of the road sign text corresponding to the road sign.
[0148] In one embodiment, step 105 may include: obtaining the feature format supported by the encoder-decoder model to obtain the target feature format; performing feature transformation on the text region features to obtain the text region features in the target feature format; and performing text encoding and decoding on the text region features in the target feature format through the encoder-decoder model to obtain the road sign text content.
[0149] The encoding and decoding model is a Transformer model. Specifically, the text region features are converted into the target feature format supported by the Transformer, and the encoder Encoder is used to encode the text region features in the target feature format. Then, the decoder Decoder is used to generate a text sequence from the encoded text region features to obtain the road sign text content.
[0150] The advantage of the above embodiment is that it is not necessary to fine-tune the encoding and decoding model in advance based on the road sign text. Instead, it is only necessary to dynamically adjust the format of the text region features during the inference stage, which can also save the cost of multiple fine-tuning of the encoding and decoding model due to the diverse forms of road sign text, and has high universality.
[0151] In one example, step 105 may include:
[0152] (1) Flatten the text region features into a vector. Specifically, a three-dimensional text region feature is unfolded into a one-dimensional vector along the number of channel bits. That is, a feature map of (C, H, W) is converted into a vector of (1, C * H * W). Then, a linear projection is performed through a fully connected layer FC, and the operation is to multiply the feature vector by a weight matrix to map the flattened feature vector to a new low-dimensional or high-dimensional space.
[0153] (2) Perform positional encoding on the flattened vector, and add the obtained positional encoding and the flattened vector together as the input of the Transformer. The subsequent work is the processing process of the general Transformer's Encoder and Decoder;
[0154] (3) After a series of calculations of the input and positional encoding, Masked Multi-Head Attention, Encoder-Decoder Multi-Head Attention, and feed-forward neural network in the Decoder, the output is obtained. The output will be input into an MLP layer to complete the linear transformation of the output of the transformer result. After inputting into the MLP, it will be input into another MLP structure. The two MLP structures are different. The first MLP is to complete batch normalization and focus on the distribution characteristics of the training batch dataset. The second MLP is to focus on the feature mapping of its own data. The final MLP output enters a softmax function for classification, and the softmax outputs the probability that each class is the road sign text. The maximum probability is taken to complete the prediction and recognition of the road sign text, and finally the road sign text content is obtained.
[0155] Based on the above embodiments, the present application can at least achieve the following beneficial effects: (1) The context information inside and around the text in the image can be utilized to improve the recognition accuracy, especially in the presence of noise, illumination changes, occlusion, etc. The position information of the text area provided by the backbone network helps subsequent text recognition to better focus on the text content and reduce background interference. (2) The combination of the robustness of the backbone network to scale changes and the sequence modeling ability of the encoder-decoder model can better handle various complex situations in natural scenes, such as texts of different sizes, orientations, fonts, and partially blurred or occluded texts. The attention mechanism of the encoder-decoder model can ignore irrelevant parts and focus on the text information, further improving the robustness. (3) The image detection model has high feature expression ability, solves the problems of complex backgrounds, different types of road sign texts, small-sized road sign texts, occlusion and deformation, and improves the accuracy of text recognition by introducing the encoder-decoder model, especially in complex scenes where the image is tilted, deformed, and has illumination changes.
[0156] Please refer to Figure 7 , the embodiments of the present application also provide a text detection device, which can implement the above text detection method. Figure 7 It is a block diagram of the module structure of the text detection device provided by the embodiments of the present application. The device includes:
[0157] An image acquisition module 701, configured to acquire an image containing a text area to obtain an initial text image; wherein, the text area includes road sign text.
[0158] A feature extraction module 702, configured to extract features from the initial text image through the backbone network of a preset image detection model to obtain text image features; wherein, the backbone network includes at least two sequentially connected convolutional layers and a feature extraction layer based on a spatial attention mechanism, and the convolutional kernels of any two convolutional layers are different.
[0159] A region detection module 703, configured to detect the text area from the text image features through the detection head network of the image detection model to obtain text area position information.
[0160] A feature segmentation module 704, configured to segment the text image features according to the text area position information to obtain text area features.
[0161] A text recognition module 705, configured to recognize the text from the text area features to obtain the road sign text content corresponding to the road sign text.
[0162] A data determination module 706, configured to determine the text area position information and the road sign text content as road sign text detection data.
[0163] It should be noted that the specific implementation of this text detection device is basically the same as the specific embodiments of the above text detection method, and will not be elaborated here.
[0164] An embodiment of this application also provides an electronic device, which includes: a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing the connection and communication between the processor and the memory. When the program is executed by the processor, the above text detection method is realized. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0165] Please refer to Figure 8 , Figure 8 , which shows the hardware structure of an electronic device in another embodiment. The electronic device includes:
[0166] A processor 801, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of this application;
[0167] A memory 802, which can be implemented in forms such as a Read Only Memory (ROM), a static storage device, a dynamic storage device, or a Random Access Memory (RAM). The memory 802 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 802, and the processor 801 is used to call and execute the text detection method of the embodiments of this application;
[0168] An input / output interface 803, which is used to realize information input and output;
[0169] A communication interface 804, which is used to realize the communication interaction between this device and other devices. It can communicate through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0170] A bus 805, which transmits information between various components of the device (such as the processor 801, the memory 802, the input / output interface 803, and the communication interface 804);
[0171] Among them, the processor 801, the memory 802, the input / output interface 803, and the communication interface 804 realize the communication connection between each other inside the device through the bus 805.
[0172] The embodiment of the present application also provides a storage medium, which is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the above-mentioned text detection method.
[0173] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0174] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation to the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0175] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation to the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.
[0176] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0177] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0178] In the description of the present application and the above-mentioned accompanying drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0179] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B may be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c may be single or multiple.
[0180] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0181] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0182] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0183] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0184] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall fall within the scope of the rights of the embodiments of the present application.
Claims
1. A text detection method, characterized in that: The method comprises: Acquire an image containing a text area to obtain an initial text image; wherein the text area contains road sign text; The initial text image is subjected to feature extraction through a backbone network of a preset image detection model to obtain text image features; wherein the backbone network includes at least two sequentially connected convolutional layers and a feature extraction layer based on a spatial attention mechanism, and the convolution kernels of any two convolutional layers are different; Performing text area detection on the text image features through the detection head network of the image detection model to obtain text area position information; Performing feature segmentation on the text image features according to the text region position information to obtain text region features; Performing text recognition on the text area features to obtain the road sign text content corresponding to the road sign text; The text area position information and the road sign text content are determined as road sign text detection data.
2. The method according to claim 1, characterized in that The extracting features of the initial text image through the backbone network of the preset image detection model to obtain text image features includes: Performing feature extraction on the initial text image through a first convolutional layer to obtain reference text image features output by the first convolutional layer; wherein the first convolutional layer is the first convolutional layer of at least two convolutional layers; Extracting the reference text image features output by the previous convolution layer through the second convolution layer to obtain the reference text image features output by the second convolution layer; wherein the second convolution layer is each convolution layer after the first convolution layer of at least two convolution layers; splicing the reference text image features output by the first convolutional layer and the reference text image features output by the second convolutional layer to obtain target reference text image features; The feature extraction layer performs spatial attention calculation on the target reference text image feature to obtain the text image feature.
3. The method according to claim 2, characterized in that The step of performing spatial attention calculation on the target reference text image feature through the feature extraction layer to obtain the text image feature includes: Perform channel average pooling on the target reference text image features through the feature extraction layer to obtain a first spatial description feature; Performing maximum pooling on the target reference text image features through the feature extraction layer to obtain a second spatial description feature; Perform channel splicing of the first spatial description feature and the second spatial description feature through the feature extraction layer to obtain a comprehensive spatial description feature; Performing convolution calculation on the comprehensive spatial description feature through the feature extraction layer to obtain a spatial attention feature; The spatial attention feature is nonlinearly activated through the feature extraction layer to obtain the text image feature.
4. The method according to claim 3, characterized in that The convolution calculation is performed on the comprehensive spatial description feature by the feature extraction layer to obtain the spatial attention feature, including: The comprehensive spatial description feature is standardized by the feature extraction layer to obtain a standard comprehensive spatial description feature; The standard comprehensive spatial description feature is nonlinearly activated by the feature extraction layer to obtain a first spatial attention feature; Performing a 1×1 convolution on the first spatial attention feature through the feature extraction layer to obtain a second spatial attention feature; Standardizing the second spatial attention feature through the feature extraction layer to obtain a standard second spatial attention feature; The standard second spatial attention feature is nonlinearly activated through the feature extraction layer to obtain a spatial attention feature.
5. The method according to any one of claims 2 to 4, characterized in that: The backbone network also includes a feature optimization layer, and the feature optimization layer is connected to the feature extraction layer; After performing spatial attention calculation on the target reference text image feature through the feature extraction layer to obtain the text image feature, the method further includes: Performing convolution calculation on the text image features output by the feature extraction layer through the feature optimization layer to obtain initial text image optimization features; Splitting the initial text image optimization feature through the feature optimization layer to obtain a first split text image optimization feature and a second split text image optimization feature; Performing spatial attention calculation on the second split text image optimization feature through the feature optimization layer to obtain a text image deep attention feature; The first split text image optimization feature and the text image deep attention feature are spliced through the feature optimization layer to obtain a text image splicing attention feature; Performing convolution calculation on the text image splicing attention feature through the feature optimization layer to obtain the target text image optimization feature; The target text image optimization feature is determined as the text image feature.
6. The method according to claim 5, characterized in that The step of performing spatial attention calculation on the second split text image optimization feature through the feature optimization layer to obtain a text image deep attention feature includes: Performing a first spatial attention calculation on the second split text image optimization feature through the first spatial attention sublayer of the feature optimization layer to obtain a first text image deep attention feature; Through the second spatial attention sublayer of the feature optimization layer, a second spatial attention calculation is performed on the first text image deep attention feature to obtain the text image deep attention feature.
7. The method according to any one of claims 1 to 4, characterized in that: The performing text recognition on the text area feature to obtain the road sign text content corresponding to the road sign text includes: Get the feature format supported by the codec model and get the target feature format; Performing feature conversion on the text region feature to obtain the text region feature in the target feature format; The text area features of the target feature format are encoded and decoded by the encoding and decoding model to obtain the text content of the road sign.
8. A text detection device, characterized in that: The device comprises: An image acquisition module, used to acquire an image containing a text area to obtain an initial text image; wherein the text area contains road sign text; A feature extraction module, used to extract features of the initial text image through a backbone network of a preset image detection model to obtain text image features; wherein the backbone network includes at least two convolutional layers and a feature extraction layer based on a spatial attention mechanism connected in sequence, and the convolution kernels of any two convolutional layers are different; A region detection module, used to perform text region detection on the text image features through the detection head network of the image detection model to obtain text region position information; A feature segmentation module, used to perform feature segmentation on the text image features according to the text region position information to obtain text region features; A text recognition module, used to perform text recognition on the text area features to obtain the road sign text content corresponding to the road sign text; The data determination module is used to determine the text area position information and the road sign text content as road sign text detection data.
9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the text detection method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the text detection method according to any one of claims 1 to 7 is implemented.