Natural scene text detection and recognition method and related device

By acquiring image data in natural scenes and determining the target application scenarios, selecting and optimizing text detection and recognition models, the problems of unstable performance and insufficient adaptability of text detection and recognition in the prior art are solved, and efficient and accurate text detection and recognition are achieved.

CN120198904APending Publication Date: 2025-06-24SHENZHEN POWER SUPPLY BUREAU
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510296644.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing natural scene text detection methods are instable due to factors such as resolution, change in direction, occlusion, and lighting, resulting in unstable detection performance, poor universality and high missed detection rate; the text recognition methods face complex fonts, occlusion, low resolution, lighting changes and special characters, and their recognition capabilities are insufficient, their adaptability to complex conditions is weak, and their accuracy is difficult to meet actual needs.

Method used

By obtaining the initial image data and determining its corresponding target application scenarios, the initial text detection and initial text recognition models are accurately selected from the pre-trained model library, and the preset training data set is used to train and optimize the model in a targeted manner, combining a variety of feature extraction methods and detection algorithms to achieve accurate acquisition of text areas and optimization of recognition results.

Benefits of technology

High-precision text detection and accurate recognition are achieved in complex natural scenarios, effectively improving the efficiency, accuracy and robustness of text detection and recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198904A_ABST
    Figure CN120198904A_ABST
Patent Text Reader

Abstract

The invention discloses a natural scene text detection and recognition method and a related device. The method comprises the following steps: acquiring initial image data and a target application scene corresponding to the initial image data; obtaining an initial text detection model from a first pre-training model library and obtaining an initial text recognition model from a second pre-training model library according to the initial image data and the target application scene; training the initial text detection model according to a preset first training data set to obtain a target text detection model, and training the initial text recognition model according to a preset second training data set to obtain a target text recognition model; preprocessing the initial image data to obtain target image data; inputting the target image data into a target text detection model to obtain a target text area; inputting the target text area into a target text recognition model to obtain a first text; and performing result optimization on the first text to obtain a second text. By adopting the method and the device, the accuracy of natural scene text detection and recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of text detection and recognition, and in particular, to a method and related device for natural scene text detection and recognition. Background Art

[0002] In the current era of rapid digital and intelligent development, natural scene text detection and recognition technology plays a key role in many fields such as robotics, industrial automation, image search, instant translation, automotive assistance, and sports video analysis.

[0003] Existing text detection methods include two categories: classical machine learning and deep learning. Among them, classical deep learning methods specifically include the sliding window method, the connected component method, etc., and deep learning methods specifically include the bounding box regression method, the segmentation method, the hybrid method, etc. Existing text recognition methods also include two categories: classical machine learning and deep learning. In the early stage, features were extracted based on Convolutional Neural Networks (CNN) and combined with Non-Maximum Suppression (NMS) to predict words or a fully connected network and the n-gram method were used to recognize characters. Subsequently, inspired by speech recognition, methods based on Connectionist Temporal Classification (CTC) were developed, or methods based on the attention mechanism were used to automatically learn and enhance decoding features. Some methods also used an image correction module to process irregular text.

[0004] Due to factors such as the resolution, variable orientation, occlusion, and illumination of natural scene text, the detection performance of existing text detection methods is unstable, the generality is poor, and the miss detection rate is high. In the face of complex fonts, occlusion, low resolution, illumination changes, and special characters, etc., the recognition ability of text recognition methods is insufficient, the adaptability to complex conditions is weak, and the accuracy is difficult to meet the actual needs. Therefore, how to improve the accuracy of natural scene text detection and recognition has become an urgent problem to be solved. Summary of the Invention

[0005] The embodiments of the present application provide a method and related device for natural scene text detection and recognition. By obtaining initial image data and determining the corresponding target application scenario, accurately selecting an initial text detection model and an initial text recognition model from a pre-trained model library according to the target application scenario and image characteristics, then using a preset training data set to perform targeted training and optimization on the models. The trained text detection model accurately obtains the text region by combining multiple feature extraction methods and detection algorithms, and the trained text recognition model determines the recognition result according to the text region. Finally, the recognition result is optimized based on a dictionary library and a language model to obtain accurate text, realizing high-precision detection and accurate recognition of text in complex natural scenes, and effectively improving the efficiency, accuracy and robustness of text detection and recognition.

[0006] In a first aspect, the embodiments of the present application provide a method for natural scene text detection and recognition, including:

[0007] Obtain initial image data and determine the target application scenario corresponding to the initial image data;

[0008] Obtain an initial text detection model from a first pre-trained model library according to the initial image data and the target application scenario, and obtain an initial text recognition model from a second pre-trained model library according to the initial image data and the target application scenario;

[0009] Train the initial text detection model according to a preset first training data set to obtain a target text detection model, and train the initial text recognition model according to a preset second training data set to obtain a target text recognition model;

[0010] Preprocess the initial image data to obtain target image data;

[0011] Input the target image data into the target text detection model to obtain a target text region;

[0012] Input the target text region into the target text recognition model to obtain a first text;

[0013] Optimize the result of the first text to obtain a second text.

[0014] In a second aspect, the embodiments of the present application provide a device for natural scene text detection and recognition, including:

[0015] An image acquisition module, configured to obtain initial image data and determine the target application scenario corresponding to the initial image data;

[0016] A model selection module, configured to obtain an initial text detection model from a first pre-trained model library according to the initial image data and the target application scenario, and obtain an initial text recognition model from a second pre-trained model library according to the initial image data and the target application scenario;

[0017] A model training module, configured to train the initial text detection model according to a preset first training data set to obtain a target text detection model, and train the initial text recognition model according to a preset second training data set to obtain a target text recognition model;

[0018] An image preprocessing module, configured to preprocess the initial image data to obtain target image data;

[0019] A text detection module, configured to input the target image data into the target text detection model to obtain a target text region;

[0020] A text recognition module, configured to input the target text region into the target text recognition model to obtain a first text;

[0021] A result optimization module, configured to optimize the result of the first text to obtain a second text.

[0022] In a third aspect, an embodiment of the present application provides an electronic device, including: a processor, a memory, a communication interface, and one or more programs, wherein the above one or more programs are stored in the above memory and are configured to be executed by the above processor, and the above programs include instructions for executing the steps in the first aspect of the embodiment of the present application.

[0023] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program for electronic data exchange, and the computer program enables a computer to execute some or all of the steps described in the first aspect of the embodiment of the present application.

[0024] In a fifth aspect, an embodiment of the present application provides a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute some or all of the steps described in the first aspect of the embodiment of the present application. The computer program product may be a software installation package.

[0025] It can be seen that by adopting the embodiment of the present application, the following beneficial effects are achieved:

[0026] By implementing the embodiments of the present application, initial image data is obtained, and the target application scenario corresponding to the initial image data is determined; an initial text detection model is obtained from the first pre-trained model library according to the initial image data and the target application scenario, and an initial text recognition model is obtained from the second pre-trained model library according to the initial image data and the target application scenario; the initial text detection model is trained according to a preset first training dataset to obtain a target text detection model, and the initial text recognition model is trained according to a preset second training dataset to obtain a target text recognition model; the initial image data is preprocessed to obtain target image data; the target image data is input into the target text detection model to obtain a target text region; the target text region is input into the target text recognition model to obtain a first text; and the first text is optimized in results to obtain a second text. It can be seen that by considering the characteristics of the initial image data and the target application scenario to select and train the models, the target text detection model and the target text recognition model are more in line with the actual application requirements, effectively improving the accuracy of text detection and recognition. Description of the Drawings

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the background art, the drawings required to be used in the embodiments of the present application or the background art will be described below.

[0028] Figure 1 is a schematic flowchart of a method for natural scene text detection and recognition provided by an embodiment of the present application;

[0029] Figure 2 is a schematic diagram of a scene of image pixel segmentation provided by an embodiment of the present application;

[0030] Figure 3 is a schematic diagram of a scene of a method for natural scene text detection and recognition provided by an embodiment of the present application;

[0031] Figure 4 is a system architecture diagram of a natural scene text detection and recognition system provided by an embodiment of the present application;

[0032] Figure 5 is a server architecture diagram provided by an embodiment of the present application;

[0033] Figure 6 is a schematic structural diagram of a natural scene text detection and recognition device provided by an embodiment of the present application;

[0034] Figure 7 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments

[0035] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts belong to the scope of protection of this application.

[0036] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products or devices.

[0037] Referring to "embodiment" herein means that a specific feature, structure or characteristic described in connection with the embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.

[0038] The relevant content, concepts, meanings, technical problems, technical solutions, beneficial effects, etc. involved in the embodiments of this application will be described below.

[0039] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a natural scene text detection and recognition method provided by an embodiment of this application. The method includes but is not limited to the following steps:

[0040] S101. Obtain initial image data and determine the target application scenario corresponding to the initial image data.

[0041] In the embodiments of this application, an image acquisition device can be used to obtain initial image data. For example, a digital camera, a mobile phone camera, a surveillance camera, a vehicle-mounted camera, etc. can be used to capture images in a natural scene, and the captured images are used as the initial image data. Image data can also be read from existing data storage media. For example, image data can be read from data storage media such as hard disks, optical discs, and USB flash drives to obtain the initial image data.

[0042] The initial image data is image data containing different scenes, texts of different fonts and sizes. By performing text detection and text recognition on the initial image data, the texts recorded in the initial image data can be obtained. The initial image data includes not only text information, but also background information, color information, texture information, the size and resolution of the image, image storage format information, etc. Considering these information during text detection and text recognition can better and more accurately recognize the texts in the image.

[0043] In the embodiments of the present application, it is also necessary to determine the target application scenario corresponding to the initial image data, that is, the natural scene. Herein, the target application scenario refers to the specific actual application context corresponding to the initial image data, which includes aspects such as the environment, field, and expected usage purpose of the captured image. For example, the target application scenario can be a traffic scene, in which images containing texts such as traffic signs and indication plates can be obtained, or a commercial scene, in which images containing texts such as store signs and product labels can be obtained.

[0044] During the text detection and text recognition process, since the text features and environmental conditions vary in different scenes, that is, factors such as font, size, color, distribution density, illumination, and background complexity may all affect the accuracy of text detection and text recognition. By analyzing the target application scenario, more appropriate initial text detection and initial text recognition models can be selected more precisely to cope with these differences and improve the accuracy of detection and recognition.

[0045] S102. Obtain an initial text detection model from the first pre-trained model library according to the initial image data and the target application scenario, and obtain an initial text recognition model from the second pre-trained model library according to the initial image data and the target application scenario.

[0046] In the embodiments of the present application, the first pre-trained model library refers to a library storing multiple pre-trained models for text detection, including models such as PMTD, CRAFT, and EAST. Among them, PMTD (Progressive Multi-text Detection) adopts a progressive detection strategy and can well detect multi-scale, multi-directional, and complexly distributed texts; CRAFT (Character Region Awareness For Text detection) is based on character region awareness and can well detect irregular text regions; EAST (Efficient and Accurate Scene Text detection) is an end-to-end model with high detection efficiency, accuracy, and speed.

[0047] In the embodiments of the present application, the second pre-trained model library refers to a library that stores multiple pre-trained models for text recognition, including models such as CLOVA, ASTER, and CRNN. Among them, CLOVA (Closed-LOop Visual Assistant) is based on the Transformer architecture and is usually used to recognize long texts and complex semantic texts; ASTER (An Attentional SceneText Recognizer with Flexible Rectification) combines the attention mechanism and deformable convolution and is often used to correct and recognize irregular texts; CRNN (Convolutional Recurrent Neural Network) combines convolutional neural networks and recurrent neural networks and is widely used for handwritten and printed text recognition.

[0048] In specific embodiments, by analyzing the initial image data and the target application scenario, the specific characteristics of the text in the initial image data can be determined, and since there are differences in the adaptability of different pre-trained models to the scenario characteristics, the scenario characteristics of the target application scenario can be determined. Thus, the most suitable initial text detection model can be selected from the first pre-trained model library according to the specific characteristics of the text and the scenario characteristics, and the most suitable initial text recognition model can be selected from the second pre-trained model library. According to the initial image data and the target application scenario, the selected initial model can better adapt to the actual task requirements, avoiding the problem of poor performance of using a general model in a specific scenario. At the same time, the initial model can be trained and optimized to improve the accuracy, efficiency, and robustness of text detection and recognition.

[0049] Optionally, step S102 above, obtaining the initial text detection model from the first pre-trained model library according to the initial image data and the target application scenario, may specifically include the following steps:

[0050] A201. Determine the first characteristics corresponding to the initial image data; the first characteristics include at least one of the following: text distribution density, text type, text direction, text occlusion degree, image resolution;

[0051] A202. Determine the second characteristics corresponding to the target application scenario; the second characteristics include at least one of the following: illumination environment, scenario domain;

[0052] A203. Analyze each first pre-trained model in the first pre-trained model library to determine the fitness of each first pre-trained model in the first pre-trained model library corresponding to the first characteristics and the second characteristics, and obtain multiple fitness values; the fitness values represent the degree of adaptation of the first pre-trained model to the initial image data and the target application scenario.

[0053] A204. Select the maximum value from the multiple fitness values, and determine the first pre-trained model corresponding to the maximum value as the initial text detection model.

[0054] In a specific embodiment, the first characteristics corresponding to the initial image data can be determined, where the first characteristics include at least one of the following: text distribution density, text type, text direction, text occlusion degree, image resolution, etc. The text distribution density refers to the proportion of the text area in the image to the entire image area and the spacing between texts. The text type refers to types such as printed text, handwritten text, and artistic text. The text direction refers to various directions such as horizontal, vertical, and inclined. The text occlusion degree refers to the situation where the text may be partially occluded by other objects, and the occlusion degree will affect the complete detection of the text by the model. The image resolution refers to the number of pixels in the image, and the resolution determines the detail level of the image. A high-resolution image can provide clearer text and object details.

[0055] Next, determine the second characteristics corresponding to the target application scenario, where the second characteristics include at least one of the following: lighting environment, scene field, etc. The lighting environment refers to environments such as strong light, weak light, backlight, and uneven lighting. Different lighting environments will make the text in the image present different visual effects. For example, there may be reflections under strong light, and the text may be blurred under weak light. The scene field refers to different fields such as transportation, commerce, education, and medical care. Different scene fields have texts with different professional characteristics.

[0056] Analyze each first pre-trained model in the first pre-trained model library, and the fitness corresponding to each first pre-trained model in the first pre-trained model library for the first characteristics and the second characteristics can be determined, obtaining multiple fitness values, where the fitness value represents the degree of adaptation of the first pre-trained model to the initial image data and the target application scenario. Specifically, test data sets for testing the performance of the first pre-trained model can be obtained. These data sets simulate the first characteristics of the initial image data and the second characteristics of the target application scenario, and use the performance metrics of each first pre-trained model on these test data sets as a reference for the fitness value. Among them, the performance metrics can be metrics such as accuracy, recall rate, and F1 value. The better the performance metric value evaluated by the first pre-trained model, the higher the fitness value of the first pre-trained model.

[0057] Select the maximum value from the multiple fitness values, and determine the first pre-trained model corresponding to the maximum value as the initial text detection model. This initial text detection model is the text detection model in the first pre-trained model library that is most suitable for the current initial image data and the target application scenario.

[0058] In the embodiments of the present application, the initial text recognition model is also obtained from the second pre-trained model library according to the initial image data and the target application scenario. When obtaining the initial text recognition model, first, determine the first characteristics of the initial image data, such as the text distribution density, text type, text orientation, text occlusion degree, and image resolution, and at the same time determine the second characteristics of the target application scenario, such as the lighting environment and scenario field corresponding to the target application scenario. Then, for each second pre-trained model in the second pre-trained model library, analyze its adaptation degree to the first characteristics of the initial image data and the second characteristics of the target application scenario, so as to obtain a plurality of fitness values, and these fitness values reflect the adaptation of each model to the current initial image data and the target application scenario. Finally, select the maximum value from these fitness values, and determine the second pre-trained model corresponding to the maximum value as the initial text recognition model. Through this series of steps similar to those for obtaining the initial text detection model, it can be ensured that the selected initial text recognition model is more suitable for accurately recognizing the text in the current initial image data and meets the actual requirements of the target application scenario.

[0059] S103. Train the initial text detection model according to a preset first training data set to obtain a target text detection model, and train the initial text recognition model according to a preset second training data set to obtain a target text recognition model.

[0060] In the embodiments of the present application, the preset first training data set refers to an image data set used to train the initial text detection model, which may include various standard data sets, such as ICDAR series data sets, COCO-Text data sets, SynthText data sets, etc. The preset first training data set contains images in various natural scenarios, that is, images with different text distribution densities, text types, text orientations, text occlusion degrees, and different image resolutions. At the same time, each image in the preset first training data set accurately annotates the text area therein, marking information such as the specific position and range of the text, so that the initial text detection model can learn how to accurately detect these text areas during the training process, improving the detection ability and accuracy of the model.

[0061] In the embodiments of the present application, the preset second training dataset refers to a dataset of images and corresponding text annotations prepared for training the initial text recognition model. It includes diverse natural scene images, and the text therein has rich variations, including different font styles, language types, vocabulary contents, etc. Each image in this dataset not only has the image itself but also accurately annotates the text region and specific content of the text therein. By allowing the initial text recognition model to learn on this dataset, the model can learn the characteristics of characters, the relationships between characters, semantic information, etc., thereby improving the recognition ability of various texts, being able to accurately recognize and output the text content in the input image, and enabling the trained target text recognition model to have higher recognition accuracy and robustness in practical applications.

[0062] In a specific embodiment, the initial text detection model will learn the training data in the preset first training dataset, continuously adjust the parameters of the model itself. By repeatedly processing the images in the preset first training dataset, the model can learn the feature patterns of text in various different situations, thereby improving the detection ability of text regions, and finally obtaining the target text detection model. The initial text recognition model will learn the training data in the preset second training dataset. By analyzing the features of the image and the corresponding text content, it learns information such as the shape, structure, and semantics of characters, as well as the relationship between the features of the image and the characters. Finally, it can accurately recognize the text content according to the image features, and thus can obtain the target text recognition model.

[0063] Optionally, the above step S103, training the initial text detection model according to the preset first training dataset to obtain the target text detection model, may specifically include the following steps:

[0064] A301. Perform data augmentation processing on the preset first training dataset to obtain the target first training dataset; the data augmentation processing includes at least one of the following: cropping, rotation, scaling, flipping, adding noise;

[0065] A302. Divide the target first training dataset based on a preset division ratio to obtain a first training set and a first validation set;

[0066] A303. Initialize the parameters of the initial text detection model to obtain the first text detection model; the parameter initialization includes: weight initialization, bias initialization;

[0067] A304. Train the first text detection model according to the first training set to obtain the second text detection model;

[0068] A305. Verify the second text detection model according to the first verification set, determine the evaluation result of the second text detection model on a preset evaluation metric, and obtain a target evaluation result; the preset evaluation metric includes at least one of the following: detection recall rate, detection precision, detection speed, and H-mean value;

[0069] A306. If the target evaluation result meets a preset performance condition, use the second text detection model as the target text detection model;

[0070] A307. Otherwise, determine model adjustment parameters;

[0071] A308. Adjust the second text detection model according to the model adjustment parameters to obtain a third text detection model;

[0072] A309. Use the third text detection model as the target text detection model.

[0073] The preset division ratio refers to the proportional relationship that is preset before dividing the target first training data set into a first training set and a first verification set. The setting of this ratio needs to consider various factors, such as the size of the data set, the complexity of the model, etc.

[0074] The preset performance condition refers to a set of criteria that are determined in advance before evaluating the second text detection model and are used to measure whether the model performance meets the standard. These criteria are set according to the performance requirements of the text detection model in the actual application scenario and are usually determined based on the preset evaluation metrics. For example, when the preset evaluation metrics include detection recall rate, detection precision, detection speed, and H-mean value, the preset performance condition can stipulate that the detection recall rate should be not less than 80%, and the detection precision should reach more than 85%. Only when the target evaluation result of the second text detection model on these preset evaluation metrics meets the corresponding preset performance condition, is the model performance considered qualified and can be used as the target text detection model; if not, the model needs to be adjusted and optimized.

[0075] In specific embodiments, by performing data augmentation on a preset first training dataset, the diversity of the dataset can be expanded, enabling the model to learn text features in more different forms, so as to improve the adaptability and robustness of the model to text images in different scenarios. Among them, the data augmentation process may include at least one of the following: cropping, rotation, scaling, flipping, adding noise, etc. It should be noted that for the cropping operation, a part of the image can be randomly selected to enable the model to adapt to texts in different positions; rotation can simulate the situation of texts at different angles in the image; scaling changes the size of the text to train the model's detection ability for texts of different scales; flipping includes horizontal flipping and vertical flipping to increase the variation of the data; adding noise simulates interference factors that may occur in the actual scenario, such as noise during shooting. Through these operations, a target first training dataset can be obtained, providing richer data for subsequent training.

[0076] Next, based on a preset division ratio, the target first training dataset is divided to obtain a first training set and a first validation set. Among them, the first training set is used for training the model to enable the model to learn the patterns and rules of text detection, and the first validation set is used to evaluate the performance of the model during the training process to prevent the model from overfitting. Then, the parameters of the initial text detection model are initialized. This is mainly to initialize the weights and biases in the model. The weights determine the connection strength between each neuron in the model, and the biases are used to adjust the activation threshold of the neurons. By reasonably initializing the weights and biases, the model can be in a better state at the beginning of training, which helps to accelerate the training convergence speed and improve the training effect. After the parameters of the initial text detection model are initialized, a first text detection model can be obtained.

[0077] The first text detection model is trained using the first training set to obtain a second text detection model. The first training set includes image data and the corresponding real annotation information of the image. This real annotation information is the position of the text area. During training, the first text detection model continuously adjusts its own parameters according to the input image data and real annotation information through the backpropagation algorithm, so that the text area position predicted by the model gets closer and closer to the real annotation information, and thus a second text detection model is obtained.

[0078] The second text detection model is verified using the first validation set to determine the evaluation result of the second text detection model on preset evaluation metrics. Among them, the preset evaluation metrics include at least one of the following: detection recall rate, detection precision, detection speed, H-mean value, etc. The detection recall rate is used to measure the ability of the model to correctly detect all text areas, the detection precision is used to measure the accuracy of the model's detection results, the detection speed is used to reflect how fast the model processes images, and the H-mean value can comprehensively consider the metrics of recall rate and precision. By evaluating the performance of the model through the preset evaluation metrics, a target evaluation result can be obtained.

[0079] Compare the target evaluation result with the preset performance conditions. If the performance of the second text detection model on the preset evaluation metrics meets the preset performance requirements, for example, the detection accuracy reaches the preset detection accuracy threshold set in advance, and the recall rate is not lower than the preset recall rate threshold, etc., then it can be considered that the second text detection model has achieved good performance, and the second text detection model can be used as the target text detection model for actual text detection tasks. Otherwise, if the target evaluation result does not meet the preset performance conditions, it indicates that there may be some problems with the second text detection model and it needs to be adjusted. Therefore, it is necessary to determine the model adjustment parameters and adjust the second text detection model according to the determined model adjustment parameters. The adjustment methods can include modifying the model structure (such as increasing or decreasing the number of layers, adjusting the number of neurons, etc.), changing the training parameters (such as the learning rate, the number of iterations, etc.), or adjusting the weights and biases of the model. After adjustment, the third text detection model is obtained. Using the adjusted third text detection model as the target text detection model can be used for subsequent text detection tasks.

[0080] Optionally, the above step A307, otherwise, determining the model adjustment parameters, specifically may include the following steps:

[0081] B301. Determine the test data set; the test data set includes: text image data with different word lengths and different aspect ratios, and the true labels of each text image data;

[0082] B302. Input the test data set into the second text detection model to obtain the model prediction result and the attention weight map; the attention weight map characterizes the degree of attention of the second text detection model to different regions of the image during the processing process;

[0083] B303. Construct a confusion matrix according to the model prediction result and the true label; the rows of the confusion matrix represent the true classes, and the columns represent the predicted classes;

[0084] B304. Determine the model error type according to the confusion matrix; the model error type includes at least one of the following: false detection, missed detection;

[0085] B305. Determine the model error reason according to the model error type and the attention weight map;

[0086] B306. Determine the model parameters to be adjusted and the adjustment coefficients corresponding to the model parameters to be adjusted according to the model error type and the model error reason;

[0087] B307. Determine the model adjustment parameters according to the model parameters to be adjusted and the adjustment coefficients.

[0088] In a specific embodiment, to accurately analyze the performance of the second text detection model and find the direction for adjustment, a test dataset can be used, which includes: text image data with different word lengths and aspect ratios, and the ground truth labels for each text image data.

[0089] Input the test dataset into the second text detection model to obtain the model prediction results. An attention mechanism can be introduced into the second text detection model. When obtaining the model prediction results, an attention weight map can also be obtained. Among them, the attention weight map characterizes the degree of attention of the second text detection model to different regions of the image during the processing. Through the attention weight map, the key attention regions of the model when detecting text can be understood.

[0090] Construct a confusion matrix based on the prediction results of the second text detection model and the ground truth labels in the test dataset. Among them, the confusion matrix is a two-dimensional table, where the rows represent the true classes, that is, the actual text region situations, such as the existence or non-existence of text, and the columns represent the predicted classes, that is, the text region situations predicted by the model. Through the confusion matrix, the classification situation of the model in different classes can be intuitively seen. For example, how many regions where text actually exists are correctly detected, and how many regions where text does not exist are misjudged as having text, etc.

[0091] Based on the confusion matrix, the model error types can be determined. Among them, the model error types include at least one of the following: false detection and missed detection. False detection means that the model detects a region where there is actually no text as having text, while missed detection means that the model fails to detect a region where text actually exists. Then, based on the model error types and the attention weight map, the reasons for the model errors can be determined. For example, if there is a false detection situation, by observing the attention weight map, it can be found that it may be that the model gives too much attention to some non-text regions (such as patterns and noises in the background), resulting in misidentifying them as text; if it is a missed detection problem, it may be that the attention weight map shows that the model pays insufficient attention to some text regions, or the features of these text regions are quite different from the features learned by the model, making the model fail to correctly detect them.

[0092] Based on the model error types and the reasons for the errors, determine the model parameters to be adjusted and the corresponding adjustment coefficients for these parameters. For example, since the weight setting of a certain convolutional layer in the model is unreasonable, resulting in being overly sensitive to non-text features and causing false detection problems in the model, then the weight parameters of the convolutional layer can be adjusted, and the adjustment coefficient can be determined. This adjustment coefficient is used to control the amplitude of parameter adjustment. Based on the model parameters to be adjusted and the adjustment coefficients, the model adjustment parameters can be determined, which are used to adjust the second text detection model to improve the performance of the model, reduce errors such as false detection and missed detection, and make it more accurate in detecting text regions.

[0093] Optionally, the above step B304, determining the model error type according to the confusion matrix, may specifically include the following steps:

[0094] C301. Determine the target true positive, target false positive, and target false negative according to the confusion matrix; the target true positive represents the number of text regions detected by the model; the target false positive represents the number of non-text regions detected as text regions by the model; the target false negative represents the number of text regions not detected by the model;

[0095] C302. Determine the accuracy of the model detection according to the target true positive and the target false positive, and obtain the target accuracy value;

[0096] C303. Determine the recall rate of the model detection according to the target true positive and the target false negative, and obtain the target recall rate;

[0097] C304. When the target accuracy value is lower than the preset accuracy threshold, determine that the model error type is misdetection;

[0098] C305. When the target recall rate is lower than the preset recall rate threshold, determine that the model error type is missed detection.

[0099] In a specific embodiment, the target true positive, target false positive, and target false negative can be determined according to the confusion matrix. Among them, the target true positive represents the number of text regions detected by the model, the target false positive represents the number of non-text regions detected as text regions by the model, and the target false negative represents the number of text regions not detected by the model.

[0100] Determine the accuracy of the model detection according to the target true positive and the target false positive, and obtain the target accuracy value. The target accuracy value represents the proportion of text regions that are truly text among the text regions detected by the model, and it can be obtained by dividing the number of target true positives by the sum of the number of target true positives and the number of target false positives. Determine the recall rate of the model detection according to the target true positive and the target false negative, and obtain the target recall rate. The target recall rate represents the ability of the model to detect actual text regions, and it can be obtained by dividing the number of target true positives by the sum of the number of target true positives and the number of target false negatives.

[0101] The preset accuracy threshold is an accuracy standard set in advance based on actual needs and application scenarios. When the calculated target accuracy value is lower than the preset accuracy threshold, it means that when predicting text areas, the model misjudges more non-text areas as text areas. At this time, it can be determined that the model's error type is false detection, that is, the model is insufficient in distinguishing text areas from non-text areas, and the model needs to be adjusted to improve its accuracy. Similarly, the preset recall rate threshold is also a pre-set recall rate standard. When the target recall rate is lower than the preset recall rate threshold, it means that the model has failed to detect more actual text areas, that is, there are missed detections. Therefore, the model's ability to capture text areas needs to be improved, and the model needs to be optimized accordingly to improve its recall rate.

[0102] S104: Preprocess the initial image data to obtain target image data.

[0103] In an embodiment of the present application, preprocessing includes at least one of the following: grayscale, binarization, noise removal, etc., and can also unify the image size and data distribution to improve image quality, as well as normalization, cropping, rotation, flipping, adding noise and other processing to simulate complex scenes, improve the accuracy of text detection and recognition, and enhance the model's adaptability and robustness to text images in different scenarios.

[0104] In a specific embodiment, the initial image data is usually a raw image obtained from various channels, which may contain various interference factors and characteristics that are not conducive to subsequent processing. Therefore, it is necessary to improve the image quality through preprocessing to improve the accuracy and efficiency of subsequent text detection and recognition. By performing preprocessing operations such as grayscale, binarization, and noise removal, the initial image data is optimized to obtain target image data, and the optimized target image data is more suitable for subsequent text detection and recognition processing.

[0105] S105: Input the target image data into the target text detection model to obtain a target text area.

[0106] In a specific embodiment, after completing the preprocessing of the initial image data and obtaining the target image data, the target image data can be input into a target text detection model. The target text detection model is a text detection model obtained after training and optimization based on a preset first training data set, which can be used for text detection and locating text areas in images.

[0107] When inputting the target image data into the target text detection model, the model will perform a series of processing and analysis on the image, including extracting various feature information from the image, such as edge information, texture information, and color information. It should be noted that even if the target image data undergoes processing such as grayscaling, the change in grayscale values can also reflect certain features, etc. Then, based on these extracted features, combined with the text patterns and rules learned during the training process, the model can determine which regions in the image may be text regions and finally determine the target text region.

[0108] Among them, the target text region refers to the part of the image that contains text, and this part can be the image part including single characters, words, sentences, paragraphs. The target text region is further input into the target text recognition model to identify the specific text content therein. In this way, the preliminary text detection of the target text detection model can improve the recognition accuracy and efficiency of the text.

[0109] Optionally, the above step S105 of inputting the target image data into the target text detection model to obtain the target text region may specifically include the following steps:

[0110] A501: Extract features from the target image data through a preset rotation-invariant feature extraction method to obtain a target feature vector;

[0111] A502: Extract features from the target image data through a preset rotation convolution filter to obtain a target feature map;

[0112] A503: Perform channel splicing and fusion on the target feature vector and the target feature map to obtain a target image feature;

[0113] A504: Detect the corner points of the text region according to the target image feature through a preset detection algorithm to obtain m text region corner points; m is an integer greater than 0;

[0114] A505: Group the m text region corner points to obtain n text region corner point groups; n is a positive integer less than m; each text region corner point group includes at least one text region corner point;

[0115] A506: Determine the target text region according to the n text region corner point groups.

[0116] Among them, the preset rotation-invariant feature extraction method refers to a class of methods that can still stably and accurately extract representative features when the image rotates. Using this type of method can effectively solve the problem of feature changes caused by image rotation and ensure that consistent feature information can be obtained at different rotation angles. Common preset rotation-invariant feature extraction methods include Scale-Invariant Feature Transform (SIFT) and Speeded-Up Robust Features (SURF), etc.

[0117] Among them, the preset rotation convolution filter refers to a convolution filter used to process rotated images. Its core feature is that when performing convolution operations on the image, it can adapt to the rotation changes of the image and capture feature information at different rotation angles. For example, the rotatable convolution kernel is a typical preset rotation convolution filter. It introduces additional rotation parameters, enabling the convolution kernel to perform convolution operations at different rotation angles. During the calculation process, the convolution kernel dynamically adjusts its own direction according to the rotation of the image, so as to better extract rotation-invariant features.

[0118] Among them, the preset detection algorithms include a variety of algorithms for detecting corner points of text regions. Common ones are Harris corner detection algorithm, Shi-Tomasi corner detection algorithm, etc.

[0119] In a specific embodiment, when processing target image data, first use the preset rotation-invariant feature extraction method to extract features from the target image data to obtain a target feature vector. The target feature vector represents the digital description of the image and includes key information such as the content and structure of the image.

[0120] Then, use the preset rotation convolution filter to extract features from the target image data, and then obtain a target feature map. Among them, the target feature map is a two-dimensional matrix, and the value of each element in this two-dimensional matrix represents the feature response intensity of the image at the corresponding position. It shows the feature distribution of different regions in the image and provides spatial feature information for detecting text regions.

[0121] Further, channel concatenation fusion is performed on the target feature vector and the target feature map to obtain the target image feature. Channel concatenation fusion refers to integrating different types of feature information contained in the target feature vector and the target feature map for better subsequent text region detection. Specifically, it can ensure the consistency of the target feature vector and the target feature map in the spatial dimension. For example, the sizes in height and width need to be the same. If the sizes are inconsistent, operations such as interpolation or cropping can be performed on the target feature vector or the target feature map. After the spatial dimension matching is completed, for example, the target feature vector is expanded into a tensor form with the same spatial size as the target feature map, enabling it to be concatenated with the target feature map in the channel dimension. In this way, the expanded target feature vector and the target feature map can be connected along the channel dimension to form a new feature tensor. This new feature tensor contains the global feature information represented by the target feature vector and the local spatial feature information contained in the target feature map. At this time, the channel concatenation fusion is completed, and the target image feature is obtained. This image feature is more comprehensive and rich, which helps to detect the text region more accurately.

[0122] The corner points of the text region are detected based on the target image feature by a preset detection algorithm. Corner points are points with obvious feature changes in the image. In the text region, corner points are usually located at the edges, corners, etc. of the text. By detecting the corner points, the approximate position of the text region can be initially determined. After detection, m text region corner points are obtained, where m is an integer greater than 0. Since in the actual image, the text may be composed of multiple independent text blocks, and each text block has its own independent set of corner points, after obtaining m text region corner points, grouping is required. The corner points belonging to the same text block are grouped together, and finally n text region corner point groups are obtained. Specifically, the spatial distance, angular relationship, etc. between the corner points can be used for grouping. For example, if two corner points are close and the direction of the line connecting them is consistent with the arrangement direction of the text, then they can be grouped into the same group.

[0123] The target text region can be determined based on the n text region corner point groups. For each text region corner point group, by connecting the corner points within the group, the approximate contour of the text block can be outlined, and the region enclosed by this contour is the target text region. In this way, the region where the text is located is accurately detected from the target image data, providing a basis for subsequent text recognition tasks.

[0124] Optionally, step S105 above, inputting the target image data into the target text detection model to obtain the target text region, may specifically include the following steps:

[0125] B501. Perform pixel segmentation on the target image data to obtain multiple pixel data;

[0126] B502. Generate a text score map based on the multiple pixel data; each pixel in the text score map corresponds to a value, and this value represents the probability that the pixel belongs to text.

[0127] B503. Compare each pixel value in the text score map with a preset pixel threshold to obtain a target comparison result.

[0128] B504. Mark the pixels corresponding to the target comparison result that are higher than the preset pixel threshold as text blocks to obtain multiple text blocks.

[0129] B505. Perform connected component analysis on the multiple text blocks to obtain multiple initial text regions; the initial text region includes at least one text block.

[0130] B506. Determine the target text region according to the multiple initial text regions through a preset regression method.

[0131] Among them, the preset pixel threshold refers to a critical value preset during the text detection process for determining whether a pixel belongs to a text region. In the generated text score map, each pixel has a value representing the probability that it belongs to text. By comparing this value with the preset pixel threshold, pixels that may belong to text and pixels that belong to the background can be initially distinguished. This threshold is set by comprehensively considering various factors, such as the characteristics of the image, the features of the text, and the performance of the target text detection model.

[0132] Among them, the preset regression method refers to a class of methods used to optimize and adjust the initial text region to determine the final target text region. The preset regression method may include: linear regression, logistic regression, support vector regression, etc.

[0133] In a specific embodiment, pixel segmentation of the target image data can obtain multiple pixel data. An image is essentially composed of individual pixel points, and pixel segmentation is to split the entire image into multiple independent pixel data. Based on the multiple pixel data obtained by segmentation, a text score map is generated. Among them, each pixel in the text score map corresponds to a specific value, and this value represents the probability that the pixel belongs to text. This probability can be obtained by analyzing the features of each pixel through the target text detection model. For example, by analyzing factors such as the color, brightness, and distribution of surrounding pixels of the pixel, it is determined whether the pixel belongs to text or non - text. If the features of a certain pixel are highly similar to the text features learned by the model, then its corresponding score will be higher, that is, the probability of belonging to text is greater; otherwise, the score is lower.

[0134] Please refer to Figure 2 , Figure 2 which is a schematic diagram of a scenario of image pixel segmentation provided by an embodiment of the present application. AsFigure 2 As shown, given an image filled with a fine grid, each small square represents a pixel. There is black text "Turn left here" in the grid. During the pixel segmentation process, the pixels of these text parts and the pixels of the background part will be distinguished. For each pixel grid, it corresponds to a value, which represents the probability that the pixel belongs to the text. For example, the 30% marked next to the area of the text "Turn left here" indicates that the probability that the pixels in this area belong to the text is 30%, while the 0% marked next to the background area indicates that the probability that the pixels in the background area belong to the text is 0.

[0135] Compare each pixel value in the text score map with a preset pixel threshold to obtain a target comparison result, which indicates the relationship between the score of each pixel and the threshold. According to the target comparison result, mark the pixels with scores higher than the preset pixel threshold as text blocks, and obtain multiple text blocks. After obtaining multiple text blocks, perform connected component analysis on the multiple text blocks to obtain multiple initial text regions. A connected component refers to a region composed of connected pixels. For example, a complete text region is usually composed of multiple connected text blocks. Through connected component analysis, the connected text blocks are grouped into an initial text region, and each initial text region contains at least one text block, thus initially dividing the text in the image into multiple relatively independent regions.

[0136] The target text region can be determined based on multiple initial text regions through a preset regression method. The position, shape, and other information of the initial text region are further optimized and adjusted through the preset regression method. For example, the precise boundary of the text region can be predicted through a regression model, some misjudged regions can be removed, and adjacent regions belonging to the same text can be merged, etc., so as to obtain an accurate target text region, providing a basis for subsequent text recognition tasks.

[0137] S106. Input the target text region into the target text recognition model to obtain a first text.

[0138] In the embodiment of the present application, the target text recognition model is obtained by training and optimizing the initial text recognition model using a preset second training dataset, and it can convert the text information in the image into editable text. When the target text region is input into the target text recognition model, the model can extract features from the input target text region image, extract feature information such as the shape, structure, and strokes of the characters from the image, and perform sequence modeling on the extracted feature information to form an ordered sequence to capture the order and context relationship between the characters. According to the features and sequence information learned by the model, predict the most likely character sequence to obtain the recognition result.

[0139] The target text recognition model outputs a first text, which is the recognition result of the actual text content in the target text area by the target text recognition model. This target recognition result can be used in applications such as information extraction, document processing, and image content understanding.

[0140] Optionally, step S106 above, inputting the target text area into the target text recognition model to obtain a first text, may specifically include the following steps:

[0141] A601. Locate the target text area, determine the rotation deviation between the target text area and a preset horizontal position, and obtain a target rotation deviation;

[0142] A602. Determine a target adjustment parameter according to the target rotation deviation;

[0143] A603. Rotate and correct and normalize the size of the target text area according to the target adjustment parameter to obtain a standard text area;

[0144] A604. Extract the character feature information of the standard text area and determine the first text according to the character feature information.

[0145] Among them, the preset horizontal position refers to a preset reference position, which represents the position direction that the text should be in when it is horizontal and without tilt, and it can ensure that different target text areas can be accurately corrected to the same horizontal direction.

[0146] In a specific embodiment, the target text area is located, the rotation deviation between the target text area and the preset horizontal position is determined to obtain a target rotation deviation, and a target adjustment parameter is determined according to the target rotation deviation. The target adjustment parameter is used for subsequent rotation correction operations on the target text area, and it is directly related to the target rotation deviation and can be information such as the rotation angle and direction.

[0147] The target text area is rotationally corrected according to the target adjustment parameter to restore it to a horizontal or preset standard direction. At the same time, the size of the target text area can also be normalized to adjust the target text area to a unified size. For example, the target text area is scaled or cropped to a fixed height and width. After the target text area is rotationally corrected and its size is normalized, a standard text area can be obtained, and this standard text area is more suitable for text recognition.

[0148] Extract the character feature information from the standard text area, extract the features such as the shape, structure, and strokes of the characters from the standard text area. According to the extracted character feature information, and based on the character feature information, the character sequence can be determined, and finally the first text is obtained. For example, the target text recognition model can match the character features with the character templates learned during the training process, so as to recognize each character and combine them into a complete text content in sequence.

[0149] Please refer to Figure 3 , Figure 3 which is a schematic diagram of the scenario of a natural scene text detection and recognition method provided by an embodiment of the present application. As Figure 3 shown, an application interface for natural scene text detection and recognition is provided. The application interface adopts a simple and intuitive design style, which is convenient for users to operate. It is mainly divided into a menu bar, an image display area, a recognition area, and other auxiliary function areas, providing users with a convenient natural scene text detection and recognition experience.

[0150] The application interface includes a menu bar. Users can select the image files stored locally through the file function, supporting common image formats such as JPEG, PNG, etc. After clicking, a file selection dialog box will pop up. Users can locate the required file and confirm to load the image into the natural scene text detection and recognition system. Users can adjust some key parameters of the system through the settings function, such as the preset edit distance threshold, the preset pixel threshold, etc. By reasonably setting these parameters, the recognition effect can be optimized according to different image qualities and recognition requirements.

[0151] The application interface also includes an image display area, which occupies a large part of the interface and is used to display the original image opened by the user and the image after text detection marking. After loading the image, users can perform operations such as zooming and panning on the image to view the image details and text areas more clearly. At the same time, when the system completes text detection, it will mark the detected target text area with rectangular frames of different colors on the image, allowing users to intuitively understand the position and scope of the text. For example, the left image frame in the image display area shows an initial image, the content of which is a traffic sign for a left turn, and there are Chinese characters "Turn left here" on the sign. It is an original image to be recognized without processing, and the main body is in an inclined state. The right image frame in the image display area shows the preprocessed image. Compared with the initial image, it becomes clearer after operations such as grayscale conversion, noise reduction, binarization, and rotation correction, improving the image quality to make subsequent text recognition more accurate.

[0152] The application interface also includes an identification area. After the user confirms that the image is loaded correctly, clicking this button will cause the system to start detecting and identifying the text in the image. A progress bar for identification will be displayed below the button, providing real-time feedback on the progress of the identification. After the identification is completed, the optimized second text will be displayed in this area. The user can perform operations such as copying and editing the identification result, facilitating the application of the result to other scenarios. For example, the identification area shows "Identification result: Turn left here", indicating that after performing text recognition on the preprocessed image, the system has successfully extracted the text content on the traffic sign.

[0153] The application interface also includes other auxiliary function areas, which include the image being processed. For example, currently object 1 is being processed, while object 2 is pre-stored as a historical record in other auxiliary function areas. The user can click to select different images or relevant data content to be recognized.

[0154] S107. Optimize the result of the first text to obtain the second text.

[0155] In the embodiments of the present application, after the first text is obtained by identifying the target text area, due to various factors such as image quality, font style, and lighting conditions, the first text may have problems such as spelling mistakes, improper grammar, unclear semantics, and non-standard formats. To improve the text quality, the result of the first text needs to be optimized to obtain the second text.

[0156] In specific embodiments, it is possible to correct recognition errors, remove extra spaces, format the output, etc., to improve the accuracy and usability of the recognition result. It is also possible to use auxiliary information such as language models and dictionaries for error correction and result optimization. Output the final recognition result to a specified format or interface, such as JSON, XML, etc., for subsequent applications.

[0157] Optionally, the above step S107, optimizing the result of the first text to obtain the second text, may specifically include the following steps:

[0158] A701. Obtain the text in the preset dictionary library to get the standard text; the preset dictionary library pre-stores the text corresponding to common words and professional terms.

[0159] A702. Determine the edit distance between the first text and the standard text based on the preset edit distance algorithm to obtain the target edit distance.

[0160] A703. Determine the difference between the target edit distance and the preset edit distance threshold to obtain the target difference.

[0161] A704. Analyze the text attributes of the first text through the preset language model; the text attributes include: grammar structure, vocabulary collocation, and semantic coherence.

[0162] A705. Determine the correction of the first text according to the text attribute and the target difference to obtain a third text;

[0163] A706. Format the third text based on a preset formatting rule to obtain the second text.

[0164] Among them, the preset edit distance algorithm refers to an algorithm used to measure the degree of difference between two strings, which quantifies this difference by calculating the minimum number of operations required to convert one string into another. The preset edit distance algorithm can be the Levenshtein Distance algorithm, etc. Among them, the preset edit distance threshold is a preset value used to determine whether the difference between the first text and the standard text is within an acceptable range. The setting of this threshold can consider factors such as the requirements of the application scenario, the characteristics of the language, and the expected accuracy of text recognition.

[0165] Among them, the preset language model includes models for analyzing text attributes (such as grammatical structure, lexical collocation, and semantic coherence). The preset language model can include N-gram models, long short-term memory networks (LSTM), gated recurrent unit (GRU) models, BERT (Bidirectional Encoder Representations from Transformers) models, etc.

[0166] Among them, the preset formatting rule refers to a set of predefined specifications used to format text to make the text more standardized, unified, and easy to read in terms of format. The rule can include capitalization rules, punctuation usage rules, paragraph format rules, number format rules, etc.

[0167] In a specific embodiment, in order to optimize the first text, first obtain the text from a preset dictionary library to obtain the standard text. The preset dictionary library is a collection that stores in advance the corresponding texts of common words and professional terms. For example, when processing texts in the medical field, the dictionary library will contain various medical professional terms, such as "coronary atherosclerosis", etc.; in daily text processing, there will be common words such as "apple", "car", etc.

[0168] Use the preset edit distance algorithm to calculate the edit distance between the first text and the standard text to obtain the target edit distance, which represents the degree of difference between the first text and the standard text. Compare the calculated target edit distance with the preset edit distance threshold to obtain the target difference. If the target difference is large, it indicates that the first text and the standard text are significantly different and there may be many errors; otherwise, the difference is small and within the acceptable range.

[0169] Analyze the text attributes of the first text through a preset language model. The text attributes include grammatical structure, lexical collocation, semantic coherence, etc. Through the text attributes, it can be checked whether the sentence conforms to the grammatical rules, such as whether the subject, predicate, and object are reasonably collocated, whether the combination of words is natural and appropriate, whether the semantics is smooth, and whether the logic is reasonable, etc.

[0170] Comprehensively consider the text attributes and the target difference, and correct the first text. If the target difference is large and there are problems with the text attributes, such as grammar errors, inappropriate lexical collocations, etc., modify the first text according to the standard text and language rules to obtain the third text.

[0171] Perform formatting processing on the third text based on the preset formatting rules. After formatting, the second text is finally obtained, and the accuracy, standardization, and readability of the second text have been significantly improved.

[0172] In summary, by implementing the embodiments of the present application, initial image data is obtained, and the target application scenario corresponding to the initial image data is determined; an initial text detection model is obtained from the first pre-trained model library according to the initial image data and the target application scenario, and an initial text recognition model is obtained from the second pre-trained model library according to the initial image data and the target application scenario; the initial text detection model is trained according to the preset first training data set to obtain a target text detection model, and the initial text recognition model is trained according to the preset second training data set to obtain a target text recognition model; the initial image data is preprocessed to obtain target image data; the target image data is input into the target text detection model to obtain a target text region; the target text region is input into the target text recognition model to obtain the first text; the result of the first text is optimized to obtain the second text. It can be seen that by considering the characteristics of the initial image data and the target application scenario to select and train the model, the target text detection model and the target text recognition model are more in line with the actual application requirements, effectively improving the accuracy of text detection and recognition.

[0173] Please refer to Figure 4 , Figure 4 which is the system architecture diagram of a natural scene text detection and recognition system provided by the embodiments of the present application. The natural scene text detection and recognition system includes: a data input unit, a text detection unit, a text recognition unit, and a data output unit.

[0174] Among them, the data input unit can obtain image data in a natural scene from various channels, such as the pictures taken in real time by a camera, stored image files, etc. These images contain various complex scenes, such as billboards on the street, text pages in books, texts displayed on the screen of electronic products, etc. The data input unit can perform a series of preprocessing operations on the collected original image data, such as grayscale conversion, binarization, noise removal, etc., to improve the accuracy and efficiency of subsequent processing. After preprocessing, target image data is obtained for subsequent text detection and recognition.

[0175] The text detection unit is used to accurately locate the area where the text is located from the preprocessed target image data. The text recognition unit is used to convert the image information in the target text area determined by the text detection unit into editable text content to obtain the first text.

[0176] The data output unit can perform multi-faceted optimization processing on the first text to improve the quality and accuracy of the text, and then output the optimized second text in a suitable way, such as displaying it on the screen, saving it as a text file, transmitting it to other systems for further processing, etc., to meet the requirements of different application scenarios.

[0177] Please refer to Figure 5 , Figure 5 which is a server architecture diagram provided by an embodiment of the present application. As Figure 5 shown, this architecture includes a terminal device 501 and a remote server 502. A direct connection can be established between the terminal device 501 and the remote server 502 through wired communication, or an indirect communication connection can be established through wireless communication. Among them, the terminal device 501 can be various types of devices, such as a smart phone, a tablet computer, a laptop computer, a desktop computer, etc. The remote server 502 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0178] The terminal device 501 can collect data through various built-in sensors (such as cameras, microphones, etc.). In the natural scene text detection and recognition system, the terminal device 501 can use the camera to capture an image containing text as the image data to be processed. The terminal device 501 provides a user interface for the user to operate application programs through this interface, such as selecting an image to upload, viewing the processing results, etc. The terminal device 501 communicates with the remote server 502 through a network (such as Wi-Fi, mobile data network, etc.), uploads the collected image data to the remote server 502, and receives the processing results returned by the remote server 502.

[0179] The remote server 502 can be used to execute the natural scene text detection and recognition method of this application. For the image data uploaded from the terminal device 501, it can use a pre-trained text detection and recognition model to process, locate the text area in the image and convert it into editable text. After completing the data processing, the remote server 502 returns the recognized text content as the recognition result to the terminal device 501 for the user to view and use.

[0180] Please refer to Figure 6 , Figure 6 FIG. is a schematic structural diagram of a natural scene text detection and recognition device provided by an embodiment of this application. The natural scene text detection and recognition device 600 includes: an image acquisition module 601, a model selection module 602, a model training module 603, an image preprocessing module 604, a text detection module 605, a text recognition module 606, and a result optimization module 607, where,

[0181] The image acquisition module 601 is used to acquire initial image data and determine the target application scenario corresponding to the initial image data;

[0182] The model selection module 602 is used to obtain an initial text detection model from a first pre-trained model library according to the initial image data and the target application scenario, and obtain an initial text recognition model from a second pre-trained model library according to the initial image data and the target application scenario;

[0183] The model training module 603 is used to train the initial text detection model according to a preset first training data set to obtain a target text detection model, and train the initial text recognition model according to a preset second training data set to obtain a target text recognition model;

[0184] The image preprocessing module 604 is used to preprocess the initial image data to obtain target image data;

[0185] The text detection module 605 is configured to input the target image data into the target text detection model to obtain a target text region;

[0186] The text recognition module 606 is configured to input the target text region into the target text recognition model to obtain a first text;

[0187] The result optimization module 607 is configured to optimize the result of the first text to obtain a second text.

[0188] Optionally, in terms of obtaining the initial text detection model from the first pre-trained model library according to the initial image data and the target application scenario, the model selection module 602 is further specifically configured to:

[0189] Determine a first characteristic corresponding to the initial image data; the first characteristic includes at least one of the following: text distribution density, text type, text direction, text occlusion degree, image resolution;

[0190] Determine a second characteristic corresponding to the target application scenario; the second characteristic includes at least one of the following: lighting environment, scene domain;

[0191] Analyze each first pre-trained model in the first pre-trained model library, determine the fitness of each first pre-trained model in the first pre-trained model library corresponding to the first characteristic and the second characteristic, and obtain a plurality of fitness values; the fitness value represents the adaptation degree of the first pre-trained model to the initial image data and the target application scenario;

[0192] Select the maximum value from the plurality of fitness values, and determine the first pre-trained model corresponding to the maximum value as the initial text detection model.

[0193] Optionally, in terms of training the initial text detection model according to a preset first training data set to obtain a target text detection model, the model selection module 602 is further specifically configured to:

[0194] Perform data augmentation processing on the preset first training data set to obtain a target first training data set; the data augmentation processing includes at least one of the following: cropping, rotation, scaling, flipping, adding noise;

[0195] Divide the target first training data set based on a preset division ratio to obtain a first training set and a first validation set;

[0196] Perform parameter initialization on the initial text detection model to obtain a first text detection model; the parameter initialization includes: weight initialization, bias initialization;

[0197] Train the first text detection model according to the first training set to obtain a second text detection model;

[0198] Validate the second text detection model according to the first validation set, determine the evaluation result of the second text detection model on a preset evaluation metric to obtain a target evaluation result; the preset evaluation metric includes at least one of the following: detection recall rate, detection precision, detection speed, H-mean value;

[0199] If the target evaluation result meets the preset performance condition, use the second text detection model as the target text detection model;

[0200] Otherwise, determine model adjustment parameters;

[0201] Adjust the second text detection model according to the model adjustment parameters to obtain a third text detection model;

[0202] Use the third text detection model as the target text detection model.

[0203] Optionally, in terms of determining the model adjustment parameters, the model selection module 602 is further specifically configured to:

[0204] Determine a test data set; the test data set includes: text image data with different word lengths and different aspect ratios and the ground truth labels of each text image data;

[0205] Input the test data set into the second text detection model to obtain a model prediction result and an attention weight map; the attention weight map characterizes the degree of attention of the second text detection model to different regions of the image during processing;

[0206] Construct a confusion matrix according to the model prediction result and the ground truth label; the rows of the confusion matrix represent the true classes and the columns represent the predicted classes;

[0207] Determine the model error type according to the confusion matrix; the model error type includes at least one of the following: false detection, missed detection;

[0208] Determine the model error reason according to the model error type and the attention weight map;

[0209] Determine the model parameters to be adjusted and the adjustment coefficients corresponding to the model parameters to be adjusted according to the model error type and the model error reason;

[0210] Determine the model adjustment parameters according to the model parameters to be adjusted and the adjustment coefficients.

[0211] Optionally, in terms of determining the model error type according to the confusion matrix, the model selection module 602 is further specifically configured to:

[0212] Determine a target true positive, a target false positive, and a target false negative according to the confusion matrix; the target true positive represents the number of text regions detected by the model; the target false positive represents the number of non-text regions detected as text regions by the model; the target false negative represents the number of text regions not detected by the model;

[0213] Determine the accuracy of the model detection according to the target true positive and the target false positive, and obtain a target accuracy value;

[0214] Determine the recall rate of the model detection according to the target true positive and the target false negative, and obtain a target recall rate;

[0215] When the target accuracy value is lower than a preset accuracy threshold, determine that the model error type is misdetection;

[0216] When the target recall rate is lower than a preset recall threshold, determine that the model error type is missed detection.

[0217] Optionally, in terms of inputting the target image data into the target text detection model to obtain a target text region, the text detection module 605 is further specifically configured to:

[0218] Extract features from the target image data through a preset rotation-invariant feature extraction method to obtain a target feature vector;

[0219] Extract features from the target image data through a preset rotation convolution filter to obtain a target feature map;

[0220] Perform channel splicing and fusion on the target feature vector and the target feature map to obtain a target image feature;

[0221] Detect text region corner points according to the target image feature through a preset detection algorithm to obtain m text region corner points; m is an integer greater than 0;

[0222] Group the m text region corner points to obtain n text region corner point groups; n is a positive integer less than m; the text region corner point group includes at least one text region corner point;

[0223] Determine the target text region according to the n text region corner point groups.

[0224] Optionally, in terms of inputting the target image data into the target text detection model to obtain a target text region, the text detection module 605 is further specifically configured to:

[0225] Perform pixel segmentation on the target image data to obtain a plurality of pixel data;

[0226] Generate a text score map based on the plurality of pixel data; each pixel in the text score map corresponds to a value, and this value represents the probability that the pixel belongs to text;

[0227] Compare each pixel value in the text score map with a preset pixel threshold to obtain a target comparison result;

[0228] Mark the pixels corresponding to those higher than the preset pixel threshold in the target comparison result as text blocks to obtain a plurality of text blocks;

[0229] Perform connected component analysis on the plurality of text blocks to obtain a plurality of initial text regions; the initial text region includes at least one text block;

[0230] Determine the target text region based on the plurality of initial text regions through a preset regression method.

[0231] Optionally, in the aspect of inputting the target text region into the target text recognition model to obtain the first text, the text recognition module 606 is further specifically configured to:

[0232] Locate the target text region, determine the rotation deviation between the target text region and a preset horizontal position to obtain a target rotation deviation;

[0233] Determine a target adjustment parameter based on the target rotation deviation;

[0234] Perform rotation correction and size normalization on the target text region according to the target adjustment parameter to obtain a standard text region;

[0235] Extract the character feature information of the standard text region, and determine the first text according to the character feature information.

[0236] Optionally, in the aspect of optimizing the result of the first text to obtain the second text, the result optimization module 607 is further specifically configured to:

[0237] Obtain the text in a preset dictionary library to obtain a standard text; the preset dictionary library pre-stores texts corresponding to common words and professional terms;

[0238] Determine the edit distance between the first text and the standard text based on a preset edit distance algorithm to obtain a target edit distance;

[0239] Determine the difference between the target edit distance and a preset edit distance threshold to obtain a target difference;

[0240] Analyze the text attributes of the first text through a preset language model; the text attributes include: syntactic structure, lexical collocation, and semantic coherence;

[0241] Determine the correction of the first text according to the text attributes and the target difference to obtain a third text;

[0242] Format the third text based on preset formatting rules to obtain the second text.

[0243] The natural scene text detection and recognition device 600 described in this application can obtain initial image data and determine the target application scenario corresponding to the initial image data; obtain an initial text detection model from a first pre-trained model library according to the initial image data and the target application scenario, and obtain an initial text recognition model from a second pre-trained model library according to the initial image data and the target application scenario; train the initial text detection model according to a preset first training data set to obtain a target text detection model, and train the initial text recognition model according to a preset second training data set to obtain a target text recognition model; preprocess the initial image data to obtain target image data; input the target image data into the target text detection model to obtain a target text region; input the target text region into the target text recognition model to obtain a first text; optimize the result of the first text to obtain a second text. It can be seen that by considering the characteristics of the initial image data and the target application scenario to select and train the model, the target text detection model and the target text recognition model are more in line with the actual application requirements, effectively improving the accuracy of text detection and recognition.

[0244] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of an electronic device provided by an embodiment of this application. The electronic device may include a processor, a memory, a communication interface, and one or more programs. The processor, the memory, and the communication interface may be interconnected through a bus; the above one or more programs are stored in the above memory and are configured to be executed by the above processor; in the embodiment of this application, the above programs include instructions for performing the following steps:

[0245] Obtain initial image data and determine the target application scenario corresponding to the initial image data;

[0246] Obtain an initial text detection model from a first pre-trained model library according to the initial image data and the target application scenario, and obtain an initial text recognition model from a second pre-trained model library according to the initial image data and the target application scenario;

[0247] Train the initial text detection model according to a preset first training data set to obtain a target text detection model, and train the initial text recognition model according to a preset second training data set to obtain a target text recognition model;

[0248] Preprocess the initial image data to obtain target image data;

[0249] Input the target image data into the target text detection model to obtain a target text region;

[0250] Input the target text region into the target text recognition model to obtain a first text;

[0251] Optimize the result of the first text to obtain a second text.

[0252] The electronic device described in this application can obtain initial image data and determine the target application scenario corresponding to the initial image data; obtain an initial text detection model from a first pre-trained model library according to the initial image data and the target application scenario, and obtain an initial text recognition model from a second pre-trained model library according to the initial image data and the target application scenario; train the initial text detection model according to a preset first training data set to obtain a target text detection model, and train the initial text recognition model according to a preset second training data set to obtain a target text recognition model; preprocess the initial image data to obtain target image data; input the target image data into the target text detection model to obtain a target text region; input the target text region into the target text recognition model to obtain a first text; optimize the result of the first text to obtain a second text. It can be seen that by considering the characteristics of the initial image data and the target application scenario to select and train the model, the target text detection model and the target text recognition model are more in line with the actual application requirements, effectively improving the accuracy of text detection and recognition.

[0253] An embodiment of this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program for electronic data exchange, and the computer program enables a computer to execute some or all of the steps of any method described in the above method embodiments. The above computer includes an electronic device.

[0254] An embodiment of this application also provides a computer program product. The computer program product includes a non-transitory computer-readable storage medium storing a computer program. The computer program is operable to enable a computer to execute some or all of the steps of any method described in the above method embodiments. The computer program product can be a software installation package, and the above computer includes an electronic device.

[0255] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by relevant hardware instructed by a computer program. This program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The foregoing storage medium includes: various media such as ROM or random access memory RAM, magnetic disk, or optical disc that can store program codes.

[0256] The steps of the methods or algorithms described in the embodiments of the present application can be implemented in a hardware manner or by a processor executing software instructions. The software instructions can be composed of corresponding software modules. The software modules can be stored in RAM, flash memory, ROM, EPROM, electrically erasable programmable read-only memory (EEPROM), registers, hard disk, removable hard disk, CD-ROM, or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC. In addition, the ASIC can be located in a terminal device or a management device. Of course, the processor and the storage medium can also exist as discrete components in the terminal device or the management device.

[0257] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the embodiments of the present application can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0258] Each device and product described in the above embodiments includes various modules / units, which can be software modules / units, hardware modules / units, or can be partially software modules / units and partially hardware modules / units. For example, for each device and product applied to or integrated into a chip, each module / unit it includes can be implemented in the form of hardware such as circuits. Or, at least some of the modules / units can be implemented in the form of software programs that run on the processor integrated inside the chip, and the remaining (if any) part of the modules / units can be implemented in the form of hardware such as circuits; for each device and product applied to or integrated into a chip module, each module / unit it includes can be implemented in the form of hardware such as circuits. Different modules / units can be located in the same component (such as a chip, a circuit module, etc.) or different components of the chip module. Or, at least some of the modules / units can be implemented in the form of software programs that run on the processor integrated inside the chip module, and the remaining (if any) part of the modules / units can be implemented in the form of hardware such as circuits; for each device and product applied to or integrated into a terminal device, each module / unit it includes can be implemented in the form of hardware such as circuits. Different modules / units can be located in the same component (such as a chip, a circuit module, etc.) or different components inside the terminal device. Or, at least some of the modules / units can be implemented in the form of software programs that run on the processor integrated inside the terminal device, and the remaining (if any) part of the modules / units can be implemented in the form of hardware such as circuits.

[0259] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the embodiments of the present application. It should be understood that the above is only the specific embodiments of the embodiments of the present application and is not used to limit the protection scope of the embodiments of the present application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the embodiments of the present application shall be included in the protection scope of the embodiments of the present application.

Claims

1. A natural scene text detection and recognition method, characterized in that: The method comprises: Acquire initial image data, and determine a target application scenario corresponding to the initial image data; Acquire an initial text detection model from a first pre-trained model library according to the initial image data and the target application scenario, and acquire an initial text recognition model from a second pre-trained model library according to the initial image data and the target application scenario; The initial text detection model is trained according to a preset first training data set to obtain a target text detection model, and the initial text recognition model is trained according to a preset second training data set to obtain a target text recognition model; Preprocessing the initial image data to obtain target image data; Inputting the target image data into the target text detection model to obtain a target text area; Inputting the target text area into the target text recognition model to obtain a first text; The first text is optimized to obtain a second text.

2. The method according to claim 1, characterized in that The obtaining an initial text detection model from a first pre-trained model library according to the initial image data and the target application scenario includes: Determine a first characteristic corresponding to the initial image data; the first characteristic includes at least one of the following: text distribution density, text type, text direction, text occlusion degree, and image resolution; Determine a second characteristic corresponding to the target application scenario; the second characteristic includes at least one of the following: lighting environment, scene field; Analyze each first pre-trained model in the first pre-trained model library, determine the fitness of each first pre-trained model in the first pre-trained model library corresponding to the first characteristic and the second characteristic, and obtain multiple fitnesses; the fitnesses represent the degree of adaptation of the first pre-trained model to the initial image data and the target application scenario; A maximum value is selected from the multiple fitness values, and a first pre-trained model corresponding to the maximum value is determined as the initial text detection model.

3. The method according to claim 1, characterized in that The initial text detection model is trained according to the preset first training data set to obtain a target text detection model, including: Performing data enhancement processing on the preset first training data set to obtain a target first training data set; the data enhancement processing includes at least one of the following: cropping, rotating, scaling, flipping, and adding noise; Dividing the target first training data set based on a preset division ratio to obtain a first training set and a first validation set; Initializing parameters of the initial text detection model to obtain a first text detection model; the parameter initialization includes: weight initialization and bias initialization; Training the first text detection model according to the first training set to obtain a second text detection model; Verifying the second text detection model according to the first verification set, determining an evaluation result of the second text detection model on a preset evaluation indicator, and obtaining a target evaluation result; the preset evaluation indicator includes at least one of the following: detection recall rate, detection accuracy, detection speed, and H-mean value; If the target evaluation result meets the preset performance condition, the second text detection model is used as the target text detection model; Otherwise, determine the model tuning parameters; Adjust the second text detection model according to the model adjustment parameter to obtain a third text detection model; The third text detection model is used as the target text detection model.

4. The method according to claim 3, characterized in that The step of determining the model adjustment parameters comprises: Determine a test data set; the test data set includes: text image data of different word lengths and different aspect ratios and a true label of each text image data; Inputting the test data set into the second text detection model to obtain a model prediction result and an attention weight map; the attention weight map represents the degree of attention paid by the second text detection model to different regions of the image during the processing; Construct a confusion matrix based on the model prediction results and the true labels; the rows of the confusion matrix represent the true categories, and the columns represent the predicted categories; Determine the model error type according to the confusion matrix; the model error type includes at least one of the following: false detection and missed detection; Determine the cause of the model error according to the model error type and the attention weight map; Determine the model parameter to be adjusted and the adjustment coefficient corresponding to the model parameter to be adjusted according to the model error type and the model error cause; The model adjustment parameter is determined according to the model parameter to be adjusted and the adjustment coefficient.

5. The method according to claim 4, characterized in that Determining the model error type according to the confusion matrix includes: Determine the target true positives, target false positives and target false negatives according to the confusion matrix; the target true positives represent the number of text regions detected by the model; the target false positives represent the number of non-text regions detected as text regions by the model; the target false negatives represent the number of text regions that the model fails to detect; Determine the accuracy of model detection according to the target true positive examples and the target false positive examples to obtain a target accuracy value; Determine the recall rate of the model detection according to the target true examples and the target false negative examples to obtain a target recall rate; When the target accuracy value is lower than a preset accuracy threshold, determining the model error type as the false detection; When the target recall rate is lower than a preset recall rate threshold, the model error type is determined to be the missed detection.

6. The method according to any one of claims 1 to 5, characterized in that: The step of inputting the target image data into the target text detection model to obtain the target text area includes: Performing feature extraction on the target image data by using a preset rotation invariant feature extraction method to obtain a target feature vector; Extracting features from the target image data by using a preset rotation convolution filter to obtain a target feature map; Perform channel splicing and fusion on the target feature vector and the target feature map to obtain target image features; Detecting text area corner points according to the target image features through a preset detection algorithm to obtain m text area corner points; m is an integer greater than 0; The m text area corner points are grouped to obtain n text area corner point groups; n is a positive integer less than m; the text area corner point group includes at least one text area corner point; The target text area is determined according to the n text area corner point groups.

7. The method according to any one of claims 1 to 5, characterized in that: The step of inputting the target image data into the target text detection model to obtain the target text area includes: Performing pixel segmentation on the target image data to obtain a plurality of pixel data; Generate a text score map according to the plurality of pixel data; each pixel in the text score map corresponds to a value, and the value represents the probability that the pixel belongs to the text; Compare each pixel value in the text score map with a preset pixel threshold to obtain a target comparison result; Marking pixels corresponding to pixels higher than the preset pixel threshold in the target comparison result as text blocks to obtain multiple text blocks; Performing connected domain analysis on the multiple text blocks to obtain multiple initial text regions; the initial text region includes at least one text block; The target text region is determined according to the multiple initial text regions by using a preset regression method.

8. The method according to claim 1, characterized in that The step of inputting the target text area into the target text recognition model to obtain a first text comprises: Positioning the target text area, determining a rotation deviation between the target text area and a preset horizontal position, and obtaining a target rotation deviation; Determining a target adjustment parameter according to the target rotation deviation; Performing rotation correction and size normalization on the target text area according to the target adjustment parameters to obtain a standard text area; Character feature information of the standard text area is extracted, and the first text is determined according to the character feature information.

9. The method according to claim 1, characterized in that The step of optimizing the result of the first text to obtain the second text includes: Obtaining text from a preset dictionary library to obtain a standard text; the preset dictionary library pre-stores texts corresponding to common words and professional terms; Determine the edit distance between the first text and the standard text based on a preset edit distance algorithm to obtain a target edit distance; Determine the difference between the target edit distance and a preset edit distance threshold to obtain a target difference; Analyzing text attributes of the first text by using a preset language model; the text attributes include: grammatical structure, vocabulary collocation and semantic coherence; Determine to modify the first text according to the text attribute and the target difference value to obtain a third text; The third text is formatted based on a preset formatting rule to obtain the second text.

10. A natural scene text detection and recognition device, characterized in that: The natural scene text detection and recognition device comprises: An image acquisition module, used to acquire initial image data and determine a target application scenario corresponding to the initial image data; A model selection module, configured to obtain an initial text detection model from a first pre-trained model library according to the initial image data and the target application scenario, and to obtain an initial text recognition model from a second pre-trained model library according to the initial image data and the target application scenario; A model training module, used to train the initial text detection model according to a preset first training data set to obtain a target text detection model, and to train the initial text recognition model according to a preset second training data set to obtain a target text recognition model; An image preprocessing module, used for preprocessing the initial image data to obtain target image data; A text detection module, used for inputting the target image data into the target text detection model to obtain a target text area; A text recognition module, used for inputting the target text area into the target text recognition model to obtain a first text; The result optimization module is used to optimize the result of the first text to obtain a second text.

Citation Information

Cited By

  • Hair body generation method and system based on joint diffusion-adversarial training model

    CN121147349A

  • File analysis method, data enhancement method, equipment, medium and product

    CN121351793A

  • File analysis method, data augmentation method, device, medium, and product

    CN121351793B