A photoelectric reconnaissance character recognition method, system, device and medium

Through deep learning algorithms and specific network structures, the character recognition model of photoelectric reconnaissance equipment is optimized, which solves the problem of insufficient recognition accuracy in complex backgrounds and dynamic environments, and achieves efficient and accurate character recognition and real-time processing capabilities.

CN119649380BActive Publication Date: 2025-05-23CHANGCHUN UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411727239.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2025-05-23
Estimated Expiration
2044-11-28

AI Technical Summary

Technical Problem

In existing photoelectric reconnaissance equipment, character recognition methods are insufficient in complex backgrounds and dynamic environments, and are prone to missed recognition, and the computing resources are consumed, which affects real-time.

Method used

Deep learning algorithm is adopted, and the PP-LCNet backbone network and SVTR network are combined with CML strategy and SE module to optimize the robustness and recognition accuracy of the character recognition model, and improve the recognition ability in complex backgrounds and dynamic environments.

Benefits of technology

It improves the character recognition accuracy and real-time processing capabilities of photoelectric reconnaissance equipment in dynamic scenarios, reduces the phenomenon of misidentification and misidentification, and meets the timeliness requirements of reconnaissance tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119649380B_ABST
    Figure CN119649380B_ABST
Patent Text Reader

Abstract

The invention discloses a photoelectric reconnaissance character recognition method, system, device and medium, and relates to the technical field of optical character recognition. The method comprises: collecting photoelectric reconnaissance image data to be recognized; inputting the photoelectric reconnaissance image data to be recognized into a character recognition model, and obtaining a fused feature map through a PP-LCNet backbone network; inputting the fused feature map into two predictors, respectively obtaining a prediction probability map and a prediction threshold map, optimizing the prediction probability map by adopting a mutual learning strategy CML based on knowledge distillation, generating a binary image according to the optimized prediction probability map and the prediction threshold map, and then obtaining an image containing a required text area; inputting the image containing the required text area into a text recognition module, outputting a text label matrix, and obtaining a character recognition result of the photoelectric reconnaissance image data to be recognized according to the text label matrix; the method improves the accuracy of character recognition in complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of optical character recognition, and in particular to a photoelectric reconnaissance character recognition method, system, equipment and medium. Background Art

[0002] Optoelectronic reconnaissance equipment collects large amounts of real-time target video data through visible light and infrared video bands, and uses character recognition to help analyze the battlefield environment, monitor enemy movements, identify targets, etc.

[0003] At present, the OCR character recognition technology of military optoelectronic reconnaissance videos is divided into three categories: OCR based on template matching, OCR based on feature extraction, and OCR based on deep learning; OCR based on template matching recognizes characters by matching characters in the image with predefined templates. This method is more effective for simple fonts, but its recognition accuracy is greatly affected by font style and image noise. In optoelectronic reconnaissance equipment, OCR based on template matching can be used to recognize characters in fixed formats (such as vehicle numbers, equipment identification, etc.), and is suitable for scenes with less battlefield background interference. However, in military environments, characters are usually affected by lighting, angles, and background complexity. The template matching method is difficult to adapt to character shape changes, image rotation, blur, etc., resulting in poor matching results. OCR based on feature extraction recognizes characters by extracting features such as strokes and contours. This method can be implemented through machine learning algorithms such as SVM (support vector machine) and KNN (nearest neighbor algorithm). In optoelectronic reconnaissance equipment, feature extraction ORC can recognize specific logos of different fonts and shapes to assist in positioning and tracking. However, this method is highly dependent on the details of the characters. In dynamic scenes or when the character shape is obscured, the effectiveness of feature extraction will be reduced, and it is easy to miss the recognition. In addition, due to the complexity of the feature extraction process, the large consumption of computing resources, and the slow processing speed in high-resolution video streams, it is not conducive to real-time reconnaissance needs. Based on deep learning, OCR automatically extracts complex features of characters by training a large number of image samples, and has higher recognition accuracy and generalization ability. This method can extract characters through network structures such as CNN and LSTM, and post-process with appropriate language models and attention mechanisms, so as to achieve higher recognition accuracy in complex scenes. In optoelectronic reconnaissance equipment, this method is widely used to identify fast-moving characters in dynamic video data, such as battlefield vehicles, identifiers, numbers on equipment, etc. However, the training and reasoning of deep learning models require a lot of computing resources, especially when processing high-resolution videos in military optoelectronic reconnaissance, which is easy to cause delays and affect real-time performance.

[0004] In short, the current recognition method not only increases the workload, but the existing character recognition method is also sensitive to environmental factors such as lighting changes, background complexity, and weather conditions, resulting in a decrease in the accuracy of character recognition. It is impossible to maintain efficient and accurate recognition performance in a dynamic environment, especially in visible light and infrared video scenes used by optoelectronic reconnaissance equipment. When characters are affected by strong background interference and light changes, it is often difficult to effectively distinguish target characters from noise, which leads to misrecognition or missed recognition. Summary of the invention

[0005] In view of the shortcomings of the prior art, such as insufficient recognition accuracy due to background interference and prone to missed recognition in dynamic scenes, the present invention proposes an optoelectronic reconnaissance character recognition method, system, equipment and medium, which enhances the robustness of character recognition in complex backgrounds and dynamic environments by utilizing a deep learning algorithm, thereby improving the character recognition capability of optoelectronic reconnaissance equipment in dynamic scenes; thereby solving the problems existing in the prior art.

[0006] A photoelectric reconnaissance character recognition method comprises the following steps:

[0007] Collect photoelectric reconnaissance image data to be identified;

[0008] The photoelectric reconnaissance image data to be recognized is input into the character recognition model, and the image data is subjected to preliminary feature extraction through the convolution layer in the PP-LCNet backbone network; the extracted features are subjected to multiple depth-separable convolutions to extract the features of each input channel, and the features of each input channel are pooled through the global average pooling layer to obtain the global features; the global features are input into the convolution layer to extract the feature map, and the feature map is subjected to convolution operations of different scales to obtain feature maps of different sizes, and the feature maps of different scales are spliced ​​along the channel dimension through the Concat operation to obtain a fused feature map; the fused feature map is input into two predictors to obtain a prediction probability map and a prediction threshold map respectively, and a binary image is generated according to the prediction probability map and the prediction threshold map, thereby obtaining an image containing the required text area;

[0009] The image containing the required text area is input into the text recognition module, and the feature map is extracted again through the PP-LCNet backbone network. The feature map is input into the text recognition network SVTR, and a text label matrix is ​​output. According to the text label matrix, the character recognition result of the optoelectronic reconnaissance image data to be identified is obtained.

[0010] Furthermore, the extracted features are subjected to multiple depth-wise separable convolutions to extract features of each input channel, specifically including further linearly combining the features of each input channel through point-by-point convolution and introducing an attention mechanism module SE to weight the features of each channel; wherein the SE module is introduced to weight the features of each channel, specifically including the following steps:

[0011] Perform global average pooling on the input feature map to obtain each global channel descriptor;

[0012] According to each global channel descriptor, the channel is weighted and adjusted through two fully connected layers, and these weight coefficients are weighted channel by channel with the original input feature map to obtain the output feature map.

[0013] Furthermore, the two predictors are each composed of a 3×3 convolutional layer with a step size of 2 and two deconvolutional layers with a step size of 2; one of the predictors is used to generate a prediction probability map, and the other predictor is used to generate a prediction threshold map.

[0014] Furthermore, the prediction probability map is optimized by using a mutual learning strategy CML based on knowledge distillation, which specifically includes the following steps:

[0015] Calculating the output feature graphs corresponding to the Teacher model and the Student model, specifically including: inputting the predicted probability graph into the Teacher model to generate the corresponding output feature graph; inputting the same predicted probability graph into the Student model to generate the output feature graph of the Student network;

[0016] Calculate the true label loss Loss gt ; The true label loss includes the probability map l p , binary mapping l b and threshold map l t Loss gt The expression is:

[0017] Loss gt (T out ,gt)=l p (S out ,gt)+αl b (S out ,gt)+βl t (S out ,gt),

[0018] in α and β are hyperparameters, T out is the output distribution of the teacher model, gt is the true label, S out is the output distribution of the student model;

[0019] Calculate the DML loss and use the KL divergence loss to measure the difference between the output distributions of the two Student models, which is expressed as:

[0020]

[0021] Among them, S1 pout is the output distribution of student model 1, S2 pout is the output distribution of student model 2;

[0022] Calculate the distillation loss, expressed as:

[0023] Loss distill =γl p (S out ,f dila (T out ))+l b (S out ,f dila (T out )),

[0024] Among them l p and l b are binary cross entropy loss and dice loss respectively, γ is a hyperparameter, and f dila is the expansion function;

[0025] The total CML loss is then expressed as:

[0026] Loss total =Loss gt +Loss dml +Loss distill ,

[0027] By setting a loss function based on the true label, the Teacher model can generate a more accurate feature distribution. Under the guidance of multiple losses, the parameters of the Student model are continuously updated until the output of the Student model is most similar to the output features of the Teacher model, and an optimized prediction probability map is obtained.

[0028] Furthermore, the feature map is extracted through the PP-LCNet backbone network, the feature map is input into the SVTR network, and a text label matrix is ​​output, which specifically includes the following steps:

[0029] Extract feature maps through the PP-LCNet backbone network;

[0030] The feature map is input into the SVTR network, and the embedding operation Patch Embedding is used to transform the feature map to obtain character components;

[0031] By embedding the position module in the character component, a character component containing position information is obtained;

[0032] The character components containing position information are subjected to three-stage feature extraction at different scales, and the extracted feature maps are output; each stage includes mixing blocks, which are used to fuse the position information with the character features; the first and second stages both use the merging operation, and the third stage uses the combining operation for downsampling;

[0033] The extracted feature map is passed through the fully connected FC layer to output a text label matrix.

[0034] Furthermore, two loss functions, CTC Loss and Attention Loss, were used to train the character recognition model.

[0035] The present invention also includes a photoelectric reconnaissance character recognition system, comprising:

[0036] An acquisition module, used for acquiring optoelectronic reconnaissance image data to be identified;

[0037] The recognition module is used to input the photoelectric reconnaissance image data to be recognized into the character recognition model, and perform preliminary feature extraction on the image data through the convolution layer in the PP-LCNet backbone network; extract the features of each input channel through multiple depth-separable convolutions, and pool the features of each input channel through the global average pooling layer to obtain global features; input the global features into the convolution layer to extract feature maps, and perform convolution operations on the feature maps of different scales to obtain feature maps of different sizes, and concatenate the feature maps of different scales along the channel dimension through the Concat operation to obtain a fused feature map; input the fused feature maps into two predictors to obtain a prediction probability map and a prediction threshold map respectively, and generate a binary image according to the prediction probability map and the prediction threshold map, thereby obtaining an image containing the required text area;

[0038] The text label matrix generation unit is used to input the image containing the required text area into the text recognition module, extract the feature map again through the PP-LCNet backbone network, input the feature map into the text recognition network SVTR, output a text label matrix, and obtain the character recognition result of the optoelectronic reconnaissance image data to be recognized based on the text label matrix.

[0039] The present invention also includes an optoelectronic reconnaissance character recognition computer device, comprising: a memory, a processor and a computer program stored in the memory, and the processor implements the steps of the optoelectronic reconnaissance character recognition method when executing the computer program.

[0040] The present invention also includes a readable storage medium, wherein the readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the steps of the photoelectric reconnaissance character recognition method are executed.

[0041] The present invention provides a method, system, device and medium for photoelectric reconnaissance character recognition, which has the following features:

[0042] Beneficial effects:

[0043] The present invention selects PP-LCNet as the backbone network of the character recognition model to process complex text detection and recognition tasks; at the same time, in the text recognition tasks faced with complex backgrounds, different fonts, low contrast or noise interference, the CML strategy is used to optimize the generated prediction probability map to ensure that the boxed area in the generated text box is as accurate as possible, thereby improving the accuracy of the text detection area, solving the problem in the prior art that it is difficult to effectively distinguish target characters from noise, which leads to misrecognition or missed recognition; the method uses a deep learning algorithm to enhance the robustness of character recognition in complex backgrounds and dynamic environments, improves the character recognition ability of optoelectronic reconnaissance equipment in dynamic scenes, ensures that character recognition responds quickly in video streams, and improves real-time processing capabilities to meet the timeliness requirements of reconnaissance tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 Flow chart of the photoelectric reconnaissance character recognition method in an embodiment of the present invention;

[0045] Figure 2 This is a CML network framework diagram in an embodiment of the present invention;

[0046] Figure 3 This is a structural diagram of PFHead in an embodiment of the present invention;

[0047] Figure 4 This is a diagram of a text recognition network structure in an embodiment of the present invention;

[0048] Figure 5 This is a PP-LCNet network structure diagram in an embodiment of the present invention;

[0049] Figure 6 This is a diagram of the SVTR network structure in an embodiment of the present invention;

[0050] Figure 7 A flow chart of a data mining solution in an embodiment of the present invention;

[0051] Figure 8 Schematic diagram of a multi-scale training strategy in an embodiment of the present invention. DETAILED DESCRIPTION

[0052] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0053] The present invention proposes a photoelectric reconnaissance character recognition method, such as Figure 1 As shown, the method specifically comprises the following steps:

[0054] S1. Collect optoelectronic reconnaissance image data to be identified.

[0055] S2, text detection module: The text detection module is responsible for locating the text area in the image. This module uses PP_LCNet as the feature extraction network and combines the CML strategy to enhance the fusion and prediction accuracy of spatial and sequence features.

[0056] S2.1. In the feature extraction part, we first perform preliminary feature extraction through a 7×7 convolution kernel and a convolution layer with a step size of 2, and output a feature map of size 16×112×112. The convolution operation at this stage can effectively extract local features.

[0057] S2.2, the feature map will go through a 3×3 depthwise separable convolution. This convolution operation extracts the features of each input channel through depthwise convolution, and then further linearly combines them through point-by-point convolution, increasing the number of output channels. After this convolution, the size of the feature map becomes 64×56×56. On this basis, a Squeeze-and-Excitation (SE) module is added to weight the features of each channel to highlight important channel features. Specifically, the input feature map is globally averaged pooled to obtain each global channel descriptor, and then the channel weighting is adjusted through two fully connected layers. These weight coefficients are weighted channel by channel with the original feature map to obtain a 256×14×14 feature map.

[0058] S2.3, the feature map undergoes a 5×5 depthwise separable convolution operation to further extract local features and increase the model’s expressiveness, outputting a feature map of size 512×7×7. In the final stage of feature extraction, all features are pooled through a global average pooling layer to obtain the mean of each channel, resulting in a 512-dimensional feature vector as a global feature representation.

[0059] S2.4. In order to enhance the fitting ability of the network, a 1280-dimensional 1×1 convolution (1x1 Conv) layer is then added after the global average pooling. This convolution layer further enhances the representation ability of the model by increasing the feature dimension.

[0060] S2.5, the feature map will be further processed through a fully connected layer, and finally a 1280-dimensional feature map will be obtained. In order to combine information of different scales, the feature map will be subjected to convolution operations of different sizes to obtain feature maps of different sizes. Through the Concat operation, these feature maps of different scales are spliced ​​along the channel dimension to obtain a fused feature map.

[0061] S2.6. Input the fused feature map into two predictors, each of which consists of a 3×3 convolutional layer with a stride of 2 and two deconvolutional layers with a stride of 2; the first predictor generates a prediction probability map for predicting the probability of the text area; the second predictor generates a prediction threshold map for generating the threshold of the text area. In order to improve the prediction accuracy, the prediction probability map is further processed by CML (mutual learning strategy), and the prediction probability map after CML processing and the previously obtained threshold map are input into a binarization module. A binary image is generated according to the prediction probability and threshold to obtain the final text detection result. Finally, the detected text area is expanded and scaled to further refine the boundary of the text area, thereby generating the final detection label.

[0062] Among them, the present invention adopts the mutual learning strategy CML based on knowledge distillation to further process the prediction probability graph. The core idea of ​​CML combines the traditional standard distillation of Teacher guiding Student with the DML mutual learning between Teacher and Student networks, which allows the Student network to learn from each other while the Teacher network provides guidance. Figure 2 As shown in the figure, when calculating the knowledge distill loss of StudentModel and TeacherModel, the network adds the KL divergence loss KL div loss to make the feature distribution of the feature maps output by the two close, thereby further improving the accuracy and stability of the model. The training process includes the following steps.

[0063] ① Initial input: First, the input feature map is sent to the Teacher model and the Student model respectively. GT Lable: The ground truth label is used to guide the standard of the Teacher model output, ensuring that the Teacher model is as close to the correct answer as possible when generating features.

[0064] ② Calculate the output feature maps (ResponseMaps) of the Teacher model and the Student model:

[0065] Teacher model prediction: Input the original image into the Teacher model to generate the corresponding output feature map.

[0066] Prediction of the Student model: Input the same original image into the Student model to generate the output feature map of the student network. At this time, the output feature map of the Student model may be significantly different from the output feature map of the Teacher model. The KL divergence is used to measure the difference between the two, thereby guiding the student model to be closer to the prediction of the teacher model.

[0067] ③Calculate various losses:

[0068] (1) Calculate the true label loss Loss gt : The true label loss is the probability mapping l p , binary mapping l b And the threshold map l t The loss is composed of gt The expression is as follows, where α and β are hyperparameters with default values ​​of 5 and 10.

[0069] Loss gt (T out ,gt)=l p (S out ,gt)+αl b (S out ,gt)+βl t (S out ,gt);

[0070] Among them, T out is the output distribution of the teacher model, gt is the true label, S out is the output distribution of the student model.

[0071] (2) Calculate DML loss: Use KL divergence loss to measure the difference between the output distributions of the two Student models. The specific expression is as follows:

[0072]

[0073] Among them, S1 pout is the output distribution of student model 1, S2 pout is the output distribution of student model 2.

[0074] (3) Calculate the distillation loss: The distillation loss reflects the supervision of the teacher model on the student model. The expression of the distillation loss is as follows, where l p and l b are binary cross entropy loss and dice loss respectively, γ is a hyperparameter with a default value of 5, and f dila is the expansion function.

[0075] Loss distill =γlp (S out ,f dila (T out ))+l b (S out ,f dila (T out ));

[0076] (4) Calculate the total CML loss. The loss function used is as follows:

[0077] Loss total =Loss gt +Loss dml +Loss distill .

[0078] ④ Combined with GT Label to guide optimization: GT Label provides supervision guidance for the Teacher model. By setting a loss function based on the true label, the Teacher model can generate a more accurate feature distribution. In other words, by combining the supervision signal generated by GTLabel, the Student model can learn the feature distribution of the Teacher model while maintaining its own accuracy and stability.

[0079] ⑤Update the parameters of the Student model: Under the guidance of multiple losses, the parameters of the Student model are continuously updated until the output of the Student model is close enough to the Teacher model to achieve the training goal.

[0080] In CML, in order to better capture features of different scales and complexities and improve the adaptability to diverse text scenes, the present invention uses PFHead (multi-branch fusion Head structure) to generate a probability map (text area feature map) to predict the probability that each pixel in the image belongs to the text area. Figure 3 As shown in the figure, after the first transposed convolution, PFHead performs upsampling and transposed convolution respectively. The upsampled output is convolved through 3x3 convolution to obtain the output result, which is then cascaded with the result of the transposed convolution branch and passed through a 1x1 convolution layer. Finally, the result of the 1x1 convolution is added to the result of the transposed convolution to obtain the final output probability map. This structure can better capture features of different scales and complexities by introducing multiple parallel branches, and improve its adaptability to diverse text scenarios.

[0081] S3, text recognition module, such as Figure 4As shown in the figure, in the text recognition module, the input image is first preprocessed to ensure the image quality and information integrity. Then, the image will be preliminarily recognized by the PP-LC Net backbone network and the SVTR network respectively. The PP-LCNet backbone network can quickly extract feature information from the image with its light weight and high efficiency, especially when dealing with complex backgrounds or different fonts, showing good adaptability. The network gradually extracts more abstract features through multi-layer convolution and pooling operations, making the subsequent text recognition process more accurate. At the same time, the SVTR network further enhances the recognition ability of text by combining the mechanism of local feature and global feature extraction. Its unique structure can effectively capture the details of character strokes, thereby improving the recognition accuracy of various text styles.

[0082] By improving the activation function and adjusting the network depth and width, the PP-LCNet network can cover a larger accuracy range. The entire network structure is as follows Figure 5 As shown. In addition, in order to improve the fitting ability of the network, H-Swish in the original ReLU is adopted, which can significantly improve the accuracy with a slight increase in inference time. In order to ensure a better speed and accuracy, it is found through experiments that the closer to the tail of the network, the better the SE module effect, so the present invention adds the SE module to the block near the tail of the network, and the two layers of activation functions in the SE module are ReLU and H-Sigmoid respectively. In addition, in PP-LCNet, the output dimension of the network after GAP is small, and directly connecting the final classification layer will lose the combination of features. In order to make the network have a stronger fitting ability, a 1x1 conv with a size of 1280 is connected to the final GAP layer, which will increase the model size without increasing the inference time.

[0083] The image processing process of the PP-LCNet backbone network is as follows: Figure 5 As shown, the specific steps include:

[0084] (1) The input image is subjected to preliminary feature extraction and dimensionality reduction through a 3×3 standard convolution and H-Swish activation function.

[0085] (2) In order to efficiently extract spatial features while reducing the number of parameters and computation, depthwise separable convolution is used. Depthwise convolution: Perform independent convolution on each input channel to extract spatial information. Pointwise convolution: Mix information from different channels through 1×1 convolution operations. All activation functions use H-Swish. At the same time, in order to improve the feature expression capability, the SE module dynamically adjusts the channel weights to enhance key information. SE has two fully connected layers, using ReLU and H-Sigmoid activation functions respectively; global average pooling further compresses the spatial dimensions (height and width) into a single feature vector, retaining important information between channels.

[0086] (3) A 1×1 convolution operation is performed before the fully connected layer to improve the feature expression capability. The number of channels is adjusted to 1280 to provide stronger feature expression capabilities, and the activation function still uses H-Swish.

[0087] (4) The last fully connected layer has an output dimension of the number of character categories + 1.

[0088] SVTR network: The recognition module is based on the text recognition algorithm SVTR optimization. The SVTR network flow chart is as follows: Figure 6 As shown in the figure, SVTR no longer uses the RNN structure. By introducing the Transformers structure, it can more effectively mine the contextual information of text images, thereby improving the text recognition ability.

[0089] The image processing process of the SVTR network specifically includes the following steps:

[0090] 1) Input processing: After inputting the image text with a size of H*W*3, it is obtained through Patch Embedding The character components of the image. The patch embedding here uses two 3×3 convolutions with a stride of 2 to perform overlapping patch embeddings.

[0091] 2) Position Embedding: Through the Position Embedding module, each character component is given position information to maintain its spatial relationship in the image, and then this information is passed to the Mix Block module for further processing.

[0092] 3) Next, three stages of feature extraction are performed at different scales, such as Figure 6As shown in the figure, each stage consists of a series of mixing blocks, merging or combing. The Merging operation is used in the first and second stages, and the Combining operation is used in the third stage for downsampling to reduce the number of parameters and reduce the computational cost. The Combining operation in the third stage first globally pools the dimension to 1, and then passes through the fully connected layer, nonlinear activation layer and dropout layer, and the characters are further compressed into a feature sequence. The Combing operation is used here instead of the Merging operation because in the final stage of the network, the height of the sample size may drop to a very small value (for example, 1). At this time, continuing to use the Merging operation may completely lose this part of the information.

[0093] Mixing Block module: The extracted position information is fused with the character features so that the model can use the specific position information of the characters in this text. The fused features are nonlinearly transformed to further enhance the expressiveness of the features. The Mixing Block module obtains a richer and more useful text representation, providing stronger feature support for the recognition task. Local feature extraction (Local Mixing): In this stage, the model adopts a local feature extraction strategy. The attention map of each token only responds to the pixels in its central area, where the height of the local box is 7 and the width is 11. This mechanism ensures that the model can effectively capture the detailed features of the character strokes. Subsequently, the local features are processed by a feedforward neural network (FFN). Global feature extraction (Global Mixing): After completing the local feature extraction, the next step is the global feature extraction stage, which uses a normal Transformer encoder to comprehensively analyze the overall information. The combination of local feature extraction and global feature extraction enables the model to understand the input image more comprehensively.

[0094] 4) Finally, the feature size The feature map of is passed through the fully connected FC layer to obtain the character sequence. For English recognition, the number of character categories N is set to 37 (numbers + letters + blank characters), and for Chinese recognition, N is set to 6625 (numbers + letters + simplified characters + traditional characters + special symbols + blank characters).

[0095] In order to continuously optimize the network performance, the model uses two loss functions, CTC Loss (Connecti onist Temporal Classification Loss, CTC Loss) and Attention Loss, during the training process; CTC Loss is used to align the predicted probability distribution with the target text sequence to optimize the model parameters; Attention Loss is used to gradually optimize the accuracy of each character generation; Among them, CTC Loss is mainly used to process sequence data, which can effectively solve the alignment problem in text recognition, so that the model can still learn effectively when the input sequence length does not match the target output length. By introducing CTC Loss, the network can autonomously learn how to map the input features to the correct character sequence. At the same time, Attention Loss further improves the model's ability to capture important information by focusing on the key parts of the input features. This mechanism enables the model to automatically focus on the most relevant areas when processing complex texts, thereby improving the accuracy and stability of recognition. Through the combined use of these two loss functions, the model is continuously optimized in each iteration, improving the accuracy and efficiency of text recognition. Finally, after multiple processing by the PP-LCNet backbone network and the SVTR network, high-quality text output is generated, as shown in Table 1.

[0096] Table 1 Text output

[0097]

[0098] S3.1, identify network optimization, such as Figure 7 shown.

[0099] (1) Data mining scheme: In order to improve the efficiency of model training and reduce the risk of overfitting, the present invention introduces DF (DataFilter) into the training process to make full use of unlabeled data. Through the semi-supervised learning strategy, the DF method can mine the hidden information in the data, effectively improve the generalization ability of the model, and reduce the risk of overfitting. Specifically, it is as follows: first, a low-precision model is obtained by rapid training with a small amount of data, and the low-precision model is used to predict tens of millions of data, and samples with a confidence level greater than 0.95 are removed. This part is considered to be redundant samples that are ineffective in improving the accuracy of the model. Secondly, PP-LCNet is used as a high-precision model to predict the remaining data, and samples with a confidence level less than 0.15 are removed. This part is considered to be difficult to identify or of poor quality. Using this strategy, tens of millions of training data are streamlined to millions, and the model training time is reduced from 2 weeks to 5 days, which significantly improves the training efficiency.

[0100] (2) Multi-Scale Training Strategy Multi-Scale.

[0101] In order to enhance the model's ability to recognize texts of different scales, a multi-scale training strategy is adopted, such as Figure 8 As shown in the figure. Through multi-scale transformation of training data, the model can learn to adapt to multiple text scales, improving its robustness in practical applications. The dynamic scale training strategy is to randomly resize the height of the input image during training to enhance the robustness of the recognition model when used in end-to-end series. During training, each iter randomly selects one of the three heights (32, 48, 64) for resizing.

[0102] This invention solves the limitations of traditional OCR in terms of generalization and expansion capabilities, especially in the recognition efficiency of superimposed characters in airborne videos. An LLM fine-tuning training interface was designed and optimized through experimental data, which significantly improved the OCR recognition accuracy in specific scenarios. The screen recording and video recording accuracy rates reached 93.56% and 99.86% respectively. Compared with the current mainstream OCR algorithms, this invention achieves a performance improvement of more than 5%, providing an innovative solution for complex text analysis needs.

[0103] Based on the same inventive concept, the present invention also proposes a photoelectric reconnaissance character recognition system, comprising:

[0104] The acquisition module is used to acquire the optoelectronic reconnaissance image data to be identified.

[0105] The recognition module is used to input the optoelectronic reconnaissance image data to be recognized into the character recognition model, and perform preliminary feature extraction on the image data through the convolution layer in the PP-LCNet backbone network; extract the features of each input channel through multiple depth-separable convolutions, and pool the features of each input channel through the global average pooling layer to obtain the global features; input the global features into the convolution layer to extract the feature map, and perform convolution operations on the feature map at different scales to obtain feature maps of different sizes, and concatenate the feature maps of different scales along the channel dimension through the Concat operation to obtain a fused feature map; input the fused feature map into two predictors to obtain a prediction probability map and a prediction threshold map respectively, and generate a binary image according to the prediction probability map and the prediction threshold map, thereby obtaining an image containing the required text area.

[0106] The text label matrix generation unit is used to input the image containing the required text area into the text recognition module, extract the feature map through the PP-LCNet backbone network, input the feature map into the SVTR network, output a text label matrix, and obtain the character recognition result based on the text label matrix.

[0107] The present invention also provides an optoelectronic reconnaissance character recognition computer device, comprising: a memory, a processor and a computer program stored in the memory, and the processor implements the steps of the optoelectronic reconnaissance character recognition method when executing the computer program.

[0108] The present invention also provides a readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, they are used to execute the steps of the photoelectric reconnaissance character recognition method.

[0109] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A photoelectric reconnaissance character recognition method, characterized in that: The following steps are involved: Collect photoelectric reconnaissance image data to be identified; The photoelectric reconnaissance image data to be recognized is input into the character recognition model, and the image data is initially extracted through the convolution layer in the PP-LCNet backbone network; the extracted features are subjected to multiple depth-separable convolutions to extract the features of each input channel, and the features of each input channel are pooled through the global average pooling layer to obtain the global features; The global features are input into the convolution layer to extract the feature map, and the feature map is subjected to convolution operations of different scales to obtain feature maps of different sizes. The feature maps of different scales are spliced ​​along the channel dimension through the Concat operation to obtain a fused feature map; the fused feature map is input into two predictors to obtain a prediction probability map and a prediction threshold map respectively, and the prediction probability map is optimized by using the mutual learning strategy CML based on knowledge distillation, and a binary image is generated according to the optimized prediction probability map and prediction threshold map, thereby obtaining an image containing the required text area; wherein the prediction probability map is optimized by using the mutual learning strategy CML based on knowledge distillation, and specifically includes the following steps: Calculate the output feature map corresponding to the teacher model and the student model; Calculate the true label loss ; The true label loss includes the probability mapping , binary mapping and threshold mapping Loss: , in and is a hyperparameter, is the output distribution of the teacher model, is the true label, is the output distribution of the student model; Calculate the DML loss and use the KL divergence loss to measure the degree of difference between the output distributions of the two student models, which is expressed as: , in, is the output distribution of student model 1, is the output distribution of student model 2; Calculate the distillation loss, expressed as: , in and are binary cross entropy loss and dice loss respectively, is a hyperparameter, is the expansion function; The total CML loss is then expressed as: , Under the guidance of multiple losses, the parameters of the teacher model are continuously updated until the output of the teacher model is most similar to the output characteristics of the student model, and then the optimized prediction probability map is obtained; The image containing the required text area is input into the text recognition module, and the feature map is extracted through the PP-LCNet backbone network. The feature map is input into the text recognition network SVTR, and a text label matrix is ​​output. According to the text label matrix, the character recognition result of the optoelectronic reconnaissance image data to be recognized is obtained.

2. The photoelectric reconnaissance character recognition method according to claim 1, characterized in that: The extracted features are subjected to multiple depth-separable convolutions to extract features of each input channel, specifically including further linear combination of the features of each input channel through point-by-point convolution and introducing an attention mechanism module SE to weight the features of each channel; The SE module is introduced to weight the features of each channel, which specifically includes the following steps: Perform global average pooling on the input feature map to obtain each global channel descriptor; According to each global channel descriptor, the channel is weighted and adjusted through two fully connected layers, and these weight coefficients are weighted channel by channel with the original input feature map to obtain the output feature map.

3. The photoelectric reconnaissance character recognition method according to claim 1, characterized in that: The two predictors are composed of a 3×3 convolutional layer with a stride of 2 and two deconvolutional layers with a stride of 2; one of the predictors is used to generate a prediction probability map, and the other predictor is used to generate a prediction threshold map.

4. The photoelectric reconnaissance character recognition method according to claim 1, characterized in that: The feature map is extracted through the PP-LCNet backbone network, the feature map is input into the SVTR network, and a text label matrix is ​​output, which specifically includes the following steps: Extract feature maps through the PP-LCNet backbone network; The feature map is input into the SVTR network, and the embedding operation Patch Embedding is used to transform the feature map to obtain character components; By embedding the position module in the character component, a character component containing position information is obtained; The character components containing position information are subjected to three-stage feature extraction at different scales, and the extracted feature maps are output; each stage includes mixing blocks, which are used to fuse the position information with the character features; the first and second stages both use the merging operation, and the third stage uses the combining operation for downsampling; The extracted feature map is passed through the fully connected FC layer to output a text label matrix.

5. The photoelectric reconnaissance character recognition method according to claim 1, characterized in that: Two loss functions, CTC Loss and Attention Loss, are used to train the character recognition model.

6. An optoelectronic reconnaissance character recognition system, characterized in that: include: An acquisition module, used for acquiring optoelectronic reconnaissance image data to be identified; The recognition module is used to input the photoelectric reconnaissance image data to be recognized into the character recognition model, and perform preliminary feature extraction on the image data through the convolution layer in the PP-LCNet backbone network; extract the features of each input channel through multiple depth-separable convolutions, and pool the features of each input channel through the global average pooling layer to obtain the global features; The global features are input into the convolution layer to extract the feature map, and the feature map is subjected to convolution operations of different scales to obtain feature maps of different sizes. The feature maps of different scales are spliced ​​along the channel dimension through the Concat operation to obtain a fused feature map; the fused feature map is input into two predictors to obtain a prediction probability map and a prediction threshold map respectively, and the prediction probability map is optimized by using the mutual learning strategy CML based on knowledge distillation, and a binary image is generated according to the optimized prediction probability map and prediction threshold map, thereby obtaining an image containing the required text area; wherein the prediction probability map is optimized by using the mutual learning strategy CML based on knowledge distillation, and specifically includes the following steps: Calculate the output feature map corresponding to the teacher model and the student model; Calculate the true label loss ; The true label loss includes the probability mapping , binary mapping and threshold mapping Loss: , in and is a hyperparameter, is the output distribution of the teacher model, is the true label, is the output distribution of the student model; Calculate the DML loss and use the KL divergence loss to measure the degree of difference between the output distributions of the two student models, which is expressed as: , in, is the output distribution of student model 1, is the output distribution of student model 2; Calculate the distillation loss, expressed as: , in and are binary cross entropy loss and dice loss respectively, is a hyperparameter, is the expansion function; The total CML loss is then expressed as: , Under the guidance of multiple losses, the parameters of the teacher model are continuously updated until the output of the student model is most similar to the output features of the teacher model, and then the optimized prediction probability map is obtained; The text label matrix generation unit is used to input the image containing the required text area into the text recognition module, extract the feature map through the PP-LCNet backbone network, input the feature map into the SVTR network, output a text label matrix, and obtain the character recognition result based on the text label matrix.

7. An optoelectronic reconnaissance character recognition computer device, characterized in that: include: A memory, a processor and a computer program stored in the memory, wherein the processor implements the steps of the photoelectric reconnaissance character recognition method according to any one of claims 1 to 5 when executing the computer program.

8. A readable storage medium, characterized in that: The readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, they are used to execute the steps of the photoelectric reconnaissance character recognition method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Medical record character recognition method and system based on deep learning

    CN117218672A

  • OCR (Optical Character Recognition) method, device and equipment based on scanning pen and medium

    CN118429976A