A picture recognition method and system for a scanning pen

Through superpixel generation, semantic segmentation, stroke extraction and reconstruction, glyph matching and context correction technologies, the recognition problems of traditional OCR in scenes such as complex backgrounds and text distortion are solved, and higher recognition accuracy and fluency are achieved.

CN119832576BActive Publication Date: 2025-06-10深圳市英得尔实业有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510302904.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-10
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

Traditional OCR technology performs poorly in complex backgrounds, uneven lighting, distorted text, handwritten and artistic characters, and has low recognition accuracy.

Method used

A picture recognition method for scanning pen is adopted to achieve accurate recognition of complex text through superpixel generation, semantic segmentation, stroke extraction and reconstruction, glyph matching and context correction technologies.

Benefits of technology

It effectively solves the recognition problem of traditional OCR in complex scenarios, improves the recognition accuracy and fluency, and provides more accurate and reliable recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832576B_ABST
    Figure CN119832576B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computer vision technology, and particularly relates to a picture recognition method and system for a scanning pen. The method includes the following steps: obtaining an original image; generating superpixels for the original image to obtain a superpixel set; training a semantic segmentation network based on adding a superpixel perception layer to a pre-constructed initial U-Net network model to obtain a semantic segmentation model; using the semantic segmentation model and the superpixel set to predict the superpixel semantic categories and construct a superpixel semantic map to obtain a superpixel semantic map; extracting stroke segments from the superpixel semantic map and constructing a stroke connection map to obtain a stroke connection map; performing stroke constraints on the stroke connection map and performing stroke reconstruction to obtain a reconstructed stroke set; generating a stroke structure diagram according to the reconstructed stroke set to obtain a stroke structure diagram. The present invention realizes more accurate and reliable recognition results of the scanning pen through computer vision technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular, to a method and system for picture recognition for a scanning pen. Background Art

[0002] The initial optical character recognition (OCR) was mainly based on technologies such as template matching and feature extraction. Early OCR systems could only recognize a limited number of fonts and characters and had high requirements for image quality. Later, the introduction of methods such as statistical pattern recognition and structural analysis improved the recognition performance and generalization ability of OCR systems. However, traditional OCR technologies still face challenges in some specific scenarios. For example, in cases of complex backgrounds, uneven lighting, and distorted text, the recognition rate is still low, and it is difficult to handle handwritten and artistic characters. The main reasons for these problems are as follows:

[0003] Complex background and uneven lighting: These factors interfere with the extraction and recognition of characters, making it difficult to extract features and thus reducing the recognition accuracy.

[0004] Text distortion and deformation: Distorted and deformed text changes the shape and structure of characters, rendering traditional shape-based recognition methods ineffective. The diverse fonts and irregular shapes of handwritten and artistic characters further exacerbate this problem. Summary of the Invention

[0005] Based on this, it is necessary to provide a method and system for picture recognition for a scanning pen to solve at least one of the above technical problems.

[0006] To achieve the above object, a method for picture recognition for a scanning pen includes the following steps:

[0007] Step S1: Obtain an original image; generate superpixels for the original image to obtain a superpixel set; train a semantic segmentation network based on adding a superpixel perception layer to a pre-constructed initial U-Net network model to obtain a semantic segmentation model; use the semantic segmentation model and the superpixel set to perform superpixel semantic category prediction and construct a superpixel semantic map to obtain a superpixel semantic map;

[0008] Step S2: Extract stroke segments from the superpixel semantic map and construct a stroke connection map to obtain a stroke connection map; perform stroke constraints on the stroke connection map and perform stroke reconstruction to obtain a reconstructed stroke set; generate a stroke structure map according to the reconstructed stroke set to obtain a stroke structure map;

[0009] Step S3: Obtain font samples; construct a multi-scale stroke structure graph library based on the font samples, and construct a glyph feature library to obtain a multi-scale stroke structure graph library and a glyph feature library; calculate the graph similarity of the multi-scale stroke structure graph library and the glyph feature library to obtain a character similarity matrix; perform glyph matching based on the character similarity matrix to obtain a candidate character set;

[0010] Step S4: Obtain the previous text; perform vector context information fusion on the candidate character set and the previous text to obtain a fused feature vector; perform optimal character prediction on the fused feature vector to obtain a predicted character; splice the predicted character into the previous text to construct a corrected character sequence, and obtain a corrected character sequence; output the text result of the corrected character sequence to obtain a display result.

[0011] The present invention obtains a clear original image through the camera of the scanning pen, generates superpixels using the SLIC algorithm, converts pixel-level processing into superpixel-level processing, effectively reduces the computational complexity, and retains the edge information of the image. Subsequently, a U-Net semantic segmentation model with a superpixel perception layer is constructed and trained. This model fuses the local image features extracted by the convolutional layer and the global features at the superpixel level (color, texture, position, shape), and through the weighted cross-entropy loss function and data augmentation strategy, significantly improves the segmentation accuracy and robustness of the model for text regions. The finally generated superpixel semantic map can precisely present the category (text, background, noise) of each pixel in the image, providing accurate input for subsequent stroke extraction. Extracting the superpixels marked as "text" from the superpixel semantic map excludes background and noise interference, improving the efficiency and accuracy of subsequent processing. Then, the stroke skeleton is extracted through Otsu binarization and Zhang-Suen thinning algorithm, and the thinned pixels are connected into stroke segments using connected component analysis. A stroke connection graph is constructed based on these stroke segments, and the connection possibility between stroke segments is predicted through a logistic regression model, considering multiple features such as direction similarity, distance, and gray value difference. Most importantly, stroke constraint conditions (length, angle, loop detection) are defined, and the Kruskal algorithm is used to solve the weighted minimum spanning tree. On the premise of meeting the constraint conditions, the stroke segments are reconstructed into complete strokes, ensuring the rationality and accuracy of the reconstructed strokes. The finally generated stroke structure diagram completely describes the stroke composition and layout of the text, providing key input for glyph matching. A multi-scale stroke structure diagram library and glyph feature library containing various fonts, font sizes, and scales are constructed, ensuring the coverage of the font library and the adaptability to input texts of different sizes. When calculating the graph similarity, the best matching scale is selected through area similarity calculation, avoiding the error caused by directly comparing stroke structure diagrams of different scales. At the same time, local feature vectors of important stroke structures (horizontal, vertical, left-falling stroke, right-falling stroke, dot, etc.) are defined and extracted, and the overall structure similarity is calculated in combination with the graph edit distance algorithm, realizing a more refined and accurate similarity calculation. Finally, by setting a similarity threshold and selecting the Top N characters, a candidate character set containing the most likely recognition results is generated, providing reliable input for subsequent context fusion and final recognition. The pre-identified previous text information and candidate character information are fully utilized to improve the recognition accuracy and fluency. The candidate characters are encoded into feature vectors through a pre-trained character embedding model (such as BERT) to capture the semantic information of the characters. At the same time, the pre-trained Transformer language model is used to encode the previous text to obtain a context vector containing semantic information. Most crucially, the attention mechanism is used to fuse the candidate character feature vectors and the context vectors, enabling the model to focus on the previous text information most relevant to the current candidate character, thereby more accurately determining which candidate character is the most suitable.Finally, the character with the highest probability is selected as the predicted character through the fully connected layer and the softmax layer, and the predicted character is concatenated into the previous text to construct a complete corrected character sequence, which is output in a user-readable format, realizing the accurate conversion from image to text. Therefore, the present invention provides a picture recognition method for a scanning pen, which effectively solves the disadvantages of traditional OCR in the recognition of complex backgrounds, uneven illumination, distorted text, handwritten and artistic characters through the combination of semantic pre-segmentation, stroke reconstruction, glyph matching and context correction technologies, and further improves the recognition accuracy through context correction. The present invention has great advantages in the application scenario of a scanning pen and can provide more accurate and reliable recognition results.

[0012] Preferably, step S1 includes the following steps:

[0013] Step S11: Obtain the original image through the scanning pen camera; generate superpixels for the original image to obtain a superpixel set;

[0014] Step S12: Extract superpixel features according to the superpixel set and the original image to obtain a superpixel feature vector;

[0015] Step S13: Obtain the training data set; add a superpixel perception layer to the pre-constructed initial U-Net network model, and train the semantic segmentation network according to the training data set to obtain a semantic segmentation model;

[0016] Step S14: Input the superpixel feature vector into the semantic segmentation model to predict the superpixel semantic category, and obtain the superpixel semantic category;

[0017] Step S15: Construct a superpixel semantic map according to the superpixel set and the superpixel semantic category to obtain a superpixel semantic map.

[0018] The present invention obtains the original image of the text to be recognized through the camera of the scanning pen, providing input data for subsequent processing. The key lies in the automatic focusing of the camera to ensure a clear image. Then, the SLIC algorithm is used to segment the original image into superpixels. Compared with directly processing pixel-level images, superpixels provide local region information of the image, grouping pixels with similar colors and spatial positions, reducing the computational complexity of subsequent processing, while retaining the edge information of the image, laying a foundation for subsequent feature extraction and semantic segmentation. Superpixel generation is a good compromise between computational efficiency and retaining image details. Rich feature vectors, including color, texture, position, and shape features, are extracted for each superpixel. These features describe the visual attributes of superpixels from multiple perspectives, providing a basis for distinguishing different categories (text, background, noise) in subsequent semantic segmentation. For example, color features can distinguish text and background of different colors, texture features can distinguish the texture differences between text regions and background regions, position features can help the model learn that text usually appears in specific regions of the image, and shape features can help distinguish text strokes with specific shapes. These features are combined into a high-dimensional feature vector, enabling the semantic segmentation model to learn and predict based on these features. A prepared and labeled training dataset is provided as samples for the model to learn. Then, a pre-built U-Net model is utilized, with its encoder-decoder structure and skip connections, to effectively extract image features and generate pixel-level segmentation results. Most importantly, a superpixel-aware layer is added to incorporate the superpixel features extracted in step S12 into the U-Net model. This enables the model to utilize not only the local image features extracted by convolutional layers but also the global features at the superpixel level, improving the segmentation accuracy of the model for text regions. Through the weighted cross-entropy loss function (assigning higher weights to text categories) and data augmentation, the recognition ability and robustness of the model for text regions are further improved. The model outputs the probabilities that a superpixel belongs to the "text", "background", or "noise" category based on the input superpixel feature vector, and takes the category with the highest probability as the prediction result. This step converts the superpixel-level feature information into semantic information, laying a foundation for subsequent construction of the superpixel semantic map. The semantic category information of the superpixels is mapped back to the original image to construct the superpixel semantic map. The semantic map intuitively shows the category (text, background, noise) of each pixel in the image. Compared with directly outputting the category of each superpixel, the superpixel semantic map provides a finer segmentation result, being able to more accurately locate the text region and providing a more accurate input for subsequent stroke extraction. At the same time, the situation of overlapping boundaries of multiple superpixels is processed to ensure the accuracy of the semantic map.

[0019] Preferably, step S13 includes the following steps:

[0020] Step S131: Perform data augmentation on the training dataset to obtain an augmented training image set;

[0021] Step S132: Integrate the pre-constructed initial U-Net network model with a superpixel perception layer according to the superpixel feature vector to obtain a U-Net model with an integrated superpixel perception layer;

[0022] Step S133: Define a weighted cross-entropy loss function according to the enhanced training image set to obtain a weighted cross-entropy loss function;

[0023] Step S134: Train a semantic segmentation network according to the enhanced training image set, the weighted cross-entropy loss function, and the U-Net model with an integrated superpixel perception layer to obtain a semantic segmentation model.

[0024] The present invention generates more training samples by applying various data augmentation operations (random rotation, scaling, cropping, flipping, color jitter) to the original training dataset. This effectively expands the scale and diversity of the training dataset, improving the generalization ability and robustness of the model. Data augmentation can simulate various image changes that may occur in real scenarios (e.g., text at different angles, lighting, and sizes), enabling the model to adapt to these changes and thus achieve better recognition results in practical applications. The key is that all augmentation operations preserve the annotation information of superpixels, ensuring the effectiveness of the augmented data. Incorporating superpixel features into the U-Net model is the key to improving text segmentation accuracy. By adding a superpixel perception layer after each convolutional layer in the U-Net encoder, the model can not only learn the local image features extracted by the convolutional layer but also utilize the global features at the superpixel level (color, texture, position, shape). The superpixel perception layer maps the superpixel features to the same dimension as the output of the convolutional layer through a fully connected layer and uses a channel attention mechanism to fuse the two types of features. This enables the model to comprehensively understand the image information and more accurately distinguish the text region and the background region, thereby improving the segmentation accuracy. A weighted cross-entropy loss function is defined to measure the difference between the model's prediction results and the true annotations. The key lies in assigning different weights to different categories (text, background, noise). Since text recognition is the main objective, the "text" category is assigned the highest weight (0.6), which makes the model pay more attention to the prediction accuracy of the text region during training, thus further improving the text segmentation accuracy. This weighting strategy can effectively solve the problem of class imbalance and improve the model's recognition ability for important classes. Integrating the achievements of the previous steps, the final training of the model is carried out. Using the augmented training image set to provide rich training samples, the weighted cross-entropy loss function to guide the model optimization direction, and the U-Net model integrated with the superpixel perception layer as the learning subject. Through the Adam optimizer and the backpropagation algorithm, the model continuously adjusts its parameters to minimize the loss function and finally obtains a trained semantic segmentation model. The validation set evaluation and model saving strategy during the training process ensure the optimal performance of the final model. This step is the core of the entire semantic segmentation process, combining data, model, and optimization strategies to obtain a model that can accurately segment the text region.

[0025] Preferably, step S2 includes the following steps:

[0026] Step S21: Extract the text region from the superpixel semantic map to obtain the superpixel set of the text region;

[0027] Step S22: Extract the stroke segments according to the superpixel set of the text region and the original image to obtain the candidate stroke segment set;

[0028] Step S23: Construct a stroke connection graph for the candidate stroke segment set to obtain a stroke connection graph;

[0029] Step S24: Apply stroke constraints to the stroke connection graph and perform stroke reconstruction to obtain a reconstructed stroke set;

[0030] Step S25: Generate a stroke structure graph based on the reconstructed stroke set to obtain a stroke structure graph.

[0031] The present invention extracts the superpixels marked as "text" to form a superpixel set of the text region. The advantage of this is that the scope of subsequent processing is limited to the text region, excluding the interference of the background and noise, and improving the efficiency and accuracy of subsequent stroke extraction and reconstruction. At the same time, the boundary coordinates and pixel coordinate sets of the superpixels in the text region are recorded, providing an index for extracting the corresponding region from the original image subsequently. Stroke segments are extracted from the original image region corresponding to the superpixels in the text region. First, the Otsu algorithm is used for binarization to separate text pixels and background pixels. Then, the Zhang-Suen thinning algorithm is used to extract the single-pixel-width stroke skeleton, which is a key step in stroke extraction and can effectively extract the shape features of the strokes. Finally, the thinned pixels are connected into stroke segments through connected component analysis, and the features of each segment (center point, main direction, length, pixel coordinates) are calculated. These stroke segments are the basis for subsequent stroke reconstruction and provide nodes for constructing a stroke connection graph. Regarding the stroke segments as nodes, a stroke connection graph is constructed to represent the connection relationship between the stroke segments. The calculation of the connection possibility comprehensively considers multiple features (direction similarity, distance, average gray value difference, endpoint distance ratio) and uses a logistic regression model for prediction. This is more accurate and robust than the judgment based on a single feature. By setting the connection possibility threshold, unreasonable connections can be filtered out, reducing the computational amount of subsequent stroke reconstruction. The stroke connection graph provides a basis for subsequent stroke merging and is a key data structure for stroke reconstruction. Stroke constraint conditions (length, angle, loop detection) are defined to avoid generating unreasonable strokes. Then, the Kruskal algorithm is used to solve the weighted minimum spanning tree, and under the premise of meeting the constraint conditions, the optimal stroke connection method is found. According to the result of the minimum spanning tree, the connected stroke segments are merged into complete strokes, and the stroke features are recalculated. This step reconstructs the scattered stroke segments into complete strokes, providing a more accurate input for subsequent glyph matching. The combination of the Kruskal algorithm and stroke constraints ensures the rationality and accuracy of the reconstructed strokes. The reconstructed strokes and their spatial relationships are represented as a graph structure. The nodes represent strokes, and the edges represent the spatial relationships (intersection, distance, direction) between the strokes. The stroke structure graph completely describes the stroke composition and layout of the text and is the key input for glyph matching. Compared with directly using the reconstructed stroke set, the stroke structure graph more clearly expresses the relationship between the strokes and provides richer information for subsequent graph similarity calculation.

[0032] Preferably, step S24 includes the following steps:

[0033] Step S241: Define stroke constraints based on the candidate stroke segment set to obtain stroke constraint definition data;

[0034] Step S242: Calculate the edge weights according to the stroke connection graph to obtain the edge weights;

[0035] Solve the weighted minimum spanning tree for the edge weights to obtain the minimum spanning tree;

[0036] Step S243: Perform multi-stroke segment connection processing on the candidate stroke segment set according to the minimum spanning tree to obtain the reconstructed stroke set.

[0037] The present invention defines the constraint conditions in the stroke reconstruction process, including length constraint, angle constraint, distance constraint, and loop constraint. These constraint conditions are based on the prior knowledge of writing rules and character structures, and can effectively avoid generating unreasonable strokes. For example, the length constraint prevents the strokes from being too long, the angle constraint prevents the sudden change of stroke directions, the distance constraint prevents the strokes from being too far apart, and the loop constraint prevents the strokes from forming closed loops. These constraint conditions provide guidance for subsequent stroke merging, ensuring the rationality and accuracy of the reconstructed strokes. Defining these rules clearly as data facilitates the subsequent use and modification of algorithms. Calculate the weight of each edge in the stroke connection graph according to the connection possibility and stroke constraint conditions. The weight calculation formula comprehensively considers the connection possibility and various constraint penalty factors, so that the weight can reflect the reasonable degree of stroke connection. Then, use the Kruskal algorithm to solve the weighted minimum spanning tree. The Kruskal algorithm can find the stroke connection method with the minimum total weight under the premise of meeting the stroke constraints, that is, find the most reasonable stroke connection scheme. The result of the minimum spanning tree provides clear guidance for subsequent stroke segment merging. According to the connection relationship of the minimum spanning tree, merge the connected stroke segments into complete strokes. Specifically, it is to merge the pixel coordinate lists of the stroke segments and recalculate the features of the strokes. If a stroke segment is not connected to other segments, it is regarded as an independent stroke. This step completes the reconstruction from stroke segments to complete strokes and obtains the final reconstructed stroke set. The reconstructed stroke set is the basis for subsequent glyph matching, and its quality directly affects the recognition accuracy.

[0038] Preferably, step S3 includes the following steps:

[0039] Step S31: Obtain the font sample; extract the stroke structure diagram of the font sample and construct a multi-scale stroke structure diagram library to obtain the multi-scale stroke structure diagram library;

[0040] Step S32: Construct a glyph feature library according to the multi-scale stroke structure diagram library to obtain the glyph feature library;

[0041] Step S33: Calculate the graph similarity of the multi-scale stroke structure diagram library and the glyph feature library to obtain the character similarity matrix;

[0042] Step S34: Screen the candidate characters from the character similarity matrix to obtain the list of candidate characters and similarity scores;

[0043] Step S35: Construct a candidate character set from the candidate character and similarity score list to obtain the candidate character set.

[0044] In the present invention, by obtaining standard character sample images of multiple fonts and sizes, the coverage range of the font library is ensured. Then, using the same stroke extraction and reconstruction method as in step S2, the stroke structure diagram of each character is extracted, which ensures the consistency between the stroke structure diagrams in the font library and those of the input image. Most importantly, the stroke structure diagrams of each character are scaled at multiple scales to construct a multi-scale stroke structure diagram library. This enables the font library to adapt to input texts of different sizes and improves the robustness of recognition. The multi-scale font library is the key to processing texts of different sizes. Feature vectors are constructed for each character in the font library for subsequent similarity calculation. The feature vectors include the number of strokes, the features of each stroke (length, direction, curvature), and the connection relationships between strokes. These features describe the glyph structure of the character from multiple perspectives and provide a basis for distinguishing different characters. Storing the feature vectors together with the character encoding, font, size, and scaling ratio constructs a glyph feature library, which facilitates subsequent fast similarity calculation. Calculate the similarity between the input stroke structure diagram and each character in the font library. First, according to the size of the input stroke structure diagram, select the set of stroke structure diagrams with the closest scale from the multi-scale font library, reducing the amount of calculation. Then, use the graph edit distance algorithm to calculate the similarity. The graph edit distance algorithm can effectively measure the difference between two graph structures and is an effective method for calculating the similarity of stroke structure diagrams. Store the calculated similarity scores in a character similarity matrix, providing a basis for subsequent candidate character screening. Screen out possible candidate characters according to the similarity scores. By setting a similarity threshold, filter out characters with low similarity, reducing the amount of calculation for subsequent processing. Store the characters with similarity scores higher than the threshold and their scores in a list and sort them by score, facilitating the subsequent selection of the optimal candidate characters. Select the top N characters with the highest scores from the candidate character and similarity score list to form the candidate character set. The candidate character set contains the most likely recognition results and is the input for subsequent context fusion and final recognition. Selecting Top N instead of only the character with the highest score can improve the fault tolerance of recognition and avoid incorrect recognition results due to a single character recognition error.

[0045] Preferably, step S33 includes the following steps:

[0046] Step S331: Calculate the area similarity of the multi-scale stroke structure diagram library and perform scale matching according to the glyph feature library to obtain matching scale selection data;

[0047] Step S332: Obtain important stroke data; define the important stroke structure according to the important stroke data to obtain important stroke structure definition data;

[0048] Step S333: Extract local structure features from the multi-scale stroke structure diagram library according to the important stroke structure definition data to obtain local structure feature vectors;

[0049] Step S334: Calculate the graph similarity according to the multi-scale stroke structure diagram library, the local structure feature vectors, and the glyph feature library to obtain a character similarity matrix.

[0050] In the present invention, by calculating the area difference between the input stroke structure diagram and the stroke structure diagrams of each character in the font library at different scales, the scale with the smallest area difference is selected as the best matching scale for the character. The advantage of this is that it avoids directly comparing the stroke structure diagrams at different scales and improves the accuracy of similarity calculation. Because for stroke structure diagrams at different scales, even for the same character, features such as the number and length of strokes will have large differences, and direct comparison will lead to errors. Through scale matching, it is ensured that the subsequent similarity calculation is carried out at the same or similar scales. Important stroke structures are usually the distinguishable strokes in Chinese characters, such as horizontal strokes, vertical strokes, left-falling strokes, right-falling strokes, dots, hooks, etc. By pre-defining the judgment rules for these important stroke structures (for example, according to features such as the direction, length, and shape of the strokes), the strokes can be classified, thus more precisely describing the glyph structure of the character. This can distinguish different characters more effectively than simply using rough features such as the number and length of strokes. According to the definition of the important stroke structure, count the occurrence times, average length, average direction, and average curvature of various important stroke structures in each stroke structure diagram. These features describe the local details of the stroke structure diagram and can more accurately distinguish different characters. Combine these statistical information into local structure feature vectors, providing a more refined feature representation for the subsequent graph similarity calculation. Use the best matching scale obtained in Step S331. Then, use the graph edit distance algorithm to calculate the overall structure similarity, and use the cosine similarity of the local structure feature vectors to calculate the local feature similarity. Finally, perform a weighted average on the two similarities to obtain the final graph similarity score. This similarity calculation method that comprehensively considers the overall and local features is more accurate and robust than a single similarity calculation method. Store the calculated similarity scores in the character similarity matrix, providing a basis for subsequent candidate character screening.

[0051] Preferably, Step S333 is specifically:

[0052] Perform stroke direction judgment on the multi-scale stroke structure diagram library to obtain stroke direction judgment data; perform stroke length judgment on the multi-scale stroke structure diagram library to obtain stroke length judgment data; perform stroke shape judgment on the multi-scale stroke structure diagram library to obtain stroke shape judgment data;

[0053] Perform important stroke structure marking processing based on the stroke direction judgment data, stroke length judgment data, and stroke shape judgment data to obtain important stroke structure markings;

[0054] Calculate the similarity of important stroke structures for the multi-scale stroke structure diagram library and the glyph feature library based on the important stroke structure markings, and perform relative position calculation to obtain local feature calculation results;

[0055] Construct a feature vector for the local feature calculation results to obtain a local structure feature vector.

[0056] In the present invention, by using principal component analysis (PCA) to calculate the principal direction angle of the stroke, the orientation of the stroke can be accurately described, which is an important basis for distinguishing different stroke types (such as horizontal, vertical, left-falling, right-falling). Calculating the length of the pixel contour of the stroke can reflect the thickness and length of the stroke, which is an important basis for distinguishing different stroke types (such as dot, short horizontal, long horizontal). Calculating the stroke shape complexity (the ratio of the perimeter to the perimeter of the minimum circumscribed rectangle) can reflect the degree of curvature of the stroke, which is an important basis for distinguishing different stroke types (such as left-falling, right-falling). These three features describe the geometric attributes of the stroke from different angles and provide basic data for subsequent important stroke structure marking. According to the previously extracted stroke features, each stroke is marked as a predefined important stroke structure (horizontal, vertical, left-falling, right-falling, dot) or others. By setting clear rules (for example, the principal direction angle of the horizontal stroke is within a certain range and the length is greater than a certain threshold), the strokes can be classified. This has more semantic information than directly using the original stroke features and can more effectively describe the glyph structure of the character. The important stroke structure marking is a key step in local feature extraction and provides more refined features for subsequent similarity calculation. Count the number of each important stroke structure and calculate the quantity difference as the similarity score. This reflects the similarity degree of the important stroke compositions of the two characters. Divide the image into grids and calculate the difference in the grid areas where the center points of the same type of important strokes are located as the relative position score. This reflects the similarity degree of the spatial layouts of the important strokes of the two characters. Combine the important stroke structure similarity score and the relative position score into a feature vector, that is, the local structure feature vector. This feature vector concisely summarizes the local features of the stroke structure diagram and provides an effective feature representation for subsequent graph similarity calculation. The local structure feature vector is a bridge connecting local feature extraction and overall similarity calculation.

[0057] Preferably, step S4 includes the following steps:

[0058] Step S41: Obtain the previous text; use a pre-trained character embedding model to perform candidate character feature encoding on the candidate character set to obtain candidate character feature vectors;

[0059] Step S42: Use a pre-trained Transformer language model to perform previous text encoding on the previous text sequence to obtain context vectors;

[0060] Step S43: Perform context information fusion on the candidate character feature vectors and the context vectors to obtain fused feature vectors;

[0061] Step S44: Perform optimal character prediction on the fused feature vectors to obtain predicted characters;

[0062] Step S45: Concatenate the predicted characters into the previous text to construct a corrected character sequence to obtain a corrected character sequence;

[0063] Step S46: Perform text format conversion on the corrected character sequence to obtain a formatted text; display the formatted text to obtain a display result.

[0064] The present invention utilizes the recognized preamble text information and candidate character information. Obtaining the preamble text provides context information for subsequent context fusion, which is the key to improving the recognition accuracy. Each character in the candidate character set is encoded into a feature vector using a pre-trained character embedding model (such as the character-level embedding of BERT). The pre-trained model has been trained on large-scale text data and can capture the semantic information of characters, making similar characters have similar feature vectors. This is more effective than one-hot encoding or other simple encoding methods. It enables the model to consider both glyph similarity and semantic information simultaneously. The Transformer model can capture long-range dependencies in the text sequence and understand the semantic information of the preamble text. The context vector contains the semantic information of the preamble text and provides an important basis for subsequently selecting the most appropriate candidate character. If the preamble text is empty, it is initialized with a zero vector to ensure the integrity of the process. The attention mechanism is used for fusion, which can effectively capture the correlation between the candidate character and the context. The attention mechanism enables the model to focus on the preamble text information most relevant to the current candidate character, thereby more accurately determining which candidate character is the most appropriate. The fused feature vector contains the comprehensive information of the candidate character and the context and provides a more comprehensive basis for subsequent optimal character prediction. Through the fully connected layer and the softmax layer, the fused feature vector is converted into a probability distribution, and the character with the highest probability is selected as the predicted character. This step completes the conversion from feature information to character recognition and is the last step of the entire recognition process. The predicted character is added to the end of the preamble text sequence to construct a new corrected character sequence. The corrected character sequence contains the currently recognized character and all previously recognized characters and is the accumulation of the recognition results. The recognition result is converted into a user-readable text format and displayed.

[0065] Preferably, the present invention further provides a picture recognition system for a scanning pen, which is used to execute the picture recognition method for a scanning pen as described above. The picture recognition system for a scanning pen includes:

[0066] A semantic pre-segmentation module, which is used to obtain the original image; generate superpixels for the original image to obtain a superpixel set; train a semantic segmentation network based on the addition of a superpixel perception layer according to a pre-constructed initial U-Net network model to obtain a semantic segmentation model; use the semantic segmentation model and the superpixel set to perform superpixel semantic category prediction and construct a superpixel semantic map to obtain a superpixel semantic map;

[0067] A stroke reconstruction module, which is used to extract stroke segments from the superpixel semantic map and construct a stroke connection map to obtain a stroke connection map; perform stroke constraints on the stroke connection map and perform stroke reconstruction to obtain a reconstructed stroke set; generate a stroke structure map according to the reconstructed stroke set to obtain a stroke structure map;

[0068] A glyph matching module, configured to obtain font samples; construct a multi-scale stroke structure graph library based on the font samples, and construct a glyph feature library to obtain the multi-scale stroke structure graph library and the glyph feature library; calculate the graph similarity of the multi-scale stroke structure graph library and the glyph feature library to obtain a character similarity matrix; perform glyph matching according to the character similarity matrix to obtain a candidate character set;

[0069] A context correction module, configured to obtain the previous text; perform vector context information fusion on the candidate character set and the previous text to obtain a fused feature vector; perform optimal character prediction on the fused feature vector to obtain a predicted character; splice the predicted character into the previous text to construct a corrected character sequence, and obtain a corrected character sequence; output the text result of the corrected character sequence to obtain a display result. Description of the Drawings

[0070] Figure 1 It is a schematic flow chart of the steps of a picture recognition method for a scanning pen;

[0071] Figure 2 It is a schematic detailed implementation step flow chart of step S2 in the present invention.

[0072] The implementation, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments

[0073] The technical method of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0074] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.

[0075] It should be understood that although terms such as "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly, the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.

[0076] To achieve the above object, please refer to Figures 1 to 2 , a picture recognition method for a scanning pen, comprising the following steps:

[0077] Step S1: Obtain the original image; generate superpixels for the original image to obtain a superpixel set; train a semantic segmentation network based on the pre-constructed initial U-Net network model with a superpixel perception layer added to obtain a semantic segmentation model; use the semantic segmentation model and the superpixel set to perform superpixel semantic category prediction and construct a superpixel semantic map to obtain a superpixel semantic map;

[0078] Step S2: Extract stroke segments from the superpixel semantic map and construct a stroke connection map to obtain a stroke connection map; perform stroke constraints on the stroke connection map and perform stroke reconstruction to obtain a reconstructed stroke set; generate a stroke structure map according to the reconstructed stroke set to obtain a stroke structure map;

[0079] Step S3: Obtain font samples; construct a multi-scale stroke structure map library and a glyph feature library according to the font samples to obtain a multi-scale stroke structure map library and a glyph feature library; calculate the graph similarity of the multi-scale stroke structure map library and the glyph feature library to obtain a character similarity matrix; perform glyph matching according to the character similarity matrix to obtain a candidate character set;

[0080] Step S4: Obtain the previous text; fuse the vector context information of the candidate character set and the previous text to obtain a fused feature vector; perform optimal character prediction on the fused feature vector to obtain a predicted character; splice the predicted character into the previous text to construct a corrected character sequence to obtain a corrected character sequence; output the text result of the corrected character sequence to obtain a display result.

[0081] In the embodiment of the present invention, referring to Figure 1 shown, it is a schematic flow chart of the steps of the picture recognition method for a scanning pen of the present invention. In this example, the picture recognition method for a scanning pen includes the following steps:

[0082] Step S1: Obtain the original image; generate superpixels for the original image to obtain a superpixel set; train a semantic segmentation network by adding a superpixel perception layer to a pre-constructed initial U-Net network model to obtain a semantic segmentation model; use the semantic segmentation model and the superpixel set to perform superpixel semantic class prediction and construct a superpixel semantic map to obtain a superpixel semantic map;

[0083] In the embodiment of the present invention, the original image collected by the scanning pen camera is obtained, and then the SLIC algorithm is used to generate a superpixel set, and the image is segmented into multiple small regions. Then, the pre-constructed initial U-Net network model is used, and a superpixel perception layer is added to its encoder part, and the semantic segmentation network is trained using the training data set to obtain a semantic segmentation model. The superpixel perception layer integrates the feature information of the superpixels into the U-Net model, improving the segmentation accuracy of the model for the text region. After training, the trained semantic segmentation model is used to perform semantic class prediction on the superpixels of the original image, and the prediction results are constructed into a superpixel semantic map, where each superpixel is labeled as "text", "background", or "noise".

[0084] Step S2: Extract stroke segments from the superpixel semantic map and construct a stroke connection map to obtain a stroke connection map; perform stroke constraints on the stroke connection map and perform stroke reconstruction to obtain a reconstructed stroke set; generate a stroke structure map according to the reconstructed stroke set to obtain a stroke structure map;

[0085] In the embodiment of the present invention, the superpixel semantic map is processed to extract the superpixels in the text region and decompose them into stroke segments. Then, a stroke connection map is constructed, where the nodes of the map represent the stroke segments and the edges represent the connection possibilities between the segments. The connection possibilities are calculated from features such as the direction, distance, gray value difference, and endpoint distance ratio of the stroke segments. Next, stroke constraints (length, angle, distance, and loop constraints) and the Kruskal algorithm are applied to optimize the stroke connection map to obtain a minimum spanning tree, and the stroke segments are merged according to the minimum spanning tree to reconstruct the complete strokes, forming a reconstructed stroke set. Finally, a stroke structure map is generated according to the reconstructed stroke set for subsequent glyph matching.

[0086] Step S3: Obtain font samples; construct a multi-scale stroke structure map library and a glyph feature library according to the font samples to obtain a multi-scale stroke structure map library and a glyph feature library; calculate the graph similarity of the multi-scale stroke structure map library and the glyph feature library to obtain a character similarity matrix; perform glyph matching according to the character similarity matrix to obtain a candidate character set;

[0087] In an embodiment of the present invention, a font sample is obtained, and a multi-scale stroke structure graph library and a glyph feature library are constructed. The multi-scale stroke structure graph library contains character stroke structure graphs of different fonts, font sizes, and scaling ratios, and the glyph feature library stores the feature vectors of each character. Then, the similarity between the input stroke structure graph and each character in the glyph feature library is calculated. The similarity calculation includes sub-steps such as scale matching, definition of important stroke structures, extraction of local structure features, and graph similarity calculation. Scale matching selects the scale closest to the area of the input stroke structure graph; the definition of important stroke structures determines the stroke type based on features such as stroke direction, length, and shape; local structure feature extraction calculates the similarity and relative position relationship of important stroke structures; graph similarity calculation combines structural similarity and local feature similarity. Finally, candidate characters are screened according to the similarity scores to obtain a candidate character set.

[0088] Step S4: Obtain the previous text; perform vector context information fusion on the candidate character set and the previous text to obtain a fused feature vector; perform optimal character prediction on the fused feature vector to obtain a predicted character; splice the predicted character into the previous text to construct a corrected character sequence; perform text result output on the corrected character sequence to obtain a display result;

[0089] In an embodiment of the present invention, the previous text recognized by the reading pen is obtained. Then, the candidate character set is feature-encoded using a pre-trained character embedding model to obtain candidate character feature vectors, and the similarity scores are added to the feature vectors. At the same time, a pre-trained Transformer language model is used to encode the previous text to obtain a context vector. Next, the attention mechanism is used to fuse the candidate character feature vectors and the context vector to obtain a fused feature vector. Then, the fused feature vector is input into a classifier for prediction, and the character with the highest probability is obtained as the predicted character. The predicted character is spliced into the previous text to construct a corrected character sequence. Finally, the corrected character sequence is converted into a text format and displayed on the screen of the reading pen or transmitted to a connected device.

[0090] Preferably, step S1 includes the following steps:

[0091] Step S11: Obtain the original image through the reading pen camera; perform superpixel generation on the original image to obtain a superpixel set;

[0092] Step S12: Extract superpixel features according to the superpixel set and the original image to obtain superpixel feature vectors;

[0093] Step S13: Obtain a training data set; add a superpixel perception layer to the pre-constructed initial U-Net network model and perform semantic segmentation network training according to the training data set to obtain a semantic segmentation model;

[0094] Step S14: Input the superpixel feature vector into a semantic segmentation model for superpixel semantic class prediction to obtain the superpixel semantic class;

[0095] Step S15: Construct a superpixel semantic map based on the superpixel set and the superpixel semantic class to obtain the superpixel semantic map.

[0096] In the embodiment of the present invention, after the scanning pen is started, the built-in camera captures an image of the text to be recognized. The camera is set to the autofocus mode to ensure clear images are captured. The obtained original image is an RGB color image with a resolution of 1200x800 pixels. The SLIC (Simple Linear Iterative Clustering) algorithm is used to perform superpixel segmentation on the original image. The number of superpixels is set to 500, and the compactness parameter is set to 20. The SLIC algorithm groups image pixels into perceptually uniform superpixels through iterative clustering, and each superpixel represents a small region with similar color and spatial position. The algorithm regards the image as a five-dimensional space (Lab color space and xy coordinate space), and iteratively optimizes the superpixel boundaries by minimizing the distance between each pixel and the center of its belonging superpixel. Finally, 500 superpixels are obtained, and each superpixel contains the coordinate information of its boundary pixels.

[0097] For each superpixel generated in step S11, a series of features are extracted to describe its visual attributes. First, calculate the mean and standard deviation of the RGB color values of all pixels within each superpixel to obtain 6-dimensional color features. Second, perform convolution operations on the original image using 5 Gabor filters with different directions and 3 different scales (a total of 15), and then calculate the mean and standard deviation of the Gabor filter responses within each superpixel to obtain 30-dimensional texture features. Third, calculate the center point coordinates of each superpixel and normalize the coordinate values to the interval [0,1] to obtain 2-dimensional position features. Finally, calculate the area, perimeter, and shape factor (4π * area / perimeter²) of each superpixel to obtain 3-dimensional shape features. Concatenate the above color, texture, position, and shape features into a 41-dimensional feature vector to represent each superpixel. Finally, a 500x41 matrix is obtained, where each row represents the feature vector of a superpixel.

[0098] Prepare a text image dataset that includes various fonts, font sizes, backgrounds, and lighting conditions. Each image in the dataset is manually annotated, where each superpixel is labeled as "text", "background", or "noise". Pre-build an initial U-Net network model that includes an encoder and a decoder part. The encoder is used to extract image features, and the decoder is used to generate pixel-level segmentation results. Add a superpixel-aware layer after each convolutional layer in the encoder part of the U-Net. The superpixel-aware layer takes the 41-dimensional superpixel feature vector extracted in step S12 as input and maps it to the same dimension as the number of output channels of the convolutional layer through a fully connected layer. Then, use an attention mechanism to fuse the mapped superpixel features with the output features of the convolutional layer. The attention mechanism generates attention weights by calculating the correlation between the superpixel features and the convolutional layer features, and uses these weights to weight the convolutional layer features. Use a weighted cross-entropy loss function to train the improved U-Net model. The loss function assigns different weights according to the importance of different classes, and the "text" class is assigned a higher weight to improve the segmentation accuracy of the text region. Use the augmented training image set to train the model. The training process uses the Adam optimizer, the learning rate is set to 0.001, and the batch size is set to 32. The training process continues until the performance of the model on the validation set reaches stability.

[0099] Input the superpixel feature vector extracted in step S12 into the semantic segmentation model trained in step S13. The model outputs the probabilities that each superpixel belongs to the "text", "background", or "noise" class. Select the class with the highest probability as the predicted class of the superpixel. Finally, obtain a vector containing 500 elements, where each element represents the semantic class of a superpixel.

[0100] Create a matrix with the same size as the original image and initialize all elements to 0. According to the superpixel set generated in step S11, map the superpixel semantic classes predicted in step S14 into this matrix. For each superpixel, set the class of all pixels within its boundary to the predicted class of the superpixel. If a pixel belongs to the boundaries of multiple superpixels, select the class of the superpixel with the largest proportion as the class of this pixel. Finally, obtain a superpixel semantic map with the same size as the original image, where each pixel is labeled as "text", "background", or "noise".

[0101] Preferably, step S13 includes the following steps:

[0102] Step S131: Perform data augmentation on the training dataset to obtain an augmented training image set;

[0103] Step S132: Integrate the pre-constructed initial U-Net network model with a superpixel perception layer according to the superpixel feature vector to obtain a U-Net model with an integrated superpixel perception layer;

[0104] Step S133: Define a weighted cross-entropy loss function according to the enhanced training image set to obtain the weighted cross-entropy loss function;

[0105] Step S134: Train a semantic segmentation network according to the enhanced training image set, the weighted cross-entropy loss function, and the U-Net model with an integrated superpixel perception layer to obtain a semantic segmentation model.

[0106] In the embodiment of the present invention, a text image training data set containing annotation information is obtained, where each superpixel of each image is labeled as "text", "background", or "noise". Apply the following data augmentation operations to each image in the training data set: 1) Random rotation: Randomly rotate the image by an angle between -15 degrees and +15 degrees; 2) Random scaling: Randomly scale the image by a factor between 0.8 and 1.2; 3) Random cropping: Randomly crop a region from the image with a size ranging from 80% to 100% of the original image; 4) Random flipping: Horizontally flip the image with a probability of 50%; 5) Color jitter: Randomly adjust the brightness, contrast, saturation, and hue of the image, with an adjustment range of ±10%. Perform 5 random transformation combinations on each original image to generate 5 enhanced images, and add the enhanced images to the training data set. All augmentation operations keep the annotation information of the superpixels unchanged. Finally, an enhanced training image set containing the original images and the enhanced images is obtained, and its size is 6 times that of the original training data set.

[0107] Pre-construct an initial U-Net network model, which has an encoder and a decoder structure for pixel-level image segmentation. The encoder part contains multiple convolutional layers and downsampling layers, and the decoder part contains multiple convolutional layers and upsampling layers. Integrate a superpixel perception layer after each convolutional layer in the U-Net encoder part. The superpixel perception layer receives a 41-dimensional superpixel feature vector as input. First, use a fully connected layer to map the superpixel feature vector to the same dimension as the number of channels of the output feature map of the corresponding convolutional layer. Then, copy and expand the mapped superpixel features to the same spatial dimension as the output feature map of the convolutional layer, so that all pixels within each superpixel share the same superpixel feature. Finally, use a channel attention mechanism to fuse the expanded superpixel features with the output features of the convolutional layer. The channel attention mechanism generates channel weights by performing global average pooling, fully connected layer transformation, and sigmoid activation function operations on the superpixel features and the convolutional layer features, and then applies the channel weights to the output features of the convolutional layer to achieve weighted fusion of the features.

[0108] Define a weighted cross-entropy loss function to measure the difference between the superpixel semantic categories predicted by the model and the ground truth annotations. The formula for the loss function is as follows:

[0109] Loss = -∑_{c=1}^{C} w_c * y_c * log(p_c);

[0110] Where C represents the number of categories ("text", "background", and "noise", C = 3), w_c represents the weight of category c, y_c represents the ground truth annotation of category c (0 or 1), and p_c represents the probability of category c predicted by the model. According to the proportion of the number of pixels of each category in the augmented training image set, the weights of the "text", "background", and "noise" categories are set to 0.6, 0.3, and 0.1 respectively. This means that the model pays more attention to the prediction accuracy of the "text" category.

[0111] Use the augmented training image set generated in step S131, the weighted cross-entropy loss function defined in step S133, and the U-Net model with an integrated superpixel perception layer constructed in step S132 to train the semantic segmentation network. Divide the augmented training image set into a training set and a validation set with a ratio of 9:1. Use the Adam optimizer for training, set the learning rate to 0.0001, and set the batch size to 16. During the training process, input the images into the U-Net model, and the model outputs the semantic category probabilities of each superpixel. Calculate the weighted cross-entropy loss between the model output and the ground truth annotation, and use the backpropagation algorithm to update the model parameters. Every certain number of iterations, evaluate the model performance on the validation set and save the model with the best performance. The training process continues until the performance of the model on the validation set no longer improves or reaches the preset number of iterations. Finally, a trained semantic segmentation model is obtained.

[0112] Preferably, step S2 includes the following steps:

[0113] Step S21: Extract the text regions from the superpixel semantic map to obtain a set of text region superpixels;

[0114] Step S22: Extract stroke segments according to the set of text region superpixels and the original image to obtain a set of candidate stroke segments;

[0115] Step S23: Construct a stroke connection graph for the set of candidate stroke segments to obtain a stroke connection graph;

[0116] Step S24: Apply stroke constraints to the stroke connection graph and perform stroke reconstruction to obtain a set of reconstructed strokes;

[0117] Step S25: Generate a stroke structure graph according to the set of reconstructed strokes to obtain a stroke structure graph.

[0118] As an example of the present invention, refer to Figure 2 As shown, in this example, step S2 includes:

[0119] Step S21: Extract the text region from the superpixel semantic map to obtain a text region superpixel set;

[0120] In the embodiment of the present invention, the superpixel semantic map generated in step S1 is input. Each superpixel in the superpixel semantic map is traversed. The superpixels with the semantic category labeled as "text" are extracted to form a text region superpixel set. The boundary coordinates (e.g., the minimum bounding rectangle) of each text region superpixel and the corresponding set of pixel coordinates in the original image are recorded. The pixel values of the extracted text region superpixels are marked in the original image, and the pixel values of other regions remain unchanged, generating a new image, where the text region is displayed in the original color and the pixel values of the non-text regions are set to 0 for subsequent processing.

[0121] Step S22: Extract stroke segments according to the text region superpixel set and the original image to obtain a candidate stroke segment set;

[0122] In the embodiment of the present invention, the text region superpixel set extracted in step S21 and the original image are input. For each text region superpixel, first, the corresponding region of interest (ROI) is extracted from the original image according to its boundary coordinates. Then, the Otsu algorithm is applied to the ROI for binarization, converting the text pixels to 1 and the background pixels to 0. Next, the Zhang-Suen thinning algorithm is used to thin the binarized image to obtain a single-pixel-wide stroke skeleton. Connected component analysis is performed on the thinned image, and the 8-connected pixel sets are marked as a stroke segment. The center point coordinates, main direction (calculated by principal component analysis PCA), length, and pixel coordinate list of each stroke segment are calculated. The information of all stroke segments is stored in the candidate stroke segment set.

[0123] Step S23: Construct a stroke connection graph for the candidate stroke segment set to obtain a stroke connection graph;

[0124] In the embodiment of the present invention, the candidate stroke segment set extracted in step S22 is input. A graph structure is created, where each node represents a stroke segment. For each pair of stroke segments, the connection possibility between them is calculated. The calculation of the connection possibility is based on the following four features: 1) Direction similarity: Calculate the cosine similarity of the main direction vectors of the two stroke segments; 2) Distance: Calculate the minimum Euclidean distance between the end points of the two stroke segments; 3) Average gray value difference: Calculate the difference between the average gray values of the corresponding pixels of the two stroke segments in the original image; 4) End point distance ratio: The ratio of the minimum distance between the end points of the two stroke segments to the sum of the lengths of the two stroke segments. Normalize these four feature values to the interval [0, 1], and then use a pre-trained logistic regression model to calculate the connection possibility score. Take the connection possibility score as the weight of the edge in the graph. If the connection possibility score between two stroke segments is greater than a preset threshold (for example, 0.5), then add an edge connecting these two nodes in the graph, and the weight of the edge is the connection possibility score.

[0125] Step S24: Perform stroke constraints on the stroke connection graph and perform stroke reconstruction to obtain a reconstructed stroke set;

[0126] In the embodiment of the present invention, the stroke connection graph constructed in step S23 is input. Define stroke constraint conditions, including: 1) Length constraint: The length of the merged stroke shall not exceed a preset maximum length threshold (for example, half of the image width); 2) Angle constraint: The included angle between adjacent stroke segments shall not exceed a preset maximum angle threshold (for example, 150 degrees); 3) Loop detection: Avoid forming circular strokes. According to the weights of the edges of the stroke connection graph and the stroke constraint conditions, use the Kruskal algorithm to solve the weighted minimum spanning tree. During the construction of the minimum spanning tree, if adding an edge will cause a violation of the stroke constraint conditions, then skip this edge. According to the result of the minimum spanning tree, merge the connected stroke segments into complete strokes. After merging, recalculate the center point coordinates, main direction, length, curvature and other features of each stroke, and store this information in the reconstructed stroke set.

[0127] Step S25: Generate a stroke structure graph according to the reconstructed stroke set to obtain a stroke structure graph;

[0128] In the embodiment of the present invention, the reconstructed stroke set generated in input step S24 is taken. A graph structure is created, where each node represents a reconstructed stroke. For each pair of strokes, the spatial relationship between them is calculated. If two strokes intersect, an edge is added in the graph to connect the corresponding two nodes, and the intersection angle and position are recorded. If two strokes do not intersect but the distance between their endpoints is less than a preset threshold, an edge is also added in the graph, and the distance and direction between the two endpoints are recorded. The finally generated stroke structure graph contains nodes (representing strokes) and edges (representing the spatial relationship between strokes), and the attributes of the edges include information such as distance, direction, intersection angle, etc.

[0129] Preferably, step S24 includes the following steps:

[0130] Step S241: Perform stroke constraint definition based on the candidate stroke segment set to obtain stroke constraint definition data;

[0131] Step S242: Calculate the edge weights according to the stroke connection graph to obtain the edge weights;

[0132] Solve the weighted minimum spanning tree for the edge weights to obtain the minimum spanning tree;

[0133] Step S243: Perform multi-stroke segment connection processing on the candidate stroke segment set according to the minimum spanning tree to obtain the reconstructed stroke set.

[0134] In the embodiment of the present invention, the candidate stroke segment set is analyzed, and the following stroke constraint conditions are defined to guide the stroke reconstruction process:

[0135] 1. Length constraint: The length of the merged stroke shall not exceed half of the image width. Calculate the length of each candidate stroke segment, and set the length threshold to image width * 0.5. In the subsequent stroke merging process, if the length of the stroke obtained by merging two stroke segments exceeds this threshold, the merging is prohibited.

[0136] 2. Angle constraint: The included angle between two connected stroke segments shall not exceed 150 degrees. Calculate the main direction of each candidate stroke segment. In the subsequent stroke merging process, if the included angle between the main directions of two stroke segments is greater than 150 degrees, its connection priority is reduced. The angle calculation uses the vector dot product formula: cos(theta) = dot(v1, v2) / (norm(v1) * norm(v2)), where v1 and v2 are the main direction vectors of the two stroke segments.

[0137] 3. Distance constraint: The minimum Euclidean distance between the endpoints of two connected stroke segments shall not exceed 3 times the length of the shorter stroke segment. Calculate the minimum endpoint distance between each pair of stroke segments. If this distance exceeds 3 times the length of the shorter stroke segment, its connection priority is reduced.

[0138] 4. Loop Constraint: The stroke connection relationship shall not form a loop. During the stroke merging process, the union-find data structure is used to detect and avoid the formation of loops. If merging two stroke segments will form a loop, the merge is prohibited.

[0139] Input the stroke connection graph constructed in step S23 and the stroke constraints defined in step S241. For each edge in the connection graph, calculate its weight. The weight calculation formula is as follows:

[0140] weight = connection_probability * penalty_length * penalty_angle *penalty_distance;

[0141] Where:

[0142] connection_probability is the connection possibility score calculated in step S23;

[0143] penalty_length is the length constraint penalty factor. If the length exceeds the threshold after merging two stroke segments, then penalty_length = 0; otherwise, penalty_length = 1;

[0144] penalty_angle is the angle constraint penalty factor, penalty_angle = 1 - (angle - 150) / (180 - 150), where angle is the included angle between the main directions of the two stroke segments. If the included angle is less than 150 degrees, then penalty_angle = 1;

[0145] penalty_distance is the distance constraint penalty factor, penalty_distance = 1 - distance / (3 * min_length), where distance is the minimum distance between the endpoints of the two stroke segments, and min_length is the shorter length of the two stroke segments. If the distance is less than 3 times the length of the shorter stroke, then penalty_distance = 1.

[0146] After calculating the weights of all edges, use the Kruskal algorithm to solve the weighted minimum spanning tree. The Kruskal algorithm sorts all edges in ascending order of weight and then tries to add the edges to the spanning tree one by one. If adding an edge will cause a loop to form (detected using the union-find), then skip that edge. Repeat this process until all nodes are connected to the spanning tree.

[0147] Input the minimum spanning tree and candidate stroke segment set generated in step S242. According to the connection relationship of the minimum spanning tree, merge the connected stroke segments into complete strokes. Specifically, traverse each edge in the minimum spanning tree, and merge the pixel coordinate lists of the two stroke segments connected by the edge into a new list. Then, according to the merged pixel coordinate list, recalculate the center point coordinates, main direction, length, curvature and other features of the stroke. Store all the reconstructed stroke information in the reconstructed stroke set. If a stroke segment is not connected to any other stroke segment in the minimum spanning tree, add it as an independent stroke to the reconstructed stroke set.

[0148] Preferably, step S3 includes the following steps:

[0149] Step S31: Obtain font samples; extract stroke structure diagrams from the font samples and construct a multi-scale stroke structure diagram library to obtain a multi-scale stroke structure diagram library;

[0150] Step S32: Construct a glyph feature library according to the multi-scale stroke structure diagram library to obtain a glyph feature library;

[0151] Step S33: Calculate the graph similarity of the multi-scale stroke structure diagram library and the glyph feature library to obtain a character similarity matrix;

[0152] Step S34: Screen candidate characters from the character similarity matrix to obtain a list of candidate characters and similarity scores;

[0153] Step S35: Construct a candidate character set from the list of candidate characters and similarity scores to obtain a candidate character set.

[0154] In the embodiments of the present invention, obtain standard character font sample images including various fonts (for example, Song typeface, boldface, regular script, Times New Roman, Arial) and font sizes (for example, 12pt, 16pt, 24pt). For each character image of each font and font size, first perform binarization processing to convert character pixels to 1 and background pixels to 0. Then, use the same stroke extraction and reconstruction method as in step S2 to obtain the stroke structure diagram of the character. To construct a multi-scale stroke structure diagram library, scale the stroke structure diagram of each character by different ratios, for example, the scaling ratios are 0.5, 0.75, 1.0, 1.25, and 1.5. Store the scaled stroke structure diagrams in the multi-scale stroke structure diagram library, and record the character encoding (for example, Unicode encoding), font, font size, and scaling ratio corresponding to each stroke structure diagram.

[0155] Construct a glyph feature library using the multi-scale stroke structure library constructed in step S31. The glyph feature library stores the feature vectors of each character. For each stroke structure diagram in the multi-scale stroke structure library, extract the following features: 1) the number of strokes; 2) the length, direction, and curvature of each stroke; 3) the connection relationship between strokes (including connection type, distance, direction, and angle). Combine these features into a feature vector of a fixed length. Store the feature vector and its corresponding character code, font, font size, and scaling ratio in the glyph feature library.

[0156] Input the stroke structure diagram generated in step S2, the multi-scale stroke structure library constructed in step S31, and the glyph feature library constructed in step S32. First, according to the overall size of the stroke structure diagram (e.g., the area of the stroke bounding box), select the set of stroke structure diagrams in the multi-scale stroke structure library that is closest in scale. Then, for each selected stroke structure diagram, use the graph edit distance algorithm to calculate its similarity to the input stroke structure diagram. The graph edit distance algorithm calculates the minimum number of edit operations (e.g., adding nodes, deleting nodes, adding edges, deleting edges) required to transform one graph into another. The fewer the number of edit operations, the higher the similarity. Store the calculated similarity scores in a character similarity matrix, where the rows of the matrix correspond to the input stroke structure diagrams, the columns correspond to the characters in the glyph feature library, and the matrix element values are the corresponding similarity scores.

[0157] Input the character similarity matrix calculated in step S33. Set a similarity threshold (e.g., 0.8). Filter out the characters with similarity scores higher than the threshold, and store these characters and their corresponding similarity scores in a candidate character and similarity score list. Sort the list in descending order of similarity scores.

[0158] Input the candidate character and similarity score list selected in step S34. Select the top N characters (e.g., N = 5) with the highest scores and their corresponding similarity scores from the list to form a candidate character set. The candidate character set will be used as the input for step S4.

[0159] Preferably, step S33 includes the following steps:

[0160] Step S331: Calculate the area similarity of the multi-scale stroke structure library and perform scale matching according to the glyph feature library to obtain matching scale selection data;

[0161] Step S332: Obtain important stroke data; define the important stroke structure according to the important stroke data to obtain important stroke structure definition data;

[0162] Step S333: Extract local structure features from the multi-scale stroke structure library according to the important stroke structure definition data to obtain local structure feature vectors;

[0163] Step S334: Calculate the graph similarity according to the multi-scale stroke structure graph library, local structure feature vectors, and glyph feature library to obtain a character similarity matrix.

[0164] In the embodiment of the present invention, the inputs are the multi-scale stroke structure graph library and the glyph feature library. First, calculate the area of the input stroke structure graph. The area calculation method is: add up the areas of the minimum bounding rectangles of all strokes in the stroke structure graph. Then, for each character in the glyph feature library, calculate the area of its stroke structure graph at different scales. The area calculation method is the same as that of the input stroke structure graph. Next, calculate the absolute value of the difference between the area of the input stroke structure graph and the areas of the stroke structure graphs of each character in the glyph feature library at different scales. For each character, select the scale with the smallest absolute value of the area difference as the best matching scale of the character. Store the best matching scale and the corresponding absolute value of the area difference of each character in the matching scale selection data.

[0165] Obtain the predefined important stroke data, which contains the description information of various important stroke structures, such as: horizontal, vertical, left-falling stroke, right-falling stroke, dot, hook, etc. According to the important stroke data, define the judgment rules for important stroke structures. For example, the judgment rule for a horizontal stroke is: the main direction angle is between -15 degrees and 15 degrees, and the length is greater than 1.5 times the average stroke length; the judgment rule for a vertical stroke is: the main direction angle is between 75 degrees and 105 degrees, and the length is greater than 1.5 times the average stroke length; the judgment rule for a dot is: the length is less than 0.5 times the average stroke length. Store these judgment rules as important stroke structure definition data.

[0166] The inputs are the multi-scale stroke structure graph library and the important stroke structure definition data. For each stroke structure graph in the multi-scale stroke structure graph library, according to the rules defined in step S332, judge which important stroke structure (such as horizontal, vertical, left-falling stroke, right-falling stroke, dot, etc.) each stroke belongs to, and record the results. Then, count the number of occurrences of each important stroke structure in the stroke structure graph. In addition, calculate the average length, average direction, and average curvature of each important stroke structure. Combine these statistical information into a feature vector, that is, the local structure feature vector.

[0167] The inputs are the multi-scale stroke structure graph library, the local structure feature vector, and the glyph feature library. First, according to the matching scale selection data obtained in step S331, select the stroke structure graph of the best matching scale for each character from the multi-scale stroke structure graph library. Then, calculate the graph similarity between the input stroke structure graph and the stroke structure graphs of the best matching scale of each character in the glyph feature library. The graph similarity calculation method is as follows:

[0168] 1. Structural similarity calculation: Use the graph edit distance algorithm to calculate the structural similarity between two stroke structure graphs. The graph edit distance algorithm calculates the minimum number of edit operations (e.g., adding nodes, deleting nodes, adding edges, deleting edges) required to transform one graph into another. The fewer the edit operations, the higher the structural similarity.

[0169] 2. Local feature similarity calculation: Calculate the cosine similarity of the local structure feature vectors of two stroke structure graphs.

[0170] 3. Final similarity calculation: Perform a weighted average of the structural similarity and the local feature similarity to obtain the final graph similarity score. The weights can be set according to experience. For example, the weight of the structural similarity is 0.7, and the weight of the local feature similarity is 0.3.

[0171] Store the calculated similarity scores in the character similarity matrix. The rows of the matrix correspond to the input stroke structure graphs, the columns correspond to the characters in the glyph feature library, and the matrix element values are the corresponding similarity scores.

[0172] Preferably, step S333 is specifically:

[0173] Judge the stroke directions of the multi-scale stroke structure library to obtain stroke direction judgment data; judge the stroke lengths of the multi-scale stroke structure library to obtain stroke length judgment data; judge the stroke shapes of the multi-scale stroke structure library to obtain stroke shape judgment data;

[0174] Perform important stroke structure marking processing based on the stroke direction judgment data, the stroke length judgment data, and the stroke shape judgment data to obtain important stroke structure marks;

[0175] Calculate the important stroke structure similarities of the multi-scale stroke structure library and the glyph feature library based on the important stroke structure marks, and perform relative position calculations to obtain local feature calculation results;

[0176] Construct feature vectors for the local feature calculation results to obtain local structure feature vectors.

[0177] In the embodiments of the present invention, each stroke structure diagram in the multi-scale stroke structure diagram library is traversed. For each stroke, its main direction angle is calculated. The main direction angle calculation method is principal component analysis (PCA). The stroke pixel coordinates are regarded as two-dimensional data points, the covariance matrix is calculated, and the eigenvector corresponding to the largest eigenvalue of the covariance matrix is the main direction vector. The main direction vector is converted into an angle value. The angle value is stored in the stroke direction judgment data. Each stroke structure diagram in the multi-scale stroke structure diagram library is traversed. For each stroke, its length is calculated. The stroke length calculation method is to calculate the length of the stroke pixel contour. The length value is stored in the stroke length judgment data. Each stroke structure diagram in the multi-scale stroke structure diagram library is traversed. For each stroke, its shape complexity is calculated. The shape complexity calculation method is the ratio of the perimeter of the stroke pixel contour to the perimeter of its minimum bounding rectangle. The shape complexity value is stored in the stroke shape judgment data.

[0178] Important stroke structure marking: Based on the stroke direction judgment data, stroke length judgment data, and stroke shape judgment data, each stroke is marked as an important stroke structure. The rules are as follows:

[0179] Horizontal: The main direction angle is between -15 degrees and 15 degrees, and the length is greater than 1.5 times the average stroke length.

[0180] Vertical: The main direction angle is between 75 degrees and 105 degrees, and the length is greater than 1.5 times the average stroke length.

[0181] Left-falling stroke: The main direction angle is between -60 degrees and -30 degrees, and the shape complexity is greater than 1.2.

[0182] Right-falling stroke: The main direction angle is between 30 degrees and 60 degrees, and the shape complexity is greater than 1.2.

[0183] Dot: The length is less than 0.5 times the average stroke length.

[0184] Each stroke is marked as "horizontal", "vertical", "left-falling stroke", "right-falling stroke", "dot", or "other".

[0185] Calculation of important stroke structure similarity: For the input stroke structure diagram and the stroke structure diagrams of each character in the glyph feature library, calculate the similarity of the important stroke structures between them. First, count the number of each important stroke structure. Then, for each important stroke structure, calculate the absolute value of the difference in the number of this structure in the two stroke structure diagrams. Add up all the differences and take the average as the important stroke structure similarity score.

[0186] Relative position calculation: For the input stroke structure diagram and the stroke structure diagram of each character in the glyph feature library, calculate the relative position relationship of the important stroke structures between them. Divide the image plane into a 3x3 grid. Record the grid area where the center point of each important stroke is located. Calculate the difference between the grid areas where the center points of the same type of important strokes in the two stroke structure diagrams are located. Add up all the differences and take the average as the relative position score.

[0187] Feature vector construction: Combine the important stroke structure similarity score and the relative position score into a feature vector, that is, the local structure feature vector.

[0188] Preferably, step S4 includes the following steps:

[0189] Step S41: Obtain the previous text; use the pre-trained character embedding model to perform candidate character feature encoding on the candidate character set to obtain candidate character feature vectors;

[0190] Step S42: Use the pre-trained Transformer language model to perform previous text encoding on the previous text sequence to obtain a context vector;

[0191] Step S43: Perform context information fusion on the candidate character feature vectors and the context vector to obtain a fused feature vector;

[0192] Step S44: Perform optimal character prediction on the fused feature vector to obtain a predicted character;

[0193] Step S45: Concatenate the predicted character to the previous text to construct a corrected character sequence and obtain a corrected character sequence;

[0194] Step S46: Perform text format conversion on the corrected character sequence to obtain a formatted text; display the formatted text to obtain a display result.

[0195] In the embodiment of the present invention, the recognized previous text sequence is read from the cache of the scanning pen. If the scanning pen has just started working, the previous text sequence is empty. Use the pre-trained character embedding model (for example, the character-level embedding of BERT) to encode each character in the candidate character set to obtain candidate character feature vectors. The feature vector dimension of each candidate character is the hidden layer dimension of the pre-trained model (for example, 768). Append the similarity score of each candidate character calculated in step S3 to the corresponding feature vector to expand the feature vector dimension.

[0196] Using a pre-trained Transformer language model (e.g., BERT), encode the previous text sequence obtained in step S41 into a context vector. Specifically, input the previous text sequence into the BERT model to obtain the output hidden state. Take the output of the last time step of the hidden state as the context vector. The dimension of the context vector is the same as the dimension of the hidden layer of the BERT model (e.g., 768). If the previous text sequence is empty, initialize the context vector as a vector of all zeros.

[0197] Fuse the candidate character feature vector obtained in step S41 and the context vector obtained in step S42. The fusion method uses an attention mechanism. Specifically, use the context vector as the query, the candidate character feature vector as the key and value, and calculate the attention weights. Then, perform a weighted sum on the candidate character feature vector using the attention weights to obtain a fused feature vector. The dimension of the fused feature vector is the same as the dimension of the candidate character feature vector.

[0198] Input the fused feature vector obtained in step S43 into a classifier composed of a fully connected layer and a softmax layer. The output dimension of the fully connected layer is equal to the size of the candidate character set. The softmax layer converts the output of the fully connected layer into a probability distribution. Select the character with the highest probability as the predicted character.

[0199] Add the predicted character obtained in step S44 to the end of the previous text sequence obtained in step S41 to construct a new corrected character sequence.

[0200] Convert the corrected character sequence obtained in step S45 into a user-readable text format. For example, if the corrected character sequence contains Unicode encoding, convert it to the corresponding character. Display the formatted text result on the screen of the scanning pen or transmit it to a device connected to it.

[0201] Preferably, the present invention also provides an image recognition system for a scanning pen, which is used to execute the above-mentioned image recognition method for a scanning pen. The image recognition system for a scanning pen includes:

[0202] A semantic pre-segmentation module, which is used to obtain an original image; generate superpixels for the original image to obtain a superpixel set; train a semantic segmentation network based on the addition of a superpixel perception layer according to a pre-constructed initial U-Net network model to obtain a semantic segmentation model; use the semantic segmentation model and the superpixel set to perform superpixel semantic category prediction and construct a superpixel semantic map to obtain a superpixel semantic map;

[0203] A stroke reconstruction module, configured to extract stroke segments from a superpixel semantic map, construct a stroke connection map, and obtain a stroke connection map; perform stroke constraints on the stroke connection map, perform stroke reconstruction, and obtain a reconstructed stroke set; generate a stroke structure diagram based on the reconstructed stroke set, and obtain a stroke structure diagram;

[0204] A glyph matching module, configured to obtain font samples; construct a multi-scale stroke structure diagram library according to the font samples, and construct a glyph feature library, and obtain a multi-scale stroke structure diagram library and a glyph feature library; calculate the graph similarity of the multi-scale stroke structure diagram library and the glyph feature library, and obtain a character similarity matrix; perform glyph matching according to the character similarity matrix, and obtain a candidate character set;

[0205] A context correction module, configured to obtain previous text; perform vector context information fusion on the candidate character set and the previous text, and obtain a fused feature vector; perform optimal character prediction on the fused feature vector, and obtain a predicted character; splice the predicted character into the previous text to construct a corrected character sequence, and obtain a corrected character sequence; output the text result of the corrected character sequence to obtain a display result.

[0206] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the application document are intended to be included in the present invention.

[0207] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features invented herein.

Claims

1. A method for image recognition using a scanning pen, characterized in that: The following steps are involved: Step S1: obtaining an original image; generating superpixels for the original image to obtain a superpixel set; The pre-built initial U-Net network model is trained with a semantic segmentation network based on the superpixel perception layer to obtain a semantic segmentation model; the semantic segmentation model and the superpixel set are used to predict the superpixel semantic category and construct a superpixel semantic map to obtain a superpixel semantic map; Step S2: extracting stroke fragments from the superpixel semantic graph and constructing a stroke connection graph to obtain a stroke connection graph; performing stroke constraints on the stroke connection graph and reconstructing the strokes to obtain a reconstructed stroke set; generating a stroke structure graph based on the reconstructed stroke set to obtain a stroke structure graph; Step S3: obtaining font samples; constructing a multi-scale stroke structure library and a glyph feature library based on the font samples to obtain a multi-scale stroke structure library and a glyph feature library; performing graph similarity calculation on the multi-scale stroke structure library and the glyph feature library to obtain a character similarity matrix; performing glyph matching based on the character similarity matrix to obtain a candidate character set; Step S4: Obtain the preceding text; The candidate character set and the preceding text are vector-context-fused to obtain a fused feature vector. The feature vectors are fused to perform optimal character prediction to obtain predicted characters; the predicted characters are concatenated into the preceding text to construct a correction character sequence to obtain a correction character sequence; and the correction character sequence is output as a text result to obtain a display result.

2. The image recognition method for a scanning pen according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: obtaining an original image through a scanning pen camera; generating superpixels for the original image to obtain a superpixel set; Step S12: extracting superpixel features based on the superpixel set and the original image to obtain a superpixel feature vector; Step S13: obtaining a training data set; adding a superpixel perception layer to the pre-built initial U-Net network model, and performing semantic segmentation network training according to the training data set to obtain a semantic segmentation model; Step S14: inputting the superpixel feature vector into the semantic segmentation model to predict the superpixel semantic category to obtain the superpixel semantic category; Step S15: construct a superpixel semantic map according to the superpixel set and the superpixel semantic category to obtain a superpixel semantic map.

3. The image recognition method for a scanning pen according to claim 2, characterized in that: Step S13 includes the following steps: Step S131: performing data enhancement on the training data set to obtain an enhanced training image set; Step S132: integrating the superpixel perception layer of the pre-constructed initial U-Net network model according to the superpixel feature vector to obtain a U-Net model integrating the superpixel perception layer; Step S133: defining a weighted cross entropy loss function according to the enhanced training image set to obtain a weighted cross entropy loss function; Step S134: Perform semantic segmentation network training according to the enhanced training image set, the weighted cross entropy loss function, and the U-Net model with integrated superpixel perception layer to obtain a semantic segmentation model.

4. The image recognition method for a scanning pen according to claim 1, characterized in that: Step S2 includes the following steps: Step S21: extracting text regions from the superpixel semantic map to obtain a text region superpixel set; Step S22: extracting stroke segments according to the text area superpixel set and the original image to obtain a candidate stroke segment set; Step S23: constructing a stroke connection graph for the candidate stroke segment set to obtain a stroke connection graph; Step S24: performing stroke constraints on the stroke connection graph and reconstructing the strokes to obtain a reconstructed stroke set; Step S25: generating a stroke structure diagram according to the reconstructed stroke set to obtain a stroke structure diagram.

5. The image recognition method for a scanning pen according to claim 4, characterized in that: Step S24 includes the following steps: Step S241: defining stroke constraints according to the candidate stroke segment set to obtain stroke constraint definition data; Step S242: Calculate edge weights according to the stroke connection graph to obtain edge weights; Solve the weighted minimum spanning tree for edge weights to obtain the minimum spanning tree; Step S243: performing multi-stroke segment connection processing on the candidate stroke segment set according to the minimum spanning tree to obtain a reconstructed stroke set.

6. The image recognition method for a scanning pen according to claim 1, characterized in that: Step S3 includes the following steps: Step S31: obtaining a font sample; extracting a stroke structure graph from the font sample, and constructing a multi-scale stroke structure graph library to obtain a multi-scale stroke structure graph library; Step S32: constructing a font feature library according to the multi-scale stroke structure library to obtain a font feature library; Step S33: performing graph similarity calculation on the multi-scale stroke structure library and the character feature library to obtain a character similarity matrix; Step S34: Screening candidate characters from the character similarity matrix to obtain a list of candidate characters and similarity scores; Step S35: construct a candidate character set for the candidate characters and similarity score lists to obtain a candidate character set.

7. The image recognition method for a scanning pen according to claim 6, characterized in that: Step S33 includes the following steps: Step S331: performing area similarity calculation on the multi-scale stroke structure library, and performing scale matching according to the character feature library to obtain matching scale selection data; Step S332: Obtain important stroke data; define important stroke structures according to the important stroke data to obtain important stroke structure definition data; Step S333: extracting local structural features from the multi-scale stroke structure library according to the important stroke structure definition data to obtain a local structural feature vector; Step S334: Calculate graph similarity based on the multi-scale stroke structure library, the local structure feature vector and the glyph feature library to obtain a character similarity matrix.

8. The image recognition method for a scanning pen according to claim 7, characterized in that: Step S333 is specifically as follows: Perform stroke direction judgment on the multi-scale stroke structure library to obtain stroke direction judgment data; perform stroke length judgment on the multi-scale stroke structure library to obtain stroke length judgment data; perform stroke shape judgment on the multi-scale stroke structure library to obtain stroke shape judgment data; Performing important stroke structure marking processing according to the stroke direction judgment data, the stroke length judgment data and the stroke shape judgment data to obtain an important stroke structure mark; According to the important stroke structure marks, the important stroke structure similarity is calculated for the multi-scale stroke structure library and the character feature library, and the relative position is calculated to obtain the local feature calculation result; The feature vector is constructed based on the local feature calculation results to obtain the local structure feature vector.

9. The image recognition method for a scanning pen according to claim 1, characterized in that: Step S4 includes the following steps: Step S41: obtaining a preceding text; using a pre-trained character embedding model to encode candidate character features of a candidate character set to obtain a candidate character feature vector; Step S42: using the pre-trained Transformer language model to encode the preceding text sequence to obtain a context vector; Step S43: fusing the candidate character feature vector and the context vector with context information to obtain a fused feature vector; Step S44: performing optimal character prediction on the fused feature vector to obtain a predicted character; Step S45: splicing the predicted characters into the preceding text to construct a correction character sequence to obtain a correction character sequence; Step S46: converting the corrected character sequence into a text format to obtain a formatted text; and displaying the formatted text to obtain a display result.

10. An image recognition system for a scanning pen, characterized in that: The image recognition system for the scanning pen is used to perform the image recognition method for the scanning pen as claimed in claim 1, and the image recognition system for the scanning pen comprises: The semantic pre-segmentation module is used to obtain the original image; generate superpixels for the original image to obtain a superpixel set; train the semantic segmentation network based on the superpixel perception layer added according to the pre-built initial U-Net network model to obtain a semantic segmentation model; use the semantic segmentation model and the superpixel set to predict the superpixel semantic category, and construct a superpixel semantic map to obtain a superpixel semantic map; The stroke reconstruction module is used to extract stroke fragments from the superpixel semantic map and construct a stroke connection map to obtain a stroke connection map; to impose stroke constraints on the stroke connection map and reconstruct the strokes to obtain a reconstructed stroke set; and to generate a stroke structure map based on the reconstructed stroke set to obtain a stroke structure map; The glyph matching module is used to obtain font samples; construct a multi-scale stroke structure library based on the font samples, and construct a glyph feature library to obtain a multi-scale stroke structure library and a glyph feature library; perform graph similarity calculation on the multi-scale stroke structure library and the glyph feature library to obtain a character similarity matrix; perform glyph matching based on the character similarity matrix to obtain a candidate character set; The context correction module is used to obtain the previous text; fuse the vector context information of the candidate character set and the previous text to obtain a fused feature vector; predict the optimal character of the fused feature vector to obtain a predicted character; splice the predicted character into the previous text to construct a correction character sequence to obtain a correction character sequence; output the text result of the correction character sequence to obtain a display result.

Citation Information

Patent Citations

  • Scene text detection method based on superpixel stroke feature transformation and deep learning region classification

    CN108345850A

  • Text detection method and apparatus, and storage medium

    US20190188528A1