Water book character recognition method and system based on target detection
By using object detection-based methods, we achieved accurate localization and feature extraction of Shui script characters. Combined with semantic information processing and confidence calculation, we solved the problem of insufficient recognition compatibility in Shui script character recognition, and improved recognition accuracy and system stability.
Patent Information
- Application Number
- CN202511856316.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2045-12-10
AI Technical Summary
Existing technologies lack the ability to handle complex document backgrounds and diverse writing styles in Shui script character recognition, resulting in insufficient recognition compatibility, difficulty in adapting to dynamic reconstruction under multi-character combinations and complex contexts, and a lack of robust recognition and semantic correction capabilities.
By employing an object detection-based approach, through image preprocessing, text region recognition, character geometric deformation correction, semantic information processing, and confidence calculation, we can achieve accurate localization, feature extraction, and error correction of Shui script characters. Combined with a multi-level semantic annotation and phrase unit reconstruction mechanism, we can generate complete Shui script recognition samples.
It improves the accuracy and structural integrity of Shui script character recognition, ensures the sequential accuracy of recognition results and the integrity of data, and enhances processing efficiency and system stability.
Smart Images

Figure CN121281071A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and in particular to a method and system for recognizing Shui script characters based on object detection. Background Technology
[0002] Currently, the Shui script, a unique pictographic writing system of my country's ethnic minorities, is widely used in document collation, ethnic cultural research, and digital storage. However, its character recognition methods still have many limitations in dealing with complex document backgrounds, diverse writing styles, and the requirements of high-precision academic research. Most existing technologies rely on template matching or manual feature extraction, lacking the ability to jointly process large-scale character sets and complex semantic structures, resulting in insufficient recognition compatibility under different writing styles and mixed typesetting conditions. Methods relying solely on sequential segmentation and linear matching ignore the characteristics of Shui script characters, such as large differences in morphology and complex stroke connections, which easily leads to errors in character localization and phrase reconstruction. Furthermore, they lack a global correction mechanism for character spatial arrangement and semantic order, resulting in deficiencies in overall recognition efficiency and reliability. Existing general OCR-based solutions mostly focus on the detection of text regions and single-character recognition, using fixed recognition processes. They are difficult to adapt to dynamic reconstruction under multi-character combinations, phrase structural levels, and complex contexts, and lack robust recognition and semantic correction capabilities for the special application scenarios of Shui script. Summary of the Invention
[0003] Therefore, it is necessary for the present invention to provide a method and system for recognizing Shui script characters based on target detection, in order to solve at least one of the above-mentioned technical problems.
[0004] To achieve the above objectives, a method for recognizing Shui script characters based on object detection includes the following steps: Step S1: Acquire the Shui script image and perform image preprocessing to obtain the Shui script image to be processed; identify the text regions of the Shui script image to be processed; Step S2: Correct the geometric deformation of the characters based on the text region and identify the features of the Shui script characters; use the features of the Shui script characters to classify the characters in the Shui script image and determine the semantic information of the characters; Step S3: Reconstruct the character order based on character semantic information; identify the Shui script phrase structure in the Shui script image based on the character order, and generate a complete Shui script recognition sample; Step S4: Calculate the confidence score of the Shui script based on the complete Shui script recognition sample, and use the confidence score to correct erroneous Shui script characters and store the corrected Shui script characters.
[0005] Preferably, this specification also provides a Shui script character recognition system based on object detection, used to perform the Shui script character recognition method based on object detection as described above, the Shui script character recognition system based on object detection includes: The image preprocessing module is used to acquire Shui script images and perform image preprocessing to obtain Shui script images to be processed; and to identify the text regions in the Shui script images to be processed. The character classification module is used to correct the geometric deformation of characters based on the text region and identify the characteristics of Shui script characters; it uses the characteristics of Shui script characters to classify characters in Shui script images and determine the semantic information of the characters; The character order reconstruction module is used to reconstruct the character order based on character semantic information; and to identify the Shui script phrase structure in Shui script images based on the character order, generating complete Shui script recognition samples. The confidence calculation module is used to calculate the confidence of Shuishu based on the complete Shuishu recognition sample, and to use the confidence of Shuishu to correct erroneous Shuishu characters and store the corrected Shuishu characters.
[0006] The beneficial effects of this invention are: (1) By accurately identifying and segmenting the text region of the Shui script image and combining it with text feature correction technology, the accurate positioning and feature extraction of Shui script text were achieved, ensuring the reliability of subsequent character classification and semantic information determination.
[0007] (2) In the character semantic information processing stage, the semantic combination and structural hierarchy analysis of adjacent characters are carried out through multi-level semantic annotation and phrase unit reconstruction mechanism, which improves the integrity and logical consistency of the Shui script phrase structure generation and is suitable for diverse Shui script typesetting and writing styles.
[0008] (3) In the process of character order reconstruction and Shuishu phrase structure recognition, complete recognition samples are established by using character semantic information, and confidence calculation and weighted correction mechanism are combined to realize automatic correction of incorrectly recognized characters, thus ensuring the order accuracy and structural integrity of Shuishu recognition results.
[0009] (4) In the system security and data processing stage, the identification results are securely stored and traceable by the dynamic evaluation and storage management mechanism of the water book confidence level, ensuring the integrity, reliability and verifiability of the identification data, while improving the processing efficiency and system stability of the entire identification process. Attached Figure Description
[0010] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a flowchart illustrating the steps of a method for recognizing Shui script characters based on target detection according to the present invention. The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0011] The technical method of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this invention.
[0012] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.
[0013] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0014] To achieve the above objectives, please refer to Figure 1 This invention provides a method for recognizing Shui script characters based on object detection, the method comprising the following steps: Step S1: Acquire the Shui script image and perform image preprocessing to obtain the Shui script image to be processed; identify the text regions of the Shui script image to be processed; In one embodiment, an image of a Shui script manuscript captured by a high-definition scanner, with a resolution of 300 dpi, is used as input. First, the image is converted to grayscale, and then adaptive thresholding is used to eliminate background interference. Next, median filtering is used to remove noise, and histogram equalization is used to enhance text contrast. Subsequently, a convolutional neural network-based object detection algorithm (such as YOLOv5) is used to locate and identify candidate bounding boxes for the text regions, resulting in rectangular regions containing individual characters.
[0015] In another embodiment, assuming a 1200×800 pixel image of a water book document is acquired, 180 candidate character regions are detected after preprocessing, of which 160 are valid character regions after threshold filtering, with an average size of 40×40 pixels per character region, basically covering the entire page of text.
[0016] Step S2: Correct the geometric deformation of the characters based on the text region and identify the features of the Shui script characters; use the features of the Shui script characters to classify the characters in the Shui script image and determine the semantic information of the characters; In one embodiment, geometric correction is performed on each candidate character region, and perspective transformation is used to eliminate tilt distortion caused by the shooting angle. Then, the stroke shape is adjusted through contour extraction and affine transformation. Afterwards, the stroke direction histogram (HOG), edge skeleton features, and local texture features of the Shui script characters are extracted and fused with a convolutional neural network. The feature is then input into a classifier to complete character classification, and finally, the semantic information corresponding to the character is output.
[0017] In another embodiment, it is assumed that 160 character regions are corrected, of which 30 regions have significant tilt distortion. After geometric correction, the character feature recognition accuracy is improved from 82% to 93%. In the classification stage, a feature vector with a dimension of 256 is extracted using a three-layer CNN. After classification by a fully connected layer, 90 common water script characters are identified, with an accuracy of 91%.
[0018] Step S3: Reconstruct the character order based on character semantic information; identify the Shui script phrase structure in the Shui script image based on the character order, and generate a complete Shui script recognition sample; In one embodiment, the identified character set is ordered by position, and the character sequence is reconstructed using a left-to-right, top-to-bottom scanning rule. Combined with a language rule base, the combination relationships between consecutive characters are analyzed to recover the semantic structure of phrases or sentences.
[0019] In another embodiment, assuming that 10 out of the 160 identified characters have positional deviations, after sorting and phrase reconstruction, 25 phrase structures were successfully parsed, with each phrase containing an average of 6.4 characters, and the overall semantic coherence after reconstruction reached 95%.
[0020] Of particular importance is that step S3, which involves recognizing the structure of Shui script phrases in the Shui script image based on character order to generate a complete Shui script recognition sample, includes: The character order is used to perform semantic segmentation on the Shui script image to obtain the segmentation result; based on the segmentation result, adjacent characters in the Shui script image are combined into Shui script phrase units. In one embodiment, the input water book image is first converted to a grayscale image. The grayscale conversion process uses linear grayscale mapping, and the image pixel values are normalized to the range of 0 to 255. Subsequently, the grayscale image is subjected to two-dimensional Gaussian smoothing, with a filter kernel size of 5×5 pixels and a standard deviation of [missing value]. The threshold is set to 1.2 to suppress image noise and subtle interference. An adaptive thresholding algorithm is used to binarize the smoothed grayscale image, with a local threshold window size of 25×25 pixels and an offset of 10. Pixel values greater than the threshold are designated as foreground, and those below the threshold as background, thus generating a preliminary binary image of the character outline. Connectivity analysis is performed on the binary image to extract the minimum bounding rectangle of each connected region. The aspect ratio (rectangle length / width) and area are calculated, with aspect ratio thresholds set between 0.3 and 5, and area thresholds between 50 and 2000 pixels, eliminating noise regions that do not meet the criteria. The remaining connected regions are sorted in ascending order by the Y-coordinate of their center points to achieve row segmentation, and then sorted in ascending order by the X-coordinate to complete the in-row sequential indexing. Subsequently, the horizontal spacing between adjacent center points of characters within a row is calculated, the average spacing value for each row is statistically analyzed, and adjacent characters with a spacing less than a set threshold of 20 pixels are grouped into a phrase unit. The number of each phrase unit is assigned from left to right along the row. After the combination is completed, the minimum bounding rectangle of each phrase unit is calculated as the bounding box, and the character index sequence, row number and phrase unit number are recorded for subsequent structural analysis and sample generation.
[0021] The hierarchical relationship of the Shui script phrase unit is labeled; the complete Shui script phrase structure is generated according to the hierarchical relationship; the complete Shui script phrase structure is converted into a sample format and the complete Shui script recognition sample is output.
[0022] In one embodiment, the center point of the bounding box and the horizontal and vertical projection positions of each phrase unit are calculated to construct a spatial relationship matrix between phrase units. For each phrase unit in the matrix, the horizontal offsets to its upper and lower neighbors and its left and right neighbors are calculated. and vertical offset ,in , The rules for determining hierarchical relationships are as follows: If If it is a hierarchical relationship, then it is marked as a superior-subordinate relationship; if If the relationship is left-right, it is labeled as a left-right hierarchical relationship. This rule is followed to traverse all phrase units, forming a complete hierarchical relationship tree. Subsequently, the hierarchical tree is converted into structured coded samples. The encoding of each phrase unit includes the unit number, character sequence number, hierarchical relationship identifier (0 for top and bottom, 1 for left and right), and bounding box coordinates. All encoded information is saved as a JSON file for subsequent text recognition or system analysis.
[0023] Most importantly, the hierarchical relationship of the phrase unit annotation structure based on the Shui script includes: A unique number is assigned to each Shui script phrase unit to determine its hierarchical relationship; a hierarchical structure is established based on the hierarchical relationship; and semantic attributes are added to each Shui script phrase unit in the hierarchical structure. In one embodiment, the Shui script phrase units generated in the Shui script image are assigned unique integer numbers in a top-to-bottom, left-to-right order, starting from 1. The coordinates of the center point of the bounding box are calculated for each phrase unit. and the size of the circumscribed rectangle And calculate the average width of all phrase units. and average height Based on the horizontal offset between center points and vertical offset Determine the subordinate relationship between phrase units: If and Then the phrase unit is labeled as a subordinate relationship; if and If the elements are parallel, they are labeled as being on the same level. Based on the above hierarchical relationships, a hierarchical structure of Shui script phrase units is constructed. The root node corresponds to the phrase unit at the top of the image, and the lower-level nodes are units that have a hierarchical relationship with it vertically or horizontally. In the hierarchical structure, semantic attributes are added to each Shui script phrase unit, including the character sequence index within the phrase unit, the line number, and the bounding box coordinates. This includes identifying text types, such as independent characters, superimposed characters, or connected characters. This operation records the numbering, hierarchical relationships, and semantic attributes of each Shui script phrase unit, providing a data foundation for subsequent structural analysis and identification.
[0024] The hierarchical relationship of the structure is marked according to semantic attributes.
[0025] In one embodiment, the semantic attributes of each Shui script phrase unit in the hierarchical structure are read, including character sequence index, line number, bounding box coordinates, and text type. For each node, its spatial positional relationship with adjacent nodes is calculated, including horizontal spacing. Vertical spacing And the overlap area of the bounding boxes. If If, then mark the subordinate relationship of the lower level; if and and Then, the parallel and same-layer relationships are marked. The hierarchical structure tree is traversed, and the structural annotation results of all phrase units are recorded in the structure coding table. Each record contains the phrase unit number, parent node number, list of child node numbers, hierarchy depth, and semantic attribute fields. After annotation is completed, the structure coding table is exported as a JSON or XML file, with fields including number, parent node, list of child nodes, hierarchy depth, and semantic attributes, providing complete sample input for the water book recognition system.
[0026] Step S4: Calculate the confidence score of the Shui script based on the complete Shui script recognition sample, and use the confidence score to correct erroneous Shui script characters and store the corrected Shui script characters.
[0027] In one embodiment, the system assigns a confidence value to each recognized character and quantifies it based on the output probability of Softmax. For example, when the confidence value is below 0.7, the character is marked as needing correction, and then error correction is performed using a contextual phrase language model, replacing it with a more semantically appropriate candidate character. Finally, the corrected result is stored in a database for subsequent retrieval or comparison.
[0028] In another embodiment, assuming that among the 160 generated character samples, the system calculates an average confidence level of 0.87, with 20 characters below the threshold of 0.7. Through a phrase context-based error correction algorithm, 15 characters are correctly corrected, and the remaining 5 characters retain their original results, ultimately improving the overall recognition accuracy from 91% to 96%.
[0029] Of particular importance, step S4, which calculates the confidence level of the Shui Book based on the complete Shui Book identification sample, includes: Perform character-by-character matching on the complete Shuishu (water script) recognition sample and calculate the matching probability; calculate the confidence of phrase units in the complete Shuishu recognition sample based on the matching probability. In one embodiment, characters within each phrase unit of a complete water script recognition sample are processed sequentially according to their bounding boxes in the image. First, the image region of each character is converted to a grayscale image and binarized using a fixed threshold of 128. Pixels above the threshold are considered foreground, and the remaining pixels are considered background. Then, a pixel-level skeletonization algorithm is used to refine the strokes in the character image into skeletons of single-pixel width, and the coordinate information of the skeleton pixels is extracted. Based on this, each skeleton pixel is compared point-by-point with the corresponding skeleton of the standard template character to determine whether the pixel position offset is within two pixels. If the offset meets the requirement, the pixel is considered to have matched successfully. At the same time, grayscale differences are checked to ensure matching under similar grayscale conditions. After character-by-character processing, the matching probability of the character is determined based on the ratio of the number of successfully matched pixels to the total number of skeleton pixels of the character. Next, the matching probabilities of all characters within the phrase unit are comprehensively analyzed to obtain the phrase unit confidence score, and the corresponding number and matching result of each phrase unit are recorded to provide basic information for subsequent processing. During operation, the width and height of the bounding box for each character are measured with pixel-level precision, and the pixel offset range and grayscale threshold are strictly limited to ensure that the matching process is determined entirely based on the image pixel features.
[0030] The confidence scores of phrase units are normalized to obtain the confidence scores to be processed; the confidence scores to be processed are then weighted and accumulated to determine the confidence scores of the Shuishu (water book) system.
[0031] In one embodiment, the confidence scores of all phrase units are sorted as a whole, and the maximum and minimum values in the sequence are identified. The confidence score of each phrase unit is then mapped to a range of 0 to 1. Next, weights are assigned based on the number of characters contained in the image within each phrase unit; phrase units with more characters receive higher weights, and those with fewer characters receive lower weights, while ensuring that the sum of all weights equals the total number of characters. During the weighting process, the normalized confidence score of each phrase unit is multiplied by its weight, and the results of all phrase units are accumulated to obtain the total confidence score of the entire water-book recognition sample. The entire operation strictly follows the number of phrase units and character distribution, requiring no external model; it is calculated entirely based on pixel-level matching results and character quantity weights.
[0032] Preferably, the text regions identified in step S1 of the water script image to be processed include: Generate text candidate boxes using a pre-defined convolutional feature extractor; calculate the response value of each text candidate box, perform character confidence prediction, and output the confidence score of each text candidate box; In one embodiment, a feature extractor based on a convolutional neural network (such as a ResNet-50 backbone network) is used to generate candidate regions by sliding a window on the input water book document image. 512-dimensional convolutional features are extracted for each candidate region, and text candidate boxes are generated through a Region Proposal Network (RPN). For each candidate box, its response value is calculated and input into a fully connected layer for classification prediction, outputting the class probability of the corresponding character as the confidence score of the candidate box.
[0033] In another embodiment, assuming a 1024×768 pixel image of a water script is processed, 500 candidate boxes are generated after convolutional feature extraction. After screening, 320 candidate boxes with response values greater than a threshold of 0.6 are selected and their confidence scores are output through a fully connected layer, with an average confidence score of 0.82, a highest confidence score of 0.97, and a lowest of 0.45. Finally, candidate boxes with confidence scores below 0.6 are discarded, retaining only 300 valid candidate boxes for subsequent recognition.
[0034] Potential text regions in the water script image to be processed are filtered and identified based on confidence level; rotation correction is performed on the potential text regions to output the text regions.
[0035] In one embodiment, candidate boxes with confidence scores greater than a set threshold are identified as potential text regions based on the confidence scores. For these regions, the rotation angle is calculated using the minimum bounding rectangle method, and rotation correction is performed through affine transformation to align the text lines with the horizontal axis of the image. After correction, a regularized text region is output for subsequent character recognition.
[0036] In another embodiment, assuming there are 300 candidate boxes, 240 of them have a confidence level greater than 0.75 and are identified as potential text regions. Further detection revealed that 60 of these regions were tilted, with an average rotation angle of 8.5° and a maximum rotation angle of 17°. After correction using affine transformation, the tilt angle of all text regions was compressed to within ±1°, thus ensuring high geometric consistency of the characters in subsequent recognition stages.
[0037] Preferably, rotating and correcting the potential text region to output the text region includes: Using the center point coordinates of the potential text region as the rotation reference point, establish an affine transformation matrix; determine the text direction angle of the potential text region, write the text direction angle as a rotation parameter into the affine transformation matrix, and keep the scaling ratio and aspect ratio unchanged; In one embodiment, the bounding box of the detected potential text region is obtained, and the reference point for rotation is obtained by calculating the center position of the bounding box. Subsequently, the text orientation in the region is analyzed, for example, by using the long side direction of the minimum bounding rectangle to determine the tilt angle of the current text relative to the horizontal line. This angle is written as a rotation parameter into the affine transformation matrix, and while maintaining the scaling ratio and aspect ratio of the original region, the final rotation mapping relationship is constructed for subsequent text region correction.
[0038] In another embodiment, assuming the bounding box of the potential text region is located at the top left corner (100, 200) and the bottom right corner (260, 280) of the image coordinates, the center point is located at (180, 240). Directional analysis reveals that the directional angle of this text region is approximately 12 degrees. Constructing an affine transformation matrix based on the center point ensures that the text region will not be stretched or compressed during rotation, maintaining a stable aspect ratio.
[0039] The original pixels of the potential text region are mapped to their coordinates using an affine transformation matrix, and each original pixel is mapped to a rotated position to output the text region.
[0040] In one embodiment, an affine transformation matrix is used to map the coordinates of each pixel in the potential text region, transforming pixels at their original positions to new rotated positions. For pixels that fall into non-integer coordinates after rotation, interpolation is used for compensation to ensure that the corrected text edges are clear and the transitions are smooth. The final output is an image of a regularized text region that has been rotated to the standard orientation.
[0041] In another embodiment, assuming the text region is 160×80 pixels, after coordinate mapping, the overall average pixel offset is approximately 5 pixels, with a maximum offset of 12 pixels. Through interpolation compensation, the resulting output image remains within the 160×80 pixel resolution range, and the tilt in the text line direction is reduced from the original 12 degrees to less than 1 degree, ensuring the accuracy of the text region in subsequent segmentation and recognition stages.
[0042] Preferably, the identification of Shui script features in step S2 includes: A pre-set neural network model is used to identify the features of Shui script characters; In one embodiment, the acquired Shui script image is preprocessed, including grayscale conversion, normalization, and noise reduction. Subsequently, the processed image is input into a pre-defined neural network model, which uses the model's forward propagation process to extract features from potential text regions in the image. During training, the model has learned the stroke structure, writing habits, and common forms of Shui script characters, thus effectively identifying the main feature points and boundaries of the characters.
[0043] In another embodiment, the input water script image is assumed to be 256×256 pixels in size, containing approximately 30 characters. After preprocessing, the number of noise points is reduced by about 40%, and the character outlines are clearer. When input into the neural network, the model detects approximately 500 low-level feature points (such as edges and corners) in the first layer, and after deep convolutions, it selects approximately 120 stable high-level structural feature points. These feature points provide a reliable input foundation for subsequent convolutional and classification layers.
[0044] The neural network model includes an input layer, multiple convolutional feature extraction layers, a feature fusion layer, and a classification decision layer. The input layer receives the Shui script image of the text region. The convolutional feature extraction layer extracts multi-scale text features from the Shui script image. The feature fusion layer upsamples and weights the text features at different levels to obtain high-dimensional comprehensive features. The classification decision layer performs character-by-character feature matching on the high-dimensional comprehensive features.
[0045] In one embodiment, the input layer of the neural network receives an image of the Shui script text region and converts it into a tensor input. The convolutional feature extraction layer, composed of multiple convolutional and pooling units, captures the detailed features and overall shape of the strokes layer by layer. Subsequently, the feature fusion layer upsamples and weights the feature maps extracted from different convolutional layers, unifying the representation of local detailed features and global structural features. Finally, the classification decision layer inputs the high-dimensional integrated features into a fully connected network to complete character-by-character classification and recognition, outputting the Shui script text category corresponding to each character.
[0046] In another embodiment, the neural network is assumed to contain 5 convolutional extraction layers, 1 feature fusion layer, and 2 classification decision layers. When an image of a traditional Chinese calligraphy style containing 30 characters is input, the convolutional layers output feature maps with a total of 64, 128, 256, 512, and 1024 channels, respectively. After feature fusion, a high-dimensional feature vector of 2048 dimensions is formed. This feature vector is assigned character-by-character in the classification layer, ultimately outputting the category predictions for the 30 characters, with an average recognition accuracy exceeding 92%. In single-character classification, the recognition accuracy for characters with complex strokes is approximately 89%, while the recognition accuracy for characters with simpler strokes can reach over 95%.
[0047] Preferably, extracting multi-scale text features from Shui script images includes: The first convolutional layer of the neural network model performs a primary convolution operation on the input Shuishu image to extract low-level edge features; the low-level edge features are then input into the second convolutional layer to extract stroke features. In one embodiment, after acquiring the input water script image, it is first normalized to a fixed size (e.g., 128×128 pixels) and noise is filtered out. It is then input into the first convolutional layer of the neural network, where a set convolutional kernel (e.g., 3×3 size) extracts character edge information to obtain a low-level edge feature map. This feature map mainly reflects the outline and basic direction of the characters. Based on this, the low-level edge features are input into a second convolutional layer to further capture local features such as stroke direction, thickness, and intersection points, laying the foundation for subsequent shape feature extraction.
[0048] In another embodiment, the input water script image is assumed to be a 128×128 pixel character image containing approximately 15 characters. After processing by a first convolutional layer with 32 kernels, the output feature map size is 64×64, yielding approximately 131,000 low-level edge feature points. After processing by a second convolutional layer with 64 kernels, the feature map is reduced to 32×32, extracting approximately 65,000 stroke-related feature points. The low-level edge features are mainly concentrated in the outline region of the characters, while the stroke features highlight the local structures such as vertical strokes, horizontal strokes, and folded strokes within the characters, providing distinguishable information for subsequent shape analysis.
[0049] The stroke features are input into the third convolutional layer to extract character shape features; the character shape features are then input into the fourth convolutional layer to extract high-level semantic features.
[0050] In one embodiment, after the stroke features are output by the second convolutional layer, they are input into the third convolutional layer. Convolution kernels with a larger receptive field are used to globally capture the stroke combinations, thereby obtaining character shape features that reflect the overall outline and structural layout of the characters. Subsequently, the character shape features are input into the fourth convolutional layer for further abstraction processing to extract high-level features that can represent the semantic units of the Shui script, such as common character radicals or compound stroke patterns, to support subsequent character recognition and phrase construction.
[0051] In another embodiment, it is assumed that the third convolutional layer uses 128 convolutional kernels to perform convolution and pooling operations on the input 32×32 feature map, obtaining a character shape feature map of size 16×16, with approximately 32,768 high-level nodes in total; then it is input into the fourth convolutional layer, where the number of convolutional kernels is 256, and the output feature map is of size 8×8, with 16,384 nodes in total. It is found in the experiment that the third convolutional layer can distinguish approximately 90% of the overall character shapes, for example, distinguishing "日" and "目" through shape features; on this basis, the fourth convolutional layer further captures the semantic composite structures in the Shui script, enabling characters with relatively high similarity to be effectively distinguished, and its classification accuracy is increased by approximately 7%.
[0052] Preferably, in step S2, using the Shui script character features to classify the Shui script images and determining the character semantic information includes: Using the Shui script character features to extract the text structure features, performing sample annotation according to the text structure features, and dividing the text structure of the Shui script images to determine the independent stroke regions and combined stroke regions; In one embodiment, the input Shui script text images are preprocessed, including grayscale conversion, binarization, and noise point removal, to obtain clear stroke contours. Subsequently, convolution feature extraction operators are used to extract the edge and local morphological features of the text, and combined with the connected component analysis method, a preliminary stroke region division is obtained. According to the spatial adjacency between strokes, the strokes are divided into independent stroke regions (such as single horizontal, vertical, and dot strokes) and combined stroke regions (such as the "十" character or "口" shape formed by the connection of horizontal and vertical strokes). On this basis, sample annotation is completed manually or semi-automatically, providing a basis for subsequent stroke morphology recording and combination analysis.
[0053] In another embodiment, assuming that the input Shui script image size is 256×256 pixels, a total of 320 connected stroke segments are detected, of which about 210 are identified as independent stroke regions and 110 are divided into combined stroke regions. During further annotation, 75 "horizontal" strokes, 68 "vertical" strokes, 41 "dot" strokes, and 26 "slash and捺" strokes are recorded in the independent stroke regions; 24 "cross" structures, 18 "square" frames, and 5 "pin" character stacked structures are identified in the combined stroke regions. The obtained annotation set can provide stable data support for subsequent stroke attribute and character semantic extraction.
[0054] In the independent stroke region, record the isolated stroke form; in the combined stroke region, identify the connection relationship of adjacent strokes; determine the stroke attribute label according to the isolated stroke form; In one embodiment, for the independent stroke region, geometric form features (such as stroke length, width, inclination angle, endpoint position) are extracted one by one and recorded as the isolated stroke form. For the combined stroke region, by calculating the Euclidean distance and angular relationship between stroke endpoints, the connection method of adjacent strokes (such as intersection, parallel, closed) is identified. On this basis, an attribute label is assigned to each isolated stroke, such as "horizontal - short", "vertical - long", "dot - slanting右下", etc.
[0055] In another embodiment, assuming that a total of 200 isolated strokes are collected in the independent stroke region, where the length ranges from 5 to 40 pixels and the inclination angle distribution is within the range of ±75°. After attribute classification, 82 horizontal strokes (of which short horizontal strokes account for about 60%), 70 vertical strokes (of which long vertical strokes account for about 55%), 28 dot strokes, and 20 oblique strokes are obtained. For the combined stroke region, a total of 95 groups of stroke connection relationships are detected, among which the cross type (such as the "cross" structure) accounts for about 38%, the closed type (such as the "square" structure) accounts for about 22%, the parallel type (such as the "double vertical" structure) accounts for about 15%, and the rest are complex mixed types. This attribute and connection information provide constraint conditions for subsequent character structure synthesis.
[0056] Match the stroke type of the preset Shui script stroke library according to the adjacent stroke connection relationship; use the stroke attribute label and stroke type to gradually synthesize the complete character structure and determine the character semantic information.
[0057] In one embodiment, the stroke attribute and connection relationship are matched with the preset Shui script stroke library. The stroke library pre - stores common stroke types and their combination rules, for example, "horizontal + vertical" can form "十", and "vertical + square" can form "中". After the matching is completed, the dynamic synthesis algorithm is used to gradually synthesize the complete character structure, and according to the semantic rules of Shui script characters, the corresponding character semantic information is output.
[0058] In another embodiment, it is assumed that the stroke library contains about 50 common stroke types, covering basic units such as horizontal strokes, vertical strokes, left-falling strokes, right-falling strokes, dots, and turns, as well as their common combinations. In the test samples, a total of 120 stroke combinations were recognized. After hierarchical combination, 85 complete characters were formed. Among them, the characters directly formed by the combination of independent strokes accounted for about 65%, and the characters obtained by parsing complex combination areas accounted for about 35%. Among the finally recognized 85 characters, the correct rate of semantic matching was about 92%, and typical Shui script symbols such as "sacrifice" and "water" could be completely recognized. This method proves that the combination strategy of using stroke attributes and library matching has high accuracy and adaptability.
[0059] Preferably, in the area of independent strokes, recording the isolated stroke form includes: In the area of independent strokes, determine the starting coordinate and ending coordinate of the stroke; calculate the stroke direction angle according to the starting coordinate and ending coordinate, and determine the stroke curvature; In one embodiment, after preprocessing the input Shui script image, first extract the stroke contour line in the area of independent strokes. In this area, determine the starting coordinate of the stroke through the contour endpoint detection method and the ending coordinate . In the extracted stroke sequence, each independent stroke has a unique pair of endpoints. For example, in a test sample, a total of 50 independent strokes were detected, and the range of the starting and ending coordinates is in the pixel interval. Then, calculate the stroke direction angle according to the endpoint coordinates , and then obtain the stroke curvature value by fitting the curvature change of the curve. After calculation, the curvature of a straight-line stroke is close to 0, while the curvature of an arc stroke is generally greater than 0.15. Thus, the basic geometric feature parameters of each stroke are obtained.
[0060] In another embodiment, it is assumed that a total of 40 independent strokes are detected in a Shui script image, and their starting coordinates are evenly distributed in the pixels, and the ending coordinates are evenly distributed in the pixels. The calculated range of the stroke direction angle is , among which there are about 22 straight-line strokes and about 18 curved strokes. Among the further calculated curvature values, the minimum value is 0.02, the maximum value is 0.35, and the average value is about 0.12. It can be seen that there are many strokes with obvious arcs in this sample, and these features can provide criteria for subsequent stroke attribute classification.
[0061] Record the endpoint features according to the stroke curvature, and identify the pen tip closing state; calculate the size of the closing gap based on the pen tip closing state, identify the direction of the closing gap, and record the isolated stroke form.
[0062] In one embodiment, based on the acquired stroke geometric features, the geometric features of the stroke endpoints (such as the distance between endpoints and the local curvature of the endpoints) are first recorded according to the difference between the curvature and the endpoint coordinates. When a stroke is detected to be connected end-to-end with a distance less than a preset threshold (e.g., 5 pixels), the stroke is determined to be in a closed state; otherwise, it is determined to be unclosed. For strokes in a closed state, the size of the closure gap is further calculated, i.e., the residual distance between the endpoints, and the direction of the closure gap is identified by the direction vector of the endpoint coordinate difference. For example, in a set of 30 closed strokes, the gap size is distributed in the range of [1.5, 4.8] pixels, and the gap direction is mainly concentrated in the horizontal and diagonal directions. Finally, based on the closure features and gap parameters, the isolated stroke shape is labeled as "straight-line closed", "arc-shaped closed", or "unclosed isolated stroke".
[0063] In another embodiment, assuming 50 independent strokes are detected, 28 are closed and 22 are open. In the closed strokes, the average gap size is 3.2 pixels, with an approximate directional distribution: horizontal 40%, vertical 30%, and diagonal 30%. In the open strokes, the average distance between endpoints is 15 pixels; some isolated strokes exhibit a distinctly long straight line shape, while others show slight curvature. Quantitative recording of the size and direction of the closed gaps provides auxiliary feature data for subsequent stroke library matching and character structure assembly.
[0064] Preferably, in the area of combined strokes, identifying the connection relationship between adjacent strokes includes: In the area of combined strokes, detect the coordinates of the endpoints of the combined strokes, calculate the distance between the endpoints, and when the distance between the endpoints is lower than a preset distance threshold, record the connection position of the endpoints and determine that the connection method of the endpoints is end-to-end connection. In one embodiment, after preprocessing the combined stroke regions in the input Shui script image, the endpoint coordinates of each stroke are first extracted. Then, the Euclidean distance between the endpoints of adjacent strokes is calculated. When the distance is less than a preset threshold (e.g., 5 pixels), the endpoint pair is recorded as an end-to-end connection, and the connection method is marked. For example, in an image, 20 pairs of combined strokes are detected, of which 12 pairs satisfy the end-to-end connection condition, with endpoint distances between [2.1, 4.8] pixels. Recording the coordinate information of these endpoints can provide a preliminary connection basis for subsequent stroke library matching and character structure assembly.
[0065] In another embodiment, assuming 25 pairs of adjacent stroke endpoints are detected in the combined stroke region, and the calculated endpoint distance range is [1.5, 6.0] pixels, 18 pairs are less than a threshold of 5 pixels and are marked as end-to-end connections. The remaining 7 pairs do not meet the condition and are not marked as connections. In this way, the endpoint connection features of combined strokes can be quantified, and a reference can be provided for the next step of determining intersection or enclosing relationships.
[0066] If two strokes of a combined stroke have an intersection point at a non-endpoint, record the coordinates and angle of the intersection point and mark it as an intersecting connection; In one embodiment, for strokes in a combined stroke region that are not yet connected end-to-end, the intersection of stroke lines at non-endpoint locations is analyzed. When an intersection point is detected, its coordinates (x, y) and the angle α between the strokes are recorded. If the intersection angle is within a reasonable range (e.g., 20°~160°), the stroke connection method is labeled as an intersecting connection. For example, in a Shui script image, a total of 10 sets of intersecting strokes were detected, and the coordinates of the intersection points were... Pixel range, intersection angle range is .
[0067] In another embodiment, assuming that 12 sets of intersection points are detected in the combined strokes, the coordinate range of which is... Pixels, with cross angle distribution as Eight of the intersection angles, ranging from 30° to 150°, are marked as intersection connections. The remaining four groups are not marked due to their intersection angles being too small or too large. By recording the characteristics of the intersection points, the spatial topological relationships between strokes can be further constructed, providing constraint information for character structure reconstruction.
[0068] Detect whether the strokes of a combination of strokes form an inner and outer enclosure relationship, identify closed regions, and mark them as enclosing connections; In one embodiment, the topological structure of the combined stroke regions is analyzed to determine whether the strokes form closed regions (withinward and outward enclosing relationships). For closed regions, their boundary coordinates and the sequence of strokes they contain are recorded and labeled as enclosing connections. For example, in a Shui script image, three closed regions were detected, each consisting of 2-4 strokes, with boundary coordinates at... Pixel range.
[0069] In another embodiment, assuming four closed regions are detected in the combined stroke region, with the number of strokes in each closed region being 2, 3, 3, and 5 respectively, and the coordinate range of the closed boundary is... Pixels. By using annotations to enclose and connect strokes, the parts that form a complete character structure in the combined strokes can be clearly identified, providing topological constraints for the step-by-step assembly of strokes.
[0070] End-to-end connections, cross connections, and enclosing connections are considered as adjacent stroke connections.
[0071] In one embodiment, end-to-end connections, cross connections, and enclosing connections are combined to construct a complete adjacent connection relationship for the combined strokes. This connection relationship can be represented as a graph structure, where strokes are nodes and connection types are edge attributes, for use in subsequent stroke library matching and character structure assembly.
[0072] In another embodiment, assume that in a certain Shui script image, 18 pairs of end-to-end connections are detected, 8 pairs of cross connections are detected, and 4 enclosed connections are detected. The constructed adjacent stroke connection relationship graph has 30 nodes and 30 edges (some edges are multiple edges), and the annotation types include "end-to-end", "cross", and "enclosed". Through this graph structure, accurate stroke connection information can be provided for subsequent step-by-step character structure restoration.
[0073] Preferably, reconstructing the character order based on the character semantic information in step S3 includes: Extracting the character position index based on the character semantic information, and performing spatial coordinate sorting to obtain the sorting result; decomposing the strokes in the sorting result to obtain the character stroke set; In one embodiment, semantic analysis is performed on the characters in the input text or image, and the spatial position index (such as line number, column number, or pixel coordinates) of each character is extracted. Then, the characters are spatially sorted according to the character position index, and the sorting result is generated from top to bottom and from left to right. For each sorted character region, the strokes inside the character are decomposed to obtain the stroke set of the character. Taking the character "水" as an example, the extracted character position index is , and the decomposed stroke set is {dot, horizontal, left-falling stroke, right-falling stroke}. This stroke set can be used for subsequent stroke order analysis and reconstruction.
[0074] In another embodiment, assume that the input document contains 5 characters, and the extracted character position indexes are respectively . After sorting by spatial coordinates, the sorting sequence obtained is the 1st, 2nd, 3rd, 4th, and 5th characters. The number of strokes decomposed for each character region is [4, 5, 3, 6, 4] respectively, forming the corresponding character stroke set. Through this sorting and decomposition, a complete stroke set can be established for each character, providing a preliminary data basis for order reconstruction.
[0075] Align the character stroke set with the preset character semantic labels, and correct the stroke order of the character stroke set according to the standard stroke order in the character semantic labels to reconstruct the character order.
[0076] In one embodiment, the character stroke set is matched with the previously established character semantic label library. When matching, features such as stroke type, start and end coordinates, stroke length, and relative position relationship are considered. After alignment, according to the standard stroke order defined in the character semantic labels, the stroke order of the character stroke set is corrected. For example, for the character "水", the standard stroke order is {dot, horizontal, left-falling stroke, right-falling stroke}. If the original stroke set order is {horizontal, dot, left-falling stroke, right-falling stroke}, then the correct order {dot, horizontal, left-falling stroke, right-falling stroke} is obtained after correction. After this step, the character stroke set not only maintains integrity but also conforms to semantics and writing norms, providing a basis for subsequent character reconstruction.
[0077] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.
[0078] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.
Claims
1. A method for recognizing Shui script characters based on object detection, characterized in that, Includes the following steps: Step S1: Obtain the Shui script image and perform image preprocessing to obtain the Shui script image to be processed; Identify the text regions in the water script image to be processed; Step S2: Correct the geometric deformation of the characters based on the text region and identify the features of the Shui script, including: A pre-set neural network model is used to identify the features of Shui script characters; The neural network model includes an input layer, multiple convolutional feature extraction layers, a feature fusion layer, and a classification decision layer. The input layer receives the Shui script image of the text region. The convolutional feature extraction layer extracts multi-scale text features from the Shui script image. The feature fusion layer upsamples and weights the text features at different levels to obtain high-dimensional comprehensive features. The classification decision layer performs character-by-character feature matching on the high-dimensional comprehensive features. Character classification of Shui script images is performed using the characteristics of Shui script characters, and the semantic information of the characters is determined, including: The structural features of the characters in the Shui script are extracted, and the samples are labeled according to the structural features. The character structure of the Shui script image is divided to determine the independent stroke regions and the combined stroke regions. In the independent stroke area, record the shape of isolated strokes; in the combined stroke area, identify the connection relationship between adjacent strokes; determine the stroke attribute label based on the shape of isolated strokes. Match the stroke types of the preset Shui script stroke library according to the connection relationship between adjacent strokes; use stroke attribute tags and stroke types to piece together the complete character structure step by step to determine the semantic information of the character; Step S3: Reconstruct the character order based on character semantic information; identify the Shui script phrase structure in the Shui script image based on the character order, and generate a complete Shui script recognition sample. Step S3, reconstructing the character order based on character semantic information, includes: Based on the semantic information of the characters, the character position index is extracted and sorted by spatial coordinates to obtain the sorting result; the strokes in the sorting result are decomposed to obtain the character stroke set; Align the character stroke set with the preset character semantic tags, and correct the stroke order of the character stroke set according to the standard stroke order in the character semantic tags in order to reconstruct the character order; Step S4: Calculate the confidence score of the Shui script based on the complete Shui script recognition sample, and use the confidence score to correct erroneous Shui script characters and store the corrected Shui script characters.
2. The method for recognizing Shui script characters based on target detection according to claim 1, characterized in that, The text regions identified in step S1 of the water script image to be processed include: Generate text candidate boxes using a pre-defined convolutional feature extractor; calculate the response value of each text candidate box, perform character confidence prediction, and output the confidence score of each text candidate box; Potential text regions in the water script image to be processed are filtered and identified based on confidence level; rotation correction is performed on the potential text regions to output the text regions.
3. The method for recognizing Shui script characters based on target detection according to claim 2, characterized in that, The potential text region is rotated and corrected to output the text region, which includes: Using the center point coordinates of the potential text region as the rotation reference point, establish an affine transformation matrix; determine the text direction angle of the potential text region, write the text direction angle as a rotation parameter into the affine transformation matrix, and keep the scaling ratio and aspect ratio unchanged; The original pixels of the potential text region are mapped to their coordinates using an affine transformation matrix, and each original pixel is mapped to a rotated position to output the text region.
4. The method for recognizing Shui script characters based on target detection according to claim 1, characterized in that, Extracting multi-scale text features from Shui script images includes: The first convolutional layer of the neural network model performs a primary convolution operation on the input Shuishu image to extract low-level edge features; the low-level edge features are then input into the second convolutional layer to extract stroke features. The stroke features are input into the third convolutional layer to extract character shape features; the character shape features are then input into the fourth convolutional layer to extract high-level semantic features.
5. The method for recognizing Shui script characters based on target detection according to claim 1, characterized in that, In the area of independent strokes, the form of isolated strokes is recorded as follows: In an independent stroke area, determine the starting and ending coordinates of the stroke; calculate the stroke direction angle based on the starting and ending coordinates to determine the stroke curvature; Record the endpoint features based on the curvature of the strokes to identify the closed state of the stroke tip; calculate the size of the closed gap based on the closed state of the stroke tip, identify the direction of the closed gap, and record the shape of isolated strokes.
6. The method for recognizing Shui script characters based on target detection according to claim 1, characterized in that, In the area of combined strokes, identifying the connection relationship between adjacent strokes includes: In the area of combined strokes, detect the coordinates of the endpoints of the combined strokes, calculate the distance between the endpoints, and when the distance between the endpoints is lower than a preset distance threshold, record the connection position of the endpoints and determine that the connection method of the endpoints is end-to-end connection. If two strokes of a combined stroke have an intersection point at a non-endpoint, record the coordinates and angle of the intersection point and mark it as an intersecting connection; Detect whether the strokes of a combination of strokes form an inner and outer enclosure relationship, identify closed regions, and mark them as enclosing connections; End-to-end connections, cross connections, and enclosing connections are considered as adjacent stroke connections.
7. A Shui script character recognition system based on target detection, characterized in that, For performing the object detection-based Shui script character recognition method as described in claim 1, the object detection-based Shui script character recognition system comprises: The image preprocessing module is used to acquire Shui script images and perform image preprocessing to obtain Shui script images to be processed; and to identify the text regions in the Shui script images to be processed. The character classification module is used to correct the geometric deformation of characters based on the text region and identify the characteristics of Shui script characters; it uses the characteristics of Shui script characters to classify characters in Shui script images and determine the semantic information of the characters; The character order reconstruction module is used to reconstruct the character order based on character semantic information; and to identify the Shui script phrase structure in Shui script images based on the character order, generating complete Shui script recognition samples. The confidence calculation module is used to calculate the confidence of Shuishu based on the complete Shuishu recognition sample, and to use the confidence of Shuishu to correct erroneous Shuishu characters and store the corrected Shuishu characters.
Citation Information
Patent Citations
Water book character recognition method based on CNN structure neural network
CN110348280A
Female book character recognition method and system based on convolutional neural network
CN116524522A
Mental language text image recognition wrong character correction method and system
CN120235736A
Ancient character image recognition and semantic analysis method
CN120472471A
Vectorization Chinese character graph generation method based on large model
CN121010668A
Cited By
Online monitoring and early warning method for printed matter printing content
CN122090461A