A calligraphy stroke intelligent extraction system and method
By using spatial transformation network registration and semantic segmentation methods in Chinese character stroke extraction, adding a priori information of strokes is solved, and the problem of extraction errors caused by the lack of priori information in the priori technology is solved, and high-precision and high-efficiency stroke extraction is achieved.
Patent Information
- Application Number
- CN202211544199.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-03
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-12-03
AI Technical Summary
The prior art lacks sufficient prior information in the extraction of Chinese characters, resulting in errors in the extraction.
The spatial transformation network registration method is used to add a priori information of strokes, and the calligraphy strokes are extracted through the stroke disassembly unit based on semantic segmentation.
High-precision extraction of calligraphy strokes is achieved, the extraction accuracy is improved, and the extraction response speed within 100ms is achieved under the GPU acceleration.
Smart Images

Figure CN115984866B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a calligraphy stroke intelligent extraction system and method. Background Art
[0002] Calligraphy is a traditional Chinese culture. Calligraphic characters are diverse in form and complex in structure. Strokes are a crucial component of each character. For intelligent technology research and product development, such as calligraphy recognition, calligraphy evaluation, and automatic calligraphy generation, strokes are a key element in improving accuracy. Stroke segmentation has always been a key research focus and challenge in intelligent calligraphy technology.
[0003] Most existing research on Chinese character stroke extraction begins with the image morphological features of strokes and radicals. Stroke extraction is generally achieved directly by analyzing the direction or other shape features of the strokes. With the development of deep learning, some studies have attempted to use deep learning methods to obtain morphological features, detect stroke intersections, and obtain character skeletons. Furthermore, to achieve better extraction results, reference templates are used during the extraction process to further improve extraction accuracy. For example, application publication number CN 104156721 A proposes a "Method for Offline Chinese Character Stroke Extraction Based on Template Matching." This method uses skeleton extraction to help identify each stroke within an offline Chinese character, improving the accuracy of stroke extraction.
[0004] However, existing extraction methods mainly achieve stroke extraction by performing morphological analysis on Chinese character strokes. Due to the lack of sufficient prior information in stroke extraction, stroke extraction is prone to errors. Summary of the Invention
[0005] At least one of the purposes of the present invention is to provide a calligraphy stroke intelligent extraction system and method to overcome the problems existing in the above-mentioned prior art, which adopts the spatial transformation network alignment method to add prior information of strokes, and extracts calligraphy strokes through a stroke decomposition unit based on semantic segmentation, thereby achieving high-precision extraction of calligraphy strokes.
[0006] In order to achieve the above objectives, the technical solutions adopted by the present invention include the following aspects.
[0007] A calligraphy stroke intelligent extraction system includes: an input module, a preprocessing module, a stroke extraction module and an output module, wherein the input module is connected to the preprocessing module, the preprocessing module is connected to the stroke extraction module, and the stroke extraction module is connected to the output module.
[0008] The input module is used to obtain the copied character image.
[0009] The preprocessing module is used to preprocess the acquired copied character image.
[0010] The stroke extraction module is used to register the pre-processed copied character image and disassemble the copied character image to obtain a single stroke image of the copied character.
[0011] The output module includes a result integration unit, and the output module is used to output the acquired single stroke image of the copied character.
[0012] Preferably, it further comprises a reference character annotation database, wherein the reference character annotation database stores reference character annotation information, and the reference character annotation information includes reference character component structure annotation information and stroke outline annotation information.
[0013] Preferably, the pre-processing module includes a frame and calligraphy segmentation unit and a calligraphy correction unit; the frame and calligraphy segmentation unit is used to segment the copied character image, and the calligraphy correction unit is used to correct the segmented copied character image.
[0014] Preferably, the stroke extraction module includes a spatial transformation network registration unit and a semantic segmentation-based stroke decomposition unit.
[0015] The spatial transformation network registration unit is used to register the preprocessed copied character image to obtain reference stroke data, and the reference stroke data is input into the semantic segmentation-based stroke disassembly unit as prior information of the preprocessed copied character image; the semantic segmentation-based stroke disassembly unit is used to disassemble the preprocessed copied character image to obtain a single stroke image.
[0016] A method for intelligently extracting calligraphy strokes, comprising the following steps:
[0017] Step 1: Get the image of the copied character;
[0018] Step 2: Preprocess the obtained copied character image;
[0019] Step 3: Register the pre-processed copied character image, and then decompose the copied character image to obtain a single stroke image;
[0020] Step 4: Integrate and output the acquired single stroke images.
[0021] Preferably, the preprocessing process of the copied character image includes: segmenting the copied character image through a border and calligraphy character segmentation unit, inputting the segmented copied character image into a calligraphy character image correction unit, and correcting the copied character image.
[0022] Preferably, the copy character image segmentation and correction steps include: inputting the copy character image into the improved Deeplabv3+ image semantic segmentation model to segment the copy character image writing area and the border auxiliary line area; based on the segmented copy character image writing area, using an adaptive image binarization algorithm to binarize the copy character image to obtain a copy character binary image; based on the segmented copy character image border auxiliary line area, using a Hough transform straight line detection algorithm to detect the border auxiliary line; according to the detected border auxiliary line, using a correction matrix to unify the angle and size of the copy character binary image to obtain a corrected binarized image.
[0023] Preferably, the process of registering the preprocessed copied character image and then disassembling the copied character image includes: using a spatial transformation network registration unit to register the preprocessed copied character image to obtain registered reference stroke data; using the registered reference stroke data as prior information of the strokes of the preprocessed copied character image, inputting it into a stroke disassembly unit based on semantic segmentation, and disassembling a single stroke image of the copied character image.
[0024] Furthermore, the process of registering the pre-processed copied character image using the spatial transformation network registration unit includes:
[0025] Obtain reference character annotation information from a reference character annotation database;
[0026] The reference information processing module is used to sequentially connect and fill the outline point data marked in the reference character annotation information to obtain a single stroke image of the reference character, and the single stroke image of the reference character is processed into 7 types of image data according to the category;
[0027] The 7 types of image data and the pre-processed copied character image are input into the deep space transformation model to obtain the image after the 7 types of reference character image data are aligned and transformed as the reference character alignment transformation data.
[0028] Furthermore, the disassembly process of the stroke disassembly unit based on semantic segmentation includes:
[0029] Input the image data after registration transformation of the 7 categories of reference character image data and the pre-processed copy character image into the stroke semantic segmentation model to obtain the intermediate layer features of the stroke semantic segmentation and the semantic segmentation results of the 7 categories of strokes of the copy character;
[0030] The intermediate layer features of stroke semantic segmentation, the stroke semantic segmentation results and the aligned single stroke images of the reference character are input into the single stroke extraction model, and the single stroke images of the copied character are extracted in sequence according to the reference stroke order through the single stroke extraction model.
[0031] In summary, due to the adoption of the above technical solution, the present invention has at least the following beneficial effects:
[0032] The present invention adopts the method of spatial transformation network registration to add prior information of strokes, and extracts the strokes of the copied characters through a stroke semantic segmentation model and a single stroke extraction model, which can improve the accuracy of stroke extraction and achieve high-precision extraction of strokes; and, since the present invention adopts a deep learning model, with GPU acceleration, it can achieve an extraction response speed of within 100ms, and in the stroke extraction process, it has high accuracy, high extraction rate and high adaptability.
[0033] By setting up a pre-processing module, stroke extraction can be performed directly on the captured image, reducing the difficulty of stroke extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a structural diagram of a calligraphy stroke intelligent extraction system according to an exemplary embodiment of the present invention.
[0035] Figure 2 It is a schematic diagram of the segmentation and correction of copied characters image according to an exemplary embodiment of the present invention.
[0036] Figure 3 2 is a schematic diagram of spatial transformation network registration according to an exemplary embodiment of the present invention.
[0037] Figure 4 4 is a schematic diagram of a stroke decomposition process based on semantic segmentation according to an exemplary embodiment of the present invention.
[0038] Figure 5 4 is a schematic diagram of the classification of Chinese character strokes according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0039] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments to make the purpose, technical solutions and advantages of the present invention more clearly understood. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0040] like Figure 1 As shown, the calligraphy stroke intelligent extraction system of an exemplary embodiment of the present invention includes: an input module, a reference character annotation database, a preprocessing module, a stroke extraction module and an output module, the input module is connected to the preprocessing module, the preprocessing module is connected to the stroke extraction module, the reference character annotation database is connected to the stroke extraction module, and the stroke extraction module is connected to the output module.
[0041] The input module is used to obtain the image of the copied character (the calligraphy practice work of copying the reference character);
[0042] The preprocessing module is used to preprocess the acquired copy character image; the preprocessing module includes a frame and calligraphy character segmentation unit and a calligraphy character correction unit, the frame and calligraphy character segmentation unit is used to segment the copy character image, and the calligraphy character correction unit is used to correct the segmented copy character image;
[0043] The reference character annotation database stores reference character annotation information, which includes reference character component structure annotation information and stroke outline annotation information;
[0044] The stroke extraction module is used to register the pre-processed copied character image and disassemble the copied character image to obtain a single stroke image of the copied character; the stroke extraction module includes a spatial transformation network registration unit and a semantic segmentation-based stroke disassembly unit. The spatial transformation network registration unit is used to register the pre-processed copied character image to obtain reference stroke data after registration transformation. The reference stroke data after registration transformation is input into the semantic segmentation-based stroke disassembly unit as prior information of the pre-processed copied character image; the semantic segmentation-based stroke disassembly unit is used to disassemble the pre-processed copied character image to obtain a single stroke image;
[0045] The output module is used to output the acquired single stroke image of the copied character; the output module includes an integration unit, and the result integration unit is used to integrate the disassembled single stroke image of the copied character to generate structured single stroke image data.
[0046] The method for intelligently extracting calligraphy strokes according to an exemplary embodiment of the present invention comprises the following steps:
[0047] Step 1: Obtain the image of the copied characters. The image of the copied characters can be obtained by taking a photo, scanning, etc.
[0048] Step 2: Preprocess the acquired copied character image to obtain a copied character binary image;
[0049] The preprocessing process of the copied character image includes: segmenting the copied character image through the border and calligraphy character segmentation unit, inputting the segmented copied character image into the calligraphy character correction unit, correcting the copied character image, and obtaining a corrected binary image;
[0050] refer to Figure 2The copy character image segmentation and correction steps include: inputting the copy character image into the improved Deeplabv3+ image semantic segmentation model, and segmenting the copy character image writing area and the border auxiliary line area; the copy character image writing area refers to the main calligraphy handwriting area of the copy character in the auxiliary frame (excluding interfering handwriting and stains), and the border auxiliary line refers to the general auxiliary frame of calligraphy writing. During the training process, these auxiliary frames are standardized to a width of 5 to 10 pixels; since the border auxiliary line is relatively thin, the present invention adds a shallow semantic segmentation branch on the basis of the existing Deeplabv3+, which improves the image semantic segmentation model's ability to segment finer and smaller categories; based on the segmented copy character image writing area, an adaptive image binarization algorithm is used to binarize the copy character image to obtain a copy character binary image; based on the segmented copy character image border auxiliary line area, a Hough transform straight line detection algorithm is used to detect the border auxiliary line, and according to the detected border auxiliary line, the copy character binary image is unified in angle and size through a correction matrix to obtain a corrected binarized image.
[0051] Step 3: Register the pre-processed copied character image, and then decompose the copied character image to obtain a single stroke image;
[0052] A spatial transformation network registration unit is used to register the copied character image to obtain the registered reference stroke data; the registered reference character strokes and the pre-processed copied character strokes are similar in stroke position and stroke shape, and the registered reference character strokes are used as prior information of the pre-processed copied character image strokes and input into a semantic segmentation-based stroke decomposition unit, which decomposes the copied character image into a single stroke image.
[0053] refer to Figure 3 ,The registration process of the copied character image includes:
[0054] Acquire reference character annotation information from a reference character annotation database, and input the acquired reference character annotation information into a reference information processing module;
[0055] The reference character annotation information includes the reference character component structure annotation information and the stroke contour annotation information. The reference character component structure annotation information and the stroke contour annotation information are annotated through open-source image annotation tools such as Labelme, and the annotation results are saved in a relational database such as mysql for the stroke extraction module to extract. In the stroke contour annotation information, each stroke contour is represented by a set of coordinate values and is attached with a text label. In the component structure annotation information, each radical structure is represented by a set of coordinate values and is attached with a text label. For example, for the character "仰", using Labelme, a series of points are used to mark the outline of each individual stroke along the edges of the strokes such as "left-falling stroke", "vertical stroke", "left-falling stroke", "vertical hook", "horizontal fold hook", and "vertical stroke". Each stroke contour is represented by a set of coordinate values and is attached with a text label. According to the general structure of Chinese characters, such as "仰" being a left-middle-right structure, the entire stroke is divided into three parts, each part is represented by a set of coordinate values and is attached with a text label.
[0056] [[ID=D3]]The reference information processing module connects and fills the contour point data marked in the reference character annotation information in sequence to obtain the single-stroke image of the reference character, and classifies the obtained single-stroke image of the reference character.
[0057] In a preferred embodiment, to simplify the registration difficulty, 31 basic strokes (due to the high similarity of some strokes, it can also be considered 25 basic strokes empirically) can be manually divided into 7 categories according to the following rules:
[0058] (1) Minimize the probability of strokes intersecting during writing within the same category;
[0059] (2) Balance the number of strokes used during writing among different categories as much as possible;
[0060] (3) Maximize the structural similarity of strokes within the same category and minimize the structural similarity of strokes between different categories.
[0061] Figure 5 The result of dividing 31 Chinese character basic strokes into 7 categories according to the above rules is shown, and it is used as a stroke classification standard.
[0062] The reference information processing module processes the single-stroke image of the reference character into 7 types of image data according to the Figure 5 shown categories, and inputs the processed 7 types of image data and the binary image of the traced character after correction into the deep space transformation model to obtain the reference stroke data after registration transformation.
[0063] The core of the deep space transformation model of the present invention is the image registration model based on TpsStn. The deep space transformation model of the present invention simultaneously establishes the registration relationship between the 7 categories of stroke images and the copied characters, and obtains the image after the registration transformation of the 7 categories of reference character image data as the reference character registration transformation data.
[0064] refer to Figure 4 ,The disassembly process of the copied character image includes:
[0065] The image data after registration with the seven categories of reference characters and the rectified copy characters were input into the Deeplabv3+ image semantic segmentation model. An inter-category attention mechanism was implemented within each block of the convolution process to improve the utilization of data information across different stroke categories. The Deeplabv3+ image semantic segmentation model obtained the 32×32 block output data from the convolution process on the stroke semantic segmentation model as intermediate-layer features for stroke semantic segmentation, along with the semantic segmentation results for the seven categories of strokes in the copy characters. The intermediate-layer features, stroke semantic segmentation results, and the single-stroke images after registration with the reference characters were then input into the single-stroke extraction model, a simple Unified Network (UNet) architecture. This model sequentially extracted the single-stroke images of the copy characters according to the reference stroke order.
[0066] Step 4: The obtained single stroke images of the copied characters are integrated through the result integration unit to generate structured single stroke image data (for example, it can be stored as a JSON format stroke image data file), and the single stroke image data is output through the output module.
[0067] The above description is only a detailed description of the specific embodiments of the present invention, and does not limit the present invention. Various substitutions, modifications and improvements made by those skilled in the relevant art without departing from the principles and scope of the present invention should be included in the scope of protection of the present invention.
Claims
1. A calligraphy stroke intelligent extraction system, It is characterized in that include: An input module, a preprocessing module, a stroke extraction module and an output module, wherein the input module is connected to the preprocessing module, the preprocessing module is connected to the stroke extraction module, and the stroke extraction module is connected to the output module; The input module is used to obtain the copied character image; The preprocessing module is used to preprocess the acquired copied character image; The stroke extraction module is used to register the pre-processed copy character image and disassemble the copy character image to obtain a single stroke image of the copy character; The output module includes a result integration unit, and the output module is used to output the acquired single stroke image of the copied character; The stroke extraction module includes a space transformation network registration unit and a stroke disassembly unit based on semantic segmentation; The spatial transformation network registration unit is used to register the preprocessed copied character image to obtain reference stroke data, and the reference stroke data is input into the stroke disassembly unit based on semantic segmentation as prior information of the preprocessed copied character image; the stroke disassembly unit based on semantic segmentation is used to disassemble the preprocessed copied character image to obtain a single stroke image.
2. The calligraphy stroke intelligent extraction system according to claim 1, It is characterized in that It also includes a reference character annotation database, which stores reference character annotation information, including reference character component structure annotation information and stroke outline annotation information.
3. The calligraphy stroke intelligent extraction system according to claim 1, It is characterized in that The preprocessing module includes a frame and calligraphy segmentation unit and a calligraphy correction unit; the frame and calligraphy segmentation unit is used to segment the copied character image, and the calligraphy correction unit is used to correct the segmented copied character image.
4. A method for intelligent extraction of calligraphy strokes, It is characterized in that The calligraphy stroke intelligent extraction method comprises the following steps: Step 1: Get the image of the copied word; Step 2: pre-processing the acquired copy character image; Step 3: register the pre-processed copy character image, and then decompose the copy character image to obtain a single stroke image; Step 4: Integrate and output the acquired single stroke images; In step 3, the process of registering the preprocessed copied character image and then disassembling the copied character image includes: using a spatial transformation network registration unit to register the preprocessed copied character image to obtain registered reference stroke data; using the registered reference stroke data as prior information of the strokes of the preprocessed copied character image, inputting it into a stroke disassembly unit based on semantic segmentation, and disassembling a single stroke image of the copied character image.
5. The intelligent calligraphy stroke extraction method according to claim 4, It is characterized in that In step 2, the preprocessing process of the copied character image includes: segmenting the copied character image through a frame and calligraphy character segmentation unit, inputting the segmented copied character image into a calligraphy character image correction unit, and correcting the copied character image.
6. The intelligent calligraphy stroke extraction method according to claim 5, It is characterized in that The copy character image segmentation and correction steps include: inputting the copy character image into the improved Deeplabv3+ image semantic segmentation model to segment the copy character image writing area and the border auxiliary line area; based on the segmented copy character image writing area, using an adaptive image binarization algorithm to binarize the copy character image to obtain a copy character binary image; based on the segmented copy character image border auxiliary line area, using a Hough transform straight line detection algorithm to detect the border auxiliary lines, and according to the detected border auxiliary lines, using a correction matrix to unify the angle and size of the copy character binary image to obtain a corrected binary image.
7. The intelligent calligraphy stroke extraction method according to claim 4, It is characterized in that The process of registering the pre-processed copied character image using the space transformation network registration unit includes: Acquire reference character annotation information from a reference character annotation database; The reference information processing module is used to sequentially connect and fill the contour point data marked in the reference character marking information to obtain a single stroke image of the reference character, and the single stroke image of the reference character is processed into 7 types of image data according to categories; The 7 types of image data and the preprocessed copied character image are input into the deep space transformation model to obtain the image after the 7 types of reference character image data are aligned and transformed as the reference character alignment transformation data.
8. The intelligent calligraphy stroke extraction method according to claim 7, It is characterized in that The disassembly process of the stroke disassembly unit based on semantic segmentation includes: Input the image data after registration transformation of the 7 types of reference character image data and the preprocessed copy character image into the stroke semantic segmentation model to obtain the intermediate layer features of the stroke semantic segmentation and the semantic segmentation results of the 7 types of strokes of the copy character; The intermediate layer features of stroke semantic segmentation, the results of stroke semantic segmentation and the aligned single stroke images of the reference characters are input into the single stroke extraction model, and the single stroke images of the copied characters are extracted in sequence according to the reference stroke order through the single stroke extraction model.
Citation Information
Patent Citations
Off-line Chinese character stroke extraction method based on template matching
CN104156721A
RMB sequence number identification method
CN102800148A
Inscription restoration method based on contour feature description of Chinese character image
CN105069766A