Stroke splitting method and device

By extracting vector contour data from Chinese character font files, performing skeleton extraction and curvature analysis, and combining it with Bezier curve fitting, the problems of low accuracy and efficiency in Chinese character stroke segmentation in the existing technology are solved, and high-precision and efficient stroke segmentation is achieved.

CN120635909APending Publication Date: 2025-09-12VIVO MOBILE COMM (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510793021.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies have problems with low accuracy and efficiency in Chinese character stroke segmentation, especially edge detection methods that tend to ignore detailed features of Chinese characters and rely on training data.

Method used

By extracting vector outline data from the Chinese character font file, performing skeleton extraction and curvature analysis, combining Bezier curve fitting, segmenting the stroke image and extracting geometric features, the stroke data is generated.

Benefits of technology

The accuracy and efficiency of Chinese character stroke segmentation are improved, the dependence on training data is avoided, and the accuracy and fluency of stroke extraction are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635909A_ABST
    Figure CN120635909A_ABST
Patent Text Reader

Abstract

The invention discloses a stroke splitting method and device. Belongs to the technical field of computer fonts. The method comprises the steps that vector contour data are extracted from font files of Chinese characters, and the vector contour data comprise shape information and structure information of the Chinese characters; based on the vector contour data, performing skeleton extraction and curvature analysis on the contour of the Chinese character to obtain a stroke unit of the Chinese character; converting the vector contour data into a character image, and segmenting the character image based on a stroke unit to obtain a stroke image of each stroke in the Chinese character; and extracting geometric features from the stroke image, and generating stroke data of the Chinese character based on the geometric features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer font technology, and more particularly to a stroke splitting method and device thereof. Background Art

[0002] Vector font formats like TrueType (TTF) are widely used in computer display and printing due to their scalability and smoothness. Typically, applications such as handwriting recognition, automatic font generation, and calligraphy analysis require a deep understanding of the structure and stroke characteristics of Chinese characters, necessitating the separation of Chinese character strokes.

[0003] In the existing technology, image processing methods such as edge detection are usually used to extract Chinese character strokes. This method easily ignores many detailed features of Chinese characters and relies on training data for model training, which has low efficiency and accuracy. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a stroke splitting method and device thereof, which improves the accuracy and efficiency of Chinese character stroke splitting.

[0005] In a first aspect, an embodiment of the present application provides a stroke splitting method, which includes: extracting vector outline data from a font file of a Chinese character, the vector outline data including shape information and structural information of the Chinese character; converting the vector outline data into a character image, performing skeleton extraction and curvature analysis on the character image to obtain the stroke unit of the Chinese character; segmenting the character image based on the stroke unit to obtain the stroke image of each stroke in the Chinese character; extracting geometric features from the stroke image, and generating the stroke data of the Chinese character based on the geometric features.

[0006] In second aspect, an embodiment of the present application provides a stroke splitting device, which includes: a first generation unit, for extracting vector contour data from a font file of a Chinese character, wherein the vector contour data includes shape information and structure information of the Chinese character; based on the vector contour data, performing skeleton extraction and curvature analysis on the contour of the Chinese character to obtain the stroke unit of the Chinese character; converting the vector contour data into a character image, and segmenting the character image based on the stroke unit to obtain a stroke image of each stroke in the Chinese character; extracting geometric features from the stroke image, and generating the stroke data of the Chinese character based on the geometric features.

[0007] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.

[0008] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in the first aspect above are implemented.

[0009] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the method described in the first aspect.

[0010] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the method described in the first aspect.

[0011] In an embodiment of the present application, the font file is first parsed to generate vector outline data including the shape and structure information of the Chinese character; then, based on the vector outline data, the outline of the Chinese character is subjected to skeleton extraction and curvature analysis to obtain the stroke unit of the Chinese character; then, the vector outline data is converted into a character image, and the character image is segmented based on the stroke unit to obtain the stroke image of each stroke in the Chinese character; finally, geometric features are extracted from the stroke image, and the stroke data of the Chinese character is generated based on the geometric features. In the above process, by combining the analysis of the vector outline data with image processing, and utilizing methods such as skeleton extraction and curvature analysis, complex Chinese character strokes can be accurately decomposed, avoiding dependence on training data, thereby improving the accuracy and efficiency of stroke extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 This is a flow chart of the stroke splitting method provided in an embodiment of the present application;

[0013] Figure 2A Schematic diagram of an application scenario of the stroke splitting method provided in an embodiment of the present application;

[0014] Figure 2B Schematic diagram of an application scenario of the stroke splitting method provided in an embodiment of the present application;

[0015] Figure 3A Schematic diagram of an application scenario of the stroke splitting method provided in an embodiment of the present application;

[0016] Figure 3B Schematic diagram of an application scenario of the stroke splitting method provided in an embodiment of the present application;

[0017] Figure 4A Schematic diagram of an application scenario of the stroke splitting method provided in an embodiment of the present application;

[0018] Figure 4BSchematic diagram of an application scenario of the stroke splitting method provided in an embodiment of the present application;

[0019] Figure 5A Schematic diagram of an application scenario of the stroke splitting method provided in an embodiment of the present application;

[0020] Figure 5B Schematic diagram of an application scenario of the stroke splitting method provided in an embodiment of the present application;

[0021] Figure 6A Schematic diagram of an application scenario of the stroke splitting method provided in an embodiment of the present application;

[0022] Figure 6B Schematic diagram of an application scenario of the stroke splitting method provided in an embodiment of the present application;

[0023] Figure 7 Schematic diagram of the structure of the stroke splitting device provided in an embodiment of the present application;

[0024] Figure 8 is a schematic structural diagram of an electronic device provided in an embodiment of the present application;

[0025] Figure 9 It is a schematic diagram of the hardware structure of an electronic device suitable for implementing the embodiments of the present application. DETAILED DESCRIPTION

[0026] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.

[0027] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0028] The stroke splitting method and device provided by the embodiments of the present application are described in detail below with reference to the accompanying drawings through specific embodiments and their application scenarios.

[0029] Please refer to Figure 1, which shows one of the flow charts of the stroke splitting method provided in the embodiment of the present application. The stroke splitting method provided in the embodiment of the present application can be applied to electronic devices. In practice, the electronic devices can be smartphones, tablet computers, laptop computers, wearable devices, etc.

[0030] The process of the stroke splitting method provided in the embodiment of the present application includes the following steps:

[0031] Step 101: extract vector outline data from a Chinese character font file. The vector outline data includes shape information and structure information of the Chinese character.

[0032] In this embodiment, a font file is a file used to store character shape and appearance information. Font files may include, but are not limited to, font files in vector font formats such as TrueType Font (TTF). TTF is a vector font format described based on Bezier curves. It contains vector data such as character outlines and curves, accurately describes character shapes, and offers advantages such as scalability and smoothness.

[0033] Vector outline data is a vector-based data format that contains the shape and structure information of Chinese characters. It consists of a series of points and their connections, accurately representing the geometric shape and structure of characters. In TTF font files, this data defines the character outlines using mathematical methods such as Bezier curves. This data is resolution-independent and can be scaled infinitely without distortion.

[0034] In practice, the font file can be loaded first, and the various data therein can be identified and processed, including but not limited to: font header table, maximum value description table, glyph positioning table, and glyph data table. Among them, the font header table contains global information of the font, such as version number, coordinate system, bounding box parameters, etc. The maximum value description table is used to record the maximum parameters in the font, such as the maximum number of contour points, the maximum number of components of a composite glyph, etc. The glyph positioning table is used to store the offset position of each glyph in the glyph data table, so as to facilitate the rapid positioning of glyph data. The glyph data table contains vector information that constitutes the glyph. During the reading process, the global attribute parameters of the font can be extracted at the same time, which may include font name, version information, character set range, and measurement information, etc., to provide the necessary basic data for subsequent processing.

[0035] After successfully loading the font file, the character encoding mapping relationship in the font file can be parsed. The specific steps are as follows: First, parse the character mapping table, read the sub-tables in the character mapping table for different platforms and encoding methods, such as the sub-table for Unicode (Unicode), to ensure the correct mapping of Chinese characters. Then traverse the mapping records in the character mapping table, extract the glyph index corresponding to the Unicode of each Chinese character, generate an index table for character retrieval and positioning, and obtain the mapping relationship between the two. Through the above mapping relationship, the corresponding glyph data can be efficiently and accurately located according to Unicode, ensuring the precise association between Chinese characters and glyphs, and realizing the glyph index of Chinese characters.

[0036] After obtaining the glyph index of a Chinese character, the glyph data table can be accessed based on the glyph index to obtain the vector information of the Chinese character. By parsing the vector information, such as data extraction and reconstruction, the vector outline data of each Chinese character can be obtained.

[0037] In some optional implementations of this embodiment, vector outline data of Chinese characters may be generated through the following sub-steps:

[0038] Sub-step 1011: extracting the Bezier curve control point information of the Chinese characters from the font file.

[0039] Specifically, the glyph data table can be accessed through the above-mentioned glyph index to obtain the vector information of the Chinese character. From this vector information, the Bezier curve control point information of the Chinese character is extracted. The Bezier curve is a parametric curve widely used in computer graphics to represent smooth curve shapes. The Bezier curve is defined by a set of control points, which determine the shape and direction of the curve. In the font file, the Bezier curve control point information includes the coordinates and type of the control points, which are used to accurately describe the curve part of the character outline. The coordinates of the control points determine the position of the curve in the plane, while the type of control points, such as the starting point, end point, intermediate control point, etc., affects the mathematical expression and shape of the Bezier curve.

[0040] Sub-step 1012: determining the outline topology structure information of the Chinese character based on the Bezier curve control point information.

[0041] The contour topology structure describes the spatial relationship and connection method between the various parts of the character contour, including the number of contours, the starting point and end point of each contour, the nesting relationship between contours, and the distinction between inner and outer contours. Contour topology information is used to describe the contour composition of the Chinese character glyph, as well as the relative position and hierarchical relationship between these contours. In practice, the number of contours contained in the Chinese character glyph can be determined by parsing the Bezier curve control point information, and the starting point index and end point index of each contour can be determined. For complex glyphs that may contain multiple contours, the nesting and inclusion relationship between contours can be determined.

[0042] Sub-step 1013, based on the Bezier curve control point information and the contour topology structure information, determine the contour point information of the Chinese character, the curve type of each contour segment in the Chinese character, and the topological relationship between the contour segments in the Chinese character.

[0043] Contour points are the basic units that make up the outline of a Chinese character. Each contour point has specific location coordinates and type. Contour point information includes its coordinates and type. Contour point types can be linear or curved. Linear contour points define straight segments, while curved contour points define curved segments. Curved contour points are typically associated with control points of a Bezier curve. Contour point information accurately describes the shape of a Chinese character.

[0044] A Chinese character may include one or more contour segments. Each contour segment in a Chinese character may refer to a sub-segment in the glyph outline of the Chinese character. The curve type may be a straight line or a curve. The curve may further include a quadratic Bezier curve, a cubic Bezier curve, etc. The curve type specifies the mathematical characteristics and drawing method of each sub-segment in the glyph outline of the Chinese character. Different curve types correspond to different mathematical expressions and numbers of control points, which are used to accurately describe the geometric shape of the character outline. For example, a straight line segment is defined by two straight line points, while a cubic Bezier curve is defined by a start point, an end point, and two intermediate control points.

[0045] The topological relationships between the contour segments of a Chinese character describe the connections and inclusion relationships between the sub-segments within the character's glyph outline. Specifically, these relationships include the relative positions of contours, such as proximity, intersection, and separation; nesting relationships, such as whether one contour is contained within another; and contour orientations, such as clockwise or counterclockwise. Topological relationships are crucial for correctly understanding and parsing character structure, especially when dealing with complex characters, such as those with holes or multiple nested layers.

[0046] In practice, for each contour segment of a Chinese character, the curve type of that contour segment can be determined based on the type and number of its control points. For example, if a contour segment has only two control points, namely the start and end points, it is determined to be a straight line segment; if there are two intermediate control points, it is determined to be a cubic Bezier curve segment. At the same time, based on the contour topology information, the contour point coordinates and contour point type of each contour point can be determined. The contour point coordinates of a straight line contour point are directly determined by the start and end points of the straight line segment, while the contour point coordinates of a curved contour point can be determined by the control points of the Bezier curve. Finally, combined with the contour topology information, the topological relationship between each contour segment can be clarified, including the nesting relationship of contours, the adjacent relationship, and the distinction between inner and outer contours.

[0047] Sub-step 1014: generating vector contour data based on the contour point information, curve type, and topological relationship.

[0048] Specifically, the determined contour point information, curve types, and topological relationships can be integrated into a standardized vector glyph data structure. This data structure should contain the contour point sequence, curve type information, and contour relationships. In this way, complete vector contour data can be generated, ensuring that it accurately describes the shape and structure of Chinese characters.

[0049] By extracting the control point information of the Bezier curve and parsing it to obtain the contour topology, key data describing the Chinese character outline is obtained, providing a foundation for subsequent processing. This information determines the contour point information, curve type, and topological relationships, enabling precise identification and classification of the various components of the Chinese character outline, including straight and curved segments and their connections. Finally, this information is integrated to form vector contour data, preserving the precise contour and structural features of the original font, ensuring the accuracy and integrity of the vector contour data and providing high-quality data support for subsequent stroke extraction and analysis.

[0050] Step 102: Based on the vector outline data, skeleton extraction and curvature analysis are performed on the outline of the Chinese character to obtain the stroke unit of the Chinese character.

[0051] In this embodiment, skeleton extraction refers to simplifying the outline of a Chinese character into a skeleton line of a single pixel width to reflect the main structure of the Chinese character strokes. Specifically, the vector outline data of the Chinese character can be processed by a skeleton extraction algorithm to gradually shrink the vector outline into a skeleton line of a single pixel width. The skeleton extraction algorithm may include but is not limited to an algorithm based on distance transformation, an algorithm based on morphology, and an algorithm based on mean curvature flow contraction. Through skeleton extraction, the main structure of the Chinese character strokes is effectively extracted, the complexity of the outline is simplified, and the key topological information is retained, ensuring that the basic direction and connectivity of the strokes are accurately reflected. In this process, the vector outline data can correspond one-to-one with the skeleton data corresponding to the skeleton line.

[0052] Curvature analysis refers to calculating the curvature of each point on the skeleton line. For each point in the skeleton line, its curvature quantifies the degree to which the curve deviates from a straight line at that point. By performing curvature analysis on the skeleton line, inflection points, intersection points, etc. in the skeleton line can be identified. On this basis, the skeleton line can be segmented into sub-segments, and adjacent segments can be merged through direction consistency clustering to form independent stroke units. Among them, a stroke unit refers to a relatively independent stroke part segmented from the vector contour of a Chinese character, which contains information such as the shape and direction of the stroke and is the basis for subsequent stroke analysis and processing. Ideally, a stroke unit is the most basic writing unit of a Chinese character stroke and is the smallest connected part that makes up a Chinese character.

[0053] In some optional implementation manners of this embodiment, the stroke units of a Chinese character can be obtained through the following sub-steps:

[0054] Sub-step 1021: Process the vector contour data through a morphological topology analysis algorithm to obtain the skeleton line of the Chinese character. Specifically, a morphological topology analysis algorithm, such as a skeleton extraction method based on mean curvature flow contraction, can be used to process the continuous vector contour of the Chinese character and convert the complex Chinese character contour into a single-pixel-width topological skeleton network, that is, the skeleton line. As an example, see Figure 2A , the skeleton line of the character "维" is shown in reference numeral 201. As another example, see Figure 2B , the skeleton line of the character "沃" is shown in reference numeral 203.

[0055] Sub-step 1022: Determine the curvature of each point in the skeleton line.

[0056] Sub-step 1023: Based on the curvature, determine the feature points in the skeleton line. Here, a feature point is a point with significant geometric features in a curve or shape, such as an endpoint, a starting point, an ending point, an inflection point, an intersection point, etc. These points play an important role in shape analysis and pattern recognition and can be used for segmenting curves, identifying shape features, etc.

[0057] Sub-step 1024: Based on the feature points, segment the skeleton line into multiple sub-segments. Here, a sub-segment is a continuous curve after the skeleton line is segmented by the feature points, and each sub-segment corresponds to a part of the stroke. Based on the feature points, a continuous skeleton line can be segmented into several smaller sub-segments according to specific rules or feature points.

[0058] Sub-step 1025: Cluster the sub-segments, and segment the skeleton line based on the clustering results to obtain the stroke units of the Chinese character. Among them, clustering is used to group the segmented sub-segments according to rules such as spatial proximity relationship, direction consistency, and connection characteristics to identify independent stroke units. Using hierarchical clustering can avoid over-segmentation or incorrect merging of strokes, and improve the rationality and stability of stroke segmentation. In this process, contour data and skeleton line data can be comprehensively utilized for collaborative segmentation.

[0059] Sub-step 1026: Use the Bezier curve fitting algorithm to smooth the stroke units to update the stroke units. Among them, the Bezier curve fitting algorithm is an algorithm used to fit discrete points or line segments into smooth Bezier curves. Bezier curves have good mathematical properties and controllability, and can change the shape of the curve by adjusting the control points. Through the Bezier curve fitting algorithm, the segmented stroke units can be smoothed to eliminate noise and discontinuity, making the stroke lines more smooth and natural. Continue to refer to Figure 2A and Figure 2B , the stroke units of the character "维" are shown in reference numeral 202; the stroke units of the character "沃" are shown in reference numeral 204.

[0060] By extracting the skeleton line of the Chinese character, the complex contour can be simplified into the central axis that can reflect the main structural features, greatly reducing the data complexity while retaining the key topological information. By calculating the curvature of each point in the skeleton line and determining the feature points, the important geometric feature positions in the stroke can be accurately identified, such as endpoints, inflection points, and intersection points, etc. These feature points provide the key basis for subsequent stroke segmentation. By cutting the skeleton line into multiple sub-segments through the feature points and clustering them, the parts belonging to the same stroke can be merged into independent stroke units, effectively solving the problem that it is difficult to separate the overlapping strokes, and improving the accuracy of stroke segmentation. By using the Bezier curve fitting algorithm to smooth the stroke units, noise and discontinuity can be eliminated, making the stroke lines more in line with the characteristics of actual writing, and enhancing the quality and usability of the stroke units. Thus, by combining technical means such as skeleton extraction, curvature analysis, and Bezier curve fitting, the accurate splitting and optimization of Chinese character strokes are achieved, improving the accuracy and effect of stroke extraction, and providing high-quality data support for subsequent applications such as Chinese character recognition, font generation, and calligraphy analysis.

[0061] Step 103: Convert the vector contour data into a character image, and segment the character image based on the stroke units to obtain the stroke images of each stroke in the Chinese character.

[0062] In this embodiment, the character image is a binary bitmap generated by rasterizing vector contour data, which retains the pixel-level form of the character. In practice, the scan line algorithm can be used to map the vector path represented by the vector contour data into a binary image, and anti-aliasing technology can be used to eliminate edge jaggedness to obtain the final character image. As an example, the character image of the character "维" is shown in Figure 3A shown, and the character image of the character "沃" is shown in Figure 3B shown.

[0063] In this embodiment, the stroke image is an image containing independent stroke regions, which can be obtained by performing stroke-level image segmentation on the character image. Specifically, it can be executed according to the following steps:

[0064] First, use the iterative closest point algorithm to align each stroke unit with the character image in terms of spatial position, and eliminate the deviation caused by factors such as rendering.

[0065] Next, perform Fourier transform on the aligned stroke units to extract their frequency domain features. Using Fourier descriptors, evaluate the consistency between the stroke units and the Chinese character contour in terms of shape and structure. Through shape matching, the possible errors between the stroke units and the character image can be corrected to ensure the accuracy of subsequent segmentation.

[0066] Next, perform a structure thinning algorithm on the character image. Adopt an adaptive edge processing method to process the character boundary to eliminate noise interference. The adaptive kernel function can dynamically adjust parameters according to the local features of the image, so as to retain the details of the strokes while removing noise. After the thinning process, the character image may have breaks or detail loss. For this reason, a morphological compensation mechanism needs to be started. By setting a dynamic threshold, reconstruct the topological structure of the character to compensate for the parts that may be lost during the thinning process. The dynamic threshold is adjusted according to the characteristics of the strokes to ensure the best compensation effect. This mechanism can restore the topological integrity of the strokes, maintain the connectivity and overall form of the strokes.

[0067] Next, the topological separation method can be used to analyze the optimized character image. By detecting connected regions and separation features, generate independent stroke regions, where each region represents an independent stroke. Thus, the stroke-level segmentation of the character image can be achieved, and the stroke images of each stroke in the Chinese character can be obtained.

[0068] In some optional implementation manners of this embodiment, after obtaining the stroke images of each stroke in the above Chinese character, shape matching verification can also be performed on them, including the following steps:

[0069] In the first step, the stroke image is matched with the standard stroke library to obtain the similarity between the stroke image and the standard stroke. Specifically, the standard stroke library of Chinese characters can be combined to compare the shape of each segmented stroke with the standard strokes in the pre-established standard stroke library. In practice, the two methods of key points and scale-invariant feature transformation can be combined for comparison. In the similarity evaluation process, Hausdorff distance and shape context matching algorithm can be used. Hausdorff distance is used to evaluate the maximum difference between stroke shapes and reflect the degree of deviation of the overall shape; the shape context matching algorithm is based on the local and global features of the strokes, calculates the matching score, and carefully evaluates the accuracy of the segmentation results. Through these quantitative indicators, the similarity between the segmented strokes and the standard strokes can be effectively measured.

[0070] In the second step, if the similarity is lower than the threshold, the image segmentation parameters are adjusted, and the character image is re-segmented based on the adjusted image segmentation parameters. In practice, a reasonable threshold can be set based on the similarity evaluation results to judge the consistency of the strokes. For strokes with similarity lower than the threshold, it is considered that there may be segmentation errors or anomalies. Then, the key parameters of the stroke segmentation algorithm are dynamically adjusted. The adjusted parameters may include the segmentation threshold, the size of the segmentation area, the segmentation step size, etc. After the parameters are adjusted, the stroke segmentation process is re-executed, and the accuracy of the segmentation results is gradually improved through iterative optimization. The iterative process is terminated when the similarity of the stroke matching results reaches a preset standard to ensure that the segmentation results meet the requirements. At the same time, to balance the computational cost, a maximum number of iterations is set. When the number of iterations reaches the upper limit, the iteration is terminated even if the similarity does not meet the standard.

[0071] By performing shape matching verification on the stroke segmentation results and iteratively optimizing the segmentation process based on the shape matching verification results, the accuracy and reliability of the segmentation results are ensured. Figure 4A As shown in the figure, the character has only 3 outline data in the TTF font file, but actually has 11 strokes. Figure 4B As shown, the character has 4 outline data in the TTF font file, but actually has 7 strokes.

[0072] In some optional implementations of this embodiment, the vector outline data may be converted into a character image through the following sub-steps:

[0073] Sub-step 1031 : extracting contour point information from the vector contour data. The contour point information includes contour point coordinates and contour point types. The contour point types include straight line contour points and curve contour points.

[0074] Line contour points are the endpoints of a line segment, connected by a straight line. In vector graphics, line contour points define the position and orientation of a line segment. Curve contour points are key points used to define the shape of a curve, typically associated with Bezier curves. Each curve contour point consists of a position coordinate and one or more control points that determine the curvature and direction of the curve.

[0075] In practice, vector contour data can be decoded to extract key geometric features, including the coordinates of each contour point and a flag indicating the type of contour point. By parsing the mathematical representation of the vector contour, the spatial properties of the vertex, its position within the contour structure, and its connection to adjacent vertices can be obtained.

[0076] Sub-step 1032, based on the contour point information, reconstructs the contour of the Chinese character to obtain a reconstructed contour path. The reconstruction process includes: connecting adjacent straight line contour points; interpolating the curve contour points based on the Bezier curve interpolation algorithm; and connecting the beginning and end of the curve contour segment to be closed.

[0077] In practice, contour points can be traversed and the contour point type of each contour point determined based on the flag bit. For linear sections, adjacent straight contour points can be directly connected to construct precise straight line segments, fully preserving the contour's linear characteristics. For curved sections, a curve interpolation algorithm can be used to generate smooth curved paths based on control point parameters. A quadratic Bezier curve model is used to ensure the smoothness and accuracy of the curved segments, faithfully reproducing the curve characteristics of the original vector contour. To ensure the integrity and correctness of the contours, the closure and directionality of the paths can be manipulated. For contours that should be closed, each contour's endpoint is connected to its starting point to form a complete closed path. If any paths are not closed, adjustments must be made to fill in the missing connections. Contour direction determines the contour type based on the direction in which the path is drawn. Typically, outer contours are drawn counterclockwise to define the glyph's outer boundary; inner contours are drawn clockwise to represent the glyph's blank area. By correctly identifying contour direction, this solution applies the even-odd rule when filling, ensuring accurate rendering of the graphics.

[0078] Sub-step 1033 , mapping the reconstructed contour path to the pixel space to generate a character image.

[0079] Specifically, a rasterization conversion engine can be used to map the reconstructed vector contour path to pixel space. Through a scan conversion algorithm, the reconstructed vector path is converted into a black and white binary bitmap format to achieve the conversion from vector space to pixel space. During the mapping process, anti-aliasing technology can be used to ensure the details of the contour and the clarity of the edges, avoiding the jagged effect caused by pixelation. In addition, the rasterization accuracy parameters can be adjusted to balance the rendering speed and image quality to ensure the geometric shape of the character image is accurate. In addition, the distortion and distortion that may occur during the conversion process can be monitored and corrected to ensure the consistency of the character image with the original vector data.

[0080] By extracting contour point information, we can accurately capture the shape features of characters. Reconstructing the contour ensures the integrity and accuracy of the character outline. Using a rasterization conversion engine, we map the reconstructed contour path to pixel space, enabling the generated character image to accurately reflect the shape of the vector outline, providing high-quality input for subsequent image processing and analysis. This process effectively preserves the character's detailed features, ensuring the clarity and accuracy of the character image.

[0081] In some optional implementations of this embodiment, stroke segmentation can be performed directly on the vector outline data to more accurately utilize the mathematical properties of the outline. This not only avoids the accuracy loss caused by image conversion but also improves the efficiency of stroke segmentation. Furthermore, operating directly on the vector outline data also facilitates preserving the vector characteristics of the strokes, facilitating subsequent scaling, deformation, and other processing.

[0082] Step 104: extracting geometric features from the stroke image and generating stroke data of the Chinese character based on the geometric features.

[0083] In this embodiment, geometric features are features that characterize the shape of a stroke in a stroke image. Specifically, they may include, but are not limited to, at least one of the following: length, width, direction, angle, curvature, etc. Geometric features can be extracted using image feature extraction. By extracting geometric features from a stroke image, key stroke information can be obtained, providing a foundation for subsequent stroke data generation.

[0084] For each stroke image, the geometric features extracted from it can be analyzed and post-processed to obtain structured data, which can be used as stroke data. This stroke data can provide data support for subsequent applications such as Chinese character recognition, font generation, and calligraphy analysis.

[0085] In some optional implementations of this embodiment, the stroke data of a Chinese character may be generated by the following sub-steps:

[0086] Sub-step 1041 : determining the central axis of the stroke in the stroke image, and the bilateral distances from each point on the central axis to the stroke outline boundary.

[0087] In practice, an algorithm based on the principle of minimum enclosure can be adopted to calculate the minimum circumscribed region of each stroke unit and establish a boundary envelope model of the stroke. This process accurately describes the overall shape contour and spatial range of the stroke. Then, using the method of edge detection combined with path tracking, traverse pixel by pixel along the boundary of the stroke to capture the subtle changes and curvature features of the boundary. After completing the boundary tracking, convert the obtained boundary point information into a sequence of geometric feature points conforming to the Cartesian coordinate system. Each feature point contains clear horizontal and vertical coordinates and is located on a two-dimensional plane. These sequences of geometric feature points completely describe the contour shape and spatial position of the stroke.

[0088] After obtaining the precise contour of the stroke, the generation of the central axis can be carried out. Specifically, based on the stroke contour data, a central axis reflecting the main structure of the stroke can be generated as a benchmark for calculating the geometric parameters of the stroke. Then, based on the extracted set of stroke contour coordinate points, adopt a spatial partitioning algorithm to partition the stroke region and generate a spatial segmentation network of the stroke. After that, identify the key nodes in the spatial segmentation network, including endpoints, bifurcation points, and intersection points. Next, adopt an optimal path algorithm to find the main path running through the entire stroke. In order to improve the accuracy and geometric continuity of the central axis, it is necessary to implement the spatial coupling optimization of the central axis and the contour, and use the dynamic alignment technology to match and adjust the initial central axis and the stroke contour. By adjusting the position, direction, and shape of the central axis, the best match between the central axis and the contour is achieved, ensuring that the central axis is accurately located at the center of the stroke. As an example, the central axis of the strokes in the stroke image of the character "维" can be seen in Figure 5A As shown, the central axis of the strokes in the stroke image of the character "沃" can be seen in Figure 5B As shown.

[0089] After obtaining the optimized central axis, the distances from each point on the central axis to the boundary of the stroke contour can be calculated. Establish a two-way ranging model. For each discrete point on the central axis, calculate the shortest distances from its left and right sides to the boundary of the stroke contour to obtain the width information of the stroke at that point. Among them, the bilateral distance is used to describe the left and right distances from each point on the central axis to the boundary of the stroke contour.

[0090] Sub-step 1042, based on the bilateral distance, determine the width information corresponding to each point. Among them, the width information is used to represent the thickness of the stroke at different positions. For each point on the central axis, the width information corresponding to that point can be obtained by adding the bilateral distances from that point to the boundary of the stroke contour.

[0091] Sub-step 1043: Gaussian filtering is performed on the width information to smooth the width information. A Gaussian filter can be used to smooth the width information. The size and standard deviation of the Gaussian filter can be selected based on actual needs. Gaussian filtering can reduce noise and fluctuations in the width information, resulting in smoother width changes.

[0092] Sub-step 1044: derive geometric features based on the central axis and the updated width information. Here, the central axis and the updated width information can be combined to form the geometric features of the stroke. The geometric features can be represented as a sequence of coordinate points of the central axis and a corresponding sequence of width information.

[0093] By determining the central axis and bilateral distances of a stroke image, we can accurately describe the stroke's center direction and thickness variations. By applying Gaussian filtering to the width information, we can reduce the influence of noise and detail, making width variations smoother and resulting in a more stable feature description. By forming geometric features based on the central axis and updated width information, we can comprehensively describe the shape and size characteristics of the stroke. This effectively extracts the key geometric features of the stroke, providing high-quality data support for subsequent stroke recognition, classification, and analysis.

[0094] In some optional implementations of this embodiment, after obtaining the stroke data, the stroke data may be compressed, stored, and output. Specifically, the following steps may be performed:

[0095] The first step is to structure the data. First, the stroke information for each Chinese character is structured. For each character, its character encoding value, such as Unicode, can be stored to uniquely identify it for easy retrieval and mapping. Its stroke index, which assigns a unique index value to each stroke to determine its order and number, can also be stored. Stroke coordinates and width information can also be stored to accurately describe the spatial position and morphological characteristics of the strokes. This data can then be organized into a unified structured format for easier storage and processing.

[0096] The second step is data compression. To improve storage efficiency, structured stroke data can be compressed. For coordinate data, differential encoding can be used to calculate the coordinate differences between adjacent points. Only the differential values ​​can be stored to reduce data redundancy. Furthermore, the precision of the coordinate data can be appropriately reduced, and the coordinate values ​​can be quantized to reduce data size while still meeting accuracy requirements. For width information, predictive encoding can be used to compress the data by leveraging its continuity and recording only the width change.

[0097] Step 3: Data storage. Specifically, an efficient binary data format can be adopted to uniformly store character encoding, stroke index, and the coordinates and widths of strokes, ensuring data compactness and readability. As an example, the final Chinese character splitting result of the character "维" can be seen in Figure... Figure 6A As shown, the final Chinese character splitting result of the character "沃" can be seen in Figure... Figure 6B as shown.

[0098] The method provided in the above embodiments of the present application first parses the font file to generate vector contour data including the shape information and structural information of Chinese characters; then, based on the vector contour data, performs skeleton extraction and curvature analysis on the contours of Chinese characters to obtain the stroke units of Chinese characters; then converts the vector contour data into character images, segments the character images based on the stroke units to obtain the stroke images of each stroke in the Chinese character; finally, extracts geometric features from the stroke images and generates the stroke data of Chinese characters based on the geometric features. In the above process, by combining the parsing of vector contour data and image processing, and using methods such as skeleton extraction and curvature analysis, complex Chinese character strokes can be accurately decomposed, avoiding the dependence on training data, thereby improving the accuracy and efficiency of stroke extraction.

[0099] It should be noted that for the stroke splitting method provided in the embodiments of the present application, the execution subject can be a stroke splitting device. In the embodiments of the present application, the stroke splitting method executed by the stroke splitting device is used as an example to illustrate the stroke splitting device provided in the embodiments of the present application.

[0100] As Figure 7 shown, the stroke splitting device 700 in this embodiment includes: a first generation unit 701 for extracting vector contour data from the font file of a Chinese character, where the vector contour data includes the shape information and structural information of the Chinese character; a processing unit 702 for performing skeleton extraction and curvature analysis on the contour of the Chinese character based on the vector contour data to obtain the stroke units of the Chinese character; a segmentation unit 703 for converting the vector contour data into a character image and segmenting the character image based on the stroke units to obtain the stroke images of each stroke in the Chinese character; a second generation unit 704 for extracting geometric features from the stroke images and generating the stroke data of the Chinese character based on the geometric features.

[0101] In some optional implementations of this embodiment, the first generation unit 701 is further configured to: extract the Bezier curve control point information of the Chinese character from the font file; determine the outline topology structure information of the Chinese character based on the Bezier curve control point information; determine the outline point information of the Chinese character, the curve type of each outline segment in the Chinese character, and the topological relationship between the outline segments in the Chinese character based on the Bezier curve control point information and the outline topology structure information; and generate the vector outline data based on the outline point information, the curve type, and the topological relationship. Thus, the precise outline and structural features of the original font are preserved, ensuring the accuracy and integrity of the vector outline data and providing high-quality data support for subsequent stroke extraction and analysis.

[0102] In some optional implementations of this embodiment, the processing unit 702 is further configured to: process the vector contour data using a morphological topology analysis algorithm to obtain the skeleton line of the Chinese character; determine the curvature of each point in the skeleton line; determine the feature points in the skeleton line based on the curvature; divide the skeleton line into a plurality of sub-segments based on the feature points; cluster the sub-segments, and segment the skeleton line based on the clustering results to obtain the stroke units of the Chinese character; and smooth the stroke units using a Bezier curve fitting algorithm to update the stroke units. Thus, by combining skeleton extraction, curvature analysis, and Bezier curve fitting techniques, accurate segmentation and optimization of Chinese character strokes are achieved, the accuracy and effect of stroke extraction are improved, and high-quality data support is provided for subsequent applications such as Chinese character recognition, font generation, and calligraphy analysis.

[0103] In some optional implementations of this embodiment, the segmentation unit is further configured to: extract contour point information from the vector contour data, the contour point information including contour point coordinates and contour point types, the contour point types including straight contour points and curved contour points; reconstruct the contour of the Chinese character based on the contour point information to obtain a reconstructed contour path, the reconstruction process including: connecting adjacent straight contour points; interpolating curved contour points based on a Bezier curve interpolation algorithm; connecting the beginning and end of the curved contour segment to be closed; and mapping the reconstructed contour path to pixel space to generate a character image. By extracting the contour point information, the shape features of the character can be accurately obtained. By reconstructing the contour, the integrity and accuracy of the character contour are ensured. By mapping the reconstructed contour path to pixel space through a rasterization conversion engine, the generated character image can accurately reflect the shape of the vector contour, providing high-quality input for subsequent image processing and analysis. This process effectively preserves the detailed features of the character and ensures the clarity and accuracy of the character image.

[0104] In some optional implementations of this embodiment, the device further includes a verification unit configured to: perform shape matching on the stroke image with a standard stroke library to determine a similarity between the stroke image and the standard strokes; if the similarity is lower than a threshold, adjust image segmentation parameters, and re-segment the character image based on the adjusted image segmentation parameters. By performing shape matching verification on the stroke segmentation results and iteratively optimizing the segmentation process based on the shape matching verification results, the accuracy and reliability of the segmentation results are ensured.

[0105] In some optional implementations of this embodiment, the second generation unit is further configured to: determine the central axis of the stroke in the stroke image and the bilateral distances between each point on the central axis and the stroke outline boundary; determine width information corresponding to each point based on the bilateral distances; perform Gaussian filtering and smoothing on the width information to update the width information; and obtain the geometric features based on the central axis and the updated width information. This effectively extracts the key geometric features of the stroke, providing high-quality data support for subsequent stroke recognition, classification, and analysis.

[0106] The device provided by the above-mentioned embodiment of the present application first parses the font file to generate vector contour data including the shape information and structural information of the Chinese character; then, based on the vector contour data, skeleton extraction and curvature analysis are performed on the contour of the Chinese character to obtain the stroke unit of the Chinese character; then, the vector contour data is converted into a character image, and the character image is segmented based on the stroke unit to obtain the stroke image of each stroke in the Chinese character; finally, geometric features are extracted from the stroke image, and the stroke data of the Chinese character is generated based on the geometric features. In the above process, by combining the analysis of the vector contour data and image processing, and utilizing methods such as skeleton extraction and curvature analysis, complex Chinese character strokes can be accurately decomposed, avoiding dependence on training data, thereby improving the accuracy and efficiency of stroke extraction.

[0107] The stroke splitting device in the embodiment of the present application can be an electronic device or a component in the electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices other than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a mobile Internet device (Mobile Internet Device, MID), an augmented reality (Augmented Reality, AR) / virtual reality (Virtual Reality, VR) device, a robot, a wearable device, an ultra-mobile personal computer (Ultra-Mobile Personal Computer, UMPC), a netbook or a personal digital assistant (Personal Digital Assistant, PDA), etc. It can also be a server, a network attached storage (Network Attached Storage, NAS), a personal computer (Personal Computer, PC), a television (Television, TV), a teller machine or a self-service machine, etc., and the embodiment of the present application does not make specific limitations.

[0108] The stroke splitting device in the embodiment of the present application can be a device having an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0109] The stroke splitting device provided in the embodiment of the present application can achieve Figure 1 To avoid repetition, the various processes implemented in the method embodiment are not described here.

[0110] Alternatively, as Figure 8 As shown, an embodiment of the present application also provides an electronic device 800, including a processor 801 and a memory 802, wherein the memory 802 stores a program or instruction that can be run on the processor 801, and when the program or instruction is executed by the processor 801, the various steps of the above-mentioned stroke splitting method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0111] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0112] Figure 9 A schematic diagram of the hardware structure of an electronic device implementing an embodiment of the present application.

[0113] The electronic device 900 includes but is not limited to components such as a radio frequency unit 901 , a network module 902 , an audio output unit 903 , an input unit 904 , a sensor 905 , a display unit 906 , a user input unit 907 , an interface unit 908 , a memory 909 , and a processor 910 .

[0114] Those skilled in the art will understand that the electronic device 900 may also include a power source (such as a battery) to power each component, and the power source may be logically connected to the processor 910 through a power management system, thereby implementing functions such as charging, discharging, and power consumption management through the power management system. Figure 9 The electronic device structure shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently, which will not be repeated here.

[0115] Among them, the processor 910 is used to extract vector outline data from the font file of the Chinese character, and the vector outline data includes the shape information and structure information of the Chinese character; convert the vector outline data into a character image, perform skeleton extraction and curvature analysis on the character image, and obtain the stroke unit of the Chinese character; segment the character image based on the stroke unit to obtain the stroke image of each stroke in the Chinese character; extract geometric features from the stroke image, and generate the stroke data of the Chinese character based on the geometric features.

[0116] By combining the analysis of vector contour data and image processing, and using methods such as skeleton extraction and curvature analysis, complex Chinese character strokes can be accurately decomposed, avoiding dependence on training data, thereby improving the accuracy and efficiency of stroke extraction.

[0117] In some optional implementations of this embodiment, the processor 910 is further configured to extract Bezier curve control point information of the Chinese character from the font file; determine the outline topology structure information of the Chinese character based on the Bezier curve control point information; determine the outline point information of the Chinese character, the curve type of each outline segment in the Chinese character, and the topological relationship between the outline segments in the Chinese character based on the Bezier curve control point information and the outline topology structure information; and generate the vector outline data based on the outline point information, the curve type, and the topological relationship. This preserves the precise outline and structural features of the original font, ensures the accuracy and integrity of the vector outline data, and provides high-quality data support for subsequent stroke extraction and analysis.

[0118] In some optional implementations of this embodiment, the processor 910 is further configured to process the vector contour data using a morphological topology analysis algorithm to obtain the skeleton line of the Chinese character; determine the curvature of each point in the skeleton line; determine the feature points in the skeleton line based on the curvature; divide the skeleton line into a plurality of sub-segments based on the feature points; cluster the sub-segments, and segment the skeleton line based on the clustering results to obtain the stroke units of the Chinese character; and smooth the stroke units using a Bezier curve fitting algorithm to update the stroke units. Thus, by combining skeleton extraction, curvature analysis, and Bezier curve fitting techniques, accurate segmentation and optimization of Chinese character strokes are achieved, the accuracy and effect of stroke extraction are improved, and high-quality data support is provided for subsequent applications such as Chinese character recognition, font generation, and calligraphy analysis.

[0119] In some optional implementations of this embodiment, the processor 910 is further configured to extract contour point information from the vector contour data, the contour point information including contour point coordinates and contour point types, the contour point types including straight contour points and curved contour points; reconstruct the contour of the Chinese character based on the contour point information to obtain a reconstructed contour path, the reconstruction process including: connecting adjacent straight contour points; interpolating the curved contour points based on a Bezier curve interpolation algorithm; connecting the beginning and end of the curved contour segment to be closed; and mapping the reconstructed contour path to pixel space to generate a character image. By extracting the contour point information, the shape features of the character can be accurately obtained. By reconstructing the contour, the integrity and accuracy of the character contour are ensured. By mapping the reconstructed contour path to pixel space through a rasterization conversion engine, the generated character image can accurately reflect the shape of the vector contour, providing high-quality input for subsequent image processing and analysis. This process effectively preserves the detailed features of the character and ensures the clarity and accuracy of the character image.

[0120] In some optional implementations of this embodiment, the processor 910 is further configured to perform shape matching on the stroke image and a standard stroke library to determine a similarity between the stroke image and the standard strokes; if the similarity is lower than a threshold, the image segmentation parameters are adjusted, and the character image is re-segmented based on the adjusted image segmentation parameters. The accuracy and reliability of the segmentation results are ensured by performing shape matching verification on the stroke segmentation results and iteratively optimizing the segmentation process based on the shape matching verification results.

[0121] In some optional implementations of this embodiment, the processor 910 is further configured to determine the central axis of the stroke in the stroke image and the bilateral distances between each point on the central axis and the stroke outline boundary; determine the width information corresponding to each point based on the bilateral distances; perform Gaussian filtering and smoothing on the width information to update the width information; and obtain the geometric features based on the central axis and the updated width information. This effectively extracts the key geometric features of the stroke, providing high-quality data support for subsequent stroke recognition, classification, and analysis.

[0122] It should be understood that in an embodiment of the present application, the input unit 904 may include a graphics processing unit (GPU) 9041 and a microphone 9042, and the graphics processor 9041 processes the image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 906 may include a display panel 9061, and the display panel 9061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 907 includes a touch panel 9071 and at least one of other input devices 9072. The touch panel 9071 is also called a touch screen. The touch panel 9071 may include two parts: a touch detection device and a touch controller. Other input devices 9072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and an operating stick, which will not be repeated here.

[0123] The memory 909 can be used to store software programs and various data. The memory 909 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 909 may include a volatile memory or a non-volatile memory, or the memory 909 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM), and a direct memory bus random access memory (DRRAM). The memory 909 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.

[0124] Processor 910 may include one or more processing units. Optionally, processor 910 integrates an application processor and a modem processor. The application processor primarily handles operations related to the operating system, user interface, and application programs, while the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into processor 910.

[0125] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned stroke splitting method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0126] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0127] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned stroke splitting method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0128] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0129] An embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the above-mentioned stroke splitting method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0130] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0131] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0132] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

Claims

1. A stroke splitting method, characterized in that: The method comprises: Extracting vector outline data from a Chinese character font file, wherein the vector outline data includes shape information and structure information of the Chinese character; Based on the vector outline data, skeleton extraction and curvature analysis are performed on the outline of the Chinese character to obtain the stroke unit of the Chinese character; Converting the vector outline data into a character image, segmenting the character image based on the stroke unit to obtain a stroke image of each stroke in the Chinese character; Geometric features are extracted from the stroke image, and stroke data of the Chinese character is obtained based on the geometric features.

2. The method according to claim 1, characterized in that The step of extracting vector outline data from the Chinese character font file includes: Extracting Bezier curve control point information of the Chinese character from the font file; Determining the outline topological structure information of the Chinese character based on the Bezier curve control point information; Determining, based on the Bezier curve control point information and the contour topology structure information, contour point information of the Chinese character, a curve type of each contour segment in the Chinese character, and a topological relationship between contour segments in the Chinese character; The vector contour data is generated based on the contour point information, the curve type and the topological relationship.

3. The method according to claim 1, characterized in that The step of performing skeleton extraction and curvature analysis on the outline of the Chinese character based on the vector outline data to obtain the stroke unit of the Chinese character includes: Processing the vector outline data using a morphological topology analysis algorithm to obtain the skeleton line of the Chinese character; determining the curvature of each point in the skeleton line; determining feature points in the skeleton line based on the curvature; Based on the feature points, the skeleton line is divided into a plurality of sub-segments; Clustering the sub-segments, and segmenting the skeleton lines based on the clustering results to obtain the stroke units of the Chinese character; A Bezier curve fitting algorithm is used to smooth the stroke unit to update the stroke unit.

4. The method according to claim 1, wherein The converting the vector outline data into a character image comprises: Extracting contour point information from the vector contour data, wherein the contour point information includes contour point coordinates and contour point types, and the contour point types include straight line contour points and curve contour points; Based on the contour point information, the contour of the Chinese character is reconstructed to obtain a reconstructed contour path, wherein the reconstruction includes: connecting adjacent straight contour points; interpolating the curve contour points based on a Bezier curve interpolation algorithm; and connecting the beginning and end of the curve contour segment to be closed; The reconstructed contour path is mapped to pixel space to generate a character image.

5. The method according to claim 1, wherein After obtaining the stroke image of each stroke in the Chinese character, the method further includes: Performing shape matching on the stroke image and a standard stroke library to obtain a similarity between the stroke image and the standard stroke; If the similarity is lower than a threshold, the image segmentation parameters are adjusted, and the character image is re-segmented based on the adjusted image segmentation parameters.

6. A stroke splitting device, characterized in that: The device comprises: A first generating unit is configured to extract vector outline data from a Chinese character font file, wherein the vector outline data includes shape information and structure information of the Chinese character; a processing unit, configured to perform skeleton extraction and curvature analysis on the outline of the Chinese character based on the vector outline data to obtain the stroke units of the Chinese character; a segmentation unit, configured to convert the vector outline data into a character image, and segment the character image based on the stroke unit to obtain a stroke image of each stroke in the Chinese character; The second generating unit is configured to extract geometric features from the stroke image and generate the stroke data of the Chinese character based on the geometric features.

7. The device according to claim 6, characterized in that The first generating unit is further configured to: Extracting Bezier curve control point information of the Chinese character from the font file; Determining the outline topological structure information of the Chinese character based on the Bezier curve control point information; Determining, based on the Bezier curve control point information and the contour topology structure information, contour point information of the Chinese character, a curve type of each contour segment in the Chinese character, and a topological relationship between contour segments in the Chinese character; The vector contour data is generated based on the contour point information, the curve type and the topological relationship.

8. The device according to claim 6, characterized in that The processing unit is further configured to: Processing the vector outline data using a morphological topology analysis algorithm to obtain the skeleton line of the Chinese character; determining the curvature of each point in the skeleton line; determining feature points in the skeleton line based on the curvature; Based on the feature points, the skeleton line is divided into a plurality of sub-segments; Clustering the sub-segments, and segmenting the skeleton lines based on the clustering results to obtain the stroke units of the Chinese character; A Bezier curve fitting algorithm is used to smooth the stroke unit to update the stroke unit.

9. The device according to claim 6, characterized in that The segmentation unit is further configured to: Extracting contour point information from the vector contour data, wherein the contour point information includes contour point coordinates and contour point types, and the contour point types include straight line contour points and curve contour points; Based on the contour point information, the contour of the Chinese character is reconstructed to obtain a reconstructed contour path, wherein the reconstruction includes: connecting adjacent straight contour points; interpolating the curve contour points based on a Bezier curve interpolation algorithm; and connecting the beginning and end of the curve contour segment to be closed; The reconstructed contour path is mapped to pixel space to generate a character image.

10. The device according to claim 6, characterized in that The device further comprises a verification unit, configured to: Performing shape matching on the stroke image and a standard stroke library to obtain a similarity between the stroke image and the standard stroke; If the similarity is lower than a threshold, the image segmentation parameters are adjusted, and the character image is re-segmented based on the adjusted image segmentation parameters.

Citation Information

Patent Citations

  • Character image vectorization method and system based on framework instruction

    CN103942552A

  • Chinese font and font library generation method based on deep learning and component splicing

    CN112784531A

  • Handwritten Chinese text segmentation method based on Snake

    CN113723413A