A method and apparatus for handwritten Chinese character stroke decomposition based on iterative adaptation and matching algorithms
By using iterative adaptation and matching algorithms, the problem of key point detection and matching in the stroke segmentation of handwritten Chinese characters is solved, achieving efficient and accurate stroke segmentation, which is suitable for robot writing systems.
Patent Information
- Application Number
- CN202310512711.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-09
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2043-05-09
AI Technical Summary
Existing contour-based stroke decomposition methods struggle to accurately detect key points and match stroke decomposition points when dealing with handwritten Chinese characters, requiring manual intervention and correction, which affects the decomposition results.
A handwritten Chinese character stroke segmentation method based on iterative adaptation and matching algorithms is adopted. The method involves image preprocessing, key point detection using iterative adaptation point method, construction of minimum weight cost function for concave point matching, and optimization of stroke segmentation using Hungarian matching algorithm.
It achieves high-accuracy stroke decomposition without human intervention, preserves the stylistic features of handwritten Chinese characters, and is suitable for robot writing systems.
Smart Images

Figure CN116959002B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of calligraphy digitization, specifically relating to a method and apparatus for decomposing handwritten Chinese character strokes based on iterative adaptation and matching algorithms. Background Technology
[0002] High-accuracy stroke segmentation technology for handwritten Chinese characters is one of the core technologies for calligraphy digitization. Current research methods mainly focus on two types: contour-based methods and refinement-based methods. However, in robotic writing systems, in order to preserve the stylistic features of the target text as much as possible, the following problems have been encountered in the research of stroke segmentation based on handwritten Chinese character contours:
[0003] (1) Traditional outline-based stroke segmentation methods have a high accuracy rate for key point detection in standard fonts because the outline polygons of the whole character are relatively regular. However, it is difficult to detect key points for handwritten fonts with variable outline curvature. Furthermore, the accuracy of key point detection greatly affects the overall segmentation effect of the algorithm.
[0004] (2) Traditional outline-based stroke decomposition methods mainly analyze the outline of handwritten Chinese characters, calculate features such as outline curvature, outline inflection point, and outline concave point, and then determine the key points of stroke decomposition. However, they lack the necessary constraints on whether the stroke decomposition points match, making it difficult to match. The final stroke decomposition point pairs often require manual intervention and secondary correction to obtain more accurate strokes. Summary of the Invention
[0005] To address the shortcomings of existing technologies, reduce the difficulty of detecting key points in handwritten Chinese character stroke segmentation due to contour curvature variations, and constrain the matching of stroke segmentation point pairs to improve the accuracy and efficiency of handwritten Chinese character stroke segmentation, this invention adopts the following technical solution:
[0006] The handwritten Chinese character stroke decomposition method based on iterative adaptation and matching algorithms includes the following steps:
[0007] Step 1: Preprocess the handwritten Chinese character image to obtain character edges;
[0008] Step 2: Detect key points of handwritten Chinese characters based on iterative adaptive point method (DP) using the pixel points on the character edges to obtain a set of pixel coordinate key points;
[0009] Step 3: Analysis of key points for decomposing handwritten Chinese characters; by using the concavity and convexity of the vertices of the geometric shape formed by the key points of pixel coordinates, concave points are used as feature points for decomposing handwritten Chinese characters.
[0010] Step 4: Construct the minimum weight cost function. Based on the Minkowski distance, cosine distance and average stroke width between concave points, calculate the minimum weight for matching the stroke splitting feature points of handwritten Chinese characters.
[0011] Step 5: Based on Hungarian matching (HMA), handwritten Chinese character strokes are split. Based on the minimum weight cost function, the concave point set is matched to obtain the matching relationship. For unmatched concave points, the concave point set is traversed. Based on the minimum weight, the augmenting path between other unmatched concave points is found. The matching relationship on the augmenting path is reversed to complete the update of the matching relationship. The set of matched concave points is used as the stroke splitting point set.
[0012] Step 6: Based on the stroke intersection type, search for stroke splitting point pairs from the stroke splitting point set and split the strokes of the handwritten Chinese character.
[0013] Further, step 1 includes the following steps:
[0014] Step 1.1: Convert the original image to grayscale;
[0015] Step 1.2: Perform HSV (Hue, Saturation, Value) color space conversion on the grayscale image;
[0016] Step 1.3: Perform median filtering on the HSV-converted image;
[0017] Step 1.4: Binarize the image after median filtering;
[0018] Step 1.5: Use the Canny edge detection algorithm to identify character edges, which can identify character edges as close as possible to the actual edges of the original image.
[0019] Further, step 2 includes the following steps:
[0020] Step 2.1: For the input image point set P(i, x, y) m×3 Sort the pixels according to the scanning direction (clockwise or counterclockwise) around the character outline. Here, i represents the connected component index, x and y represent the horizontal and vertical coordinates of the key point, m represents the number of pixels, and ×3 represents the corresponding i, x, and y columns.
[0021] Step 2.2: Connect the coordinates of the arranged pixels pairwise to obtain the generalized curve S(P) describing the character outline image. m×3 ), and the generalized curve is processed according to a certain step size t. step Divide the curve into segments, and denote the first and last two pixels of each segment as P. start and P end Then the curve segment is denoted as
[0022] Step 2.3: Set the starting point P of the current piecewise curve start and termination point P end The line connecting the two chords is denoted as a chord. By traversing each pixel P in the current segment, the distance chord of the curve in that segment is calculated. The pixel P with the largest distance mid The maximum distance is denoted as d. max ;
[0023] Step 2.4: If d max Less than a pre-given distance threshold d threshold Then the straight line segment As a curve An approximation, that is, retaining only the coordinates of point P. start and P end The number of key points k = k + 1, k is initially 0, the algorithm ends, and we jump to step 2.6;
[0024] Step 2.5: If d max Greater than a pre-given distance threshold d threshold Then use P mid Divide the original curve into two segments. and The two curve segments are respectively regarded as curves Enter the update information and continue with steps 2.3 and 2.4;
[0025] Step 2.6: Obtain the final set of retained pixel coordinate keypoints P(i, x, y). k×3 .
[0026] Furthermore, step 3 includes the following steps:
[0027] Step 3.1: Based on the existing keypoint set P(i, x, y) k×3 Connect all key points sequentially to construct the topological polygon T(i, x, y). k×3 ;
[0028] Step 3.2: Let the topological polygon be T(i, x, y). k×3 The coordinates of the nth key point in the image are t. n (i, x, y), through t n-1 (i, x, y) and t n+1 (i, x, y), calculate the in-vector v of the current point. in and out vector v out The calculation formula is as follows:
[0029] v in =t n(i, x, y)-t n -1(i, x, y)=(x n y n )
[0030] v out =t n+1 (i, x, y)-t n (i, x, y) = (x n+1 y n+1 )
[0031] Step 3.3: Calculate the vector cross product:
[0032] A t =v in ×v out =(x n y n )×(x n+1 y n+1 )=(x n y n+1 -x n+1 y n )=|v in |·|v out |·sinθ
[0033] Get the current key point t n (i, x, y) are given by the input vector v in and out vector v out Let A be the area of the parallelogram formed by combining adjacent sides. t θ represents the angle formed by the input vector and the output vector;
[0034] Step 3.4: According to the geometric meaning of the vector cross product, if A t If A is greater than 0, then the current point is a convex point. t If the value is less than 0, then the current point is a concave point. By judging the concave points, we can obtain the set of concave points.
[0035] Furthermore, step 4 includes the following steps:
[0036] Step 4.1: Define the Minkowski distance to describe the matching of concave points in a stroke, where two points belong to the concave point set C(i, x, y). a×3 concave point c n (i, x, y) and c n+1 The Minkowski distance between (i, x, y) in 2D space is represented as follows:
[0037]
[0038] Where i represents the connected component index, x and y represent the horizontal and vertical coordinates of the key point, respectively, and w1 and w2 are the weights for calculating the Mink distance by balancing the Manhattan distance and Euclidean distance.
[0039] Step 4.2: Define the cosine similarity describing the matching of concave points in the connected components of a stroke. Let the input and output vectors of each concave point be v. in and v out The resulting vector space is V in and V out The input / output vector of the nth concave point is represented as and The ingress and egress vectors of the concave point are represented as follows:
[0040]
[0041]
[0042]
[0043]
[0044] Where c n (i, x, y) = t n (i, x, y), c n+1 (i, x, y) = t' n+1 (i, x, y), c n (i, x, y) ∈ C(i, x, y) a×3 , t n (i, x, y) ∈ T(i, x, y) k×3 ;
[0045] Step 4.3: Define the cosine distance describing the matching of concave points of the stroke, where the current concave point c is... n The input vector of (i, x, y) and the next concave point c n+1 The outgoing vector of (i, x, y) The cosine distance in 2D space is represented as follows:
[0046]
[0047] Step 4.4: Average stroke width is W average The weighted cost D(c) includes the Minkowski distance and the cosine distance. n c n+1 ) cost for:
[0048]
[0049] The concave point matching is transformed into a matching with the minimum weight cost of a bipartite graph.
[0050] Furthermore, step 5 includes the following steps:
[0051] Step 5.1: Encode all indentations and initialize the count variable count = 1;
[0052] Step 5.2: If the count variable count is less than the set number of vertices N, then proceed to step 5.3. If the count variable count is greater than the number of vertices N (here N = 2a, which is twice the number of concave points), then the splitting algorithm ends.
[0053] Step 5.3: Based on the minimum weight cost function, match the set of concave points to obtain the matching edges between concave points; for unmatched concave points, use depth-first search to find the augmenting path of the current vertex. If an augmenting path is found, invert the matched edges and unmatched edges between concave points on the augmenting path to make all points on the augmenting path saturated. At this time, the count variable count = count + 1, and jump to step 5.2.
[0054] Furthermore, in step 5.2, finding the augmenting path to the current vertex includes the following steps:
[0055] Step 5.3.1: Select an unmatched concave point, traverse the Minkowski distances between this concave point and other concave points, and then calculate the input vector to this concave point. Outgoing vector with smaller included angle The corresponding concave points are matched based on the minimum cost function. When a matching concave point is found, step 5.3.2 is executed; when no matching concave point is found, step 5.3.4 is executed.
[0056] Step 5.3.2: Iterate through and calculate the Minkowski distance between the concave point and other concave points. Based on the geometric distance, find the concave point with a shorter Minkowski distance to the concave point and execute step 5.3.3.
[0057] Step 5.3.3: Calculate the concave point corresponding to the outgoing vector that has a smaller angle with the incoming vector of the concave point, and then proceed to step 5.3.4;
[0058] Step 5.3.4: Based on the Minkowski distance and the angle between the ingress and egress vectors, calculate the concave point with the smallest comprehensive cost weight, increment the number of matched points count by 1, and proceed to step 5.2; if no matching point is found, proceed to step 5.3.5.
[0059] Step 5.3.5: Use depth-first search to find the augmenting path of the current unmatched concave point. Traverse all points and find the unmatched concave point with the smallest weight to the currently selected concave point as the matching point. An augmenting path is formed between the two. Execute step 5.3.6.
[0060] Step 5.3.6: Invert each of the unmatched and matched edges on the augmenting path, and add the matched edges to the matching relationship. This will increase the number of matched points. Mark the corresponding unmatched points as matched points and execute step 5.2.
[0061] Furthermore, in step 6, the stroke splitting point pair search is performed based on pairwise matching of isolated point pairs, including the following steps:
[0062] Step 6.1.1: Determine the direction of the traversal search, which can be clockwise or counterclockwise. Encode the isolated matching point pairs found in sequence according to the traversal direction.
[0063] Step 6.1.2: Visit any matching point in the traversal order, and directly return its matching point pair to obtain the set M of adjacent point pairs used for the search. I .
[0064] Furthermore, in step 6, based on pairwise matching triangular point pairs, a stroke splitting point pair search is performed, including the following steps:
[0065] Step 6.2.1: Determine the direction of the traversal search, which can be clockwise or counterclockwise. Encode the found triangular matching point pairs sequentially as p1(i, x1, y1), p2(i, x2, y2), and p3(i, x3, y3) according to the traversal direction. The initial set of these triangular matching point pairs is T.
[0066] T[(p1,p2),(p2,p3),(p3,p1)]
[0067] Step 6.2.2: When the sequential search reaches a matching point p1(x1, y1), the directed cycle G1 is obtained by traversing each matching point in turn.
[0068] G1[p1→p2→p3→p1]
[0069] Step 6.2.3: Define its directed neighbor as the 3rd node of the directed cycle, and obtain the final neighbor pair M1 of p1(x1, y1):
[0070] G1[p1→p2→p3→p1]→M1(p1, p3)
[0071] Step 6.2.4: Repeat the above steps for the other two points;
[0072] Step 6.2.5: Solve to obtain the set of adjacent point pairs M for the triangular matching point pair set T. T :
[0073] T[(p1, p2), (p2, p3), (p3, p1)]→M T[(p1, p3), (p2, p1), (p3, p2)].
[0074] A handwritten Chinese character stroke splitting device based on an iterative adaptation and matching algorithm includes a memory and one or more processors. The memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the handwritten Chinese character stroke splitting method based on the iterative adaptation and matching algorithm.
[0075] The advantages and beneficial effects of this invention are as follows:
[0076] This invention relates to a method and apparatus for handwritten Chinese character stroke decomposition based on an iterative adaptive matching algorithm. The method involves preprocessing handwritten Chinese characters into images, detecting key points of the handwritten Chinese character outline using the iterative adaptive point method (DP), obtaining key points for stroke decomposition, and analyzing these key points to obtain key point features that conform to the logic of handwritten Chinese character decomposition. This transforms the concave point matching problem into a minimum weight cost matching problem in a bipartite graph, thus achieving stroke decomposition point matching for handwritten Chinese characters. This invention can obtain decomposed strokes with high accuracy and without manual intervention by analyzing the search direction of the entire character's strokes, enabling the preservation of stylistic features of the target text to be copied in a robotic writing system. Attached Figure Description
[0077] Figure 1 This is a flowchart of the handwritten Chinese character splitting method in an embodiment of the present invention.
[0078] Figure 2a In this embodiment of the invention, the outline of Chinese characters based on DP is in 2tste p A generalized contour curve within the range.
[0079] Figure 2b This is a schematic diagram of the first segment of the contour curve processing of Chinese characters based on DP in an embodiment of the present invention.
[0080] Figure 2c This is a schematic diagram of the first stage of the contour curve of segment II in the Chinese character contour processing based on DP in an embodiment of the present invention.
[0081] Figure 2d This is a schematic diagram of the second stage of the contour curve of the second segment of Chinese character contour processing based on DP in an embodiment of the present invention.
[0082] Figure 2e This is a schematic diagram of the third stage of the contour curve of segment II in the Chinese character contour processing based on DP in an embodiment of the present invention.
[0083] Figure 2f This is a diagram showing the completed contour curve processing of the first and second segments of the Chinese character outline based on DP in an embodiment of the present invention.
[0084] Figure 3a It is one of the schematic diagrams of connecting key point sets into topological polygons in the embodiments of the present invention.
[0085] Figure 3b It is the second schematic diagram of connecting key point sets into topological polygons in the embodiments of the present invention.
[0086] Figure 3c It is one of the schematic diagrams of generating concave point sets (hollow) from key point sets (solid) through concavity and convexity judgment in the embodiments of the present invention.
[0087] Figure 3d It is the second schematic diagram of generating concave point sets (hollow) from key point sets (solid) through concavity and convexity judgment in the embodiments of the present invention.
[0088] Figure 4 It is the flowchart of stroke splitting of handwritten Chinese characters based on the Hungarian matching HMA in the embodiments of the present invention.
[0089] Figure 5a It is the schematic diagram of the Minkowski distance of the weight matrix based on calculating concave point matching in the embodiments of the present invention.
[0090] Figure 5b It is the schematic diagram of the cosine distance of the weight matrix based on calculating concave point matching in the embodiments of the present invention.
[0091] Figure 6 It is the schematic diagram of the matching process of the concave point set in the embodiments of the present invention. This is a schematic diagram of the handwritten Chinese character splitting device in an embodiment of the present invention. Detailed Implementation
[0100] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0101] like Figure 1 As shown, the handwritten Chinese character stroke decomposition method based on iterative adaptation and matching algorithm includes the following steps:
[0102] Step 1: Preprocessing the handwritten Chinese character image to obtain character edges, including the following steps:
[0103] Step 1.1: Convert the original image to grayscale;
[0104] Step 1.2: Perform HSV (Hue, Saturation, Value) color space conversion on the grayscale image;
[0105] Step 1.3: Perform median filtering on the HSV-converted image;
[0106] Step 1.4: Binarize the image after median filtering;
[0107] Step 1.5: Use the Canny edge detection algorithm to identify character edges, which can identify character edges as close as possible to the actual edges of the original image.
[0108] Step 2: Using the pixel points on the character edges, perform key point detection on the outline of handwritten Chinese characters based on the iterative adaptive point method (DP) to obtain a set of pixel coordinate key points, such as... Figures 2a to 2f As shown, it includes the following steps:
[0109] Step 2.1: Let the initial number of keypoints be k = 0. For the input image point set P(i, x, y) m×3 Sort the pixels; arrange the coordinates of the pixels in order according to the scanning direction (clockwise or counterclockwise) around the character outline;
[0110] Step 2.2: Connect the coordinates of the arranged pixels pairwise to obtain the generalized curve S(P) describing the character outline image. m×3 ), and the generalized curve is processed according to a certain step size t. step Divide the curve into segments, and denote the first and last two pixels of each segment as P. start and P end Then the curve segment is denoted as
[0111] Step 2.3: Set the starting point P of the current piecewise curve start and termination point P end The line connecting the two chords is denoted as a chord. By traversing each pixel P in the current segment, the distance chord of the curve in that segment is calculated. The pixel P with the largest distance mid The maximum distance is denoted as d. max ;
[0112] Step 2.4: If d max Less than a pre-given distance threshold d threshold Then the straight line segment As a curve An approximation, that is, retaining only the coordinates of point P. start and P end The number of key points k = k + 1, the algorithm ends, and we jump to step 2.6;
[0113] Step 2.5: If d max Greater than a pre-given distance threshold d threshold Then use P mid Divide the original curve into two segments. and Let these two curve segments be respectively treated as curves Enter the update information and continue with steps 2.3 and 2.4;
[0114] Step 2.6: Obtain the final set of retained pixel coordinate keypoints P(i, x, y). k×3 .
[0115] Step 3: Analyze the key points of handwritten Chinese character stroke breakdown; analyze the key point set P(i, x, y). k×3 Concavity / convexity is determined based on the established topological polygon T(i, x, y). k×3 The concavity and convexity of each vertex ultimately yields a set of concave points C(i, x, y) containing the stroke splitting points. a×3 ,like Figures 3a to 3d As shown, it includes the following steps:
[0116] Step 3.1: Based on the existing keypoint set P(i, x, y) k×3 Connect all key points sequentially to construct the topological polygon T(i, x, y). k×3 ;
[0117] Step 3.2: Let the topological polygon be T(i, x, y). k×3 The coordinates of the nth key point in the image are t. n Given (i, x, y), the calculation formula is as follows:
[0118] v in =tn (i, x, y)-t n-1 (i, x, y) = (x n y n )
[0119] v out =t n+1 (i, x, y)-t n (i, x, y) = (x n+1 y n+1 )
[0120] via t n-1 (i, x, y) and t n+1 (i, x, y) is used to calculate the infeed vector v of the current point. in and out vector v out ;
[0121] Step 3.3: Calculate using the vector cross product formula:
[0122] A t =v in ×v out =(x n y n )×(x n+1 y n+1 )=(x n y n+1 -x n+1 y n )=|v in |·|v out |·sinθ
[0123] Get the current key point t n (i, x, y) are given by the input vector v in and out vector v out Let A be the area of the parallelogram formed by combining adjacent sides. t θ represents the angle formed by the input vector and the output vector;
[0124] Step 3.4: According to the geometric meaning of the vector cross product, if A t If A is greater than 0, then the current point is a convex point. t If the value is less than 0, then the current point is a concave point. By judging the concave points, we can obtain the set of concave points.
[0125] In subsequent steps, the concave point set C(i, x, y) will be... a×3 The concave points in the graph are matched pairwise to obtain the final stroke splitting point set S(i, x, y). q×3 'a' represents the number of concave points, and 'q' represents the number of splitting points.
[0126] S(i, x, y) q×3∈C(i, x, y) a×3 ∈T(i, x, y) k×3 ≡P(i, x, y) k×3 .
[0127] Step 4: Calculate the weights for matching feature points of stroke decomposition in handwritten Chinese characters, including the following steps:
[0128] Step 4.1: Define the Minkowski distance to describe the matching of concave points in a stroke. Two points belonging to the concave point set C(i, x, ..., ...) are considered as a single point. y ) a×3 concave point c n (i, x, y) and c n+1 The Minkowski distance between (i, x, y) in 2D space is represented as follows:
[0129]
[0130] Where w1 and w2 are the weights used to calculate the Mink distance by balancing the Manhattan distance and Euclidean distance metrics;
[0131] Step 4.2: Define the cosine similarity describing the matching of concave points in the connected components of a stroke. Let the input and output vectors of each concave point be v. in and v out The resulting vector space is V in and V out The input / output vector of the nth concave point is represented as and The ingress and egress vectors of the concave point are represented as follows:
[0132]
[0133]
[0134]
[0135]
[0136] Where c n (i, x, y) = t n (i, x, y), c n+1 (i, x, y) = t' n+1 (i, x, y), c n (i, x, y) ∈ C(i, x, y) a×3 , t n (i, x, y) ∈ T(i, x, y) k×3 ;
[0137] Step 4.3: Define the cosine distance describing the matching of concave points of the stroke, where the current concave point c is... n The input vector of (i, x, y) and the next concave point c n+1 The outgoing vector of (i, x, y) The cosine distance in 2D space is represented as follows:
[0138]
[0139] Step 4.4: Average stroke width is W average The weighted cost D(c) includes the Minkowski distance and the cosine distance. n c n+1 ) cost for:
[0140]
[0141] Thus, the matching problem is transformed into a matching problem with minimum weight cost in a bipartite graph.
[0142] Step 5: Segmentation of handwritten Chinese characters based on Hungarian matching (HMA), such as... Figure 4 As shown, it includes the following steps:
[0143] Step 5.1: Encode all indentations and initialize the count variable count = 1;
[0144] Step 5.2: If the count variable count is less than the number of vertices N, continue to step 5.3. If the count variable count is greater than the number of vertices N (here N = 2a, i.e., twice the number of concave points), the matching algorithm ends, and the handwritten Chinese character strokes are decomposed based on the matching relationship;
[0145] Step 5.3: Based on the minimum weight cost function, match the set of concave vertices to obtain matching edges between them; for unmatched concave vertices, use depth-first search to find augmenting paths for the current vertex. If an augmenting path is found, invert the matched edges and unmatched edges on that path to make all points on the augmenting path saturated. At this point, the count variable count = count + 1, and proceed to step 5.2. Specifically, this includes the following steps:
[0146] Step 5.3.1: Select a concave point and iterate through the concave points c to calculate the value. x Other concave points {c1, c2, ..., c n The Minkowski distance is calculated, and then the distance with c is calculated. x input vector Outgoing vector with smaller included angle For the corresponding concave point, if a suitable matching vertex can be found, proceed to step 5.3.2; if no suitable matching vertex is found for the current vertex, proceed to step 5.3.4.
[0147] Step 5.3.2: Calculate Figure 5a Taking the matching point of concave point c1 as an example, we iterate through and calculate the Minkowski distance between concave point c1 and other concave points {c2, c3, ..., c7}. Based on the geometric distance, the concave point with the shorter Minkowski distance to concave point c1 is c7. Then we execute step 5.3.3.
[0148] Step 5.3.3: Then calculate the input vector with c1. Outgoing vector with smaller included angle The corresponding concave point, for Figure 5b Analysis shows that the input vector of c1 The smaller included angles are the outgoing vectors of c5. and the outgoing vector of c7 Proceed to step 5.3.4;
[0149] Step 5.3.4: The Minkowski distance between c1 and c7 is smaller. Therefore, the concave point c7 has the smallest combined cost weight among the concave points c1 and the other concave points {c2, c3, ..., c7}, thus obtaining the final matching point pair. Increment the number of matched points count by 1, and return to step 5.2; if no matched points are found, proceed to step 5.3.5.
[0150] Step 5.3.5: Use depth-first search to find the augmenting path to the current vertex, such as... Figure 6 As shown, the vertices of the thick box are all vertices to be matched, and the vertices of the thin box are all matched vertices. Each step is divided into left and right sides. Taking the calculation of the concave point c3 inside the box as an example, all points are traversed, and the point with the smallest weight with the currently selected concave point is taken as the matching point. Figure 6 In this case, the augmenting path is c3(right)-c2(left)-c4(right)-c5(left), and step 5.3.6 is executed;
[0151] Step 5.3.6: Invert each of the unmatched and matched edges on the augmenting path, and add the matched edges to the matching relationship. This will increase the number of matched points. Mark the corresponding unmatched points as matched points and execute step 5.2. Figure 6 In the diagram, c2 to c4 are already matched edges, and c3 to c2 and c4 to c5 are newly calculated edges to be matched. The purpose of inverting the values is to obtain pairwise matches between c3 to c2 and c4 to c5. Since c2 to c4 is already in the concave point matching set, the inversion operation is performed here to prevent duplication. At the same time, the newly matched points are added to the concave point matching set.
[0152] Step 6: Stroke segmentation search for handwritten Chinese characters. Further search is performed on the segmented character image data. Matching of stroke segmentation point pairs for different stroke intersection types is divided into two main categories: one is pairwise matching of isolated point pairs, and the other is pairwise matching of triangular point pairs, such as... Figures 7a to 7f As shown, it includes the following steps:
[0153] Step 6.1: For isolated matching points, divide them into four types of stroke intersection types: "X", "T", "L", and "Z" for searching;
[0154] Step 6.1.1: Determine the traversal search direction, either clockwise or counterclockwise. According to the traversal direction, sequentially encode the isolated matching point pairs found;
[0155] Step 6.1.2: When accessing any matching point in the traversal order, directly return its matching point pair to obtain the set M of adjacent point pairs for searching I .
[0156] Step 6.2: For triangular matching point pairs, divide this matching type into "K" type and "Y" type for searching:
[0157] Step 6.2.1: It is necessary to determine the traversal search direction, either clockwise or counterclockwise. According to the traversal direction, sequentially encode the triangular matching point pairs found as p1(i, x1, y1), p2(i, x2, y2), and p3(i, x3, y3). The initial set of this triangular matching point pair is T:
[0158] T[(p1, p2), (p2, p3), (p3, p1)]
[0159] Step 6.2.2: When sequentially searching and accessing the matching point p1(x1, y1), sequentially traverse each matching point to obtain the directed loop G1:
[0160] G[p1→p2→p3→p1]
[0161] Step 6.2.3: Define its directed adjacent point as the third node of the directed loop to obtain the final adjacent point pair M1 of p1(x1, y1): [[ID=二十九]]
[0162] G[p1→p2→p3→p1]→M1(p1, p3)
[0163] Step 6. Figure 8 As shown.
[0167] Corresponding to the aforementioned embodiments of the handwritten Chinese character stroke decomposition method based on iterative adaptation and matching algorithm, the present invention also provides embodiments of a handwritten Chinese character stroke decomposition device based on iterative adaptation and matching algorithm.
[0168] See Figure 9 The handwritten Chinese character stroke splitting device based on iterative adaptation and matching algorithm provided in this embodiment of the invention includes a memory and one or more processors. The memory stores executable code. When the one or more processors execute the executable code, they are used to implement the handwritten Chinese character stroke splitting method based on iterative adaptation and matching algorithm in the above embodiment.
[0169] The embodiment of the handwritten Chinese character stroke decomposition device based on the iterative adaptation and matching algorithm of this invention can be applied to any device with data processing capabilities, such as a computer. The device embodiment can be implemented in software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by the processor of any data processing device loading the corresponding computer program instructions from non-volatile memory into memory for execution. From a hardware perspective, such as... Figure 9 The diagram shown is a hardware structure diagram of any device with data processing capabilities, including the handwritten Chinese character stroke decomposition device based on the iterative adaptation and matching algorithm of this invention. (Except for...) Figure 9 In addition to the processor, memory, network interface, and non-volatile memory shown, any data processing device in the embodiment may also include other hardware depending on the actual function of the data processing device, which will not be described in detail here.
[0170] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0171] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the present invention according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0172] This invention also provides a computer-readable storage medium storing a program that, when executed by a processor, implements the handwritten Chinese character stroke decomposition method based on iterative adaptation and matching algorithm described in the above embodiments.
[0173] The computer-readable storage medium can be an internal storage unit of any data processing device as described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device of any data processing device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc., equipped on the device. Furthermore, the computer-readable storage medium can include both internal storage units and external storage devices of any data processing device. The computer-readable storage medium is used to store the computer program and other programs and data required by the data processing device, and can also be used to temporarily store data that has been output or will be output.
[0174] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for decomposing handwritten Chinese character strokes based on iterative adaptation and matching algorithms, characterized in that... Includes the following steps: Step 1: Preprocess the handwritten Chinese character image to obtain character edges; Step 2: Detect key points of the handwritten Chinese character outline using the pixel points on the character edge to obtain a set of pixel coordinate key points; Step 3: Analysis of key points for decomposing handwritten Chinese characters; by using the concavity and convexity of the vertices of the geometric shape formed by the key points of pixel coordinates, concave points are used as feature points for decomposing handwritten Chinese characters. Step 4: Construct the minimum weight cost function. Based on the distance between concave points and the average stroke width, calculate the minimum weight for matching the stroke segmentation feature points of handwritten Chinese characters. This includes the following steps: Step 4.1: Define the Minkowski distance to describe the matching of concave points in a stroke. Two concave points belonging to the concave point set... and exist The Minkowski distance in 3D space is represented as follows: Where i represents the connected component index, and x and y represent the horizontal and vertical coordinates of the key point, respectively. This indicates the weight given to the Manhattan distance in the Minkowski distance calculation. This indicates the weight given to the Euclidean distance in the Minkowski distance calculation; Step 4.2: Define the cosine similarity describing the matching of concave points in the connected components of a stroke. Let the in-vectors of each concave point be... and The vector space formed is and , No. The input / output vectors of the concave points are represented as follows: and The ingress and egress vectors of the concave point are represented as follows: in , , , 'a' represents the number of concave points. This represents the topological polygon created by connecting all key points sequentially from beginning to end. ×3 represents the corresponding i, x, and y columns, and k represents the number of key points. This represents the current nth key point; Step 4.3: Define the cosine similarity to describe the matching of concave points of strokes, and the current concave point. input vector and the next indentation out vector exist The cosine similarity in 3D space is represented as follows: in, Represents the input vector of the nth key point. out vector of the next concave point The angle of formation; Step 4.4: Weighting Cost Includes Minkowski distance, cosine similarity, and average stroke width. The specific formula is as follows: Transform concave point matching into minimum weight cost matching of bipartite graphs; Step 5: Handwritten Chinese character stroke decomposition. Based on the minimum weight cost function, the concave point set is matched to obtain the matching relationship. For unmatched concave points, the concave point set is traversed. Based on the minimum weight, the augmenting path between other unmatched concave points is found. The matching relationship on the augmenting path is reversed to complete the update of the matching relationship. The set of matched concave points is used as the stroke decomposition point set. Step 6: Based on the stroke intersection type, search for stroke splitting point pairs from the stroke splitting point set and split the strokes of the handwritten Chinese character.
2. The handwritten Chinese character stroke decomposition method based on iterative adaptation and matching algorithm according to claim 1, characterized in that: Step 1 includes the following steps: Step 1.1: Convert the original image to grayscale; Step 1.2: Perform HSV color space conversion on the grayscale image; Step 1.3: Perform median filtering on the HSV-converted image; Step 1.4: Binarize the image after median filtering; Step 1.5: Use an edge detection algorithm to identify character edges.
3. The handwritten Chinese character stroke decomposition method based on iterative adaptation and matching algorithm according to claim 1, characterized in that: Step 2 includes the following steps: Step 2.1: Process the input image point set Sort the pixels; arrange the coordinates of the pixels in order according to the scanning direction around the character outline, where i represents the connected component index, x and y represent the horizontal and vertical coordinates of the key point respectively, m represents the number of pixels, and ×3 represents the corresponding i, x, and y columns; Step 2.2: Connect the coordinates of the arranged pixels pairwise to obtain the generalized curves describing the character outline image. And the generalized curve is processed according to a certain step size. Divide the curve into segments, and denote the first and last two pixels of each segment as... and Then the curve segment is denoted as ; Step 2.3: Set the starting point of the current piecewise curve. and termination point The line connecting the two chords is denoted as a chord. By traversing each pixel of the current segment The distance chord of the piecewise curve is calculated. The pixel with the largest distance The maximum distance is denoted as ; Step 2.4: If Less than a pre-given distance threshold Then the straight line segment As a curve An approximation, that is, retaining only the point coordinates. and Number of key points Since k is initially 0, skip to step 2.6; Step 2.5: If Greater than a pre-defined distance threshold Then use Divide the original curve into two segments. and The two curve segments are respectively taken as curves Enter the update information and continue with steps 2.3 and 2.4; Step 2.6: Obtain the final set of retained pixel coordinate keypoints. .
4. The handwritten Chinese character stroke decomposition method based on iterative adaptation and matching algorithm according to claim 1, characterized in that: Step 3 includes the following steps: Step 3.1: Based on the existing keypoint set Connect all key points end to end to create a topological polygon. ; Step 3.2: Define the topological polygon The first in The coordinates of the key points are: ,pass and Calculate the in vector of the current point. out vector The calculation formula is as follows: Step 3.3: Calculate the vector cross product: Get the current key point From the input vector out vector The area of the parallelogram formed by combining adjacent sides. , Indicates the angle formed by the input vector and the output vector; Step 3.4: According to the geometric meaning of the vector cross product, if If the value is greater than 0, then the current point is a convex point. If the value is less than 0, then the current point is a concave point. By judging the concave points, we can obtain the set of concave points.
5. The handwritten Chinese character stroke decomposition method based on iterative adaptation and matching algorithm according to claim 1, characterized in that: Step 5 includes the following steps: Step 5.1: Encode all indentations and initialize the counting variable. ; Step 5.2: If the counting variable Less than the set number of vertices Then proceed to step 5.3, if the counter variable... Greater than the number of vertices Then the splitting ends; Step 5.3: Based on the minimum weight cost function, match the set of concave vertices to obtain the matching edges between them; for unmatched concave vertices, use depth-first search to find the augmenting path of the current vertex. If an augmenting path is found, invert the matched edges and unmatched edges between concave vertices on the augmenting path. At this time, the counter variable... Proceed to step 5.
2.
6. The handwritten Chinese character stroke decomposition method based on iterative adaptation and matching algorithm according to claim 5, characterized in that: In step 5.2, finding the augmenting path to the current vertex includes the following steps: Step 5.3.1: Select an unmatched concave point, traverse the Minkowski distances between this concave point and other concave points, and then calculate the input vector to this concave point. Outgoing vector with smaller included angle The corresponding concave points are matched based on the minimum weight cost function. When a matching concave point is found, step 5.3.2 is executed; when no matching concave point is found, step 5.3.4 is executed. Step 5.3.2: Iterate through and calculate the Minkowski distance between the concave point and other concave points. Based on the geometric distance, find the concave point with a shorter Minkowski distance to the concave point and execute step 5.3.
3. Step 5.3.3: Calculate the concave point corresponding to the outgoing vector that has a smaller angle with the incoming vector of the concave point, and then proceed to step 5.3.4; Step 5.3.4: Based on the Minkowski distance and the angle between the ingress and egress vectors, calculate the concave point with the minimum comprehensive cost weight and the number of matched points. Proceed to step 5.2; if no matching point is found, proceed to step 5.3.
5. Step 5.3.5: Use depth-first search to find the augmenting path of the current unmatched concave point. Traverse all points and find the unmatched concave point with the smallest weight to the currently selected concave point as the matching point. An augmenting path is formed between the two. Execute step 5.3.
6. Step 5.3.6: Invert each of the unmatched and matched edges on the augmenting path, and add the matched edges to the matching relationship. This will increase the number of matched points. Mark the corresponding unmatched points as matched points and execute step 5.
2.
7. The handwritten Chinese character stroke decomposition method based on iterative adaptation and matching algorithm according to claim 1, characterized in that: In step 6, the stroke splitting point pair search is performed based on pairwise matching isolated point pairs, including the following steps: Step 6.1.1: Determine the direction of the traversal search, and encode the isolated matching point pairs found in sequence according to the traversal direction; Step 6.1.2: Visit any matching point in the traversal order and directly return its matching point pair to obtain the set of adjacent point pairs used for the search. .
8. The handwritten Chinese character stroke decomposition method based on iterative adaptation and matching algorithm according to claim 1, characterized in that: In step 6, the stroke splitting point pair search is performed based on pairwise matching triangle point pairs, including the following steps: Step 6.2.1: Determine the direction of the traversal search, and encode the triangular matching point pairs found according to the traversal direction. , and The initial set of matching points for the triangle is : Step 6.2.2: When the sequential search reaches the matching point At that time, a directed cycle is obtained by traversing each matching point in turn. : Step 6.2.3: Define its directed neighbor as the 3rd node of the directed cycle, and obtain... Final adjacent point pairs : Step 6.2.4: Repeat the above steps for the other two points; Step 6.2.5: Solve to obtain the set of matching point pairs for this triangle. The set of adjacent points : 。 9. A handwritten Chinese character stroke decomposition device based on iterative adaptation and matching algorithms, characterized in that, The method includes a memory and one or more processors, wherein the memory stores executable code, and the one or more processors execute the executable code to implement the handwritten Chinese character stroke decomposition method based on the iterative adaptation and matching algorithm as described in any one of claims 1-8.
Citation Information
Patent Citations
A method for fast correction of license plate distortion in complex scene
CN109145915A
Generation method, generation system and application method of font writing stroke order
CN113657330A