Ray composition complex scene ancient Chinese character detection system and method based on human-computer interaction
Through ray composition and human-computer interaction, the edges of ancient Chinese characters in complex scenes are captured, and the text detection model is optimized, which solves the problem of inaccurate recognition of ancient Chinese characters in the existing technology, and achieves high-accurate ancient Chinese character detection.
Patent Information
- Application Number
- CN202311652758.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-07-11
AI Technical Summary
The existing ancient Chinese character detection model based on deep learning has the problem of inaccurate identification in complex scenarios, especially continuous and overlapping words, which is difficult to cover all complex scenarios through data augmentation and multi-scale strategies.
The ray composition strategy is used combined with human-computer interaction, incremental learning is performed through user feedback, and the undetected text areas are used to smear the edges of ancient Chinese characters, and the text detection model is optimized.
It improves the flexibility and accuracy of the positioning detection of ancient Chinese characters, effectively solves the problem of continuous and overlapping text recognition in complex scenarios, and significantly improves the recognition accuracy.
Smart Images

Figure CN120299055A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of ancient Chinese character recognition, and specifically to an ancient Chinese character detection system and method for complex scenes of ray composition based on human-computer interaction. Background Art
[0002] The prior art uses a text localization detection model based on deep learning to perform text localization detection of ancient Chinese characters in complex scenes, and a large amount of labeled data is required for training. For some problems such as complex and changeable scenes, cursive characters and overlapping characters, the model detection is often inaccurate, resulting in misrecognition. Since the performance of the deep learning model highly depends on the quality and quantity of the training data, the diversity, complexity of natural scenes and the diversity of fonts make it difficult to obtain sufficient and comprehensive training data. Although data augmentation, multi-scale strategies and post-processing techniques improve the robustness and accuracy of the model, solving problems from the perspective of the detection model still cannot cover all actual complex scenes. Summary of the Invention
[0003] To solve the technical problems in the above background, the present invention uses a ray composition strategy to first calculate the angle of each ray, and then aggregate the angles of multiple rays to obtain the directional information within the region; at the same time, through the feedback obtained from user interaction, incremental learning of the model is realized.
[0004] To achieve the above object, the present invention provides an ancient Chinese character detection system for complex scenes based on ray composition, including: a pre-training module, a localization module, a human-computer interaction module, a determination module and an optimization module;
[0005] The pre-training module is used to pre-train a text detection model based on the collected ancient Chinese character images;
[0006] The localization module is used to extract features and perform preliminary localization on the ancient Chinese character image to be detected by using the text detection model to obtain a feature map;
[0007] The human-computer interaction module performs ray smearing on the text area not detected in the feature map to obtain a smeared area;
[0008] The determination module is used to capture the edge of the ancient Chinese character based on the smeared area;
[0009] The optimization module is used to optimize the text detection model based on the edge of the ancient Chinese character; use the optimized text detection model to outline the boundary of the ancient Chinese character.
[0010] Preferably, the working process of the pre-training module includes:
[0011] Preprocess the captured ancient Chinese character images in complex scenes into pixel sizes of width m pixels * height n pixels to adapt to the model input; then use a convolutional neural network for model training, and the preprocessed images are fed into the convolutional neural network for forward propagation; the network generates a set of feature maps through convolutional layers and activation functions; these feature maps will then be input into the proposal function of a sliding window or region proposal network to generate a batch of candidate regions; in the post-processing stage, the candidate regions will undergo threshold screening and non-maximum suppression to remove redundant and inaccurate bounding boxes; finally, the screened and optimized candidate regions are used as the final output of the model.
[0012] Preferably, the workflow of the human-computer interaction module includes: the user selects a starting point on the human-computer interaction module and then drags it to an ending point. The internal event listening mechanism of the human-computer interaction module triggers a listening event when the user clicks or touches the screen and records the coordinates of the starting point; when the user drags and finally releases, the coordinates of the ending point are recorded, and the human-computer interaction module takes the path and direction between these two points as a ray.
[0013] Preferably, the determination module includes: a calculation unit and an edge detection unit;
[0014] The calculation unit is used to predict the inclination angle of the ancient Chinese character;
[0015] The edge detection unit is used to capture the edge of the ancient Chinese character based on the inclination angle.
[0016] Preferably, the optimization module includes: a generation unit and a feedback unit;
[0017] The generation unit is used to calculate and generate an optimal irregular quadrilateral to cover the ancient Chinese character within the smeared area and outline the boundary of the ancient Chinese character;
[0018] Feed the optimal irregular quadrilateral back to the human-computer interaction interface.
[0019] Preferably, after receiving the optimal irregular quadrilateral on the human-computer interaction interface, the user makes a judgment. If the judgment result is reasonable, it is updated to the training set to improve the character detection model.
[0020] The present invention also provides a method for detecting ancient Chinese characters in complex scenes based on ray composition. The method is applied to the above system, and the steps include:
[0021] Based on the captured ancient Chinese character images, pre-train the character detection model;
[0022] Use the character detection model to perform feature extraction and preliminary localization on the ancient Chinese character images to be detected to obtain feature maps;
[0023] Perform ray smearing on the text regions not detected in the feature map to obtain the smeared regions;
[0024] Capture the edges of ancient Chinese characters based on the smeared regions;
[0025] Optimize the text detection model based on the edges of ancient Chinese characters; and use the optimized text detection model to outline the boundaries of ancient Chinese characters.
[0026] Preferably, the method for performing the pre-training includes:
[0027] Preprocess the collected ancient Chinese character images in complex scenes into a pixel size of width m pixels * height n pixels to adapt to the model input; then use a convolutional neural network for model training, and the preprocessed images are fed into the convolutional neural network for forward propagation; the network generates a set of feature maps through convolutional layers and activation functions; these feature maps will then be input into the proposal function of a sliding window or region proposal network to generate a batch of candidate regions; in the post-processing stage, the candidate regions will undergo threshold screening and non-maximum suppression to remove redundant and inaccurate boxes; finally, the screened and optimized candidate regions are used as the final output of the model.
[0028] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0029] Through a unique ray composition strategy and human-computer interaction method, the present invention not only provides accurate directional recognition to deal with the inclination or rotation of ancient Chinese characters, but also effectively solves the difficulties brought by cursive and overlapping characters to positioning detection, making it have significant flexibility and accuracy advantages in the positioning detection of ancient Chinese characters in complex scenes, significantly different from the prior art that completely relies on deep learning text positioning detection models, and can accurately perform the positioning detection of ancient Chinese characters in complex scenes by means of the ray composition strategy and human-computer interaction method, reducing the influence of text region noise, thereby greatly improving the accuracy of text recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions of the present invention, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0031] Figure 1 It is a schematic structural diagram of the system according to an embodiment of the present invention;
[0032] Figure 2 It is a schematic flowchart of the method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0034] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0035] Embodiment 1
[0036] As Figure 1 shown, it is a schematic diagram of the system structure of this embodiment, including: a pre-training module, a positioning module, a human-computer interaction module, a determination module, and an optimization module; the pre-training module is used to pre-train a text detection model based on the collected ancient Chinese character images; the positioning module is used to extract features and perform preliminary positioning on the ancient Chinese character images to be detected by using the text detection model to obtain a feature map; the human-computer interaction module performs ray painting on the text areas not detected in the feature map to obtain a painted area; the determination module is used to capture the edges of the ancient Chinese characters based on the painted area; the optimization module is used to optimize the text detection model based on the edges of the ancient Chinese characters, and use the optimized text detection model to outline the boundaries of the ancient Chinese characters.
[0037] Next, in conjunction with this embodiment, it will be detailed how the present invention solves technical problems in real life.
[0038] Before feature extraction and preliminary positioning, the pre-training module is first used to pre-train the text detection model. The collected ancient Chinese character pictures in complex scenes are pre-processed into a pixel size of width m pixels * height n pixels (typical value is 750 pixels * 1000 pixels) to adapt to the model input. The model is trained using a convolutional neural network. The pre-processed image is fed into the convolutional neural network (CNN) for forward propagation. The network generates a set of feature maps through convolutional layers (Conv2D) and activation functions (ReLU). These feature maps will then be input into the proposal function of a sliding window or region proposal network (RPN) to generate a batch of candidate regions. In the post-processing stage, these candidate regions will undergo threshold screening (thresholding) and non-maximum suppression (NMS, using the nms function) to remove redundant and inaccurate boxes. Finally, the screened and optimized candidate regions are used as the final output of the model.
[0039] After the pre-training of the text detection model is completed, the positioning module uses the pre-trained text detection model to extract features and perform preliminary positioning on the ancient Chinese character images to be detected, obtaining a feature map.
[0040] The positioning module uses this model to extract features from the input image and achieve the preliminary positioning of the ancient Chinese character area. Run the forward propagation of the model to generate a feature map representing possible text areas. Based on this feature map, the model can generate a set of candidate boxes, each box representing a possible area containing text. The user uploads a picture containing ancient Chinese characters in the scene; once the picture is uploaded, the module automatically preprocesses it to a size of m pixels wide * n pixels high (typical values are 750 pixels * 1000 pixels) and sends it to the pre-trained detection model for feature extraction and preliminary text area positioning. After the detection model finishes running, the system marks the areas on the original image that the model thinks may contain ancient Chinese characters with a quadrilateral frame.
[0041] After that, the human-computer interaction module is used to perform ray smearing on the text areas not detected in the feature map to obtain the smeared area.
[0042] The user can draw the main strokes and their directions of the text in the text area by drawing rays. The user selects a starting point on the interface and then drags it to an end point. The internal event listening mechanism of the module, when the user clicks or touches the screen, triggers the listening event and records the coordinates of the starting point. When the user drags and finally releases, the coordinates of the end point are recorded, and the module takes the path and direction between these two points as a ray; finally, through the ray composition algorithm, the smeared area is determined.
[0043] The determination module is used to capture the edges of ancient Chinese characters based on the smeared area.
[0044] According to the length and direction of the ray, a structural element consistent with the ray direction is generated. If the ray goes from the upper left to the lower right, then we can choose an oblique rectangular structural element. Apply the dilation operation on the ray. To obtain better results, the dilation operation may need to be iterated multiple times, especially when the ray is short or the ancient Chinese character is large. The specific process includes:
[0045] First, the calculation unit calculates the expected direction of the ancient Chinese character based on multiple rays drawn by the user, so as to predict its tilt angle. First, calculate the direction of a single ray, and its direction or tilt angle θ can be calculated as follows: Then aggregate the directions of multiple rays: It should be noted that when x2 = x1, that is, the ray is vertical, then set its angle to 90° or -90° as needed.
[0046] Then the edge detection unit performs edge detection in the smeared area, and at the same time uses the direction of each ray calculated previously to optimize this process to more accurately capture the edges of ancient Chinese characters.
[0047] The first step is basic edge detection, which is accomplished by calculating the gradient magnitude of the image. The horizontal gradient is: G x = I * S x , and the vertical gradient is: C y = I * S y , where * is the convolution operation, and S x and S y are the Sobel operators in the horizontal and vertical directions respectively.
[0048] The second step is directional optimization. Based on the basic edge detection results, the edge image is dilated using the structural element generated by the determination module that is consistent with the ray direction, strengthening the edges in the ray direction and weakening the edges in other directions. Through the pixel maximum merging method, the optimized edges are combined with the original edge image.
[0049] The third step is gradient direction filtering, which only retains those edges that are close to the pre-computed ray direction. For example, if the expected direction is 45°, we can only retain those edges with directions between 40° and 50°.
[0050] The fourth step is gradient intensity adjustment, which enhances or weakens the edge intensity according to the similarity between the edge direction and the expected direction. Edges closer to the expected direction may be enhanced, while edges far from the expected direction may be weakened or deleted.
[0051] The last step is non-maximum suppression and double thresholding. To obtain clearer edges, non-maximum suppression is required to thin the edges. In the direction of the gradient, the gradient magnitudes of the current pixel and its neighbor pixels are checked. Only when the gradient magnitude of the current pixel is larger than those of its neighbor pixels on both sides is it considered a pixel on the edge. And the double threshold method is used, setting two thresholds, high and low. Pixels with gradient magnitudes higher than the high threshold are considered real edge pixels, pixels with gradient magnitudes lower than the low threshold are excluded, and pixels in between are marked as potential edge pixels to further determine and link the edges.
[0052] Finally, the optimization module optimizes the text detection model based on the edges of ancient Chinese characters, and uses the optimized text detection model to outline the boundaries of ancient Chinese characters.
[0053] The generation unit calculates and generates the optimal irregular quadrilateral to cover the ancient Chinese characters within the smeared area. First, based on the ray and edge information, the intersection points of the ray direction and the text edges are determined. The intersection points of each ray are screened. If an intersection point is on the line connecting other intersection points, the corresponding other intersection point is selected. Four intersection points are screened out to become the corner points of the irregular quadrilateral, forming a preliminary irregular quadrilateral that represents the boundary of the ancient Chinese character.
[0054] The feedback unit feeds back the optimal irregular quadrilateral to the human-computer interaction module. If the user believes that the generated bounding box is reasonable, then these data and labels are updated to the training set, and the pre-trained detection model is fine-tuned to achieve incremental learning. If the bounding box is unreasonable, the user is allowed to correct the corner points of the quadrilateral on the human-computer interface.
[0055] The user sees the generated bounding box on the human-computer interaction module and makes a judgment. Whether the representation of the bounding box is correct mainly depends on the user's satisfaction. A simple confirmation button and a correction button are provided for each bounding box on the interface. If the user clicks the confirmation button, then the image and its corresponding bounding box coordinates are marked as correct. If the correction button is clicked, the user can directly adjust the position and size of the bounding box.
[0056] The image data and its labels (coordinates and size of the bounding box) confirmed by the user are added to the training dataset, and fine-tuning is performed using the newly added or corrected data. There is no need to train from scratch, just perform a few additional iterations or epochs on the pre-trained model, thereby optimizing the pre-trained text detection model.
[0057] Finally, the optimized text detection model is used to complete the detection of ancient Chinese characters in complex environments.
[0058] Embodiment 2
[0059] As Figure 2 shown, the method flow schematic diagram of this embodiment includes the following steps:
[0060] Based on the collected ancient Chinese character images, pre-train the text detection model; use the text detection model to perform feature extraction and preliminary localization on the ancient Chinese character images to be detected to obtain feature maps; perform ray painting on the text regions not detected in the feature maps to obtain painted regions; based on the painted regions, capture the edges of ancient Chinese characters; based on the edges of ancient Chinese characters, optimize the text detection model; and use the optimized text detection model to outline the boundaries of ancient Chinese characters.
[0061] Among them, the method for pre-training includes: preprocessing the collected ancient Chinese character pictures in complex scenes into a pixel size of width m pixels * height n pixels to adapt to the model input; then using a convolutional neural network for model training, and the preprocessed images are fed into the convolutional neural network for forward propagation; the network generates a set of feature maps through convolutional layers and activation functions; these feature maps will then be input into the proposal function of the sliding window or region proposal network to generate a batch of candidate regions; in the post-processing stage, these candidate regions will undergo threshold screening and non-maximum suppression (to remove redundant and inaccurate bounding boxes; finally, the screened and optimized candidate regions are used as the final output of the model.
[0062] The embodiments described above are only descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the spirit of the design of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A complex-scene ancient Chinese character detection system based on ray composition, characterized in that Including: A pre-training module, a positioning module, a human-computer interaction module, a determination module, and an optimization module; The pre-training module is used to pre-train a text detection model based on the collected ancient Chinese character images; The positioning module is used to perform feature extraction and preliminary positioning on the ancient Chinese character image to be detected by using the text detection model, and obtain a feature map; The human-computer interaction module performs ray smearing on the text area not detected in the feature map to obtain a smeared area; The determination module is used to capture the edge of the ancient Chinese character based on the smeared area; The optimization module is used to optimize the text detection model based on the edge of the ancient Chinese character; and use the optimized text detection model to outline the boundary of the ancient Chinese character.
2. The complex scene ancient Chinese character detection system based on ray composition according to claim 1, characterized in that, The working process of the pre-training module includes: Preprocessing the collected ancient Chinese character pictures in complex scenes into a pixel size of width m pixels * height n pixels to adapt to the model input; then using a convolutional neural network for model training, and the preprocessed image is fed into the convolutional neural network for forward propagation; the network generates a set of feature maps through convolutional layers and activation functions; these feature maps will then be input into the proposal function of the sliding window or region proposal network to generate a batch of candidate regions; in the post-processing stage, the candidate regions will undergo threshold screening and non-maximum suppression to remove redundant and inaccurate boxes; finally, the screened and optimized candidate regions are used as the final output of the model.
3. The complex scene ancient Chinese character detection system based on ray composition according to claim 1, characterized in that, The working process of the human-computer interaction module includes: the user selects a starting point on the human-computer interaction module, and then drags it to an end point. The internal event listening mechanism of the human-computer interaction module triggers a listening event when the user clicks or touches the screen, and records the coordinates of the starting point; when the user drags and finally releases, the coordinates of the end point are recorded, and the human-computer interaction module takes the path and direction between these two points as a ray.
4. The ancient Chinese character detection system for complex scenes based on ray composition according to claim 1, wherein The determination module includes: a calculation unit and an edge detection unit; The calculation unit is used to predict the tilt angle of the ancient Chinese character; The edge detection unit is used to capture the edge of the ancient Chinese character based on the tilt angle.
5. The complex scene ancient Chinese character detection system based on ray composition according to claim 1, characterized in that, The optimization module includes: a generation unit and a feedback unit; The generation unit is used to calculate and generate an optimal irregular quadrilateral to cover the ancient Chinese character in the smeared area, and outline the boundary of the ancient Chinese character; Feed the optimal irregular quadrilateral back to the human-computer interaction interface.
6. The complex scene ancient Chinese character detection system based on ray composition according to claim 5, characterized in that, After receiving the optimal irregular quadrilateral on the human-computer interaction interface, the user makes a judgment. If the judgment result is reasonable, it is updated to the training set to improve the text detection model.
7. A complex scene ancient Chinese character detection method based on ray composition, the method being applied to the system according to any one of claims 1-6, characterized in that the steps Including: Pre-training a text detection model based on the collected ancient Chinese character images; Performing feature extraction and preliminary positioning on the ancient Chinese character image to be detected by using the text detection model to obtain a feature map; Performing ray smearing on the text area not detected in the feature map to obtain a smeared area; Capturing the edge of the ancient Chinese character based on the smeared area; Optimizing the text detection model based on the edge of the ancient Chinese character; and using the optimized text detection model to outline the boundary of the ancient Chinese character.
8. The method for detecting ancient Chinese characters in complex scenes based on ray composition according to claim 7, characterized in that The method for performing the pre-training includes: Preprocess the collected ancient Chinese character images in complex scenes into pixel sizes of width m pixels * height n pixels to adapt to the model input; then use a convolutional neural network for model training, and the preprocessed images are fed into the convolutional neural network for forward propagation; the network generates a set of feature maps through convolutional layers and activation functions; these feature maps will then be input into the proposal function of the sliding window or region proposal network to generate a batch of candidate regions; in the post-processing stage, the candidate regions will undergo threshold screening and non-maximum suppression to remove redundant and inaccurate boxes; finally, the screened and optimized candidate regions are used as the final output of the model.