An intelligent scanning and processing method for test papers in a marking system
By segmenting the uniform light area and shadow area in the test paper image, character characteristics are analyzed and repaired, the shadow and brightness abnormalities caused by the difference in the shooting environment of the test paper image are solved, the clarity and recognition of the image are improved, and the marking efficiency is improved.
Patent Information
- Application Number
- CN202411699242.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-26
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-11-26
AI Technical Summary
In the online marking system, due to differences in the shooting environment, the test paper images are prone to shadow, too bright or too dark abnormalities, which affects the image quality and reduces the marking efficiency.
Through contour detection and clustering technology, the test paper images are divided into uniform light areas and shadow areas, the pixel difference coefficients and character characteristics of each area are analyzed, and the character external rectangles are merged or split, and the enhanced repair coefficients of each similar character are obtained, and image enhancement and repair are performed.
Effectively distinguish between uniform light areas and shadow areas, improve the accuracy and consistency of character recognition, enhance the clarity and recognizability of the test paper images, and improve the efficiency of marking papers.
Smart Images

Figure CN119516568B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and particularly relates to a method for intelligent scanning and processing of test papers for a marking system. Background Art
[0002] With the continuous development of intelligent technologies, the traditional manual marking method has gradually been replaced by a more efficient online marking system; in the online marking system, markers can view and grade candidates' test papers anytime and anywhere, which not only improves the marking efficiency but also optimizes the resource allocation, making the entire test marking process more convenient and intelligent.
[0003] During the process of photographing, scanning, and uploading test papers, due to differences in different shooting environments, such as the direction of light, mobile phone occlusion, etc., the influence of the light source position and occluders will cause shadow areas in the test paper images, resulting in abnormal phenomena such as large shadows, overexposure, or underexposure in the uploaded test paper images. Excessively thick shadows will affect the test paper characters, reduce the quality of the test paper images, and affect the marking efficiency. Summary of the Invention
[0004] A method for intelligent scanning and processing of test papers for a marking system according to the present invention adopts the following technical solutions:
[0005] An embodiment of the present invention provides a method for intelligent scanning and processing of test papers for a marking system, and the method includes the following steps:
[0006] Obtain a plurality of test paper images, and obtain a plurality of contour pixel points of each test paper image through contour detection;
[0007] Cluster the pixel points in the test paper image to obtain a plurality of clusters, and obtain the pixel difference coefficient of each cluster according to the pixel value differences of different pixel points in each cluster; obtain the light-uniform area and the shadow area according to the pixel difference coefficients of each cluster;
[0008] Obtain the circumscribed rectangle of each character in the light-uniform area, obtain the merged character difference value after each merge of each row of characters according to the size of each circumscribed rectangle and the distance between adjacent circumscribed rectangles, obtain the final merge result of each row of characters based on the merged character difference value, and further split to obtain the split character difference value after each split of each row of characters, and finally obtain the final split result of each row of characters;
[0009] According to the characters with non-rectangular circumscribed polygons in the light-uniform area, combined with the shadow area, obtain the semi-occluded characters in the light-uniform area and the occluded characters in the shadow area; based on the distribution of the contour pixel points in the semi-occluded characters and the occluded characters, obtain the same stroke state parameter of the semi-occluded characters in the light-uniform area and the occluded characters in the shadow area, and further obtain a plurality of similar characters;
[0010] Obtain the image sharpness parameter of the uniformly illuminated area of each similar character according to the pixel difference coefficient of similar characters in the uniformly illuminated area and the gradient direction of the contour pixel points; obtain the image sharpness parameter of the shadow area of the similar characters; obtain the enhancement repair coefficient of each similar character in the test paper image through the image sharpness parameter of the similar characters and the same stroke state parameter of the semi-occluded characters in the corresponding uniformly illuminated area and the occluded characters in the shadow area.
[0011] Further, the specific steps for obtaining the pixel difference coefficient of each cluster are as follows:
[0012] The calculation method of the regional feature parameter of the k-th cluster is:
[0013]
[0014] The calculation method of the pixel difference coefficient of the k-th cluster is:
[0015]
[0016] Among them, represents the regional feature parameter of the k-th cluster, represents the standard deviation of the pixel values of all pixel points in the k-th cluster, represents the number of contour pixel points in the k-th cluster, represents the pixel difference coefficient of the k-th cluster, represents the maximum value of the regional feature parameters among all clusters, represents the maximum value of the differences between the pixel values of all pixel points in the k-th cluster, is a linear normalization function.
[0017] Further, the specific steps for obtaining the merged character difference value after each merge of each row of characters are as follows:
[0018] Obtain the length of the circumscribed rectangle of each character in the X-axis direction , and simultaneously record the Euclidean distance between adjacent circumscribed rectangles in the same row; according to the length of the circumscribed rectangle, and the Euclidean distance between adjacent circumscribed rectangles, respectively obtain the standard deviation of the lengths of the circumscribed rectangles of all characters in the row and the standard deviation of the Euclidean distances, denoted as the length discrete parameter value and the distance discrete parameter value;
[0019] For any whole line of characters, two adjacent bounding rectangles corresponding to the minimum Euclidean distance among the adjacent characters in the line are merged as the first merge, and the length of the bounding rectangle and the Euclidean distance between the adjacent bounding rectangles are recalculated, so as to obtain the length discrete parameter value and the distance discrete parameter value after the first merge; and so on, several merges are performed in the order of Euclidean distance from small to large, and the length discrete parameter value and the distance discrete parameter value after each merge of the whole line of characters are obtained; the merged character difference value of the whole line of characters after any merge is obtained. .
[0020] Furthermore, the step of obtaining the final merging result of each line of characters according to the merging character difference value includes the following specific steps:
[0021] For any whole line of characters, the merge result corresponding to the minimum value of the merged character difference values after all merges is selected as the final merge result of the whole line of characters.
[0022] Furthermore, the final splitting result of each line of characters is finally obtained, and the specific steps included are as follows:
[0023] For any whole line of characters, under the final merged result of the whole line of characters, select the character with the longest length in the X-axis direction for splitting in turn, and take the position of the character with the least pixels under the Y-axis coordinate point in the bounding rectangle of each character as the character segmentation area, and so on, split the bounding rectangle of the character from large to small according to the length; obtain the standard deviation of the length of the bounding rectangle of the whole line of characters after each split , and the standard deviation of the Euclidean distance , and further obtain the split character difference value , select the split result corresponding to the minimum value of the split character difference value as the final split result of the entire line of characters.
[0024] Furthermore, the step of obtaining the semi-shielded characters in the uniformly illuminated area and the shielded characters in the shadow area includes the following specific steps:
[0025] Characters in the uniformly illuminated area whose circumscribed polygons are not rectangular are regarded as semi-occluded characters. A rectangle is circumscribed to the circumscribed polygon of the semi-occluded character. The area with the maximum difference between the circumscribed rectangle and the polygon is the shaded part that is obscured, which is regarded as the obscured character.
[0026] Furthermore, the obtaining of a plurality of similar characters comprises the following specific steps:
[0027] The distance between any two contour pixels with opposite gradient directions in the same part is recorded as the width of the character stroke. The same stroke state parameters of the semi-occluded character in the uniformly illuminated area and the occluded character in the shadow area are calculated as follows:
[0028]
[0029] Among them, is the same stroke state parameter of the semi-occluded character in the uniformly illuminated area and the occluded character in the shadow area, is the character stroke width of the occluded character in the shadow area, is the character stroke width of the semi-occluded character in the uniformly illuminated area, is the same stroke influence factor of the semi-occluded character in the uniformly illuminated area and the occluded character in the shadow area; use function to normalize the result;
[0030] Set a similarity threshold. If the same stroke state parameter is greater than or equal to the similarity threshold, mark the semi-occluded character in the uniformly illuminated area and the occluded character in the shadow area as similar characters.
[0031] Furthermore, the specific method for obtaining the same stroke influence factor is as follows:
[0032] For the semi-occluded character b in the uniformly illuminated area, obtain the gradient direction of each contour pixel point of the semi-occluded character. For the character stroke of the examinee, obtain the fitting direction vector of the normal vector through the curve fitting algorithm , denoted as the change direction vector of the character stroke; for the occluded character a in the shadow area, obtain the change direction vector of the character stroke ;
[0033] By obtaining the included angle 、 of the character stroke change direction vectors of the occluded character a and the semi-occluded character b , and using it as the same stroke influence parameter of the semi-occluded character in the uniformly illuminated area and the occluded character in the shadow area, perform processing on using the linear normalization function to obtain the same stroke influence factor .
[0034] Furthermore, the specific steps for obtaining the image sharpness parameter of the uniformly illuminated area of each similar character are as follows:
[0035]
[0036] Among them, represents the image sharpness parameter of the uniformly illuminated area of the similar character in the i-th row and j-th column, n represents the length of the semi-occluded character corresponding to the similar character in the i-th row and j-th column in the X-axis direction, and m represents the length of the semi-occluded character corresponding to the similar character in the i-th row and j-th column in the Y-axis direction, Denote the pixel difference coefficient of the cluster to which the uniformly illuminated region corresponding to the similar characters in the \(i\)-th row and \(j\)-th column belongs. Denote the coordinates in the uniformly illuminated region of the similar characters in the \(i\)-th row and \(j\)-th column as The gradient direction of the contour pixel points.
[0037] Furthermore, the specific steps for obtaining the enhancement and repair coefficient of each similar character in the test paper image are as follows:
[0038]
[0039] Wherein, Is the enhancement and repair coefficient of the similar characters in the \(i\)-th row and \(j\)-th column, Denote the image sharpness parameter of the uniformly illuminated region of the similar characters in the \(i\)-th row and \(j\)-th column, Denote the image sharpness parameter of the shadow region of the similar characters in the \(i\)-th row and \(j\)-th column, Denote the same stroke state parameter of the semi-occluded characters in the uniformly illuminated region corresponding to the similar characters in the \(i\)-th row and \(j\)-th column and the occluded characters in the shadow region.
[0040] The beneficial effects of the technical solution of the present invention are as follows: Obtain the uniformly illuminated region and the shadow region to analyze the characters in different regions; According to the size and distance of the circumscribed rectangles of each character in the uniformly illuminated region, obtain the final merging result after each merging of the characters in each row, and further split to obtain the final splitting result of the characters in each row, which is beneficial to distinguish the recognition of each character under different writing habits; Obtain the same stroke state parameter according to the distribution of the contour pixel points, and then obtain several similar characters, and analyze each similar character; According to the pixel difference in the uniformly illuminated region of the similar characters and the gradient direction of the contour pixel points, obtain the image sharpness parameters of the uniformly illuminated region and the shadow region of each similar character, so as to obtain the enhancement and repair coefficient of each similar character in the test paper image, enhance and repair the characters, facilitate the detection and recognition of the test paper content, and improve the recognizability of the test paper content. Description of the Drawings
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings.
[0042] Figure 1 Is the step flowchart of a method for intelligent scanning and processing of test papers for a marking system of the present invention;
[0043] Figure 2Schematic diagram of the circumscribed rectangle of a character;
[0044] Figure 3 Schematic diagram of the shaded area and the evenly illuminated area of the character. Detailed implementation manners
[0045] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following specifically describes, with reference to the accompanying drawings and preferred embodiments, a method for intelligent scanning and processing of test papers for a marking system, including its specific implementation manners, structures, features and effects. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0047] The following specifically describes the specific solution of a method for intelligent scanning and processing of test papers provided by the present invention with reference to the accompanying drawings.
[0048] Please refer to Figure 1 , which shows a flowchart of the steps of a method for intelligent scanning and processing of test papers provided by an embodiment of the present invention. The method includes the following steps:
[0049] Step S001: Obtain a plurality of test paper images, and obtain a plurality of contour pixel points of each test paper image.
[0050] The purpose of this embodiment is to ensure the integrity of characters and images in the occluded or shaded areas through the extension trend of font symbols in the test paper images, so as to detect and recognize the test paper content. Therefore, the filled test papers need to be uploaded to the marking system by means of photographing and scanning first, and the test paper images are stored according to the exam ID numbers in the educational administration information module.
[0051] Specifically, for each uploaded test paper image, the canny operator is used to sharpen the image to improve the clarity of the image, and the Gaussian filter is used to denoise the image to obtain the preprocessed test paper image. The contour detection is performed on each test paper image through the findContours function to obtain a plurality of edge pixel points, which are used as a plurality of contour pixel points of each test paper image. A plane rectangular coordinate system is established with the upper left corner of each test paper image as the origin, and the coordinates of each pixel point data in the test paper image are , and the pixel value is ; Subsequently, any test paper image is processed.
[0052] Step S002: Cluster the pixel points in the test paper image to obtain several clusters, and obtain the pixel difference coefficient of each cluster according to the pixel value differences of different pixel points in each cluster; obtain the uniformly illuminated area and the shadow area according to the pixel difference coefficients of each cluster.
[0053] It should be noted that during the process of uploading the test paper image to the marking system, due to the influence of light and occlusions, there may be obvious shadow areas in the test paper image, resulting in unclear content in the shadow areas, or when some candidates are answering questions, due to insufficient pen ink or writing habits, etc., the handwriting is relatively light during writing, which further exacerbates the phenomenon of blurred handwriting, making the answer characters detected in the system abnormal, and thus affecting the scanning result of the test paper. In the uniformly illuminated area, the brightness difference between the written characters and the paper background is obvious, and when performing contour detection, a relatively large number of contour pixel points are detected. However, within the range of the shadow area, due to the low brightness of the shadow, the brightness difference between the written characters and the paper background is small, resulting in a relatively small number of contour pixel points detected during contour detection; therefore, by analyzing the abnormal character areas in the image, the character contour parameters are analyzed according to the change trend of the image or characters.
[0054] Specifically, for the pixel points in any test paper image, the K-means clustering algorithm is used, and the elbow method is used to determine the optimal value, and then the test paper image is clustered and divided, where the distance metric uses the absolute value of the difference between the pixel values of the pixel points, and clusters are obtained;
[0055] The calculation method of the regional feature parameter of the k-th cluster is:
[0056]
[0057] The calculation method of the pixel difference coefficient of the k-th cluster is:
[0058]
[0059] Among them, represents the regional feature parameter of the k-th cluster, represents the standard deviation of the pixel values of all pixel points in the k-th cluster, represents the number of contour pixel points in the k-th cluster, represents the pixel difference coefficient of the k-th cluster, represents the maximum value of the regional feature parameters among all clusters, represents the maximum value of the differences between the pixel values of all pixel points in the k-th cluster, is a linear normalization function, and the normalization object is the of all clusters.
[0060] It should be noted that when the standard deviation of pixel values and the number of contour pixels during contour detection are larger, it indicates that it is more likely to be the area where the examinee fills in the characters. It represents the ratio of the regional feature parameter of the k-th cluster to the regional feature parameters in all clusters. The larger the value, the more likely it is to be the area where the examinee fills in the characters.
[0061] Furthermore, traverse all clusters to obtain the pixel difference coefficient of each cluster; set a difference threshold. When the pixel difference coefficient of the k-th cluster is greater than or equal to the difference threshold, several regions composed of the pixels in the k-th cluster are recorded as uniformly illuminated regions. When the pixel difference coefficient of the k-th cluster is less than the difference threshold, several regions composed of the pixels in the k-th cluster are recorded as shadow regions; in this embodiment, the set difference threshold is 0.6, and other values can be set in other embodiments, which are not specifically limited in this embodiment.
[0062] Step S003: Obtain the bounding rectangle of each character in the uniformly illuminated region. According to the size of each bounding rectangle and the distance between adjacent bounding rectangles, obtain the merged character difference value after each merge of the characters in each row. Based on the merged character difference value, obtain the final merge result of the characters in each row, and further split to obtain the split character difference value after each split of the characters in each row, and finally obtain the final split result of the characters in each row.
[0063] It should be noted that since different examinees have different writing habits during the writing process, there will be certain personalized differences in aspects such as font selection, line spacing control, and the spacing between letters during the exam, which are not only reflected in the size and clarity of the glyphs, but also include multiple detailed aspects such as the spacing between lines and the distance between characters; due to the diversity of these writing habits, it affects the content recognition of the test paper image; therefore, when recognizing the content in the test paper image, it is necessary to first obtain the spacing and size of the font on the test paper image.
[0064] Furthermore, it should be noted that during the writing process of the examinee, the inside of the character contour is filled with pen ink, and the brightness of the character area is lower than that of the paper background part. At the same time, there are phenomena such as connected strokes or overly long or narrow character widths caused by left-right or up-down structures during the specific writing process. To avoid such phenomena, it is necessary to merge or split the bounding rectangles of adjacent characters.
[0065] Specifically, in the uniformly illuminated region, use the boundingRect function in the image algorithm to calculate the bounding rectangle of the character, as Figure 2 shown, to obtain the length of each character's bounding rectangle in the X-axis direction , and at the same time record the Euclidean distance between adjacent bounding rectangles in the same row ; according to the length of the circumscribed rectangle , and the Euclidean distance between adjacent circumscribed rectangles , respectively obtain the standard deviation of the length of the circumscribed rectangles of the whole-line characters and the standard deviation of the Euclidean distance , denoted as the length discrete parameter value and the distance discrete parameter value.
[0066] Further, for any whole-line character, merge the two adjacent circumscribed rectangles corresponding to the minimum Euclidean distance among the adjacent characters in this line as the first merge, and recalculate the length of the circumscribed rectangle and the Euclidean distance between adjacent circumscribed rectangles, so as to obtain the length discrete parameter value and the distance discrete parameter value after the first merge; and so on, perform several merges in ascending order of the Euclidean distance, and obtain the length discrete parameter value and the distance discrete parameter value after each merge of the whole-line characters; obtain the merged character difference value after any merge of this whole-line character ; select the merge result corresponding to the minimum value of the merged character difference values after all merges as the final merge result of this whole-line character.
[0067] It should be noted that due to the cursive characters generated during the candidate's filling process, the length of the circumscribed rectangle of the characters is relatively large, resulting in a relatively large merged character difference value after merging, which is not conducive to character detection. Therefore, it is necessary to split the cursive characters. During the writing process of cursive characters, there is adhesion between the pen-down of the previous character and the pen-up of the next character. The adhesive part has narrow strokes and occupies a small number of pixel points.
[0068] Further, under the final merge result of this whole-line character, successively select the circumscribed rectangles of the characters with the largest length in the X-axis direction for splitting, and use the position where the characters occupy the fewest pixel points under the Y-axis coordinate points in the circumscribed rectangle of each character as the character splitting area, and so on, split the circumscribed rectangles of the characters according to the length from large to small; after splitting, obtain the standard deviation of the length of the circumscribed rectangles of the whole-line characters in the same way as the merge operation , and the standard deviation of the Euclidean distance , and further obtain the split character difference value , select the split result corresponding to the minimum value of the split character difference value as the final split result of this whole-line character.
[0069] Step S004: Based on the non-rectangular characters of the circumscribed polygons in the light-uniform area and in combination with the shadow area, obtain the semi-occluded characters in the light-uniform area and the occluded characters in the shadow area; based on the distribution of the contour pixel points in the semi-occluded characters and the occluded characters, obtain the same stroke state parameter of the semi-occluded characters in the light-uniform area and the occluded characters in the shadow area, and further obtain several similar characters.
[0070] It should be noted that for the characters inside the shadow area, due to the low brightness of the shadow, it is difficult to obtain the contour parameters of the characters through contour detection. Therefore, when identifying the shadow area with characters, it is necessary to analyze according to the manifestations of the shadow and the characters. For the semi-occluded characters, part of the area is within the uniformly illuminated area, and the rest is within the shadow area. The characters can be analyzed through the semi-occluded characters.
[0071] Furthermore, it should be noted that under normal conditions, the circumscribed polygon of the character is a rectangle. After being occluded, the character is missing, resulting in a non-rectangular circumscribed polygon. The pixel value of the contour pixel points of the character changes most significantly in the gradient direction towards the pixel value of the paper background. The normal vector in the gradient direction is the stroke direction of the character at the position of the pixel point. Due to the existence of non-linear strokes such as flicks and strokes in the font itself and the irregularity of the examinee's font, the font strokes are not horizontal and vertical. At the same time, during the writing process, there is coherence between the same strokes. Therefore, when they belong to the same stroke, the change direction of the character strokes is similar or opposite.
[0072] Specifically, the characters whose circumscribed polygons in the uniformly illuminated area are not rectangles are regarded as semi-occluded characters. A rectangle is circumscribed around the circumscribed polygon of the semi-occluded characters. Then, the largest difference area between the circumscribed rectangle and the polygon is the occluded shadow part, which is regarded as the occluded character. As Figure 3 shown, part a is the occluded shadow part, and part b is the uniformly illuminated part;
[0073] For the semi-occluded character b in the uniformly illuminated area, obtain the gradient direction of each contour pixel point of the semi-occluded character. For the character strokes of the examinee, obtain the fitting direction vector of the normal vector through the curve fitting algorithm , which is denoted as the change direction vector of the character strokes. Similarly, for the occluded character a in the shadow area, obtain the change direction vector of the character strokes .
[0074] Furthermore, by obtaining the change direction vectors of the character strokes of the occluded character a and the semi-occluded character b 、 the included angle , and use it as the influence parameter of the same stroke between the semi-occluded character in the uniformly illuminated area and the occluded character in the shadow area. Use the linear normalization function for processing to obtain the influence factor of the same stroke , so that when, , that is, the interval range before the linear normalization function processing is , and the interval range after normalization is .
[0075] Further, the distance between any two contour pixel points with opposite gradient directions in the same part is denoted as the width of the character stroke. The calculation method for the same stroke state parameter of the semi-occluded character in the uniformly illuminated area and the occluded character in the shadow area is as follows:
[0076]
[0077] Among them, is the same stroke state parameter of the semi-occluded character in the uniformly illuminated area and the occluded character in the shadow area, is the character stroke width of the occluded character in the shadow area, is the character stroke width of the semi-occluded character in the uniformly illuminated area, is the same stroke influence factor of the semi-occluded character in the uniformly illuminated area and the occluded character in the shadow area; when is closer to 1, the a part and the b part are more similar. Use the function to normalize the result, so that when the result value is closer to 0, the value is closer to 1.
[0078] Further, a similarity threshold is set. In this embodiment, the similarity threshold is described as 0.8. If the same stroke state parameter is greater than or equal to the similarity threshold, it means that the gradient directions of the uniformly illuminated area and the shadow area are similar, and it is more likely to be the same stroke. Mark the semi-occluded character in the uniformly illuminated area and the occluded character in the shadow area as similar characters; if the same stroke state parameter is less than the similarity threshold, it means that the gradient directions of the uniformly illuminated area and the shadow area are less similar, and it is more likely not to be the same stroke, and no marking is performed.
[0079] Step S005: Obtain the image sharpness parameter of the uniformly illuminated area of each similar character according to the pixel difference coefficient and the gradient direction of the contour pixel points in the uniformly illuminated area; obtain the image sharpness parameter of the shadow area of the similar character; obtain the enhancement repair coefficient of each similar character in the test paper image through the image sharpness parameter of the similar character and the same stroke state parameter of the semi-occluded character in the uniformly illuminated area and the occluded character in the shadow area corresponding to it.
[0080] It should be noted that there are random phenomena such as the size and shape of the shadow part in the occluded area of the test paper. It is impossible to predict the area of the occluded part in the test paper and the number of characters. To ensure the integrity and clarity of the content of the test paper obtained by photographing and scanning; therefore, it is necessary to predict and lay the same row direction in the answer box through the character circumscribed rectangle obtained from the test paper answer box.
[0081] Specifically, the circumscribed rectangle of the similar characters corresponding to the semi-occluded characters in the evenly illuminated area and the occluded characters in the shadow area is used as the circumscribed rectangle of the starting character for laying. Starting from the circumscribed rectangle of the starting character, for the shadow area along the X-axis direction of the characters in the same row, based on the average Euclidean distance and the average length of all adjacent circumscribed rectangles in the final splitting result of the characters in this entire row, the circumscribed rectangles of the characters are laid until the boundary position of the answer box is reached; similarly, for the shadow area along the Y-axis direction, the laying ends when the boundary position of the answer box is reached.
[0082] The calculation method of the image clarity parameter of the evenly illuminated area of the similar characters in the j-th column of the i-th row is as follows:
[0083]
[0084] Among them, represents the image clarity parameter of the evenly illuminated area of the similar characters in the j-th column of the i-th row, n represents the length of the semi-occluded characters in the evenly illuminated area corresponding to the similar characters in the j-th column of the i-th row in the X-axis direction, m represents the length of the semi-occluded characters in the evenly illuminated area corresponding to the similar characters in the j-th column of the i-th row in the Y-axis direction, represents the pixel difference coefficient of the cluster to which the evenly illuminated area corresponding to the similar characters in the j-th column of the i-th row belongs, represents the gradient direction of the contour pixel point with coordinates in the evenly illuminated area of the similar characters in the j-th column of the i-th row.
[0085] Similarly, based on the pixel points in the occluded characters in the shadow area corresponding to the similar characters and the pixel difference coefficient of the cluster to which they belong, the image clarity parameter of the similar characters in the j-th column of the i-th row in the shadow area is obtained .
[0086] Furthermore, the calculation method of the enhancement and repair coefficient of the test paper image is as follows:
[0087]
[0088] Among them, is the enhancement and repair coefficient of the similar characters in the j-th column of the i-th row, represents the image clarity parameter of the evenly illuminated area of the similar characters in the j-th column of the i-th row, represents the image clarity parameter of the shadow area of the similar characters in the j-th column of the i-th row, is the ratio of the image clarity parameters between the evenly illuminated area and the shadow area, representing the difference parameter between the two areas, represents the same stroke state parameter of the semi-occluded characters in the evenly illuminated area corresponding to the similar characters and the occluded characters in the shadow area in the j-th column of the i-th row.
[0089] Further, multiply the pixel value of each pixel point in the shadow area of each similar character by the enhancement repair coefficient, and use the result as the enhanced pixel value of each pixel point, thereby completing the enhancement of the test paper image and realizing the intelligent scanning process of the test paper.
[0090] So far, the present invention is completed.
[0091] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A test paper intelligent scanning processing method for a test paper marking system, characterized in that: The method comprises the following steps: Acquire a number of test paper images, and obtain a number of contour pixel points of each test paper image through contour detection; The pixels in the test paper image are clustered to obtain several clusters. According to the pixel value differences of different pixels in each cluster, the pixel difference coefficient of each cluster is obtained; according to the pixel difference coefficient of each cluster, the uniform illumination area and the shadow area are obtained; Obtain the bounding rectangle of each character in the uniformly illuminated area, obtain the merged character difference value after each merge of each line of characters according to the size of each bounding rectangle and the distance between adjacent bounding rectangles, obtain the final merged result of each line of characters according to the merged character difference value, and further split each line of characters to obtain the split character difference value after each split, and finally obtain the final split result of each line of characters; According to the characters with non-rectangular circumscribed polygons in the uniformly illuminated area, combined with the shadow area, the semi-occluded characters in the uniformly illuminated area and the obscured characters in the shadow area are obtained; based on the distribution of contour pixels in the semi-occluded characters and the obscured characters, the same stroke state parameters of the semi-occluded characters in the uniformly illuminated area and the obscured characters in the shadow area are obtained, thereby obtaining a number of similar characters; According to the pixel difference coefficient of similar characters in the uniformly illuminated area and the gradient direction of the contour pixels, the image clarity parameters of the uniformly illuminated area of each similar character are obtained; the image clarity parameters of the shadow area of the similar characters are obtained; the enhanced restoration coefficient of each similar character in the test paper image is obtained through the image clarity parameters of the similar characters and the same stroke state parameters of the semi-occluded characters in the uniformly illuminated area and the occluded characters in the shadow area.
2. According to claim 1, a test paper intelligent scanning and processing method for a test paper marking system is characterized in that: The specific steps of obtaining the pixel difference coefficient of each cluster are as follows: The calculation method of the regional characteristic parameters of the kth cluster is: The pixel difference coefficient of the kth cluster is calculated as: in, represents the regional characteristic parameter of the kth cluster, represents the standard deviation of the pixel values of all pixels in the kth cluster, represents the number of contour pixels in the kth cluster, represents the pixel difference coefficient of the kth cluster, represents the maximum value of the regional characteristic parameters in all clusters, represents the maximum value of the difference between the pixel values of all pixels in the kth cluster, is a linear normalization function.
3. According to claim 1, a test paper intelligent scanning processing method for a test paper marking system is characterized in that: The specific steps of obtaining the merged character difference value after each merge of each line of characters are as follows: Get the length of the bounding rectangle of each character in the X-axis direction , and simultaneously record the Euclidean distance between adjacent bounding rectangles in the same row ; According to the length of the circumscribed rectangle , and the Euclidean distance between adjacent bounding rectangles , respectively obtain the standard deviation of the length of the circumscribed rectangle of the entire line of characters and the standard deviation of the Euclidean distance , recorded as the length discrete parameter value and the distance discrete parameter value; For any whole line of characters, two adjacent bounding rectangles corresponding to the minimum Euclidean distance among the adjacent characters in the line are merged as the first merge, and the length of the bounding rectangle and the Euclidean distance between the adjacent bounding rectangles are recalculated, so as to obtain the length discrete parameter value and the distance discrete parameter value after the first merge; and so on, several merges are performed in the order of Euclidean distance from small to large, and the length discrete parameter value and the distance discrete parameter value after each merge of the whole line of characters are obtained; the merged character difference value of the whole line of characters after any merge is obtained. .
4. According to claim 3, a test paper intelligent scanning processing method for a marking system is characterized in that: The specific steps of obtaining the final merging result of each line of characters according to the merging character difference value are as follows: For any whole line of characters, the merge result corresponding to the minimum value of the merged character difference values after all merges is selected as the final merge result of the whole line of characters.
5. According to claim 3, a test paper intelligent scanning and processing method for a test paper marking system is characterized in that: The final splitting result of each line of characters is finally obtained, and the specific steps included are as follows: For any whole line of characters, under the final merged result of the whole line of characters, select the character with the longest length in the X-axis direction for splitting in turn, and take the position of the character with the least pixels under the Y-axis coordinate point in the bounding rectangle of each character as the character segmentation area, and so on, split the bounding rectangle of the character from large to small according to the length; obtain the standard deviation of the length of the bounding rectangle of the whole line of characters after each split , and the standard deviation of the Euclidean distance , and further obtain the split character difference value , select the split result corresponding to the minimum value of the split character difference value as the final split result of the entire line of characters.
6. According to claim 1, a test paper intelligent scanning and processing method for a test paper marking system is characterized in that: The specific steps of obtaining the semi-shielded characters in the uniformly illuminated area and the shielded characters in the shadow area are as follows: Characters in the uniformly illuminated area whose circumscribed polygons are not rectangular are regarded as semi-occluded characters. A rectangle is circumscribed to the circumscribed polygon of the semi-occluded character. The area with the maximum difference between the circumscribed rectangle and the polygon is the shaded part that is obscured, which is regarded as the obscured character.
7. According to claim 1, a test paper intelligent scanning and processing method for a test paper marking system is characterized in that: The specific steps of obtaining a plurality of similar characters are as follows: The distance between any two contour pixels with opposite gradient directions in the same part is recorded as the width of the character stroke. The same stroke state parameters of the semi-occluded character in the uniformly illuminated area and the occluded character in the shadow area are calculated as follows: in, is the same stroke state parameter of the semi-occluded character in the uniformly illuminated area and the occluded character in the shadow area, is the stroke width of the blocked characters in the shadow area, is the character stroke width of the semi-occluded character in the uniformly illuminated area, is the same stroke influence factor of the semi-occluded character in the uniformly illuminated area and the occluded character in the shadow area; The function normalizes the result; A similarity threshold is set. If the state parameter of the same stroke is greater than or equal to the similarity threshold, the semi-occluded characters in the evenly illuminated area and the occluded characters in the shadow area are marked as similar characters.
8. The method for intelligent scanning and processing of examination papers for a paper marking system according to claim 7, characterized in that: The specific method of obtaining the same stroke influence factor is as follows: For the semi-occluded character b in the uniformly illuminated area, obtain the gradient direction of each contour pixel point of the semi-occluded character. For the character strokes of the examinee, obtain the fitting direction vector of the normal vector through the curve fitting algorithm. , recorded as the change direction vector of the character stroke; for the blocked character a in the shadow area, obtain the change direction vector of the character stroke ; By obtaining the character stroke change direction vector of the blocked character a and the semi-blocked character b , Angle , and used as the same stroke influence parameter of the semi-occluded characters in the uniform illumination area and the occluded characters in the shadow area. Use the linear normalization function to process and get the same stroke influence factor .
9. The method for intelligent scanning and processing of examination papers for a paper marking system according to claim 1, characterized in that: The specific steps of obtaining the image clarity parameter of the uniformly illuminated area of each similar character are as follows: in, represents the image clarity parameter of the uniformly illuminated area of the similar characters in the i-th row and j-th column, n represents the length of the semi-occluded characters in the uniformly illuminated area corresponding to the similar characters in the i-th row and j-th column in the X-axis direction, m represents the length of the semi-occluded characters in the uniformly illuminated area corresponding to the similar characters in the i-th row and j-th column in the Y-axis direction, It represents the pixel difference coefficient of the cluster of the uniformly illuminated area corresponding to the similar characters in the i-th row and j-th column. The coordinates of the similar characters in the uniform illumination area of the i-th row and j-th column are The gradient direction of the contour pixel.
10. The method for intelligent scanning and processing of examination papers for a paper marking system according to claim 1, characterized in that: The specific steps of obtaining the enhancement restoration coefficient of each similar character in the test paper image are as follows: in, is the enhancement repair coefficient of similar characters in the i-th row and j-th column, represents the image clarity parameter of the uniformly illuminated area of similar characters in the i-th row and j-th column, represents the image clarity parameter of the shadow area of the similar character in the i-th row and j-th column, Indicates the same stroke state parameter of the semi-occluded character in the uniformly illuminated area and the obscured character in the shadow area corresponding to the similar character in the i-th row and j-th column.
Citation Information
Patent Citations
Shaded character recovery device and method thereof, shaded character recognition device and method thereof
CN102208022A
Stroke width figure based method for extracting Chinese character data from image
CN104598907A