An intelligent marking content detection and recognition method and system based on deep learning
By combining deep learning and contextual semantic analysis with pixel value processing and writing models, the problem of inaccurate fuzzy text recognition has been solved, achieving high efficiency and accuracy in intelligent marking.
Patent Information
- Application Number
- CN202510211218.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-02-25
AI Technical Summary
The diversity of handwritten text, the clarity of writing, and the quality of scanning in existing technologies make it difficult to accurately recognize blurry text, resulting in inaccurate intelligent marking results.
By using deep learning-based methods, combined with contextual semantic analysis and writing models, the system identifies fuzzy text on exam papers. It utilizes techniques such as locators, pixel value analysis, threshold segmentation, and dilation erosion to accurately locate and process effective detection areas, filter out clear and fuzzy text, and determine target text through various strategies.
It improves the accuracy and rationality of fuzzy text recognition, enhances the adaptability and accuracy of intelligent marking, reduces the time for recognizing fuzzy text, and improves marking efficiency and fairness.
Smart Images

Figure CN120126146B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to an intelligent paper review content detection and recognition method and system based on deep learning. BACKGROUND
[0002] Deep learning technology has made great progress in recent years and has shown strong capabilities in image and text processing. In terms of images, deep learning architectures such as convolutional neural networks (CNN) and recurrent neural networks (RNN) can effectively extract image features and have strong representation capabilities for various complex image content, enabling better handling of handwriting recognition tasks. In terms of text processing, deep learning can provide a technical basis for intelligent paper review by deeply understanding and analyzing the semantics and structure of text.
[0003] The patent document with the Chinese patent application publication number CN113222788A discloses an intelligent paper review method, which includes the following steps: first, scanning the complaint, obtaining the complaint label and content, then magnifying or rotating the extracted label and content for correction, recognizing the scanned text and obtaining the complaint content text and confirming the keywords, corresponding the obtained keywords with the keywords in the case filing table, and filling the content into the case filing table; according to the text keywords, retrieving the relevant electronic case files from the electronic case file database and selecting whether to translate according to the operator's needs, filling the electronic case file content into the case filing table, and auditing whether the content is accurate, if the filling is accurate, the case filing table is completed, if it is not accurate, the content is modified, the modified content is synchronized to the electronic case file and the modified content is backfilled into the case filing table, and the case filing table is completed.
[0004] In the prior art, due to the diversity of handwritten text, writing clarity, and scanning quality, etc., there is a problem of fuzzy text that is difficult to accurately recognize. Simple keyword matching does not involve in-depth understanding and analysis of text semantics, making the text recognition result not comprehensive, resulting in inaccurate intelligent paper review results. SUMMARY
[0005] Therefore, the present application provides an intelligent paper review content detection and recognition method and system based on deep learning, which can solve the problem of inaccurate intelligent paper review results by identifying fuzzy text on the test paper through context semantic analysis, deep learning to construct a writing model, etc.
[0006] To achieve the above purpose, the present application provides an intelligent paper review content detection and recognition method based on deep learning, which comprises:
[0007] Determining the test paper area to be detected based on the locator;
[0008] Recognizing the area to be detected based on the pixel value, and determining the effective detection area according to the recognition result;
[0009] identify a plurality of uploaded characters in the effective detection area, match the plurality of uploaded characters with a preset character database to determine a plurality of clear characters and a plurality of ambiguous characters according to a matching result;
[0010] determine a plurality of alternative preset characters corresponding to any ambiguous character, and screen the plurality of alternative preset characters based on context semantics to determine a plurality of first preset characters;
[0011] determine a plurality of second preset characters corresponding to any ambiguous character based on the plurality of ambiguous characters, determine a target character according to the plurality of first preset characters or the plurality of second preset characters based on a judgment result, or determine a plurality of third preset characters based on a writing model determined by a deep learning algorithm and the plurality of clear characters, and determine the target character based on the plurality of first preset characters and the plurality of third preset characters;
[0012] label the ambiguous characters according to the target character to assist intelligent marking.
[0013] Further, the step of determining the effective detection area according to the recognition result comprises:
[0014] identify a plurality of actual pixel values corresponding to the to-be-detected area, and determine an initial character area based on a threshold segmentation algorithm and the plurality of actual pixel values;
[0015] adjust the initial character area based on dilation operation and erosion operation to obtain an adjusted character area;
[0016] perform projection analysis on the adjusted character area based on character arrangement characteristics to determine the effective detection area according to an analysis result.
[0017] Further, the step of determining a plurality of clear characters and a plurality of ambiguous characters according to a matching result comprises:
[0018] perform similarity matching on a plurality of uploaded characters and a plurality of preset characters in a preset character database to obtain a plurality of actual similarity values;
[0019] compare the plurality of actual similarity values with a preset similarity value to determine a plurality of clear characters or a plurality of ambiguous characters according to a comparison result.
[0020] Further, the step of determining a plurality of alternative preset characters corresponding to any ambiguous character comprises:
[0021] calculate the similarity of any ambiguous character and characters in a preset character database and sort the similarity to obtain a first similarity sequence;
[0022] Determine a plurality of candidate words corresponding to the first similarity sequence and a preset number threshold.
[0023] Further, the step of screening a plurality of candidate words based on context semantics comprises:
[0024] Locate any ambiguous word, determine its corresponding context content based on punctuation marks, and convert it into a vector form based on a word vector model to obtain an upper word vector and a lower word vector;
[0025] Calculate the average semantic similarity between a plurality of candidate words and the upper word vector and the lower word vector, respectively, to obtain a first semantic similarity;
[0026] Compare a plurality of first semantic similarities with a first preset semantic similarity, and determine a plurality of first preset words based on the comparison result.
[0027] Further, the step of determining the target word according to a plurality of first preset words or a plurality of second preset words comprises:
[0028] Calculate a plurality of second semantic similarities corresponding to a plurality of second preset words corresponding to the same ambiguous word;
[0029] Select a plurality of fourth preset words that are the same as a plurality of first preset words and a plurality of second preset words;
[0030] Determine a third semantic similarity corresponding to the fourth preset word based on the first semantic similarity and the second semantic similarity, and determine the target word based on a plurality of third semantic similarities.
[0031] Further, the step of determining the writing model based on a deep learning algorithm and a plurality of clear words comprises:
[0032] Input a plurality of clear words into a preset deep learning model to train and obtain the writing model;
[0033] Extract features of the ambiguous word to be recognized through the writing model to obtain an ambiguous word feature vector;
[0034] Match the ambiguous word feature vector with a preset word feature vector in a preset word database to determine a plurality of third preset words according to the matching result.
[0035] Further, the step of determining the target word based on a plurality of first preset words and a plurality of third preset words comprises:
[0036] Calculate the average semantic similarity between a plurality of third preset words and the upper word vector and the lower word vector, respectively, to obtain a third semantic similarity;
[0037] determine the target character based on a comparison result of the first semantic similarity and the third semantic similarity.
[0038] Further, the step of determining the test paper to-be-detected region based on the locator comprises:
[0039] determining the locator based on the preset shape and the preset pixel value;
[0040] determining the horizontal locator and the vertical locator based on the distance between the centers of the adjacent locators;
[0041] correcting the test paper based on the horizontal locator and the vertical locator to determine the to-be-detected region corresponding to the locator.
[0042] In another aspect, the present application also provides a system of the intelligent test paper content detection and recognition method based on deep learning, which comprises:
[0043] a region determination module, configured to determine the test paper to-be-detected region based on the locator, recognize the to-be-detected region based on the pixel value, and determine the effective detection region according to the recognition result;
[0044] a character recognition module, connected with the region determination module, configured to recognize a plurality of uploaded characters in the effective detection region, match the plurality of uploaded characters with a preset character database, and determine a plurality of clear characters and a plurality of fuzzy characters according to a matching result;
[0045] a screening module, connected with the character recognition module, configured to determine a plurality of candidate preset characters corresponding to any fuzzy character, screen the plurality of candidate preset characters based on context semantics, and determine a plurality of first preset characters;
[0046] a target determination module, connected with the screening module, configured to determine whether there is a same fuzzy character based on the plurality of fuzzy characters, determine a plurality of second preset characters corresponding to the same fuzzy character based on a determination result, and determine a target character according to the plurality of first preset characters or the plurality of second preset characters, or determine a writing model based on a deep learning algorithm and the plurality of clear characters, determine a plurality of third preset characters based on the writing model and the fuzzy character, and determine the target character based on the plurality of first preset characters and the plurality of third preset characters;
[0047] a labeling module, connected with the target determination module, configured to label the fuzzy character according to the target character to assist intelligent test paper review.
[0048] Compared with the prior art, the beneficial effects of the present application are that the specific area needing detection in the test paper can be accurately locked by the locator, the processing efficiency is improved, the adaptability of the method to different test paper templates is enhanced, through analysis of the pixel value, the threshold segmentation, expansion and corrosion are used to effectively remove the noise and irrelevant background information in the image, only the effective area related to the text is retained, the accuracy of subsequent text recognition is improved, the effective detection area is clearly defined, invalid calculation is avoided, and the processing efficiency is further improved, the clear and distinguishable text and the fuzzy text are quickly distinguished, which provides a basis for subsequent processing of different types of text, the quality of text recognition is preliminarily evaluated by matching with the preset text database, whether the text is clear and distinguishable is judged, which helps to improve the accuracy and reliability of the test paper reading, the context semantic screening of the candidate text is used, so that the determined first preset text is more in line with the overall context of the text in semantics, the accuracy and rationality of the fuzzy text recognition are improved, the target text is determined through multiple strategies, the accuracy of the fuzzy text recognition is improved, the writing model is learned from the clear text based on the deep learning algorithm, the writing characteristics of the text can be better understood, so that the fuzzy text can be more accurately inferred, the self-adaptability is improved, the target text is marked in the fuzzy text, the test paper reading efficiency is greatly improved, the time for recognizing the fuzzy text is reduced, and the accuracy and fairness of the test paper reading are improved.
[0049] Especially, by identifying the actual pixel value of the detection area, the information in the test paper image can be converted into a data form that can be quantified and processed, providing basic data for subsequent processing, the threshold segmentation algorithm divides the image into different regions according to the different pixel values, and preliminarily separates the region that may contain text and the background region, thereby realizing the first step of screening the text area from the complex test paper image, reducing the range that needs to be further processed, and improving the pertinence and efficiency of subsequent processing, the disconnected parts of text strokes are connected through the dilation operation, for some texts that are written roughly or have strokes broken due to scanning, the disconnected strokes can be connected into complete shapes through the dilation operation, which is beneficial to subsequent text recognition and processing, and the text structure is more complete, the redundant noise points generated by the expansion are removed through the erosion operation, the quality of the text area is improved, and the continuity and integrity of the text area are enhanced, the projection analysis is performed according to the arrangement characteristics that the text is usually arranged in lines with certain line spacing and character spacing, the text area can be better organized and segmented according to natural lines and blocks, the projection curve is obtained by accumulating the pixel values in the horizontal and vertical directions, the text line position and the text block position are accurately determined according to the characteristics of the wave crest and the wave trough of the curve, which helps to divide the text content into the correct lines and paragraphs, avoids misjudging the text of different lines as one line, ensures the accuracy and orderliness of subsequent text content recognition and processing, and improves the accuracy and reliability of the entire intelligent test paper reading process. BRIEF DESCRIPTION OF DRAWINGS
[0050] Figure 1 A flowchart of the intelligent paper review content detection and recognition method based on deep learning provided by the embodiment of the present application is shown in the figure.
[0051] Figure 2 A flowchart of determining an effective detection area in the intelligent paper review content detection and recognition method based on deep learning provided by the embodiment of the present application is shown in the figure.
[0052] Figure 3 A structure diagram of a paper to be detected area in the intelligent paper review content detection and recognition method based on deep learning provided by the embodiment of the present application is shown in the figure.
[0053] Figure 4 A structure block diagram of the intelligent paper review content detection and recognition system based on deep learning provided by the embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0054] In order to make the purpose and advantages of the present application clearer and more apparent, the present application will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.
[0055] The preferred embodiments of the present application will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present application and are not intended to limit the protection scope of the present application.
[0056] It should be noted that in the description of the present application, the terms "upper", "lower", "left", "right", "inner", "outer" and the like indicating the direction or positional relationship of the terms are based on the direction or positional relationship shown in the drawings, which is only for the convenience of description and does not indicate or imply that the device or element must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application.
[0057] In addition, it should also be noted that in the description of the present application, unless otherwise explicitly specified and limited, the terms "mounting", "connecting", "connection" should be understood broadly, for example, it can be fixed connection, or detachable connection, or integral connection; it can be mechanical connection, or electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, or the internal communication of two elements. Those skilled in the art can understand the specific meaning of the above terms in the present application according to the specific circumstances.
[0058] Please refer to Figure 1 The embodiment of the present application provides a kind of intelligent paper review content detection and recognition method based on deep learning, which comprises:
[0059] Step S100, determine the paper to be detected area based on locator;
[0060] In step S200, the to-be-detected area is recognized based on pixel values, and an effective detection area is determined according to a recognition result.
[0061] In step S300, a plurality of uploaded texts in the effective detection area are recognized, the plurality of uploaded texts are matched with a preset text database, and a plurality of clear texts and a plurality of fuzzy texts are determined according to a matching result.
[0062] In step S400, a plurality of candidate preset texts corresponding to any fuzzy text are determined, and the plurality of candidate preset texts are screened based on context semantics to determine a plurality of first preset texts.
[0063] In step S500, it is determined whether there is a same fuzzy text based on the plurality of fuzzy texts, a plurality of second preset texts corresponding to the same fuzzy text are determined based on a determination result, a target text is determined according to the plurality of first preset texts or the plurality of second preset texts, or a writing model is determined based on a deep learning algorithm and the plurality of clear texts, a plurality of third preset texts are determined based on the writing model and the fuzzy text, and the target text is determined based on the plurality of first preset texts and the plurality of third preset texts.
[0064] In step S600, the fuzzy text is labeled according to the target text to assist intelligent marking.
[0065] It can be understood that the to-be-detected area of the test paper in the embodiment of the present application can refer to the short answer question area part corresponding to the liberal arts test paper, and the answer areas are all text descriptions without other formulas or character descriptions.
[0066] It can be understood that the preset text database in the embodiment of the present application should contain various types of vocabulary, phrases and other text contents that may be involved in liberal arts short answer questions. These contents can come from teaching syllabus, teaching materials, true test questions and answers in previous years, etc. For example, for the liberal arts short answer questions of history subject, the database will contain historical event names, historical figure names, related historical terms, etc.; for the liberal arts subject of Chinese, there will be various literary genre names, rhetorical device names, words, etc. The texts in the database are usually stored in a standardized form, and each text or vocabulary has a corresponding standard representation.
[0067] Specifically, the embodiment of the present application can accurately lock the specific area that needs to be detected in the test paper through the locator, improve the processing efficiency, and enhance the adaptability of the method to different test paper templates. Through the analysis of the pixel value, the threshold segmentation, the expansion corrosion, the noise and irrelevant background information in the image are effectively removed, only the effective area related to the text is reserved, the accuracy of subsequent text recognition is improved, the effective detection area is determined, the invalid calculation is avoided, the processing efficiency is further improved, the clear and distinguishable text and the fuzzy text are quickly distinguished, which provides a basis for subsequent processing of different types of text, the quality of text recognition is preliminarily evaluated through matching with the preset text database, whether the text is clear and distinguishable is judged, which helps to improve the accuracy and reliability of the test paper reading, the context semantic screening of the candidate text is used, the determined first preset text is more in line with the overall context of the text in the semantic, the accuracy and rationality of the fuzzy text recognition are improved, the target text is determined through multiple strategies, the accuracy of the fuzzy text recognition is improved, the writing model is learned from the clear text based on the deep learning algorithm, the writing characteristics of the text can be better understood, so that the fuzzy text can be more accurately inferred, the self-adaptability is improved, the target text is marked in the fuzzy text, the test paper reading efficiency is greatly improved, the time for recognizing the fuzzy text is reduced, and the accuracy and fairness of the test paper reading are improved.
[0068] Referring to Figure 2 As shown, the step of determining the effective detection area according to the recognition result comprises:
[0069] Step S210, identifying a plurality of actual pixel values corresponding to the to-be-detected area, and determining an initial text area based on a threshold segmentation algorithm and the plurality of actual pixel values;
[0070] Step S220, adjusting the initial text area based on expansion operation and corrosion operation to obtain an adjusted text area;
[0071] Step S230, performing projection analysis on the adjusted text area based on text arrangement characteristics to determine the effective detection area according to the analysis result.
[0072] Specifically, the embodiment of the present application can convert the information in the test paper image into quantifiable and processable data form by identifying the actual pixel value of the to-be-detected region, thereby providing basic data for subsequent processing. The threshold segmentation algorithm divides the image into different regions according to the difference of pixel value, and preliminarily separates the region possibly containing characters and the background region, thereby realizing the first step of screening the region where the characters are located from the complex test paper image, reducing the range that needs to be further processed, and improving the pertinence and efficiency of subsequent processing. The disconnected parts of the character strokes are connected through the dilation operation. For some characters with rough handwriting or broken strokes caused by scanning, the disconnected strokes can be connected into complete shapes through the dilation operation, which is beneficial to subsequent character recognition and processing, and makes the character structure more complete. The redundant noise points generated by the dilation are removed through the erosion operation, thereby improving the quality of the character region and enhancing the continuity and integrity of the character region. The projection analysis is performed according to the arrangement characteristics that the characters are usually arranged in lines with certain line spacing and character spacing, which can better organize and segment the character region according to natural lines and blocks. The projection curve is obtained by accumulating the pixel values in the horizontal and vertical directions. The character line position and the character block position are accurately determined according to the characteristics of the wave crest and the wave trough of the curve, which is helpful to divide the character content into correct lines and paragraphs, avoids misjudging the characters of different lines as one line, ensures the accuracy and orderliness of subsequent character content recognition and processing, and improves the accuracy and reliability of the entire intelligent marking process.
[0073] It can be understood that the initial character region adjusted based on the dilation operation and the erosion operation in the embodiment of the present application can be adjusted by first using the dilation operation, connecting the disconnected parts of the character strokes by setting a suitable structure element (such as a 3*3 rectangular structure element), and then performing the erosion operation to remove the redundant noise points generated by the dilation, so that the character region is more clear and accurate.
[0074] It can be understood that the character arrangement characteristics in the embodiment of the present application are that the characters are usually arranged in lines with certain line spacing and character spacing.
[0075] It can be understood that the step of performing the projection analysis on the adjusted character region in the embodiment of the present application can include:
[0076] The pixel values in the horizontal direction and the vertical direction corresponding to the adjusted character region are accumulated to obtain a horizontal projection curve and a vertical projection curve.
[0077] The horizontal wave peak feature, the horizontal wave trough feature, the vertical wave peak feature and the vertical wave trough feature corresponding to the horizontal projection curve and the vertical projection curve are determined respectively, the character line position is determined based on the horizontal wave peak feature and the horizontal wave trough feature, the character block position is determined based on the vertical wave peak feature and the vertical wave trough feature, and the effective detection region is determined according to the character line position and the character block position.
[0078] It can be understood that, according to the wave peak and wave trough features of the curve, the position of the character line and the position of the character block are determined, for example, on the horizontal projection curve, the wave trough position corresponds to the blank area between the character lines, and by reasonably setting the wave trough threshold, different character line areas can be accurately divided.
[0079] Specifically, the step of determining the clear characters and the fuzzy characters according to the matching result comprises:
[0080] The similarity between the uploaded characters and the preset characters in the preset character database is matched to obtain actual similarity values;
[0081] The actual similarity values are compared with the preset similarity value to determine the clear characters or the fuzzy characters according to the comparison result.
[0082] It can be understood that, the preset similarity value in the embodiment of the application is 98%.
[0083] It can be understood that, when the actual similarity value is greater than or equal to the preset similarity value, the corresponding uploaded character is determined as a clear character, and when the actual similarity value is less than the preset similarity value, the corresponding uploaded character is determined as a fuzzy character.
[0084] Specifically, the step of determining the clear characters and the fuzzy characters according to the matching result comprises:
[0085] The similarity between the uploaded characters and the preset characters in the preset character database is matched to obtain actual similarity values;
[0086] The actual similarity values are compared with the preset similarity value to determine the clear characters or the fuzzy characters according to the comparison result.
[0087] It can be understood that, the preset number threshold in the embodiment of the application is 5.
[0088] Specifically, the step of screening the plurality of candidate preset characters based on the context semantics comprises:
[0089] Any fuzzy character is located, the context content corresponding to the fuzzy character is determined based on the punctuation symbol, and the context content is converted into a vector form based on a word vector model to obtain an upper word vector and a lower word vector;
[0090] Calculate the average semantic similarity between the upper word vector and the lower word vector of each of the candidate preset characters, respectively, to obtain a first semantic similarity;
[0091] Compare the first semantic similarity with a first preset semantic similarity, and determine the first preset character based on the comparison result.
[0092] It can be understood that the embodiment of the application determines the corresponding context content based on the punctuation symbol and converts it into a vector form based on the word vector model, which includes:
[0093] Taking any of the ambiguous characters as the center, identify the text between the ambiguous character and the nearest previous punctuation symbol as the context content, and the text between the ambiguous character and the nearest next punctuation symbol as the context content based on the preset punctuation symbol library;
[0094] Perform word segmentation processing on the determined context content and context content, respectively, to obtain a plurality of context segmentation and a plurality of context segmentation;
[0095] Convert the plurality of context segmentation and the plurality of context segmentation into a plurality of context segmentation vectors and a plurality of context segmentation vectors based on the word vector model.
[0096] It can be understood that the preset punctuation symbol library of the embodiment of the application can include a plurality of punctuation symbols determined in the prior art, such as comma, period, semicolon, etc.
[0097] It can be understood that the word segmentation processing of the embodiment of the application can be performed by a natural language processing toolkit (such as NLTK, jieba, etc.).
[0098] It can be understood that the vector conversion based on the word vector model of the embodiment of the application can be performed by looking up the corresponding vector in the word table of the word vector model for each word, and if it is an OOV (Out-Of-Vocabulary) word, the vector representation can be generated by randomly initializing the vector or based on the character vector representation method.
[0099] It can be understood that the word vector model of the embodiment of the application is, for example, Word2Vec, GloVe, or a word vector model trained for a specific field.
[0100] It can be understood that the first preset semantic similarity of the embodiment of the application is 0.9.
[0101] It can be understood that when the first semantic similarity is greater than the first preset semantic similarity, the corresponding candidate preset character is taken as the first preset character.
[0102] Specifically, the step of determining the target character according to the plurality of first preset characters or the plurality of second preset characters includes:
[0103] Calculate a plurality of second semantic similarities corresponding to a plurality of second preset character pairs corresponding to the same ambiguous character;
[0104] Select a plurality of fourth preset characters that are the same as the plurality of first preset characters and the plurality of second preset characters;
[0105] Determine a third semantic similarity corresponding to the fourth preset character based on the first semantic similarity and the second semantic similarity, and determine the target character based on a plurality of third semantic similarities.
[0106] It can be understood that the embodiments of the present application also include that if the first preset character set and the second preset character set have no common elements, they can be processed according to specific conditions, for example, according to the semantic similarity scores of the two sets, the elements in the set with a high score can be taken as the fourth preset character; or the two sets are combined and re-evaluated semantically.
[0107] It can be understood that for each element in the fourth preset character set, the embodiments of the present application obtain its corresponding third semantic similarity in a weighted average manner according to its first semantic similarity in the first preset character set and its second semantic similarity in the second preset character set.
[0108] It can be understood that the third semantic similarity of the embodiments of the present application is 0.9.
[0109] Specifically, the step of determining the writing model based on the deep learning algorithm and a plurality of clear characters includes:
[0110] Inputting a plurality of clear characters into a preset deep learning model to train and obtain the writing model;
[0111] Extracting features of the ambiguous character to be recognized through the writing model to obtain an ambiguous character feature vector;
[0112] Matching the ambiguous character feature vector with a preset character feature vector in a preset character database to determine a plurality of third preset characters according to the matching result.
[0113] It can be understood that the preset deep learning model of the embodiments of the present application can be a convolutional neural network model, a recurrent neural network model or a long short-term memory network model, etc., which is used to identify and learn the writing features of clear characters.
[0114] It can be understood that the preset character feature vector of the embodiment of the present application is a vector representation obtained by feature extraction on the characters in the preset character database, which can include information such as the structure, strokes, semantics and the like of the characters, wherein an example of feature vector extraction can be that for each preset character, it is represented in the form of an image, the image is subjected to convolution operation by using the convolution layer in the CNN, the convolution kernel slides on the image, and local features at different levels such as the edges of strokes, corners and the like are extracted; processing is performed through multiple convolution layers and pooling layers; and the feature map after pooling is converted into a fixed-length vector through the full connection layer.
[0115] Specifically, the step of determining the target character based on the plurality of first preset characters and the plurality of third preset characters comprises:
[0116] respectively calculating average semantic similarity of the plurality of third preset characters and the upper word vector and the lower word vector to obtain third semantic similarity;
[0117] determining the target character based on the comparison result of the plurality of first semantic similarities and the plurality of third semantic similarities.
[0118] It can be understood that one possible example of determining the target character in the embodiment of the present application is to select the preset character corresponding to the maximum value in the plurality of first semantic similarities and the plurality of third semantic similarities as the target character.
[0119] Specifically, the step of determining the target character based on the plurality of first preset characters and the plurality of third preset characters comprises:
[0120] determining the locator based on the preset shape and the preset pixel value;
[0121] determining the horizontal locator and the vertical locator based on the distance between the centers of adjacent locators;
[0122] correcting the test paper based on the horizontal locator and the vertical locator to determine the to-be-detected area corresponding to the locator.
[0123] It can be understood that the preset shape in the embodiment of the present application is a template shape determined according to the shape characteristics of the actual design of the locator on the test paper. For example, the locator can be designed as a simple geometric shape such as a circle, a square or a triangle, or a shape with a unique contour composed of specific lines. In actual application, the key geometric features such as the radius of the circle, the side length of the square, the side length and angle of the triangle and the like are extracted as the description parameters of the preset shape through analysis of the shape of the locator.
[0124] The preset pixel value is related to the color characteristic of the locator. When processing the test paper image, the preset pixel value is determined according to the pixel value range of the color of the locator in the corresponding color space (such as RGB, HSV, etc.). For example, if the locator is red, in the RGB color space, the pixel value range corresponding to the red color may be (R: 200-255, G: 0-50, B: 0-50), which is an example of the preset pixel value. By setting such a pixel value range, the area that may belong to the locator can be preliminarily screened out in the image, and the locator is further accurately identified in combination with the preset shape.
[0125] Referring to Figure 4 As shown in the figure, the embodiment of the present application also provides a system of the intelligent test paper content detection and recognition method based on deep learning, which comprises:
[0126] The area determination module 10 is used to determine the test paper to be detected area based on the locator, to recognize the to-be-detected area based on the pixel value, and to determine the effective detection area according to the recognition result;
[0127] The character recognition module 20 is connected with the area determination module 10 and is used to recognize a plurality of uploaded characters in the effective detection area, to match the plurality of uploaded characters with a preset character database, and to determine a plurality of clear characters and a plurality of fuzzy characters according to the matching result;
[0128] The screening module 30 is connected with the character recognition module 20 and is used to determine a plurality of candidate preset characters corresponding to any fuzzy character, to screen the plurality of candidate preset characters based on the context semantics, and to determine a plurality of first preset characters;
[0129] The target determination module 40 is connected with the screening module 30 and is used to judge whether there is the same fuzzy character based on the plurality of fuzzy characters, to determine a plurality of second preset characters corresponding to the fuzzy character same as any of the fuzzy characters based on the judgment result, to determine the target character according to the plurality of first preset characters or the plurality of second preset characters, or to determine a plurality of third preset characters based on the writing model and the fuzzy character, and to determine the target character based on the plurality of first preset characters and the plurality of third preset characters;
[0130] The labeling module 50 is connected with the target determination module 40 and is used to label the fuzzy character according to the target character, so as to assist the intelligent test paper review.
[0131] Specifically, the system of the intelligent test paper content detection and recognition method based on deep learning provided by the embodiment of the present application can execute the above-mentioned intelligent test paper content detection and recognition method based on deep learning, and achieve the same technical effect, which will not be described here.
[0132] So far, the technical solutions of the present application have been described in combination with the preferred embodiments shown in the drawings, but it is easy for those skilled in the art to understand that the protection scope of the present application is obviously not limited to these specific embodiments. Those skilled in the art can make equivalent changes or replacements to the related technical features without departing from the principles of the present application, and the technical solutions after the changes or replacements will all fall within the protection scope of the present application.
[0133] The above only describes the preferred embodiments of the present application and is not intended to limit the present application; the present application can have various changes and variations for those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A deep learning-based intelligent marking content detection and recognition method, characterized in that, The method comprises the following steps: determining a test paper detection area based on a locator; identifying the detection area based on pixel values, and determining an effective detection area according to the identification result; identifying a plurality of uploaded texts in the effective detection area, matching the plurality of uploaded texts with a preset text database, and determining a plurality of clear texts and a plurality of fuzzy texts according to the matching result; determining a plurality of candidate preset texts corresponding to any fuzzy text, screening the plurality of candidate preset texts based on context semantics, and determining a plurality of first preset texts; determining a writing model based on a deep learning algorithm and the plurality of clear texts, determining a plurality of third preset texts based on the writing model and the fuzzy text, and determining a target text based on the plurality of first preset texts and the plurality of third preset texts; annotating the fuzzy text according to the target text to assist intelligent paper marking; The step of screening the plurality of candidate preset texts based on context semantics comprises: locating any fuzzy text, determining its corresponding context content based on punctuation marks, and converting the context content into a vector form based on a word vector model to obtain an upper word vector and a lower word vector; calculating the average semantic similarity between the plurality of candidate preset texts and the upper word vector and the lower word vector respectively to obtain a plurality of first semantic similarities; comparing the plurality of first semantic similarities with a first preset semantic similarity, and determining the plurality of first preset texts based on the comparison result; The step of determining a writing model based on a deep learning algorithm and a plurality of clear texts comprises: inputting the plurality of clear texts into a preset deep learning model to train and obtain the writing model; extracting features of the to-be-identified fuzzy text through the writing model to obtain a fuzzy text feature vector; matching the fuzzy text feature vector with a preset text feature vector in a preset text database to determine a plurality of third preset texts according to the matching result; The step of determining a target text based on a plurality of first preset texts and a plurality of third preset texts comprises: calculating the average semantic similarity between the plurality of third preset texts and the upper word vector and the lower word vector respectively to obtain a plurality of third semantic similarities; determining the target text based on the comparison result of the plurality of first semantic similarities and the plurality of third semantic similarities; The preset deep learning model is a convolutional neural network model, a recurrent neural network model or a long short-term memory network model, which is used to identify and learn the writing features of clear texts; The preset text feature vector is a vector representation obtained by extracting features of the texts in the preset text database, including the structure, strokes and semantic information of the texts. One feature vector extraction example is to represent each preset text as an image, use the convolutional layer in CNN to perform convolution operation on the image, and slide the convolution kernel on the image to extract the edges and corners of the strokes. After processing through multiple convolutional layers and pooling layers, the feature map after pooling is converted into a fixed-length vector through a fully connected layer.
2. The deep learning-based intelligent marking content detection and recognition method according to claim 1, characterized in that, The step of determining an effective detection area according to the identification result comprises: identifying a plurality of actual pixel values corresponding to the detection area, and determining an initial text area based on a threshold segmentation algorithm and the plurality of actual pixel values; adjust the initial character region based on the expansion operation and the erosion operation to obtain an adjusted character region; perform projection analysis on the adjusted character region based on a character arrangement feature to determine the effective detection region according to an analysis result. 3.The deep learning based intelligent marking content detection and recognition method according to claim 2, characterized in that, The step of determining the clear characters and the fuzzy characters according to the matching result includes: matching the uploaded characters with the preset characters in the preset character database to obtain actual similarity values; comparing the actual similarity values with preset similarity values to determine the clear characters or the fuzzy characters according to a comparison result.
4. The deep learning-based intelligent marking content detection and recognition method according to claim 3, characterized in that, The step of determining the candidate preset characters corresponding to any fuzzy character includes: calculating the similarity of any fuzzy character with the characters in the preset character database and sorting to obtain a first similarity sequence; determining the candidate preset characters corresponding to any fuzzy character based on the first similarity sequence and a preset number threshold. 5.The deep learning based intelligent marking content detection and recognition method according to claim 1, characterized in that, Further includes: determining whether there are identical fuzzy characters based on the fuzzy characters, determining the second preset characters corresponding to the identical fuzzy characters based on a determination result, and determining the target character according to the first preset characters or the second preset characters; The step of determining the target character according to the first preset characters or the second preset characters includes: calculating the second semantic similarity of the second preset characters corresponding to the identical fuzzy characters; selecting the fourth preset characters identical to the first preset characters and the second preset characters; determining the third semantic similarity of the fourth preset characters based on the first semantic similarity and the second semantic similarity, and determining the target character based on the third semantic similarity. 6.The deep learning based intelligent marking content detection and recognition method according to claim 1, characterized in that, The step of determining the detection region of the test paper based on the locator includes: determining the locator based on a preset shape and a preset pixel value; determining the horizontal locator and the vertical locator based on the distance between the centers of adjacent locators; correcting the test paper based on the horizontal locator and the vertical locator to determine the detection region corresponding to the locator.
7. A system applying the intelligent marking content detection and recognition method based on deep learning according to any one of claims 1-6, characterized in that, includes: a region determination module configured to determine the detection region of the test paper based on the locator, identify the detection region based on the pixel value, and determine the effective detection region according to an identification result; a character recognition module connected to the region determination module and configured to recognize the uploaded characters in the effective detection region, match the uploaded characters with the preset character database, and determine the clear characters and the fuzzy characters according to a matching result; a screening module connected to the character recognition module and configured to determine the candidate preset characters corresponding to any fuzzy character, screen the candidate preset characters based on the context semantics, and determine the first preset characters; a target determination module connected to the screening module and configured to determine a writing model based on a deep learning algorithm and the clear characters, determine the third preset characters based on the writing model and the fuzzy characters, and determine the target character based on the first preset characters and the third preset characters; a marking module connected to the target determination module and configured to mark the fuzzy characters according to the target character to assist intelligent paper marking.
Citation Information
Patent Citations
Intelligent paper marking method
CN113222788A
Intelligent paper marking implementation method and system based on deep learning and computer program
CN110110585A
Character recognition method and system based on deep learning
CN114049641A