Multi-mode intelligent dictation and correction system and method based on dynamic calibration

The multimodal intelligent dictation and correction system based on dynamic calibration achieves dynamic calibration of handwriting through feature segmentation, layout analysis, and region mapping comparison. This improves the accuracy of recognizing illegible handwriting and stroke overlap, ensuring the accuracy and reliability of the correction process.

CN121236768AActive Publication Date: 2025-12-30BEIJING CETEN EDUCATION TECH GRP CO LTD

Patent Information

Application Number
CN202511758536.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2025-12-30
Estimated Expiration
2045-11-27

AI Technical Summary

Technical Problem

Existing technologies fail to effectively consider individual differences in students' handwriting habits, resulting in low accuracy in recognizing complex situations such as illegible handwriting and overlapping strokes. Furthermore, they fail to implement a context-based replacement mechanism, making it impossible to accurately handle ambiguous writing content.

Method used

A multimodal intelligent dictation system is adopted, which uses a dynamic calibration-based method and extraction technology, including: acquisition module, feature segmentation module, layout analysis module, recognition module and calibration module, to construct feature observation group, perform local box selection and region mapping comparison, and realize dynamic calibration of handwriting.

Benefits of technology

It improves the accuracy of recognizing complex handwriting such as illegible handwriting and overlapping strokes, and enhances the accuracy and reliability of the correction process through a context restoration replacement mechanism, thus solving the shortcomings of traditional methods in complex handwriting recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121236768A_ABST
    Figure CN121236768A_ABST
Patent Text Reader

Abstract

The invention relates to the field of character and image recognition, in particular to a multi-mode intelligent dictation and correction system and method based on dynamic calibration, and the system comprises an acquisition module for obtaining a writing test image based on voice feedback and carrying out character recognition, a feature splitting module for constructing a feature observation group and labeling a structural unit, the layout analysis module selects a recombination structure by sliding a window frame and sets a fuzzy identification tag based on the prior probability of spatial distribution features, and the identification module compares contour features through region mapping to determine difference types. And the calibration module realizes context restoration and dynamic calibration by removing fuzzy blocks, performing mapping comparison with a sample database and replacing contour features. According to the method, the recognition accuracy of complex handwriting such as scribbling and stroke adhesion is improved, and on this basis, the local image content is calibrated through a replacement mechanism based on context restoration, so that the accuracy and reliability of the correction process are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of text image recognition, and in particular to a multimodal intelligent dictation and correction system and method based on dynamic calibration. Background Technology

[0002] In current language teaching and learning assessment practices, dictation, as a fundamental and important skill training method, is widely used in classroom teaching and after-class exercises. Traditional grading mainly relies on manual completion by teachers, which suffers from problems such as low grading efficiency, long feedback cycles, and difficulty in providing personalized guidance. Currently, assisted grading systems based on optical character recognition or speech recognition have emerged.

[0003] For example, Chinese Patent Publication No. CN113486786A discloses an automatic homework correction system, which relates to the field of homework correction technology. The system includes the following steps: (1) Teacher-side setting operation: The teacher uploads the homework image and sets various parameters. Uploading the homework image mainly refers to taking pictures of the homework page by page and uploading them to the system. After receiving the uploaded homework sample, the system enters the parameter setting module, where the teacher sets the specific parameters; (2) Student-side homework upload operation: Students upload the completed homework image. They only need to upload clear, complete images with neat and upright text according to the page number of the current homework; (3) System homework correction operation: The system corrects the homework images uploaded by students by comparing them with the homework sample. This automatic homework correction system adopts electronic and image processing, which allows teachers to correct homework through mobile terminals. This method solves the problem of students answering questions on computers and avoids the drawbacks of paperless homework.

[0004] However, the following problems still exist in the existing technology: 1. Existing technologies do not take into account individual differences in students' handwriting habits and cannot dynamically calibrate according to writing characteristics, resulting in low recognition accuracy for complex situations such as illegible handwriting and strokes sticking together. 2. In the existing technology, no replacement mechanism based on context restoration is considered, and dynamic calibration of ambiguous written content cannot be achieved. Summary of the Invention

[0005] To address this, the present invention provides a multimodal intelligent dictation and correction system and method based on dynamic calibration, which overcomes the problems in the prior art that do not consider individual differences in students' handwriting habits, cannot perform dynamic calibration according to writing characteristics, resulting in low recognition accuracy for complex situations such as illegible handwriting and strokes sticking together, and do not consider a replacement mechanism based on context restoration, thus failing to achieve dynamic calibration for fuzzy writing content.

[0006] To achieve the above objectives, the present invention provides a multimodal intelligent dictation and correction system based on dynamic calibration, comprising: The acquisition module is used to acquire the pen image data based on the user's output voice feedback, and to perform character recognition on the pen image data. A feature segmentation module, which is connected to the acquisition module, is used to construct several feature observation groups based on the text recognition results of the pen image data and to label several structural units in the feature observation groups. The layout analysis module, which is connected to the feature splitting module, is used to construct a sliding window to locally select structural units in the layout observation group in order to identify several recombination structures with a tendency to combine. Based on the prior probability of the spatial distribution features of the recombination structure in the corpus database, a fuzzy identification label is set for the recombination structure. The recognition module, which is connected to the layout analysis module, is used to read the contour features in the reconstructed structure corresponding to the fuzzy recognition label, divide the local image in the reconstructed structure corresponding to the fuzzy recognition label into regions, compare the contour features in each local image with the corpus in the sample database, and determine the difference type of the local image based on the region mapping comparison result. The calibration module, connected to the recognition module, is used to remove unmatched regions in local images of local difference types, compare the remaining regions with the corpus in the sample database for region mapping, extract a predetermined number of corpus data based on sample similarity to replace the contour features in the local images one by one to determine the calibration contour features, and identify the calibration contour features.

[0007] Furthermore, the feature segmentation module is used to construct several feature observation groups based on adjacent contour features in the pen image data, and the labeled structural units in the feature observation groups include, Used to obtain character recognition results to identify unrecognized local images; Used to identify horizontally adjacent local images as feature observation groups; Used to perform pixel clustering on local images, and the resulting pixel clusters are labeled as structural units.

[0008] Furthermore, the layout analysis module performs local bounding selection on the structural units in the layout observation group to identify several recombination structures with a tendency to combine, including, Used to construct a sliding window based on the size features of the recognized characters; Used to move the sliding window horizontally to select different areas of the feature observation group and determine several structural units within the selected area as the reorganized structure.

[0009] Furthermore, the layout analysis module is used to identify several recombination structures with a tendency to combine, and to set fuzzy identification labels for the recombination structures based on the prior probability of the spatial distribution features corresponding to the recombination structures in the corpus database, including... Determine the spatial distribution characteristics of several recombinant structures, and determine the prior probability of the spatial distribution characteristics in the corpus database; If the prior probability corresponding to the recombined structure is greater than or equal to the predetermined prior probability threshold, then no fuzzy recognition label is set for the recombined structure. If the prior probability corresponding to the recombined structure is less than the predetermined prior probability threshold, then a fuzzy identification label is set for the recombined structure.

[0010] Furthermore, the layout analysis module is also used to determine the spatial distribution characteristics of the reorganized structure, including, Used to divide the sliding window into several spatial regions and determine the mapping relationship between the spatial regions and the serial numbers; Used to determine the spatial region where each structural unit is located in the reorganized structure, and to determine the sequence numbering based on the mapping relationship; This is used to determine the spatial distribution characteristics of the sequence number arrangement.

[0011] Furthermore, the identification module determines the type of difference based on the region mapping comparison results, including: Perform region mapping comparison to determine whether each region matches the corpus in the sample database; If the local matching condition is met, it is determined to be a local difference type; If the local matching condition is not met, it is determined to be a non-local difference class; The local matching condition is that there exists a single region that does not match, while all remaining regions match.

[0012] Furthermore, the identification module is used for region mapping comparison, including: The contour features are divided into several regions, and the corpus in the sample database is divided into several corresponding corpus regions. The similarity between the region and the corresponding corpus region is compared to determine the contour similarity. If the contour similarity of the regions is greater than a preset threshold, then the regions are determined to match.

[0013] Furthermore, the calibration module is used to perform region mapping comparison between the remaining region and the corpus in the sample database, and extract a predetermined number of corpora based on sample similarity to replace the contour features in the local image one by one, including, The remaining regions are compared with the corresponding regions in the corpus of the sample database to obtain the sample similarity. The corpora are sorted from largest to smallest based on sample similarity, and a predetermined number of corpora are selected as candidate corpora in descending order. The candidate corpus is used sequentially to replace the contour features in the local image.

[0014] Furthermore, the calibration module is used to determine calibration profile features, including: After the replacement of candidate texts is completed, the texts are identified, and contextual understanding is performed based on the texts to determine the semantic matching degree. The candidate corpus with the highest semantic matching degree is selected as the calibration contour feature.

[0015] On the other hand, a method for dynamic calibration-based multimodal intelligent dictation and correction applied to a dynamic calibration-based multimodal intelligent dictation and correction system is also provided, including: Acquire pen image data based on user input voice feedback, and perform text recognition on the pen image data; Based on the character recognition results of the pen image data, several feature observation groups are constructed, and several structural units in the feature observation groups are labeled. A sliding window is constructed to locally select structural units in the layout observation group to identify several recombination structures with a tendency to combine. Fuzzy identification labels are set for the recombination structures based on the prior probability of the spatial distribution features of the recombination structures in the corpus database. Read the contour features in the reconstructed structure corresponding to the fuzzy recognition label, divide the local image in the reconstructed structure corresponding to the fuzzy recognition label into regions, compare the contour features in each local image with the corpus in the sample database, and determine the difference type based on the region mapping comparison result. The corpus samples required for context reconstruction are determined based on the difference type, including: Blurred areas in the local image are removed, and the remaining areas are compared with the corpus in the sample database. Based on the sample similarity, a predetermined number of corpus data are extracted to replace the contour features in the local image one by one to determine the calibration contour features. The calibration contour features are then identified.

[0016] Compared with existing technologies, this invention acquires handwriting images based on voice feedback through an acquisition module and performs character recognition; a feature segmentation module constructs feature observation groups and labels structural units; a layout analysis module reconstructs structures by selecting and recombining them through a sliding window and sets fuzzy recognition labels based on prior probabilities of spatial distribution features; a recognition module compares contour features through region mapping to determine the type of difference; and a calibration module achieves context restoration and dynamic calibration by removing fuzzy blocks, comparing them with a sample database, and replacing contour features. This invention improves the recognition accuracy for complex handwriting such as illegible handwriting and overlapping strokes, and further enhances the accuracy and reliability of the correction process by calibrating local image content through a context-based replacement mechanism.

[0017] In particular, this invention considers a feature decomposition mechanism based on structural units and feature observation groups. In practice, the continuity and individual differences in handwriting pose challenges to feature extraction, and traditional methods struggle to effectively separate and represent complex strokes. This invention constructs feature observation groups based on adjacent contour features in the stroke image data through a feature decomposition module, and processes local images using a pixel clustering algorithm, labeling the resulting pixel clusters as structural units. This achieves a systematic decomposition and standardized representation of complex handwriting features. This mechanism not only accurately captures the detailed features of strokes but also maintains the spatial correlation between features, laying the foundation for subsequent layout analysis and recognition calibration, and improving the system's adaptability to various writing styles.

[0018] In particular, this invention considers a layout analysis mechanism based on a sliding window and prior probabilities. In reality, the spatial distribution characteristics of handwritten text are complex and varied, making it difficult for traditional methods to effectively utilize layout information for recognition. This invention constructs a sliding window through a layout analysis module to locally select structural units in the feature observation group. After determining the recombined structure, fuzzy recognition labels are set based on its prior probabilities in the corpus database, thereby achieving quantitative analysis and intelligent recognition of the spatial distribution characteristics of text.

[0019] In particular, this invention considers a dynamic calibration mechanism based on region mapping comparison and blurred block processing. In practice, handwritten images often suffer from localized smudges, blurriness, or missing features, affecting recognition accuracy. Traditional correction methods often struggle to achieve accurate repair while maintaining semantic coherence. This invention divides contour features into several regions through a recognition module, performing region mapping comparison with corpora in a sample database to determine the type of difference. The calibration module then removes blurred blocks, performs multiple rounds of region mapping comparison based on the remaining regions, and replaces contour features one by one using candidate corpora ranked by sample similarity, achieving accurate calibration of locally blurred writing. This hierarchical processing mechanism ensures both the accuracy of the correction process and semantic coherence, solving the problem of recognition errors caused by missing or interfered local features.

[0020] In particular, this invention considers a contour feature replacement mechanism based on sample similarity and semantic matching. In practice, a single feature comparison is insufficient to guarantee the accuracy of the replacement result. This invention uses a calibration module to perform region mapping comparison between the corpus in the sample database and the remaining regions to obtain sample similarity. Candidate corpora are selected for contour feature replacement based on similarity ranking. Then, the calibration contour features are determined through semantic matching calculation. This constructs a synergistic mechanism between character shape comparison and semantic analysis, ensuring the reliability and accuracy of the correction result. Attached Figure Description

[0021] Figure 1This is a schematic diagram of the structural connection of the multimodal intelligent dictation and correction system based on dynamic calibration according to an embodiment of the present invention; Figure 2 A logic block diagram for setting fuzzy identification tags according to an embodiment of the invention; Figure 3 A logic block diagram for determining the difference type in an embodiment of the invention; Figure 4 This is a logic block diagram for determining region matching in an embodiment of the invention. Detailed Implementation

[0022] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0023] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0024] It should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the term "connection" should be interpreted broadly. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can be a connection within two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0025] Please see Figure 1 As shown, Figure 1 This is a schematic diagram of the structural connection of a multimodal intelligent dictation and correction system based on dynamic calibration according to an embodiment of the present invention. The multimodal intelligent dictation and correction system based on dynamic calibration of the present invention includes: The acquisition module is used to acquire the pen image data based on the user's output voice feedback, and to perform character recognition on the pen image data. A feature segmentation module, which is connected to the acquisition module, is used to construct several feature observation groups based on the text recognition results of the pen image data and to label several structural units in the feature observation groups. The layout analysis module, which is connected to the feature splitting module, is used to construct a sliding window to locally select structural units in the layout observation group in order to identify several recombination structures with a tendency to combine. Based on the prior probability of the spatial distribution features of the recombination structure in the corpus database, a fuzzy identification label is set for the recombination structure. The recognition module, which is connected to the layout analysis module, is used to read the contour features in the reconstructed structure corresponding to the fuzzy recognition label, divide the local image in the reconstructed structure corresponding to the fuzzy recognition label into regions, compare the contour features in each local image with the corpus in the sample database, and determine the difference type of the local image based on the region mapping comparison result. The calibration module, connected to the recognition module, is used to remove unmatched regions in local images of local difference types, compare the remaining regions with the corpus in the sample database for region mapping, extract a predetermined number of corpus data based on sample similarity to replace the contour features in the local images one by one to determine the calibration contour features, and identify the calibration contour features.

[0026] Specifically, the output speech can be dictation speech, and students can write based on the output speech. The pen image data containing the students' handwriting can be collected through a user application.

[0027] There are no restrictions on the structure of the acquisition module, feature segmentation module, layout analysis module, recognition module, and calibration module. They can be composed of logic components, including field-programmable processors, computers, or microprocessors in computers.

[0028] Specifically, there are no restrictions on the method of text recognition for image data. Existing open-source text recognition models can be used, or an image processing model that can recognize text can be trained independently. This will not be elaborated further.

[0029] Specifically, the feature segmentation module is used to construct several feature observation groups based on adjacent contour features in the pen image data, and the labeled structural units in the feature observation groups include, Used to obtain character recognition results to identify unrecognized local images; Used to identify horizontally adjacent local images as feature observation groups; Used to perform pixel clustering on local images, and the resulting pixel clusters are labeled as structural units.

[0030] In practice, the local image is an image containing unrecognized characters. The purpose of pixel clustering is to identify the various parts of the contour in the local image. For example, a single stroke of a Chinese character or a continuously written radical will be classified as a cluster, and the corresponding single cluster will be labeled as a single structural unit.

[0031] Specifically, the layout analysis module performs local bounding selection on the structural units in the layout observation group to identify several recombination structures with a tendency to combine, including, Used to construct a sliding window based on the size features of the recognized characters; Used to move the sliding window horizontally to select different areas of the feature observation group and determine several structural units within the selected area as the reorganized structure.

[0032] Specifically, the layout analysis module is configured to construct a sliding window based on the average size characteristics of the recognized characters. The average height and average width are obtained by statistical analysis of the height and width of the recognized characters. Based on this, the size of the sliding window is determined by enlarging it according to a preset magnification ratio.

[0033] Understandably, the preset magnification ratio should not be too large to avoid the sliding window covering too many irrelevant structural units, introducing unnecessary feature interference, and affecting the effective identification and analysis of the target reconstructed structure. In practice, the preset magnification ratio is typically set within the range of [115%, 130%]. Please see Figure 2 As shown, Figure 2 This is a logic block diagram for setting fuzzy identification labels according to an embodiment of the invention. The layout analysis module is used to determine several recombination structures with a tendency to combine. Setting fuzzy identification labels for the recombination structures based on the prior probability of the spatial distribution features corresponding to the recombination structures in the corpus database includes... Determine the spatial distribution characteristics of several recombinant structures, and determine the prior probability of the spatial distribution characteristics in the corpus database; If the prior probability corresponding to the recombined structure is greater than or equal to the predetermined prior probability threshold, then no fuzzy recognition label is set for the recombined structure. If the prior probability corresponding to the recombined structure is less than the predetermined prior probability threshold, then a fuzzy identification label is set for the recombined structure.

[0034] Specifically, the predetermined prior probability threshold is set by those skilled in the art based on the statistical characteristics of the occurrence probability of commonly used characters. In practice, the spatial distribution characteristics of commonly used characters in the modern Chinese character list are determined, and the average occurrence probability of each spatial distribution characteristic in the Chinese dictionary is determined. The product of the average occurrence probability and a correction coefficient ranging from 0.8 to 0.95 is used as the prior probability threshold. This setting method ensures that the threshold effectively covers most common writing structures while retaining appropriate adjustment space through the correction coefficient, allowing the system to fine-tune the recognition sensitivity according to the needs of actual application scenarios. In practice, the correction coefficient is preferably 0.85.

[0035] Specifically, the layout analysis module is also used to determine the spatial distribution characteristics of the reorganized structure, including, Used to divide the sliding window into several spatial regions and determine the mapping relationship between the spatial regions and the serial numbers; Used to determine the spatial region where each structural unit is located in the reorganized structure, and to determine the sequence numbering based on the mapping relationship; This is used to determine the spatial distribution characteristics of the sequence number arrangement.

[0036] In implementation, the sliding window is divided into grids to obtain several spatial regions, each of which corresponds to a serial number. Preferably, the sliding window is divided into 16 spatial regions, each corresponding to a serial number, which can be the number 1-16.

[0037] During implementation, when structural units exist in a spatial region, the corresponding serial number of the spatial region is recorded; when they do not exist, the serial number is not recorded. The serial numbers are arranged in ascending order to obtain the serial number arrangement, which will not be elaborated further.

[0038] It is understandable that spatial distribution features reflect the distribution of various parts in Chinese characters. For blocks with messy layouts or missing features, the spatial distribution features will change and deviate from the norm, thus making the prior probability of the spatial distribution features in the corpus database low.

[0039] Please see Figure 3 As shown, Figure 3 This is a logic block diagram illustrating the determination of the difference type according to an embodiment of the invention. The identification module determines the difference type based on the region mapping comparison result, including: Perform region mapping comparison to determine whether each region matches the corpus in the sample database; If the local matching condition is met, it is determined to be a local difference type; If the local matching condition is not met, it is determined to be a non-local difference class; The local matching condition is that there exists a single region that does not match, while all remaining regions match.

[0040] Please see Figure 4 As shown, Figure 4 This is a logic block diagram for determining region matching according to an embodiment of the invention. The identification module is used to perform region mapping comparison, including: The contour features are divided into several regions, and the corpus in the sample database is divided into several corresponding corpus regions. The similarity between the region and the corresponding corpus region is compared to determine the contour similarity. If the contour similarity of the regions is greater than a preset threshold, then the regions are determined to match.

[0041] Specifically, the purpose of setting a preset threshold is to characterize the blocks that are similar to the contour features and the corpus. The preset threshold is pre-set. Several correctly identified contour features are manually selected, and the contour similarity between the contour features and the corresponding corpus is determined. The mean contour similarity is calculated, and the 85th percentile of the mean contour similarity is set as the preset threshold.

[0042] The corpus in the sample database consists of templates for single Chinese characters.

[0043] Specifically, the calibration module is used to perform region mapping comparison between the remaining region and the corpus in the sample database, and extract a predetermined number of corpora based on sample similarity to replace the contour features in the local image one by one, including, The remaining regions are compared with the corresponding regions in the corpus of the sample database to obtain the sample similarity. The corpora are sorted from largest to smallest based on sample similarity, and a predetermined number of corpora are selected as candidate corpora in descending order. The candidate corpus is used sequentially to replace the contour features in the local image.

[0044] Specifically, the preset number should not be too large in order to reduce the number of times context understanding is required later. In practice, the preset number is selected within the range [3, 5], preferably 4.

[0045] Specifically, sample similarity is the contour similarity when the remaining region is compared with the corresponding region in the corpus.

[0046] Specifically, the calibration module is used to determine calibration profile features, including: After the replacement of candidate texts is completed, the texts are identified, and contextual understanding is performed based on the texts to determine the semantic matching degree. The candidate corpus with the highest semantic matching degree is selected as the calibration contour feature.

[0047] Specifically, in context understanding, the semantic matching degree is calculated based on the overall semantic match between the text corresponding to the replaced candidate text and the already recognized text. There are no restrictions on the method for calculating semantic matching degree. Word segmentation can be performed to determine a number of keywords. After the keywords are vectorized, the cosine similarity between the keywords is calculated, and the mean cosine similarity is determined as the semantic matching degree.

[0048] When vectorizing the model, existing open-source models can be used, such as the BERT context-aware model, which can obtain the vector representation of keywords in the context. Of course, other forms can also be used, which will not be elaborated here.

[0049] Specifically, after recognizing all the text, the system can correct the correspondence between the text and the output speech, which will not be elaborated further.

[0050] This embodiment also provides a dynamic calibration-based multimodal intelligent dictation and correction method applied to a dynamic calibration-based multimodal intelligent dictation and correction system, including: Step S1: Obtain the pen image data based on the user's output voice feedback, and perform text recognition on the pen image data; Step S2: Construct several feature observation groups based on the text recognition results of the pen image data, and label several structural units in the feature observation groups; Step S3: Construct a sliding window to locally select structural units in the layout observation group to identify several recombination structures with a tendency to combine. Based on the prior probability of the spatial distribution features of the recombination structures in the corpus database, set fuzzy recognition labels for the recombination structures. Step S4: Read the contour features in the reconstructed structure corresponding to the fuzzy recognition label, divide the local image in the reconstructed structure corresponding to the fuzzy recognition label into regions, compare the contour features in each local image with the corpus in the sample database for region mapping, and determine the difference type of the local image based on the region mapping comparison result. Step S5: Remove unmatched regions in local images of local difference type, compare the remaining regions with the corpus in the sample database for region mapping, extract a predetermined number of corpus based on sample similarity, replace the contour features in the local images one by one to determine the calibration contour features, and identify the calibration contour features.

[0051] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.

Claims

1. A multi-modal intelligent dictation and correction system based on dynamic calibration, characterized in that, Comprise: A collection module is used to obtain the user's handwriting image data based on the output voice feedback, and to perform character recognition on the handwriting image data; A feature splitting module connected with the collection module is used to construct a plurality of feature observation groups based on the character recognition results of the handwriting image data, and to label a plurality of structural units in the feature observation groups; A layout analysis module connected with the feature splitting module is used to construct a sliding window to locally frame the structural units in the layout observation groups to determine a plurality of reorganization structures with combination tendency, and to set a fuzzy recognition label for the reorganization structures based on the prior probability of the corresponding spatial distribution features of the reorganization structures in the corpus database; An identification module connected with the layout analysis module is used to read the contour features in the reorganization structures corresponding to the fuzzy recognition label, to divide the local images in the reorganization structures corresponding to the fuzzy recognition label into regions, to compare the contour features in each local image with the corpus in the sample database, and to determine the difference type of the local image based on the region mapping comparison result; A calibration module connected with the identification module is used to remove the unmatched regions in the local image of the local difference type, to compare the remaining regions with the corpus in the sample database, to extract a predetermined number of corpora one by one to replace the contour features in the local image based on the sample similarity, and to determine the calibrated contour features and to identify the calibrated contour features.

2. The multi-modal intelligent dictation and grading system based on dynamic calibration as claimed in claim 1, wherein, The feature splitting module is used to construct a plurality of feature observation groups based on the adjacent contour features in the handwriting image data, and to label a plurality of structural units in the feature observation groups, including, To obtain the character recognition result to determine the un-recognized local image; To determine the horizontally adjacent local images as the feature observation groups; To perform pixel clustering on the local images, and to label the obtained pixel clustering clusters as structural units.

3. The multi-modal smart dictation and grading system based on dynamic calibration as claimed in claim 1, wherein, The layout analysis module locally frames the structural units in the layout observation groups to determine a plurality of reorganization structures with combination tendency, including, To construct a sliding window based on the size features of the recognized characters; To move the sliding window horizontally to frame different regions of the feature observation groups, and to determine a plurality of structural units in the framed regions as reorganization structures.

4. The multi-modal smart dictation and grading system based on dynamic calibration as claimed in claim 3, wherein, The layout analysis module is used to determine a plurality of reorganization structures with combination tendency, and to set a fuzzy recognition label for the reorganization structures based on the prior probability of the corresponding spatial distribution features of the reorganization structures in the corpus database, including, To determine the spatial distribution features of a plurality of reorganization structures, and to determine the prior probability of the spatial distribution features in the corpus database; If the prior probability corresponding to the reorganization structure is greater than or equal to a predetermined prior probability threshold, no fuzzy recognition label is set for the reorganization structure; If the prior probability corresponding to the reorganization structure is less than the predetermined prior probability threshold, a fuzzy recognition label is set for the reorganization structure.

5. The multi-modal smart dictation and grading system based on dynamic calibration as claimed in claim 4, wherein, The layout analysis module is also used to determine the spatial distribution features of the reorganization structure, including, To divide the sliding window into a plurality of spatial regions, and to determine the mapping relationship between the spatial regions and the serial numbers; To determine the spatial region where each structural unit in the reorganization structure is located, and to determine the serial number arrangement based on the mapping relationship; To determine the serial number arrangement as the spatial distribution features. ​ 6. The multi-modal smart dictation and grading system based on dynamic calibration as claimed in claim 1, wherein, The recognition module determines a difference type based on a region mapping comparison result, including, The region mapping comparison determines whether each region matches the corpus in the sample database; If the local matching condition is met, it is determined as a local difference type; If the local matching condition is not met, it is determined as a non-local difference type; The local matching condition is that there is a single region that does not match and the remaining regions all match.

7. The multi-modal smart dictation and grading system based on dynamic calibration as claimed in claim 6, wherein, The recognition module is used to perform region mapping comparison, including, The contour features are divided into several regions, and the corpus in the sample database is divided into corresponding several corpus regions; The region similarity comparison is performed between the region and the corresponding corpus region to determine the contour similarity; If the contour similarity of the region is greater than a preset threshold, it is determined that the region matches.

8. The multi-modal smart dictation and grading system based on dynamic calibration as claimed in claim 1, wherein, The calibration module is used to perform region mapping comparison between the remaining regions and the corpus in the sample database, and based on the sample similarity, a predetermined number of corpora are extracted one by one to replace the contour features in the local image, including, The region mapping comparison is performed between the remaining regions and the corresponding regions of the corpus in the sample database to obtain the sample similarity; The corpora are sorted from large to small according to the sample similarity, and a preset number of corpora are selected in descending order as candidate corpora; The candidate corpora are used in turn to replace the contour features in the local image.

9. The multi-modal smart dictation and grading system based on dynamic calibration as claimed in claim 1, wherein, The calibration module is used to determine the calibrated contour features, including, The text corresponding to the candidate corpus after the replacement is determined, the context understanding is performed based on the text, and the semantic matching degree is determined; The candidate corpus with the highest semantic matching degree is determined as the calibrated contour feature.

10. A method applied to the multi-modal intelligent dictation and correction system based on dynamic calibration according to any one of claims 1-9, characterized in that, Including, Obtaining the user's test image data based on the output voice feedback, and performing text recognition on the test image data; Based on the text recognition result of the test image data, a plurality of feature observation groups are constructed, and a plurality of structure units in the feature observation group are labeled; A sliding window is constructed to locally frame the structure units in the layout observation group to determine a plurality of reorganized structures with combination tendency, and a fuzzy recognition label is set for the reorganized structure based on the prior probability of the spatial distribution characteristics of the reorganized structure in the corpus database; The contour features in the reorganized structure corresponding to the fuzzy recognition label are read, the local image in the reorganized structure corresponding to the fuzzy recognition label is regionally divided, the contour features in each local image are regionally mapped and compared with the corpus in the sample database, and the difference type of the local image is determined based on the region mapping comparison result; The unmatched regions in the local image of the local difference type are removed, the remaining regions are regionally mapped and compared with the corpus in the sample database, a predetermined number of corpora are extracted based on the sample similarity, and the contour features in the local image are replaced one by one to determine the calibrated contour features, and the calibrated contour features are recognized.

Citation Information

Patent Citations

  • Automatic homework correcting system

    CN113486786A

  • Archive digitization management method and system and storage medium

    CN120014650A

  • Intelligent paper marking content detection and identification method and system based on deep learning

    CN120126146A

  • Method and system for correcting answer content based on image recognition

    CN120599636A

  • Intelligent paper marking system based on image cutting and recognition

    CN120997851A

Cited By

  • Incremental training method and system for accounting document generation

    CN121725494A