Training method of grammar correction model and grammar correction method
By directly identifying and correcting text information in images from multimodal information, a grammar correction model was trained to solve the cascading error problem caused by text recognition errors, achieving efficient grammar correction results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING FLYING ELEPHANT PLANET TECH CO LTD
- Filing Date
- 2026-01-15
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies suffer from poor correction results due to cascading errors in grammar judgment caused by text recognition errors in multimodal information to be corrected.
By acquiring text information from images, identifying target text, determining target text modification, and training a grammar correction model until the training stops, cascading errors in the text conversion process are avoided.
It enables precise correction of grammatical errors in text information directly within multimodal information, improving the accuracy and efficiency of correction.
Smart Images

Figure CN121936459A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a training method and a syntax correction method for a syntax correction model. This application also relates to a training apparatus and a syntax correction apparatus for a syntax correction model, a computing device, a computer-readable storage medium, and a computer program product. Background Technology
[0002] With the continuous development of computer technology, educational aids, intelligent writing assistants, and other technologies are placing increasingly higher demands on the grammatical accuracy of texts.
[0003] For multimodal information to be corrected, text recognition technology is typically used to convert the multimodal information into text information of the corresponding text format, and then use this text information as the basis for grammatical judgment to correct the text information. However, this method is prone to recognition errors during the text recognition process. Therefore, using the converted text information as the basis for grammatical judgment can easily lead to cascading errors in the grammatical judgment process, resulting in poor correction results.
[0004] Therefore, how to accurately correct the multimodal information to be corrected has become an urgent problem to be solved. Summary of the Invention
[0005] In view of this, embodiments of this application provide a training method for a syntax correction model and a syntax correction method. This application also relates to a training apparatus and a syntax correction apparatus for a syntax correction model, a computing device, a computer-readable storage medium, and a computer program product, to solve the aforementioned problems existing in the prior art.
[0006] According to a first aspect of the embodiments of this application, a method for training a syntax correction model is provided, comprising: Obtain an image to be processed, wherein the image to be processed includes at least one text message; Identify the image to be processed, determine the target recognition text corresponding to each text information, and determine the target modification text corresponding to each target recognition text; The grammar correction model is trained based on the target recognition text and the target modification text corresponding to each target recognition text until the training stopping condition of the grammar correction model is met.
[0007] According to a second aspect of the embodiments of this application, a syntax correction method is provided, including: Obtain an image to be processed, wherein the image to be processed includes at least one initial text to be corrected; The image to be processed is input into the syntax correction model to obtain at least one target text to be corrected and the target corrected text corresponding to each target text to be corrected. The syntax correction model is trained based on the syntax correction model training method described in the first aspect of the present application.
[0008] According to a third aspect of the embodiments of this application, a training apparatus for a syntax correction model is provided, comprising: The acquisition unit is configured to acquire an image to be processed, wherein the image to be processed includes at least one text information; The recognition unit is configured to recognize the image to be processed, determine the target recognition text corresponding to each text information, and determine the target modification text corresponding to each target recognition text; The training unit is configured to train a grammar correction model based on each target recognition text and the target modification text corresponding to each target recognition text, until the training stopping condition of the grammar correction model is met.
[0009] According to a fourth aspect of the embodiments of this application, a syntax correction apparatus is provided, comprising: The acquisition unit is configured to acquire an image to be processed, wherein the image to be processed includes at least one initial text to be corrected; The processing unit is configured to input the image to be processed into the syntax correction model to obtain at least one target text to be corrected and the target corrected text corresponding to each target text to be corrected. The syntax correction model is trained based on the syntax correction model training method described in the first aspect of the present application.
[0010] According to a fifth aspect of the embodiments of this application, a computing device is provided, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, they implement the training method and steps of the syntax correction method described above.
[0011] According to a sixth aspect of the embodiments of this application, a computer-readable storage medium is provided, which stores a computer program / instructions that, when executed by a processor, implement the training method and steps of the syntax correction method for the above-described syntax correction model.
[0012] According to a seventh aspect of the present application, a computer program product is provided, including a computer program / instruction that, when executed by a processor, implements the training method and steps of the syntax correction method for the above-described syntax correction model.
[0013] According to the training method of the grammar correction model involved in the above embodiments of this application, an image containing text information is obtained to obtain text information in image format. Based on this, the image to be processed is identified to obtain target recognition text corresponding to the text information in the image. The target recognition text is checked to obtain target modified text that does not contain grammatical errors. The grammar correction model is trained based on the target recognition text and the corresponding target modified text, so that after receiving the image to be processed, the grammar correction model can autonomously learn the text information containing grammatical errors in the image and correct the grammatically incorrect text information. According to the trained grammar correction model involved in this application, the method of obtaining the corresponding target modified text based on the image to be processed enables the grammar correction model to not only accurately identify text from images but also to perform grammatical error correction on the text in the image. It eliminates the need to convert the text information in image format into text and then use the converted text to determine whether the text contains grammatical errors, thus avoiding cascading errors in the process of sentence recognition. Attached Figure Description
[0014] Figure 1 A flowchart is shown for a training method of a syntax correction model according to an embodiment of this application; Figure 2 A flowchart illustrating a syntax correction method according to an embodiment of this application is shown. Figure 3 This paper shows a schematic diagram of the structure of a training device for a syntax correction model according to an embodiment of this application; Figure 4 This invention provides a schematic diagram of the structure of a syntax correction device according to an embodiment of the present application. Figure 5 A structural block diagram of a computing device according to an embodiment of this application is shown. Detailed Implementation
[0015] Many specific details are set forth in the following description to provide a full understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this application; therefore, this application is not limited to the specific embodiments disclosed below.
[0016] The terminology used in one or more embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the scope of one or more embodiments of this application. The singular forms “a,” “the,” and “the” used in one or more embodiments of this application and in the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” used in one or more embodiments of this application refers to and includes any or all possible combinations of one or more associated listed items.
[0017] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this application, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0018] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0019] The training method and grammar correction method of the grammar correction model involved in this application are applied to scenarios where the grammar of text existing in a multimodal form is corrected. For example, in an educational setting, to facilitate the grading of student essays, student essays are stored as images. During subsequent grading, for images containing student essays, in order to accurately and quickly determine whether there are grammatical errors in the student essays, the training method and grammar correction method of the grammar correction model involved in this application can be used to correct the grammatical errors in the student essays in the images.
[0020] With the continuous development of computer technology, in educational settings, text in multimodal form is typically converted into a text format using text recognition technology, and subsequent processing is based on this converted text. For example, in grammar correction, grammatical errors are usually analyzed and corrected within the converted text format.
[0021] However, when using the methods described above, if the text identified from the image by the character recognition technology is inaccurate, continuing to analyze and modify the identified text can easily lead to cascading errors.
[0022] In view of this, this application provides a training method for a syntax correction model and a syntax correction method. This application also relates to a training device for a syntax correction model and a syntax correction device, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0023] Figure 1 A flowchart illustrating a training method for a syntax correction model according to an embodiment of this application is shown, specifically including the following steps 102-106: Step 102: Obtain the image to be processed, wherein the image to be processed includes at least one text information.
[0024] The image to be processed can be understood as an image containing text information, which serves as input data for the grammar correction model during training. The image contains clearly legible text. For example, the text in the image could be a student's essay, and the image can be obtained by taking a photograph or scanning it. It should be understood that this application does not limit the source of the image to be processed.
[0025] Text information can be understood as text content that requires grammatical analysis. For example, text information could be student essays or answers. The text information contained in the image to be processed is stored in image format. Given this, in order for the grammar correction model to learn the text information contained in the image during training, subsequent operations train the grammar correction model to identify the target text corresponding to the text information, thereby enabling the grammar correction model to correct grammatical errors in the text information in the image. For ease of understanding, this application will be explained below.
[0026] In one specific embodiment of this application, text information in the form of an image is acquired through photography or scanning. In this application, a grammar correction model is subsequently trained based on the acquired image to be processed. This enables the trained grammar correction model to directly extract the target recognition text and the target modified text from the image to be processed, avoiding the cascading errors that occur during the process of converting the image to text and then correcting the text to obtain the corrected text.
[0027] Furthermore, to improve the accuracy of the training of the grammar correction model, this application explains the training process of the grammar correction model in the following manner.
[0028] Step 104: Identify the image to be processed, determine the target recognition text corresponding to each text information, and determine the target modification text corresponding to each target recognition text.
[0029] In this context, target-recognized text can be understood as the text corresponding to the text information in the image to be processed, determined based on the image itself. It's important to understand that the target-recognized text must be a perfect match to the text information in the image to be processed.
[0030] Target-corrected text can be understood as text that conforms to the standard syntax format after the target recognition text has been modified according to a set standard syntax format. For example, the standard syntax format may include at least one of the following: a): For inappropriate grammatical formats.
[0031] For example, if there are grammatical errors in the target recognition text, the target correction text can be the text that has been corrected to remove the grammatical errors.
[0032] b): Syntax format for incomplete components.
[0033] For example, if there are grammatical errors with missing components in the target recognition text, the target correction text can be the text that has been corrected to remove the grammatical errors with missing components.
[0034] c): For grammatical formats with incorrect word order.
[0035] For example, if there are grammatical errors such as incorrect word order in the target recognition text, the target correction text can be the text that has been corrected to remove the incorrect word order grammatical errors.
[0036] d): For custom syntax formats.
[0037] For example, if there are custom grammatical errors in the target recognition text, the target correction text can be the text that corrects the custom grammatical errors and then eliminates the custom grammatical errors.
[0038] It should be understood that the target correction text involved in this application may include at least one grammatical error; that is, it can be understood as correcting the target correction text according to at least one standard grammatical format. Furthermore, the custom grammatical format can be the grammatical format corresponding to other grammatical errors besides any of the aforementioned grammatical formats, and this application does not limit the custom grammatical format.
[0039] In one specific embodiment provided in this application, an image to be processed is identified, and all text information contained in the image is determined. For each piece of text information, a corresponding target recognition text is determined. Since grammatical errors exist in each target recognition text, a corresponding target modification text without grammatical errors is determined based on each target recognition text and standard grammatical format.
[0040] Specifically, the image to be processed is identified, and the target text corresponding to each text information is determined, including S1042-S1046: S1042. Identify each text information in the image to be processed, and obtain the initial text to be identified corresponding to each text information.
[0041] The initial text to be recognized can be understood as the initial text identified from the image to be processed using methods such as character recognition technology. It is important to understand that the initial text to be recognized may contain grammatical errors.
[0042] In one specific embodiment provided in this application, clear text information is identified in the original image to be processed, and the identified clear text information is converted into initial text to be recognized. It should be understood that each text information in the image to be processed corresponds to its own identified initial text to be recognized.
[0043] Specifically, identifying each text information in the image to be processed and obtaining the initial text to be identified corresponding to each text information includes: At least one text message is determined in the image to be processed based on a preset end character; Identify each piece of text information and obtain the initial text to be identified for each piece of text information.
[0044] The preset end character can be understood as a punctuation mark used to indicate the end of a text segment. For example, preset end characters can be periods, question marks, exclamation marks, etc. It's important to understand that the number of preset end characters is directly proportional to the recognition accuracy of the initial text to be recognized. The more preset end characters, the more initial text to be recognized can be extracted from the text information.
[0045] In one specific embodiment provided in this application, a preset end character is set, the preset end character is detected in the image to be processed, and at least one corresponding text information is extracted based on the detected preset end character. The extracted at least one text information is recognized to obtain the initial text to be recognized corresponding to each text information.
[0046] According to a specific embodiment provided in this application, the method of determining at least one text information in the image to be processed based on a preset end character improves the convenience of determining text information.
[0047] S1044. Compare each text information with the corresponding initial text to be identified to obtain the comparison results of each initial text to be identified.
[0048] In one specific embodiment provided in this application, in order to further ensure the accuracy of the identified initial text information, the identified initial text information is compared with the text information in the original image to be processed to obtain the comparison result, so as to further improve the accuracy of the initial text information based on the comparison result.
[0049] Specifically, by comparing each piece of text information with the corresponding initial text to be identified, the comparison results of each initial text to be identified are obtained, including: Determine the initial text to be identified for comparison, wherein the initial text to be identified for comparison is any one of the initial texts to be identified; Based on the text attribute information of the initial text to be identified to be compared, the initial text information corresponding to the initial text to be identified to be compared is located in the image to be processed; The initial text information is compared with the initial text to be identified to obtain the comparison result of the initial text to be identified.
[0050] The text attribute information can be understood as information used to locate the position of text information in the image to be processed. For example, based on the text attribute information of the initial text to be identified, the original text information corresponding to the initial text to be identified in the image to be processed is determined.
[0051] In one specific embodiment provided in this application, any one of the initial texts to be identified is determined as the initial text to be compared. Based on the text attribute information of the initial text to be compared, the original text information is located in the image to be processed. The initial text to be compared is compared with its corresponding original text information, for example, character-by-character comparison, punctuation-by-punctuation comparison, etc., to determine the comparison result, so that the compared initial text to be compared remains consistent with the original text information.
[0052] According to a specific implementation provided in this specification, the text attribute information corresponding to the initial text information to be identified is used to accurately locate the original text information in the original image to be processed, thereby improving the accuracy of subsequent comparison results. Then, the original text information located is compared with the initial text to be identified to obtain an accurate comparison result.
[0053] S1046. Based on the comparison results of each initial text to be identified, determine the target text to be identified for each initial text to be identified.
[0054] In one specific embodiment provided in this application, the process involves comparing each initial text to be identified with the text information in the image to be processed, obtaining the comparison result, and determining the target text to be identified based on the comparison result.
[0055] Specifically, based on the comparison results of each initial text to be identified, the target text corresponding to each initial text to be identified is determined, including: Determine the initial text to be identified to be processed, and determine the comparison result corresponding to the initial text to be identified to be processed, wherein the initial text to be identified to be processed is any one of the initial texts to be identified; If the comparison result is consistent, the initial text to be identified is determined as the target text to be identified. If the comparison result is inconsistent, the initial text to be identified is corrected, and the corrected initial text to be identified is determined as the target text to be identified.
[0056] Here, the initial text to be recognized can be understood as any one of multiple initial texts to be recognized obtained after recognizing the image to be processed. It should be understood that in this application, the initial text to be recognized is selected sequentially or randomly from multiple initial texts to be recognized, until each initial text to be recognized is obtained from multiple initial texts.
[0057] The comparison result can be understood as the result of comparing the identified text from the image with the original text in the image to obtain a comparison result regarding the consistency of information. For example, it is based on the comparison result between the text information in the original image to be processed and the initial text to be identified obtained from the image to be processed.
[0058] In one specific embodiment provided in this application, an initial text to be processed is randomly or sequentially determined from a plurality of initial texts to be recognized. The obtained initial text to be processed is compared with the text information in the original image to be processed. If the comparison result is a match, the initial text to be processed is determined as the target text to be recognized. If the comparison result is a mismatch, the initial text to be processed is corrected so that the corrected initial text to be recognized matches the text information in the image to be processed, and thus the corrected initial text to be processed that matches the text information in the image to be processed is determined as the target text to be recognized.
[0059] It should be understood that the text information in the image to be processed involved in this application is stored in the image to be processed in the image format. Under this premise, the text information in the image to be processed is compared with the initial text to be recognized, so as to determine the target recognition text that can finally represent the text format in the text information in the image to be processed.
[0060] According to a specific embodiment provided in this application, the initial text to be identified is compared with the original text information in the original image to be processed, so that the initial text to be identified is corrected based on the original text information in the original image to be processed. Further, if the comparison results are consistent, the initial text to be identified in text format is directly used as the target text to be identified, and the target text to be identified is determined specifically. If the comparison results are inconsistent, the initial text to be identified in text format is corrected to ensure that the corrected text to be identified is consistent with the text information in the original image to be processed, and thus it is used as the target text to be identified. The above method ensures strict consistency between the target text to be identified and the text information in the original image to be processed.
[0061] In one specific embodiment provided in this application, initial text to be recognized is identified from an image to be processed using methods such as character recognition (CR). The identified initial text to be recognized is compared with the original text information in the image to be processed. Based on the comparison result, the target text to be recognized is determined. In this case, for example, if the initial text to be recognized is consistent with the original text information in the image to be processed, then the text to be recognized is taken as the target text to be recognized. If the initial text to be recognized is inconsistent with the original text information in the image to be processed, then the text to be recognized is adjusted according to the original text information in the image to be processed, so that the text to be recognized is consistent with the original text information. It should be understood that the original text information in the image to be processed and each text information in the image to be processed have the same meaning as described above. Furthermore, the process of comparing the text information in the image to be processed with the initial text to be recognized can be a manual comparison or a comparison based on a large-scale language model. Similarly, when the original text information in the image to be processed is inconsistent with the initial text information to be recognized, the process of adjusting the initial text information to be recognized can also be understood as a manual adjustment. This application does not limit the specific operation methods involved above.
[0062] According to the specific embodiments described above in this application, by recognizing the text information in the image to be processed, the initially identified text depends on the original image to be processed, thus providing the original basis for subsequently determining the target text. Furthermore, the initially identified text is compared with the original text information in the original image to be processed, and the comparison result is determined based on the content of the original image to be processed. This further ensures the consistency between the text information in the original image to be processed and the target text in the conversion to target text, thereby guaranteeing the accuracy of the target text recognition.
[0063] In one specific embodiment provided in this application, the target recognition text corresponding to each text information is obtained based on the above content. On this basis, to improve the training accuracy of the grammar correction model, this application uses the following method to determine the target modification text corresponding to each target recognition text.
[0064] Specifically, determine the target modification text corresponding to each target recognition text, including: Determine an initial target recognition text, wherein the initial target recognition text is any one of the target recognition texts; Based on the set grammatical structure, the initial target recognition text is adjusted to obtain the target modified text that satisfies the set grammatical structure.
[0065] The syntax structure can be a syntax structure that conforms to a standard syntax format. For example, the syntax structure can include at least one of the following: a): Used to correct inappropriate grammatical structures.
[0066] For example, if there are grammatical errors in the target recognition text, the target recognition text can be corrected according to the grammatical structure used to adjust the grammatical errors, so as to obtain a text without grammatical errors after the grammatical errors are corrected.
[0067] b): Used to adjust grammatical structures with incomplete components.
[0068] For example, if there are grammatical errors with incomplete components in the target recognition text, the target recognition text can be corrected according to the grammatical structure used to adjust the incomplete components, so as to obtain a text without incomplete components after correcting the grammatical errors with incomplete components.
[0069] c): Used to correct grammatical structures with incorrect word order.
[0070] For example, if there are grammatical errors with incorrect word order in the target recognition text, the target recognition text can be corrected according to the grammatical structure used to adjust the incorrect word order, so as to obtain a text without incorrect word order after correcting the grammatical errors with incorrect word order.
[0071] d): Used to adjust custom syntax structures.
[0072] For example, if there are custom grammatical errors in the target recognition text, the target recognition text is corrected according to the custom grammatical structure used to adjust it, so as to obtain text without custom grammatical errors after the custom grammatical errors are corrected.
[0073] It should be understood that the target modified text involved in this application can be modified according to one or more of the grammatical structures shown above. Furthermore, the custom grammatical structure can be any other grammatical structure besides any of the above-mentioned grammatical structures, and this application does not limit the custom grammatical structure.
[0074] In one specific embodiment provided in this application, the initial target recognition text is adjusted according to the set grammatical structure to obtain the target correction text without grammatical errors, so as to facilitate the subsequent training of the grammar correction model based on the target correction text.
[0075] Step 106: Train the grammar correction model based on each target recognition text and the target modification text corresponding to each target recognition text until the training stopping condition of the grammar correction model is met.
[0076] The grammar correction model can be understood as a model trained on a large-scale language model to identify target text in image format and correct grammatical errors in the target text.
[0077] The training stopping condition for the grammar correction model can be understood as follows: the training epochs reach a set number, or all text information in the image to be processed is obtained, or the loss value of the grammar correction model reaches a loss threshold, or the loss function of the grammar correction model converges to a certain condition. This application does not impose any restrictions on the training stopping condition for the grammar correction model.
[0078] In one specific embodiment provided in this application, the image to be processed is input into the grammar correction model to obtain the output result of the grammar correction model, and the grammar correction model is trained according to the target recognition text and the target modification text corresponding to each target recognition text until the set training rounds are reached to obtain the trained grammar correction model.
[0079] Specifically, a grammar correction model is trained based on the target recognition text and the corresponding target modification text until the training stopping condition of the grammar correction model is met, including: The image to be processed is input into the syntax correction model to obtain at least one predicted recognition text output by the syntax correction model and the predicted modified text corresponding to each predicted recognition text. Based on each predicted recognized text, the predicted modified text corresponding to each predicted recognized text, each target recognized text, and the target modified text corresponding to each target recognized text, the model loss value of the syntax correction model is determined. Based on the model loss value, the syntax correction model is trained until the training stopping condition of the syntax correction model is met.
[0080] In one specific embodiment provided in this application, the image to be processed is input into a grammar correction model to obtain all text information recognized by the grammar correction model based on the image, and this information is used as the predicted recognition text. The grammar correction model then outputs the corresponding predicted modified text for each predicted recognition text.
[0081] Based on the above, the target recognition text obtained in the above manner and the predicted recognition text output by the grammar correction model are used to train the grammar correction model's ability to recognize text in images. The target modification text corresponding to each target recognition text and the predicted modification text corresponding to each predicted recognition text are then used to train the grammar correction model's ability to correct grammatical errors in the text in images. Based on training these two capabilities of the grammar correction model, the grammar correction model is obtained.
[0082] It should be understood that the training method of the grammar correction model mentioned above in this application is only one training method. This application can also train the grammar correction model in other ways until the grammar correction model learns the ability to recognize and correct text in images.
[0083] According to the specific implementation provided in this application, the grammar correction model is trained based on each predicted recognition text, the predicted modified text corresponding to each predicted recognition text, each target recognition text, and the target modified text corresponding to each target recognition text. This supervises the grammar correction model to have the ability to recognize and correct text in images. By inputting images into the grammar training model, the target recognition text and the target modified text corresponding to each target recognition text are directly output by the grammar training model. This eliminates the need for the two-step operation of first converting the text in the image into text format and then correcting it with text format information, thus avoiding cascading errors and improving the accuracy of grammatical error correction in text in images.
[0084] The following is in conjunction with the appendix Figure 2 Taking the training method of the grammar correction model provided in this application as an example in the application of grammar correction, the grammar correction method will be explained and illustrated. Among them, Figure 2 The following is a flowchart illustrating a syntax correction method according to an embodiment of this application, specifically including the following steps 202-204: Step 202: Obtain the image to be processed, wherein the image to be processed includes at least one initial text to be corrected.
[0085] Step 204: Input the image to be processed into the syntax correction model to obtain at least one target text to be corrected and the target corrected text corresponding to each target text to be corrected.
[0086] The syntax correction model is trained based on the syntax correction model training method described above.
[0087] In one specific embodiment provided in this application, the image to be processed is input into the grammar correction model trained above, and at least one target text to be corrected and the target corrected text corresponding to each target text to be corrected are obtained from the output of the grammar correction model.
[0088] To facilitate understanding, this manual uses the following examples to explain the training method and syntax correction method of the syntax correction model mentioned above.
[0089] In one specific embodiment provided in this application, the text information in the image to be processed is taken as an example of a student's essay for explanation.
[0090] We collect photos of students' essays as training data to obtain images to be processed.
[0091] The image to be processed is identified using text recognition technology, thereby converting the image into text information, namely the initial text to be recognized mentioned above in this application. The converted text information is annotated in two ways. Firstly, if the converted initial text to be recognized is inconsistent with the student essay displayed in the original image, it is marked as text requiring correction to facilitate subsequent correction processing. Secondly, if the converted initial text to be recognized is consistent with the student essay displayed in the original image, or if the corrected initial text to be recognized is consistent with the student essay displayed in the original image, it is marked as text requiring grammatical correction, namely the target modified text mentioned above in this application. Based on the above two annotations, the annotation format is as follows: Reference: Original image to be processed; Annotation results: A list of sentence granularities, with each sentence formatted as follows: { Original sentence: xxxx, (Text obtained after converting the image to text and correcting it to match the image) "Modified": xxxx (grammatically correct text used as a tag) } The list of sentence granularities can be understood as at least one piece of text information in the image to be processed, such as each sentence in a student's essay. The format of each sentence can be understood as text information containing a preset ending character.
[0092] Based on the above annotations, the image to be processed is input into the grammar correction model, and the grammar correction model is fine-tuned in a supervised manner, so that the grammar correction model has the ability to generate text information in the image based on the image and correct the grammar of the text information.
[0093] Specifically, the trained syntax correction model can be represented in the following format: Model input: Image or text information; Model expected output: [{ Original sentence: xxxx, Edit: xxxx }] According to the embodiments provided in this application, the text discovery and modification method based on multimodal input can support text titles or image titles, while eliminating the need for image-to-text conversion, thus avoiding the generation and spread of errors.
[0094] Corresponding to the above method embodiments, this application also provides embodiments of a training apparatus for a syntax correction model. Figure 3 This diagram illustrates the structure of a training apparatus for a syntax correction model according to an embodiment of this application. Figure 3 As shown, the device includes: The acquisition unit 302 is configured to acquire an image to be processed, wherein the image to be processed includes at least one text message.
[0095] The recognition unit 304 is configured to recognize the image to be processed, determine the target recognition text corresponding to each text information, and determine the target modification text corresponding to each target recognition text.
[0096] Training unit 306 is configured to train a grammar correction model based on each target recognition text and the target modification text corresponding to each target recognition text, until the training stopping condition of the grammar correction model is met.
[0097] Furthermore, the identification unit 304 is further configured as follows: Identify each text information in the image to be processed, and obtain the initial text to be identified corresponding to each text information; By comparing each piece of text information with the corresponding initial text to be identified, the comparison results of each initial text to be identified are obtained; Based on the comparison results of each initial text to be identified, the target text corresponding to each initial text to be identified is determined.
[0098] Furthermore, the identification unit 304 is further configured as follows: Determine the initial text to be identified to be processed, and determine the comparison result corresponding to the initial text to be identified to be processed, wherein the initial text to be identified to be processed is any one of the initial texts to be identified; If the comparison result is consistent, the initial text to be identified is determined as the target text to be identified. If the comparison result is inconsistent, the initial text to be identified is corrected, and the corrected initial text to be identified is determined as the target text to be identified.
[0099] Furthermore, the identification unit 304 is further configured as follows: Determine the initial text to be identified for comparison, wherein the initial text to be identified for comparison is any one of the initial texts to be identified; Based on the text attribute information of the initial text to be identified to be compared, the initial text information corresponding to the initial text to be identified to be compared is located in the image to be processed; The initial text information is compared with the initial text to be identified to obtain the comparison result of the initial text to be identified.
[0100] Furthermore, the identification unit 304 is further configured as follows: At least one text message is determined in the image to be processed based on a preset end character; Identify each piece of text information and obtain the initial text to be identified for each piece of text information.
[0101] Furthermore, the identification unit 304 is further configured as follows: Determine an initial target recognition text, wherein the initial target recognition text is any one of the target recognition texts; Based on the set grammatical structure, the initial target recognition text is adjusted to obtain the target modified text that satisfies the set grammatical structure.
[0102] Furthermore, training unit 306 is further configured as follows: The image to be processed is input into the syntax correction model to obtain at least one predicted recognition text output by the syntax correction model and the predicted modified text corresponding to each predicted recognition text. Based on each predicted recognized text, the predicted modified text corresponding to each predicted recognized text, each target recognized text, and the target modified text corresponding to each target recognized text, the model loss value of the syntax correction model is determined. Based on the model loss value, the syntax correction model is trained until the training stopping condition of the syntax correction model is met.
[0103] Corresponding to the above method embodiments, this application also provides embodiments of a syntax correction device. Figure 4 A schematic diagram of the structure of a syntax correction device according to an embodiment of this application is shown. Figure 4 As shown, the device includes: The acquisition unit 402 is configured to acquire an image to be processed, wherein the image to be processed includes at least one initial text to be corrected.
[0104] The processing unit 404 is configured to input the image to be processed into the syntax correction model to obtain at least one target text to be corrected and the target corrected text corresponding to each target text to be corrected.
[0105] The syntax correction model is trained based on the syntax correction model training method described above.
[0106] The above is a schematic scheme of a training device and a grammar correction device for a grammar correction model according to this embodiment. It should be noted that the technical solution of the training device and the grammar correction device for this grammar correction model belongs to the same concept as the technical solution of the training method and the grammar correction method for the grammar correction model described above. For details not described in detail in the technical solution of the training device and the grammar correction device for this grammar correction model, please refer to the description of the technical solution of the training method and the grammar correction method for the grammar correction model described above.
[0107] Figure 5 A structural block diagram of a computing device according to an embodiment of this application is shown. The components of the computing device 500 include, but are not limited to, a memory 510 and a processor 520. The processor 520 is connected to the memory 510 via a bus 530, and a database 550 is used to store data.
[0108] The computing device 500 also includes an access device 540, which enables the computing device 500 to communicate via one or more networks 560. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 540 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0109] In one embodiment of this application, the aforementioned components of the computing device 500 and Figure 5 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 5 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this application. Those skilled in the art can add or replace other components as needed.
[0110] The computing device 500 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 500 can also be a mobile or stationary server.
[0111] The processor 520 is used to execute the following computer program / instruction, which, when executed by the processor, implements the training method and steps of the syntax correction method for the above-mentioned syntax correction model.
[0112] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device belongs to the same concept as the training method of the syntax correction model and the technical solution of the syntax correction method described above. For details not described in detail in the technical solution of the computing device, please refer to the description of the training method of the syntax correction model and the technical solution of the syntax correction method described above.
[0113] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the training method and steps of the syntax correction method described above for the syntax correction model.
[0114] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solutions of the training method for the syntax correction model and the syntax correction method described above. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solutions of the training method for the syntax correction model and the syntax correction method described above.
[0115] An embodiment of this specification also provides a computer program product, including a computer program / instruction that, when executed by a processor, implements the training method and steps of the syntax correction method described above for the syntax correction model.
[0116] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product belongs to the same concept as the technical solution of the training method of the syntax correction model and the syntax correction method described above. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the training method of the syntax correction model and the syntax correction method described above.
[0117] The foregoing has described specific embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0118] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0119] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0120] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0121] The preferred embodiments disclosed above are merely illustrative of this application. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this application. These embodiments are selected and specifically described in this application to better explain the principles and practical applications of this application, thereby enabling those skilled in the art to better understand and utilize this application. This application is limited only by the claims and their full scope and equivalents.
Claims
1. A training method for a syntax correction model, characterized in that, include: Obtain an image to be processed, wherein the image to be processed includes at least one text message; Identify the image to be processed, determine the target recognition text corresponding to each text information, and determine the target modification text corresponding to each target recognition text; The grammar correction model is trained based on the target recognition text and the target modification text corresponding to each target recognition text until the training stopping condition of the grammar correction model is met.
2. The method as described in claim 1, characterized in that, Identify the image to be processed and determine the target text corresponding to each text information, including: Identify each text information in the image to be processed, and obtain the initial text to be identified corresponding to each text information; By comparing each piece of text information with the corresponding initial text to be identified, the comparison results of each initial text to be identified are obtained; Based on the comparison results of each initial text to be identified, the target text corresponding to each initial text to be identified is determined.
3. The method as described in claim 2, characterized in that, Based on the comparison results of each initial text to be identified, the target text corresponding to each initial text to be identified is determined, including: Determine the initial text to be identified to be processed, and determine the comparison result corresponding to the initial text to be identified to be processed, wherein the initial text to be identified to be processed is any one of the initial texts to be identified; If the comparison result is consistent, the initial text to be identified is determined as the target text to be identified. If the comparison result is inconsistent, the initial text to be identified is corrected, and the corrected initial text to be identified is determined as the target text to be identified.
4. The method as described in claim 2, characterized in that, By comparing each piece of text information with the corresponding initial text to be identified, the comparison results of each initial text to be identified are obtained, including: Determine the initial text to be identified for comparison, wherein the initial text to be identified for comparison is any one of the initial texts to be identified; Based on the text attribute information of the initial text to be identified to be compared, the initial text information corresponding to the initial text to be identified to be compared is located in the image to be processed; The initial text information is compared with the initial text to be identified to obtain the comparison result of the initial text to be identified.
5. The method as described in claim 2, characterized in that, Identify each text information in the image to be processed, and obtain the initial text to be identified corresponding to each text information, including: At least one text message is determined in the image to be processed based on a preset end character; Identify each piece of text information and obtain the initial text to be identified for each piece of text information.
6. The method according to any one of claims 1 to 5, characterized in that, Determine the target modification text corresponding to each target recognition text, including: Determine an initial target recognition text, wherein the initial target recognition text is any one of the target recognition texts; Based on the set grammatical structure, the initial target recognition text is adjusted to obtain the target modified text that satisfies the set grammatical structure.
7. The method according to any one of claims 1 to 5, characterized in that, The grammar correction model is trained based on the target recognition text and the corresponding target modification text until the training stopping condition of the grammar correction model is met, including: The image to be processed is input into the syntax correction model to obtain at least one predicted recognition text output by the syntax correction model and the predicted modified text corresponding to each predicted recognition text. Based on each predicted recognized text, the predicted modified text corresponding to each predicted recognized text, each target recognized text, and the target modified text corresponding to each target recognized text, the model loss value of the syntax correction model is determined. Based on the model loss value, the syntax correction model is trained until the training stopping condition of the syntax correction model is met.
8. A grammar correction method, characterized in that, include: Obtain an image to be processed, wherein the image to be processed includes at least one initial text to be corrected; The image to be processed is input into the syntax correction model to obtain at least one target text to be corrected and the target corrected text corresponding to each target text to be corrected. The syntax correction model is trained based on the syntax correction model training method of any one of claims 1 to 7.
9. A training device for a syntax correction model, characterized in that, include: The acquisition unit is configured to acquire an image to be processed, wherein the image to be processed includes at least one text information; The recognition unit is configured to recognize the image to be processed, determine the target recognition text corresponding to each text information, and determine the target modification text corresponding to each target recognition text; The training unit is configured to train a grammar correction model based on each target recognition text and the target modification text corresponding to each target recognition text, until the training stopping condition of the grammar correction model is met.
10. A grammar correction device, characterized in that, include: The acquisition unit is configured to acquire an image to be processed, wherein the image to be processed includes at least one initial text to be corrected; The processing unit is configured to input the image to be processed into the syntax correction model to obtain at least one target text to be corrected and the target corrected text corresponding to each target text to be corrected. The syntax correction model is trained based on the syntax correction model training method of any one of claims 1 to 7.
11. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 8.
12. A computer-readable storage medium storing a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 8.
13. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 8.