An intelligent method and system for generating error analysis documents, and a storage medium

CN117933203BActive Publication Date: 2026-09-25陕西巨微图书文化传播有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410094964.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-23
Publication Date
2026-09-25
Estimated Expiration
2044-01-23

AI Technical Summary

Technical Problem

[0005]本申请实施例提供了一种智能生成差错分析文档方法及系统、存储介质,用以解决现有技术中质检人员采用人工处理方式存在的效率低和漏检的问题

Benefits of technology

[0022]对批注进行智能化识别,提取批注中的差错信息,对差错信息进行统计,使质检人员免去了在批注统计上的经历投入,大大提高了处理的效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117933203B_ABST
    Figure CN117933203B_ABST
Patent Text Reader

Abstract

The application discloses an intelligent error analysis document generation method and system and a storage medium, relates to the technical field of electric digital data processing, and comprises the following steps: acquiring an original document; identifying errors in the original document, setting a comment in the original document at a position corresponding to the identified errors, forming a first comment-containing document; acquiring a comment added by a quality inspector in the first comment-containing document, forming a second comment-containing document; classifying and modifying the comment with a simple description in the second comment-containing document, forming a modified second comment-containing document; acquiring the comment content in the modified second comment-containing document and extracting error information in the comment content; and performing statistics on the error information to form an error analysis document. The application intelligently identifies the comment, extracts error information in the comment, and performs statistics on the error information, so that the quality inspector is relieved of the experience and input in comment statistics, and the processing efficiency is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electronic digital data processing technology, and in particular to a method and system for intelligently generating error analysis documents, and a storage medium. Background Technology

[0002] In document editing, such as producing books, newspapers, and web pages, the generation and editing of documents largely rely on manual operation. Therefore, errors are inevitable during the editing process, including typos, punctuation mistakes, and even some substantive errors. Error management involves two methods: self-checking by the document editor and inspection by quality control personnel. Because document editors often have fixed mindsets, even after spending time on self-checking, some errors may still exist in the document. Therefore, the role of quality control personnel is crucial.

[0003] Currently, when quality inspectors check documents, they read through the entire document, identify errors one by one, and mark them as annotations, creating an annotated document. Finally, they perform statistical analysis on all the annotations, summarizing information such as the types and quantities of errors found in the document, and forming an error analysis table. This error analysis table is then fed back to the document editors, serving as a reminder and incentive for them.

[0004] When using the above method, the statistics of annotations in the documents are all done manually by quality inspectors, which results in low efficiency. Summary of the Invention

[0005] This application provides a method and system for intelligently generating error analysis documents, as well as a storage medium, to solve the problems of low efficiency and missed detections caused by the manual processing methods used by quality inspectors in the prior art.

[0006] On one hand, embodiments of this application provide a method for intelligently generating error analysis documents, including:

[0007] Get the original document;

[0008] Identify errors in the original document, and set comments in the original document at the positions corresponding to the identified errors to form the first document with comments;

[0009] The annotations added by quality inspectors in the first document containing annotations are retrieved to create a second document containing annotations.

[0010] An error classification model is used to classify the comments with simple descriptions in the second annotated document, and to determine the comments with detailed descriptions that correspond to the comments with simple descriptions. The content of the comments with detailed descriptions is added to the comments with simple descriptions to form the modified second annotated document.

[0011] Obtain the annotation content from the modified second annotated document and extract the error information from the annotation content;

[0012] The error information is statistically analyzed to create an error analysis document.

[0013] On the other hand, embodiments of this application also provide an intelligent error analysis document generation system, including:

[0014] The document acquisition module is used to acquire the original document;

[0015] The first batch of annotation modules is used to identify errors in the original document and set annotations in the original document at the positions corresponding to the identified errors, thus forming the first annotated document;

[0016] The second annotation module is used to obtain the annotations added by quality inspectors in the first annotation document and form the second annotation document.

[0017] The document editing module is used to classify the annotations with simple descriptions in the second annotated document using an error classification model, determine the annotations with detailed descriptions corresponding to the annotations with simple descriptions, add the content of the annotations with detailed descriptions to the annotations with simple descriptions, and form the modified second annotated document.

[0018] The annotation extraction module is used to obtain the annotation content in the modified second annotated document and extract the error information in the annotation content;

[0019] The error analysis module is used to statistically analyze error information and generate error analysis documents.

[0020] On the other hand, embodiments of this application also provide a computer storage medium storing a plurality of computer instructions for causing a computer to execute the above-described method.

[0021] The intelligent method, system, and storage medium for generating error analysis documents disclosed in this application have the following advantages:

[0022] Intelligent identification of annotations, extraction of error information from annotations, and statistical analysis of error information free quality inspectors from the time and effort required for annotation statistics, greatly improving processing efficiency. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a flowchart illustrating an intelligent method for generating error analysis documents, as provided in an embodiment of this application. Detailed Implementation

[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0026] Figure 1 A flowchart illustrating an intelligent method for generating error analysis documents, provided in an embodiment of this application. This application provides an intelligent method for generating error analysis documents, including:

[0027] S100, retrieve the original document.

[0028] For example, a document editor can create a new document according to the requirements and add certain content to the new document to obtain the original document. Taking teaching materials as an example, when dealing with content related to real numbers, all relevant knowledge points can be listed, and the definitions and examples of these knowledge points can be shown. Additionally, some thought-provoking questions related to these knowledge points can be provided.

[0029] S110, identify errors in the original document, set comments in the original document at the positions corresponding to the identified errors, and form the first document with comments.

[0030] For example, the identified errors include a variety of possibilities. This application embodiment will explain three aspects: typos, punctuation marks, and formatting.

[0031] There are two possibilities for misspellings in word form: either the entire word is misspelled, or only a portion of the word is misspelled. For the former, each sentence in the original document can be divided into multiple real word segments. Each real word segment is converted into a corresponding word vector. Based on the word vectors of other segments besides the real word segments, predicted word segments are obtained. The distance between the word vectors of the predicted word segments and the word vectors of the real word segments is determined, and the presence of misspellings in the real word segments is then determined based on this distance.

[0032] For the latter case, each sentence in the original document can be divided into a plurality of actual word segments, each actual word segment is compared with a standard word segment in a dictionary, and whether the actual word segment has a typo is determined according to the comparison result.

[0033] When the whole word is a typo, for example, "certificate and fractions are collectively called rational numbers" appears in content related to real numbers, it is obvious that the whole "certificate" is a typo, and the correct form should be "integer". For this case, a prediction model based on LSTM (Long Short-Term Memory) network can be established, and the prediction model is trained by Tensorflow. The word segment is predicted by using word vectors of other analyses except the word segment to obtain a predicted word segment. After the predicted word segment is also converted into a corresponding word vector, the distance between the word vector of the predicted word segment and the word vector of the actual word segment can be calculated. If the distance is less than a set distance threshold, it indicates that the actual word segment and the predicted word segment are similar or even the same, and it can be considered that there is no typo in the current actual word segment. When the distance is greater than or equal to the distance threshold, it indicates that there is a great difference between the actual word segment and the predicted word segment, and it can be considered that there is a typo in the current actual word segment.

[0034] When a part of the word is a typo, for example, "certificate number and fractions are collectively called rational numbers" appears, it is obvious that "certificate" in "certificate number" is a typo, and the correct form should be "integer". For this case, since "certificate number" is not a commonly used word in the field of teaching auxiliary books, after querying in a dictionary established based on commonly used words in the field of teaching auxiliary books, if the same standard word segment is not queried, it indicates that there is a partial typo in the current word segment.

[0035] For punctuation marks, it is necessary to preset the usage requirements for punctuation marks, for example, the end of each paragraph needs to be ended with a period. Based on the usage requirements, paragraphs in the document can be segmented first, and then the punctuation marks at the end of the paragraphs are detected. If the punctuation mark is not a period, it can be determined that there is a punctuation error at this time.

[0036] For formatting, it is also necessary to preset templates for all formats in the document. For example, for a knowledge point, the method of listing the knowledge point, definition, interpretation, thinking question and the like in sequence can be adopted. Of course, there may be a variety of different format templates in a document, so all format templates need to be summarized in advance. In error analysis, the document can be divided into a plurality of paragraphs, word segments in each paragraph are determined, and the word segments are compared with keywords in each template to determine the format template corresponding to the part where the current paragraph is located. After the format template is determined, it is judged whether the current part of the document meets all the requirements in the format template, for example, whether the knowledge point, definition, interpretation and thinking question appear in sequence. If not, it is determined that a format error has occurred.

[0037] After identifying errors including typos, punctuation mistakes, and formatting errors, comments can be inserted at the locations of the errors, with the comment content inserted into the comment box. This comment content includes the commenter, the date of the comment, and the error information, with the error information directly describing the specific error.

[0038] S120: Obtain the annotations added by the quality inspectors in the first annotated document and form a second annotated document.

[0039] For example, after the computer completes the identification of errors in the original document, it can be handed over to quality inspectors. The quality inspectors can modify the comments added by the computer, or add corresponding comments for some substantive errors, such as inappropriate wording or formatting errors. These comments should be consistent in form with the comments added by the computer so that the second document with comments can be analyzed with comments of the same form.

[0040] In the embodiments of this application, after obtaining the annotations added by the quality inspector in the first annotated document, the added annotations are further analyzed to determine whether the added annotations are correct. If the added annotations are incorrect, a reminder message is sent to the quality inspector.

[0041] Specifically, since quality inspectors also add annotations based on subjective judgment, errors may occur. For example, a quality inspector might think a word in a sentence is misspelled, but after computer analysis, the location where the quality inspector added the annotation is not incorrect. In this case, it could be that the quality inspector's judgment was wrong, or that the quality inspector inserted the annotation in the wrong place. In either case, the computer will generate a warning message for that annotation. If the computer generates such a warning message, the second document containing the annotations, along with the warning message, needs to be sent to the quality inspector. After the quality inspector makes corrections and confirms that there are no errors, and the computer analysis is also successful, the annotations in the second document containing the annotations will then be analyzed.

[0042] Furthermore, when analyzing the added annotations, the error information of the annotations is extracted, the error type to which the error information belongs is determined, and when the error type matches the preset type, the content in the second annotated document corresponding to the error type is analyzed.

[0043] Specifically, annotations added by quality inspectors may address formal errors, such as typos, punctuation mistakes, and formatting issues, or they may address substantive errors. For the former, computers can analyze based on objective standards to determine the correctness of the annotation. However, for the latter, due to the highly subjective nature of annotations, it is difficult for computers to determine their correctness. Therefore, it is necessary to first determine the type of error—formal or substantive—based on the keywords in the annotation. If it is a formal error, the corresponding content is then analyzed to determine the annotation's correctness; if it is a substantive error, it is ignored.

[0044] S130, use an error classification model to classify the comments with simple descriptions in the second annotated document, determine the comments with detailed descriptions corresponding to the comments with simple descriptions, add the content of the comments with detailed descriptions to the comments with simple descriptions, and form the modified second annotated document.

[0045] For example, when quality inspectors add comments on formal or substantive errors, there may be many similar or identical errors in the first document containing the comments. The quality inspectors will add a more detailed comment on the first error they find, but in subsequent content, they will only give a simple description of the same or similar errors in the comments, such as "same as above" or "formal error". If these comments with only simple descriptions are analyzed, it is very likely to cause the analysis results to be incorrect. Therefore, before analyzing the annotations in the second annotated document, this application establishes an error classification model. This error classification model can be based on a convolutional neural network. The model is trained using a piece of content with an error and detailed annotations made by quality inspectors or a pre-prepared training set as training samples. Then, the trained error classification model is used to classify annotations with simple descriptions. After determining the annotations with detailed descriptions that correspond to the annotations with simple descriptions, the content of the annotations with detailed descriptions is added to the annotations with simple descriptions. This ensures that each annotation in the second annotated document has a detailed description, improving the accuracy of annotation analysis in subsequent processing.

[0046] S140, obtain the annotation content from the second annotated document, and extract the error information from the annotation content.

[0047] S150: Statistically analyze error information to generate an error analysis document.

[0048] For example, when compiling error information, the errors in comments added by quality inspectors are also counted. Since there may be more than one quality inspector, and the same document may need to undergo two or more quality inspections, the second document containing comments may contain comments added by multiple quality inspectors during multiple quality inspections. In this case, the comments added by the same person can be identified based on the information of the person who added the comments. Therefore, when compiling the statistics on errors in comments added by quality inspectors, the statistics are compiled separately according to the quality inspector who added the comments.

[0049] In the embodiments of this application, when statistically analyzing error information, keywords are extracted from the error information, the keywords are compared with standard information to determine the error type to which the error information belongs, and error information involving the same error type is statistically analyzed.

[0050] Specifically, the same type of error may appear multiple times in the same original document, such as multiple typos. The type of these errors can usually be determined based on the error information. For example, if the error information is "Typo exists here," it contains the keyword "typo." If the standard information also contains "typo," then the keyword and a certain standard information are the same. The error type "formal error" to which the standard information belongs can be used as the error type of the current annotation. After determining the error type of each error, the number of occurrences of the same error type can be counted, and the location of these errors in the error analysis document can be marked, such as "page x."

[0051] It should be understood that since the annotations added by quality inspectors were also analyzed above, the accuracy of the annotations also needs to be statistically analyzed and displayed in the error analysis table.

[0052] This application also provides an intelligent error analysis document generation system, the system comprising:

[0053] The document acquisition module is used to acquire the original document;

[0054] The first batch of annotation modules is used to identify errors in the original document and set annotations in the original document at the positions corresponding to the identified errors, thus forming the first annotated document;

[0055] The second annotation module is used to obtain the annotations added by quality inspectors in the first annotation document and form the second annotation document.

[0056] The document editing module is used to classify the annotations with simple descriptions in the second annotated document using an error classification model, determine the annotations with detailed descriptions corresponding to the annotations with simple descriptions, add the content of the annotations with detailed descriptions to the annotations with simple descriptions, and form the modified second annotated document.

[0057] The annotation extraction module is used to obtain the annotation content in the modified second annotated document and extract the error information in the annotation content;

[0058] The error analysis module is used to statistically analyze error information and generate error analysis documents.

[0059] This application also provides a computer storage medium storing a plurality of computer instructions for causing a computer to execute the above-described method.

[0060] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0061] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for intelligently generating error analysis documents, characterized in that, include: Get the original document; Identify errors in the original document, and set annotations in the original document at the positions corresponding to the identified errors to form a first document with annotations; Obtain the annotations added by the quality inspectors in the first annotated document, and form a second annotated document; An error classification model is used to classify the comments with simple descriptions in the second annotated document, and to determine the comments with detailed descriptions that correspond to the comments with simple descriptions. The content of the comments with detailed descriptions is added to the comments with simple descriptions to form the modified second annotated document. Obtain the annotation content from the modified second annotated document, and extract the error information from the annotation content; The error information is statistically analyzed to form an error analysis document; Among them, after obtaining the annotations added by the quality inspectors in the first annotated document, the added annotations are analyzed to determine whether the added annotations are correct. If the added annotations are incorrect, a reminder message is sent to the quality inspectors. When analyzing the added annotations, the error information of the annotations is extracted, the error type to which the error information belongs is determined, and when the error type matches the preset type, the content in the second annotated document corresponding to the error type is analyzed.

2. The method for intelligently generating error analysis documents according to claim 1, characterized in that, Errors in the original document are identified as typos. When identifying typos, each sentence in the original document is divided into multiple real word segments, each real word segment is converted into a corresponding word vector, and a predicted word segment is obtained based on the word vectors of other words besides each real word segment. The distance between the word vector of the predicted word segment and the word vector of the real word segment is determined, and the presence of typos in the real word segment is determined based on the distance.

3. The method for intelligently generating error analysis documents according to claim 1, characterized in that, Errors in the original document are identified as typos. When identifying typos, each sentence in the original document is divided into multiple real word segments. Each real word segment is compared with the standard word segment in the dictionary. Based on the comparison results, it is determined whether there are typos in the real word segment.

4. The method for intelligently generating error analysis documents according to claim 1, characterized in that, When compiling the error information, the errors in the annotations added by quality inspectors were also statistically analyzed.

5. The method for intelligently generating error analysis documents according to claim 4, characterized in that, When compiling statistics on errors in annotations added by quality inspectors, statistics were also compiled separately for each quality inspector who added the annotations.

6. The method for intelligently generating error analysis documents according to claim 1, characterized in that, When statistically analyzing the error information, keywords are extracted from the error information, the keywords are compared with standard information to determine the error type to which the error information belongs, and the error information involving the same error type is statistically analyzed.

7. A system applying the intelligent error analysis document generation method of claim 1, characterized in that, include: The document acquisition module is used to acquire the original document; The first batch of annotation modules is used to identify errors in the original document and set annotations in the original document at the positions corresponding to the identified errors, thus forming the first annotated document; The second annotation module is used to obtain the annotations added by quality inspectors in the first annotated document and form a second annotated document. The document editing module is used to classify the annotations with simple descriptions in the second annotated document using an error classification model, determine the annotations with detailed descriptions corresponding to the annotations with simple descriptions, add the content of the annotations with detailed descriptions to the annotations with simple descriptions, and form the modified second annotated document. The annotation extraction module is used to obtain the annotation content in the modified second annotated document and extract the error information in the annotation content; The error analysis module is used to statistically analyze the error information and generate an error analysis document.

8. A computer storage medium, characterized in that, The computer storage medium stores a plurality of computer instructions, which are used to cause the computer to perform the method described in any one of claims 1-6.

Citation Information

Patent Citations

  • Document proofreading method, device and equipment

    CN116502627A

  • Document response method and device, equipment and storage medium

    CN116932727A