Document proofreading support device, document proofreading support method, and document proofreading support program
The document proofing support device addresses the limitation of existing technologies by identifying semantic errors and providing calibration time information, enhancing the accuracy and efficiency of document proofreading.
Patent Information
- Application Number
- JP2021184785
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-11-12
- Publication Date
- 2025-05-08
- Estimated Expiration
- 2041-11-12
AI Technical Summary
Existing document proofing support devices primarily focus on detecting grammatical or type errors, failing to effectively identify semantic errors in documents.
A document proofing support device that includes a document analyzing unit to identify unknown sentences potentially containing semantic errors, and a display processing unit to display the components to which these unknown sentences approximate, along with the required calibration time.
Enables the detection of potential semantic errors in documents, providing users with information on the components that unknown sentences approximate and the time required for calibration, thus improving the accuracy and efficiency of document proofreading.
Smart Images

Figure 0007672955000001 
Figure 0007672955000002 
Figure 0007672955000003
Abstract
Description
[Technical field]
[0001] The present invention relates to a document proofreading support device, a document proofreading support method, and a document proofreading support program. [Background technology]
[0002] In companies, a large number of different documents are created on a daily basis. However, document creators often do not have enough time to review documents. Furthermore, business documents must be written in a way that accurately conveys technical content. Recently, it has become common for computers to perform such reviews.
[0003] The document proofreading support device of Patent Document 1 extracts inappropriate descriptions that do not conform to predetermined rules from a document, calculates the estimated revision time required to correct them, and outputs the inappropriate descriptions and the estimated revision time. The document proofreading support device compares the revision time specified by the user with the estimated revision time. If the estimated revision time is shorter, the document proofreading support device targets all inappropriate descriptions for revision. Conversely, if the estimated revision time is longer, the document proofreading support device targets only inappropriate descriptions that are of high importance for revision. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] JP 2017-41164 A Summary of the Invention [Problem to be solved by the invention]
[0005] A manuscript before proofreading may contain not only grammatical or stylistic errors such as typos, omitted characters, and spelling variations, but also semantic errors. The rules in Patent Document 1 are intended to detect grammatical or stylistic errors. A separate measure was needed to inform the user of potential semantic errors. Therefore, an object of the present invention is to detect locations in a document where semantic errors potentially exist. [Means for solving the problem]
[0006] The document proofreading support device of the present invention is , which is a criterion for detecting grammatical or typological errors in the sentence. Matches a given rule Includes or not Determine , the sentence If the sentence contains a portion that matches the rule, the sentence is a matching sentence that contains a grammatical or typological error, and if the sentence does not contain a portion that matches the rule, the sentence is a non-grammatical or typological error, but Unknown sentences that may contain semantic errors and an unknown sentence processing unit that estimates which of a plurality of components defined according to the type of the document the unknown sentence is similar to; and a display unit that displays the unknown sentence and the component to which the unknown sentence is similar. and displaying the constituents that can proofread the unknown sentence in the shortest time. A display processing unit. Other means will be described in the description of the embodiment of the invention. Effect of the Invention
[0007] According to the present invention, it is possible to detect potential semantic errors in a document. [Brief description of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram illustrating a configuration of a document proofreading support device. [Diagram 2] 1 is an example of a document. [Diagram 3] 11 is an example of rule information. [Figure 4] 11 is an example of matching sentence information. [Diagram 5] 1 is an example of unknown sentence information. [Figure 6] 13 is an example of rule-specific proofreading time information. [Figure 7] 13 is an example of proofreading time information for each matching sentence. [Figure 8] 1 is an example of document type and component information. [Figure 9] This is an example of a sentence space. [Figure 10] 11 is an example of distance information. [Figure 11] 13 is an example of component-specific importance information. [Figure 12] 11 is an example of score information. [Figure 13] This is an example of score-proofreading time conversion information. [Figure 14] 13 is an example of proofreading time information for each unknown sentence. [Figure 15] 13 is a flowchart of a document analysis process. [Figure 16] 13 is a flow chart of a matching sentence processing procedure. [Figure 17] 13 is a flow chart of an unknown sentence processing procedure. [Figure 18] 13 is an example of a calibration time display screen. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] Hereinafter, an embodiment of the present invention (referred to as "the present embodiment") will be described in detail with reference to the drawings, etc. The present embodiment is an example of extracting grammatical or typographical errors and semantic errors from business documents. The task of correcting typographical errors or expressions in a manuscript created by a document creator to produce a final draft for the purpose of printing, binding, etc. is generally called "proofreading." The present embodiment can also be used for purposes other than printing and binding. In this embodiment, the word "proofreading" is used to include "corrections" in such cases.
[0010] (terminology, etc.) A document is an electronic file that contains character strings and is a manuscript before proofreading. A sentence is a unit of continuous character strings contained in a document, separated by a period "." In this embodiment, the terms "sentence" and "paragraph" are synonymous. A rule is a specific criterion for detecting grammatical or typological errors in a sentence.
[0011] A matching sentence is a sentence that contains a match to a rule. A matching sentence may contain grammatical or typological errors. An unknown sentence is a sentence that does not contain any part that matches the rules. An unknown sentence may contain a semantic error. The reason for the name "unknown" sentence is that it is unknown whether or not it contains a semantic error. In that sense, an unknown sentence can also be said to contain a potential proofreading part. The matched sentence may also contain semantic errors. In this embodiment, the matched sentence becomes an unknown sentence once it has been proofread and no longer contains grammatical or typological errors.
[0012] The document type is a category of the document, for example, "quotation", "patent specification", "report", "specification", "minutes", "decision document", etc. A component is an item that a document usually contains, and is defined for each document type. For example, the components of a quotation are "process", "labor cost", "travel cost", and "work content".
[0013] (Configuration of document proofreading support device) 1 is a diagram explaining the configuration of a document proofreading support device 1. The document proofreading support device 1 is a general computer, and includes a central control device 11, an input device 12 such as a mouse and a keyboard, an output device 13 such as a display, a main memory device 14, and an auxiliary memory device 15. These are interconnected via a bus.
[0014] The auxiliary storage device 15 stores documents 31, rule information 32, matching sentence information 33, unknown sentence information 34, rule-specific proofreading time information 35, matching sentence-specific proofreading time information 36, document type and component information 37, distance information 38, component-specific importance information 39, score information 40, score and proofreading time conversion information 41, unknown sentence-specific proofreading time information 42, and component estimation models 43 (described in detail below).
[0015] Of these, the document 31, rule information 32, rule-specific proofreading time information 35, document type and component information 37, component-specific importance information 39, score and proofreading time conversion information 41 and component estimation model 43 are the results of information created by the user and imported by the document proofreading support device 1 into the auxiliary storage device 15. The remaining information, namely, matched sentence information 33, unknown sentence information 34, matched sentence-specific proofreading time information 36, distance information 38, score information 40 and unknown sentence-specific proofreading time information 42, were created by the document proofreading support device 1 during processing.
[0016] The document analysis unit 21, the matching sentence processing unit 22, the unknown sentence processing unit 23, and the display processing unit 24 in the main memory device 14 are programs. The central control device 11 realizes the functions of each program (described in detail later) by reading these programs from the auxiliary memory device 15 and loading them into the main memory device 14. The auxiliary memory device 15 may be configured independent of the supply and demand adjustment support device 1 (cloud).
[0017] (document) FIG. 2 is an example of document 31. Document 31 in FIG. 2 is a development-related report for a manufacturer. Document 31 includes sentences SE01 and SE02. Two rules match sentence SE01 (reference numbers 51 and 52). Therefore, sentence SE01 is a matching sentence. No rule matches sentence SE02 (reference number 53). Therefore, sentence SE02 is an unknown sentence. Note that the rules reference numbers 51 to 53 are for explanatory purposes and are not stated in document 31 itself.
[0018] (Rule information) 3 is an example of the rule information 32. In the rule information 32, a rule ID (column 101), a rule (column 102), and an importance level (column 103) are stored in association with each other. The rule ID (column 101) is an identifier that uniquely identifies a rule. The rules (field 102) are the rules described above. The importance is a relative weight among multiple rules. The user sets the importance within the range of "0<importance≦1".
[0019] (Matching sentence information) 4 is an example of the matching sentence information 33. In the matching sentence information 33, a sentence ID (column 111), a matching sentence (column 112), a rule ID (column 113), and an importance level (column 114) are stored in association with each other. The sentence ID (field 111) is an identifier that uniquely identifies a sentence, in this case a matching sentence. The matched sentence (column 112) is the matched sentence described above. The rule ID (field 113) is the same as the rule in FIG. The importance (column 114) is the same as that in FIG. 4 includes two records for sentence SE01, which correspond to rules 51 and 52 in FIG.
[0020] (Unknown sentence information) 5 is an example of the unknown sentence information 34. In the unknown sentence information 34, a sentence ID (column 121) and an unknown sentence (column 122) are stored in association with each other. The sentence ID (field 121) is an identifier that uniquely identifies a sentence, and in this case, identifies an unknown sentence. The unknown sentence (column 122) is the unknown sentence described above. 5 includes one record for sentence SE02, which corresponds to column 53 (no matching rule) in FIG.
[0021] (Proofreading time information by rule) 6 is an example of the rule-specific proofreading time information 35. In the rule-specific proofreading time information 35, a rule ID (field 131) and a proofreading time (field 132) are stored in association with each other. The rule ID (field 131) is the same as the rule ID in FIG. The proofreading time (field 132) is the time required to proofread a mistake that matches the rule. The user sets the proofreading time in seconds based on past examples.
[0022] (Proofreading time information for each matching sentence) 7 is an example of the matching sentence-specific proofreading time information 36. In the matching sentence-specific proofreading time information 36, a sentence ID (field 141) and a proofreading time (field 142) are stored in association with each other. The sentence ID (field 141) is the same as the sentence ID in FIG. The proofreading time (column 142) is the same as the proofreading time in Fig. 6, but here, the proofreading times in Fig. 6 are tallied for each matching sentence. For example, the proofreading time "390" for sentence SE01 is the sum of "360" for rule R02 and "30" for rule R04 in Fig. 6.
[0023] (Document type and component information) 8 is an example of the document type / constituent element information 37. In the document type / constituent element information 37, constituent elements 1 (column 152) to 4 (column 155) are stored in association with the document type (column 151). The document type (field 151) is the document type described above. A user sets a plurality of document types according to his / her own business. Component 1 (column 152) to Component 4 (column 155) are the components described above. The user sets any number of components for each document type. "KPI" stands for "Key Performance Indicator."
[0024] 9 is an example of the sentence space. A sentence vector is defined as a premise for the document proofreading support device 1 to perform processing using the sentence space.
[0025] (Sentence Vector) The document proofreading support device 1 converts one sentence as a character string into one sentence vector. The number of dimensions (number of elements) of the sentence vector is equal to the number of words in the word dictionary of the language of the sentence. Each element of the sentence vector is, for example, the number of times that the word appears in the sentence. Now, if the word dictionary consists of words a, b, c, d, and e, and word a appears once, word b 0 times, word c 2 times, word d 0 times, and word e appears once in the sentence, the sentence vector is "(1,0,2,0,1)". The sentence vector described here is a very simple example. The document proofreading support device 1 can create a more precise sentence vector that more accurately indicates the semantic features of the sentence by any method.
[0026] (Sentence Space) The document proofreading support device 1 can draw sentence vectors as points in a sentence space 44. The number of dimensions of the sentence space 44 is equal to the number of dimensions of the sentence vector. Each axis of the sentence space 44 indicates the number of occurrences of a particular word. The document proofreading support device 1 collects multiple sample documents (learning data) for each document type, in which all sentences are known to be grammatically, typologically, and semantically correct, converts all sentences in each sample document into sentence vectors, and draws them as "●" in the sentence space 44.
[0027] As a result, the document proofreading support device 1 creates a sentence space 44 for each document type. Each "●" in FIG. 9 corresponds to one sentence. The document proofreading support device 1 classifies these ● into clusters using a technique such as the k-means method. It is empirically known that the clusters 61a to 61d often correspond one-to-one to the components of the document type. The sentence space 44 is a space in which sentences as learning data are classified into multiple clusters.
[0028] The document proofreading support device 1 converts a certain unknown sentence into a sentence vector and draws it as a "circle" in the sentence space 44. Then, it happens that a certain circle is classified into one of the clusters 61a to 61d, while another circle is not classified into any of the clusters 61a to 61d. The sentence circle 62a is classified into the cluster 61a and describes the component "process" of the document type "quote". The sentence circle 62b is not classified into any of the clusters 61a to 61d. The sentence in question cannot be said to describe any component of the quotation, and is highly likely to include a semantic error (for example, an advertising statement that is not appropriate for the content of the quotation). Incidentally, if at least one circle is classified into the clusters of all components of a certain document type, it can be said that the document covers all the necessary description items. If there is even one cluster into which the circle is not classified, it can be said that the document lacks that component (description item).
[0029] (Component estimation model) The component estimation model 43 is a function that, when a sentence vector constituting a document of a certain document type is input, outputs the distance between the sentence vector (◯) in the sentence space 44 and each component of the document type (the center of each cluster). A component estimation model 43 exists for each document type. The component estimation model 43 may also perform a process of converting an unknown sentence into a sentence vector. The document proofreading support device 1 may update the positions and sizes of the clusters 61a to 61d in the sentence space 44 using the latest learning data at any timing and store them in the auxiliary storage device 15.
[0030] (Distance information) 10 is an example of the distance information 38. In the distance information 38, a sentence ID (column 161), an unknown sentence (column 162), a process distance (column 163), a task cost distance (column 164), a travel cost distance (column 165), and a task content distance (column 166) are stored in association with each other. The sentence ID (field 161) is the same as the sentence ID in FIG. The unknown sentence (column 162) is the same as the unknown sentence in FIG.
[0031] The process distance (column 163) is the distance between the unknown sentence (indicated by "O") and the center of the cluster 61a in the sentence space 44 (FIG. 9). The distance can be the Euclidean distance, the Mahalanobis distance, or other distance. If the distance is greater than a predetermined threshold (e.g., the radius of the cluster 61a), there is a high possibility that the unknown sentence does not describe at least the component "process" (and so on). The work cost distance (column 164) is the distance in sentence space 44 between the unknown sentence and the center of cluster 61b. The travel distance (column 165) is the distance in sentence space 44 between the unknown sentence and the center of cluster 61c. The task distance (column 166) is the distance in sentence space 44 between the unknown sentence and the centre of cluster 61d.
[0032] The distance information 38 in Fig. 10 is distance information 38 for the document type "quotation". If Fig. 10 is distance information 38 for the document type "patent specification", for example, the process distance, task distance, travel cost distance, and task content distance change to problem distance, solution method distance, claim distance, and prior art distance, respectively.
[0033] (Importance information by component) 11 is an example of the component importance information 39. In the component importance information 39, the components (column 171) and the importance (column 172) are stored in association with each other. The components (column 171) are the components described above. The importance (column 172) is a relative weight among multiple components. The user sets the importance within the range of "0<importance≦1". The document proofreading support device 1 may automatically set the importance based on the number of characters or the number of keywords in a sentence in each component of the sample document. The component-specific importance information 39 exists for each document type.
[0034] (Score information) 12 is an example of the score information 40. In the score information 40, a sentence ID (column 181), an unknown sentence (column 182), a process score (column 183), a work cost score (column 184), a travel cost score (column 185), and a work content score (column 186) are stored in association with each other. The sentence ID (field 181) is the same as the sentence ID in FIG. The unknown sentence (column 182) is the same as the unknown sentence in FIG.
[0035] The process score (column 183) is a value obtained by multiplying the process distance in FIG. 10 by the importance level in FIG. 11 that corresponds to the process. The operation cost score (column 184) is a value obtained by multiplying the operation cost distance in FIG. 10 by the importance in FIG. 11 that corresponds to the operation cost. The travel cost score (column 185) is a value obtained by multiplying the travel cost distance in FIG. 10 by the importance level in FIG. 11 that corresponds to the travel cost. The task score (column 186) is a value obtained by multiplying the task distance in FIG. 10 by the importance level in FIG. 11 that corresponds to the task. Score information 40 also exists for each document type. In the above, the score is calculated by multiplying the distance by the importance, but this is merely an example. The score may be calculated using addition, exponential calculation, etc. In short, the greater the distance and the greater the importance, the greater the score should be.
[0036] (Score / Proofreading time conversion information) 13 is an example of the score / proofreading time conversion information 41. In the score / proofreading time conversion information 41, the score (column 191) and the proofreading time (column 192) are stored in association with each other. The score (column 191) is, for example, the “process score” mentioned above, or more generally, a value calculated (e.g., multiplied) by the importance of the component to the distance between the sentence vector and the center of the cluster for that component. The proofreading time (column 192) is the time required to proofread the mistake corresponding to the score in the unknown sentence. The user sets the proofreading time in seconds based on past cases. The document proofreading support device 1 may update the proofreading time based on the time the user actually spent on proofreading.
[0037] (Proofreading time information for each unknown sentence) 14 is an example of the unknown sentence-specific proofreading time information 42. In the unknown sentence-specific proofreading time information 42, a sentence ID (column 201), an unknown sentence (column 202), a process score (column 203a), a process proofreading time (column 203b), a task cost score (column 204a), a task cost proofreading time (column 204b), a travel expense score (column 205a), a travel expense proofreading time (column 205b), a task content score (column 206a), and a task content proofreading time (column 206b) are stored in association with each other.
[0038] The sentence ID (column 201) is the same as the sentence ID in FIG. The unknown sentence (column 202) is the same as the unknown sentence in FIG. The process score (column 203a) is the same as the process score in FIG. The process proofreading time (column 203b) is the proofreading time resulting from conversion of the process score by the score / proofreading time conversion information 41 (FIG. 13). The operation cost score (column 204a) is the same as the operation cost score in FIG. The labor cost proofreading time (column 204b) is the proofreading time resulting from the score / proofreading time conversion information 41 converting the labor cost score.
[0039] The travel expense score (column 205a) is the same as the travel expense score in FIG. The travel expense proofreading time (column 205b) is the proofreading time resulting from the score / proofreading time conversion information 41 converting the travel expense score. The task content score (column 206a) is the same as the task content score in FIG. The task content proofreading time (column 206b) is the proofreading time resulting from the score / proofreading time conversion information 41 converting the task content score.
[0040] Since the proofreading time information 42 for each unknown sentence stores not only the score but also the proofreading time, the user can know how much time it will take to proofread each unknown sentence for each component.
[0041] (Processing Procedure) The processing procedures of this embodiment are described below. There are three processing procedures: a document analysis processing procedure, a matching sentence processing procedure, and an unknown sentence processing procedure.
[0042] (Document analysis process) FIG. 15 is a flowchart of the document analysis process. In step S301, the document analysis unit 21 of the document proofreading support device 1 acquires a document. Specifically, the document analysis unit 21 acquires the document 31 from the outside via the input device 12 or from the auxiliary storage device 15.
[0043] In step S302, the document analysis unit 21 acquires a character string. Specifically, the document analysis unit 21 acquires a character string from the document 31. In step S303, the document analysis unit 21 divides the character string into sentences. Specifically, the document analysis unit 21 divides the character string into a plurality of sentences, with a period "." as a separator. At this time, the document analysis unit 21 may perform morphological analysis (part of speech analysis) and dependency analysis between words.
[0044] In step S304, the document analysis unit 21 matches the sentence with the rule. Specifically, first, the document analysis unit 21 acquires any one of the unprocessed sentences. Second, the document analysis unit 21 matches the sentence with each rule in the rule information 32 (FIG. 3) and identifies all rules that match the sentence. Third, document analysis unit 21 counts the number of rules identified in “second” step S304. The count result is “0”, “1”, “2”, “3”, . . .
[0045] In step S305, the document analysis unit 21 judges whether the sentence matches the rule. Specifically, if the count result in the "third" in step S304 is "0" (step S305 "NO"), the document analysis unit 21 proceeds to step S307, otherwise (step S305 "YES"), the document analysis unit 21 proceeds to step S306.
[0046] In step S306, document analysis section 21 registers the matching sentence information 33 (FIG. 4). Specifically, document analysis section 21 creates a record in matching sentence information 33 for the sentence to be processed.
[0047] In step S307, document analysis unit 21 registers the sentence in unknown sentence information 34 (FIG. 5). Specifically, document analysis unit 21 creates a record for the sentence to be processed in unknown sentence information 34. If, in step S305, the document analysis unit 21 determines that the sentence contained in document 31 does not match a predetermined rule, in step S307, the sentence is an unknown sentence that may contain a semantic error.
[0048] Document analysis unit 21 repeats the process from step S304 onwards for each unprocessed sentence, and ends the document analysis procedure after step S306 or S307 for the last sentence. When the document analysis procedure ends, all sentences included in document 31 acquired in step S301 are sorted into matching sentence information 33 (FIG. 4) or unknown sentence information 34 (FIG. 5) and stored.
[0049] (Matching Sentence Processing Procedure) FIG. 16 is a flow chart of the matching sentence processing procedure. In step S321, the matched sentence processing unit 22 of the document proofreading support apparatus 1 acquires a matched sentence. Specifically, the matched sentence processing unit 22 acquires an arbitrary unprocessed matched sentence from the matched sentence information 33 (FIG. 4).
[0050] In step S322, the matching sentence processing unit 22 acquires the proofreading times based on the rules. Specifically, the matching sentence processing unit 22 acquires the proofreading times of all rules that match the sentence acquired in step S321 from the rule-specific proofreading time information 35 (FIG. 6).
[0051] In step S323, the match sentence processor 22 sums up the proofreading times for each sentence. Specifically, the match sentence processor 22 sums up the proofreading times acquired in step S322.
[0052] In step S324, the matched sentence processor 22 registers the matched sentence-specific proofreading time information 36 (FIG. 7). Specifically, the matched sentence processor 22 creates a record for the sentence to be processed in the matched sentence-specific proofreading time information 36. Match sentence processor 22 repeats the process of steps S321 to S324 for each unprocessed match sentence, and ends the match sentence processing procedure when there are no more unprocessed match sentences.
[0053] (Unknown sentence processing procedure) FIG. 17 is a flow chart of the unknown sentence processing procedure. In step S341, the unknown sentence processing unit 23 of the document proofreading support device 1 acquires an unknown sentence. Specifically, the unknown sentence processing unit 23 acquires an arbitrary unprocessed unknown sentence from the unknown sentence information 34 (FIG. 5).
[0054] In step S342, the unknown sentence processing unit 23 accepts the document type. Specifically, first, the unknown sentence processing unit 23 displays on the output device 13 the document 31 acquired in step S301. Secondly, the unknown sentence processing unit 23 accepts the user's input of the document type via the input device 12. The user visually checks the document 31 and determines the document type to be input. For convenience of explanation, it is assumed here that "quotation" is input. The unknown sentence processing unit 23 may automatically determine the document type based on, for example, the title of the document 31, without waiting for an input by the user.
[0055] In step S343, the unknown sentence processing unit 23 creates a sentence vector. Specifically, the unknown sentence processing unit 23 converts the sentence acquired in step S341 into a sentence vector by the above-mentioned method.
[0056] In step S344, the unknown sentence processing unit 23 creates a sentence space 44. Specifically, first, the unknown sentence processing unit 23 creates the sentence space 44 in FIG. 9, and creates a plurality of clusters using sample documents of estimates as learning data (●). Each of the created clusters corresponds to component 1 to component 4 of the document type / component information 37 (FIG. 8). The cluster here may be the smallest sphere that envelops all the ● classified into that cluster, or may be a sphere with the center of gravity of all the ● as its center and the distance from the center of gravity to the farthest ● as its radius. The unknown sentence processing unit 23 may complete this process in advance at any timing. Secondly, the unknown sentence processing unit 23 draws the sentence vector (◯) created in step S343 in the sentence space 44.
[0057] In step S345, the unknown sentence processing unit 23 judges whether the unknown sentence includes a component. Specifically, first, the unknown sentence processing unit 23 checks whether the circle drawn in “second” of step S344 exists inside any cluster. Secondly, if the circle exists inside any of the clusters (step S345 "YES"), the unknown sentence processing unit 23 proceeds to step S346, otherwise (step S345 "NO"), the unknown sentence processing unit 23 proceeds to step S347.
[0058] In step S346, the unknown sentence processing unit 23 sets the score and the proofreading time to "0." Specifically, the unknown sentence processing unit 23 sets the score and the proofreading time of the unknown sentence acquired in step S341 to "0." Here, the unknown sentence processing unit 23 determines that the unknown sentence does not require proofreading because the unknown sentence describes any of the components that are usually included in an estimate.
[0059] In step S347, the unknown sentence processing unit 23 calculates the distance. Specifically, first, the unknown sentence processing unit 23 inputs the sentence vector created in step S343 to the component estimation model 43 for the estimate. Then, the component estimation model 43 outputs the distance between the unknown sentence (◯) and the center of each cluster in the sentence space 44. The unknown sentence processing unit 23 receives this distance. Second, the unknown sentence processing unit 23 creates a record of the distance information 38 (FIG. 10) based on the distance received in “first” of step S347. In step S347, the unknown sentence processing unit 23 estimates which of a plurality of components defined according to the type of document 31 the unknown sentence is similar to.
[0060] In step S348, the unknown sentence processing unit 23 calculates a score. Specifically, first, the unknown sentence processing unit 23 multiplies the process distance of the record created in "second" of step S347 by the importance corresponding to the process in Fig. 11 to calculate a process score. The unknown sentence processing unit 23 also calculates a task cost score, a travel cost score, and a task content score in the same manner. Secondly, the unknown sentence processing unit 23 creates a record of the score information 40 (FIG. 12) based on the score calculated in the "first" of step S348.
[0061] In step S349, the unknown sentence processing unit 23 calculates the proofreading time. Specifically, the unknown sentence processing unit 23 applies the score-proofreading time conversion information 41 in Fig. 13 to the process score of the record created in "second" of step S348 to calculate the process proofreading time. The unknown sentence processing unit 23 also calculates the work cost proofreading time, the travel expense proofreading time, and the work content proofreading time in the same manner.
[0062] In step S350, the unknown sentence processing unit 23 registers the score and the proofreading time calculated in steps S346, S348, and S349 in the proofreading time information 42 (FIG. 14). Specifically, the unknown sentence processing unit 23 creates a record in the proofreading time information 42 (FIG. 14) for each unknown sentence based on the score and the proofreading time calculated in steps S346, S348, and S349.
[0063] In step S351, the display processing unit 24 of the document proofreading support device 1 displays the proofreading time. Specifically, the display processing unit 24 uses the record created in step S324 and the record created in step S350 to display the proofreading time display screen 71 (FIG. 18) on the output device 13. Thereafter, the unknown sentence processing procedure is terminated.
[0064] FIG. 18 is an example of a proofreading time display screen 71. In the matching sentence column 72, the proofreading time and importance of the matching sentence in the document 31 are displayed. In principle, the proofreading time and importance are displayed for each rule that matches the matching sentence. In the unknown sentence column 73, the score and proofreading time of the unknown sentence in the document 31 are displayed. In principle, the score and proofreading time are displayed for each unknown sentence and for each component. Now, assume that the user inputs a check mark in the selection column of a record in the matching sentence column 72 and the unknown sentence column 73. Then, the display processing unit 24 displays the document 31 in the document column 74 and highlights (for example, underlines) the selected sentence. The document 31 here is different from the document 31 in FIG. 2.
[0065] In the text field 74, sentence SE03 is an unknown sentence. The display processing unit 24 adds a speech bubble 75 to sentence SE03. The speech bubble 75 states "closest component: process." This indicates that of the distances between sentence SE03 and each component, the "process distance" is the shortest.
[0066] In this case, for example, the following is assumed. Although the document writer was trying to write sentence SE03 about the process, it is highly likely that sentence SE03 contained a semantic error as a result of a slight lack of attention. It is highly likely that the document author wrote sentence SE03 about an event that is not related to any of the components. It is possible to proofread this sentence to a description of one of the components. In that case, considering that unknown sentence SE03 is the closest to a process, the proofreading time required to proofread unknown sentence SE03 to a sentence about a process is shortest, unless the importance of the process is extremely high.
[0067] The display processing unit 24 adds a speech bubble 76 to the sentence SE21. The speech bubble 76 describes two rules that match the sentence SE21. The display processing unit 24 receives the time that the user (document creator or proofreader) can use to proofread the document 31 from the user or obtains it from the user's schedule information or the like, and displays it as available time 77. The display processing unit 24 displays the time required to proofread all sentences included in the document 31 or sentences corresponding to the input check marks (the sum of the proofreading times) as predicted proofreading time 78. The display processing unit 24 may store the results of the user proofreading the sentences in the document column 74 in the auxiliary storage device 15.
[0068] The display processing unit 24 may display, at any position on the proofreading time display screen 71, the unknown sentence that is determined not to require proofreading in step S345.
[0069] (Effects of this embodiment) The document proofreading support device of this embodiment has the following advantages. (1) The document proofreading support device can display unknown sentences that may contain semantic errors and the components of documents that are similar to the unknown sentences. (2) The document proofreading support device can quantify the approximation between the unknown sentence and each component of the document as a distance in the sentence space. (3) The document proofreading support device can accurately estimate the components of an unknown sentence by converting the unknown sentence into a sentence vector.
[0070] (4) The document proofreading support device can update the positions and sizes of the clusters by updating the learning data. (5) The document proofreading support device can accurately identify unknown sentences that do not need to be proofread. (6) The document proofreading support device can reflect the importance of each component in the distance. (7) The document proofreading support device can display the time required to proofread an unknown sentence.
[0071] The present invention is not limited to the above-described embodiment, but includes various modified examples. For example, the above-described embodiment has been described in detail to easily explain the present invention, and is not necessarily limited to those including all of the configurations described. It is also possible to replace a part of the configuration of one embodiment with the configuration of another embodiment, and it is also possible to add the configuration of another embodiment to the configuration of one embodiment. It is also possible to add, delete, or replace a part of the configuration of each embodiment with another configuration. [Explanation of symbols]
[0072] 1 Document proofreading support device 11 Central control unit 12 Input Devices 13 Output Devices 14 Main memory 15 Auxiliary storage 21 Document Analysis Section 22 Matching Sentence Processing Unit 23 Unknown sentence processing unit 24 Display processing section 31 Documents 32 Rules Information 33 Matching sentence information 34 Unknown sentence information 35 Proofreading time information by rule 36 Proofreading time information for each matching sentence 37 Document type and component information 38 Distance Information 39 Importance information by component 40 Score Information 41 Score and proofreading time conversion information 42 Proofreading time information for each unknown sentence 43 Component Estimation Model 44 Sentence Space 71 Calibration time display screen
Claims
1. determining whether a sentence contained in the document contains a portion that matches a predetermined rule that is a criterion for detecting grammatical or typographical errors in the sentence; If the sentence contains a portion that matches the rule, the sentence being a matching sentence containing a grammatical or typological error; If the sentence does not contain a match for the rule, a document analysis unit for determining the sentence as an unknown sentence that does not contain grammatical or typological errors but may contain semantic errors; an unknown sentence processing unit that estimates which of a plurality of components defined according to the type of the document the unknown sentence is similar to; a display processing unit that displays the unknown sentence and the components similar to the unknown sentence, thereby displaying the components that can be used to proofread the unknown sentence in the shortest time; A document proofreading support device comprising:
2. The unknown sentence processing unit using a component estimation model that takes the unknown sentence as an input and outputs a distance between the unknown sentence and each of the plurality of components; 2. The document proofreading support device according to claim 1,
3. The unknown sentence processing unit converting the unknown sentence into a sentence vector, inputting the converted sentence vector into the component estimation model, and obtaining the distance from the component estimation model; 3. The document proofreading support device according to claim 2, wherein:
4. The component estimation model is calculating a distance between the transformed sentence vector and each of the plurality of clusters in a space in which sentences as learning data are classified into a plurality of clusters; 4. The document proofreading support device according to claim 3,
5. The unknown sentence processing unit determining that the unknown sentence does not need to be proofread if the unknown sentence is classified into any of the plurality of clusters; 5. The document proofreading support device according to claim 4,
6. The unknown sentence processing unit calculating a score for each of the components based on the distance and an importance level defined for each of the components; The display processing unit is displaying the calculated score in association with the unknown sentence; 6. The document proofreading support device according to claim 5,
7. The unknown sentence processing unit Converting the score into the time required for proofreading; The display processing unit is displaying the reduced time in association with the unknown sentence; 7. The document proofreading support device according to claim 6,
8. The document analysis unit of the document proofreading support device includes: determining whether a sentence contained in the document contains a portion that matches a predetermined rule that is a criterion for detecting grammatical or typographical errors in the sentence; If the sentence contains a portion that matches the rule, the sentence being a matching sentence containing a grammatical or typological error; If the sentence does not contain a match for the rule, The sentence is an unknown sentence that does not contain grammatical or typological errors, but may contain semantic errors; The unknown sentence processing unit of the document proofreading support device includes: predicting which of a plurality of components defined according to the type of the document the unknown sentence is closest to; The display processing unit of the document proofreading support device includes: displaying the unknown sentence and the constituents to which the unknown sentence is similar, thereby displaying the constituents that can proofread the unknown sentence in the shortest time; A document proofreading support method comprising the document proofreading support device.
9. Computer, determining whether a sentence contained in the document contains a portion that matches a predetermined rule that is a criterion for detecting grammatical or typographical errors in the sentence; If the sentence contains a portion that matches the rule, the sentence being a matching sentence containing a grammatical or typological error; If the sentence does not contain a match for the rule, a document analysis unit for determining the sentence as an unknown sentence that does not contain grammatical or typological errors but may contain semantic errors; an unknown sentence processing unit that estimates which of a plurality of components defined according to the type of the document the unknown sentence is similar to; a display processing unit that displays the unknown sentence and the components similar to the unknown sentence, thereby displaying the components that can be used to proofread the unknown sentence in the shortest time; A document proofreading support program that functions as a proofreader.
Citation Information
Patent Citations
Document / Sentence knowledge storage device
JP1995141371A
Document analysis device and program
JP2016212533A
Document calibration support device, document calibration support method, and document calibration support program
JP2017041164A
Document revision device and program
JP2019145023A
Verification device and method and program
JP2019194789A