Document creation support device
The document creation support device assists users in selecting document components by querying and analyzing user responses, ensuring important information is included, thereby simplifying the documentation process and preserving valuable knowledge.
Patent Information
- Application Number
- JP2020175911
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2020-10-20
- Publication Date
- 2025-10-16
- Estimated Expiration
- 2040-10-20
AI Technical Summary
Existing methods fail to effectively assist users in selecting the most important components to include in a document, leading to potential omissions and misunderstandings due to varying writing skills and the difficulty in determining what information is crucial.
A document creation support device that queries users about document components, tracks answers, and generates a draft based on statistical analysis to ensure important information is included.
Facilitates easy inclusion of important information in documents, preserving knowledge that exists only in users' memories and reducing the effort required to organize complex information.
Smart Images

Figure 0007755377000001 
Figure 0007755377000002 
Figure 0007755377000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a technique for assisting users in writing documents. [Background technology]
[0002] In various research and development activities and surveys, facts and empirical rules confirmed by employees can be obtained. This tacit knowledge is extremely valuable information for the employer who employs the employees, and it is desirable to share it with many employees, not just the employees themselves.
[0003] To achieve this, the method of creating and sharing documents such as reports is often adopted. However, when writing documents, the quality of information conveyed varies depending on the writer's writing skills. In other words, it is not easy to acquire the writing skills to convey information to the reader without excess or deficiency, which can lead to omissions and misunderstandings.
[0004] To assist with this, systems are used to assist with writing by automatically checking the grammar of documents created. However, in reality, deciding what to include in a document is often more difficult than writing the document itself. In particular, when describing something that consists of multiple components, providing detailed descriptions of all the components is extremely time-consuming and difficult to read. For this reason, it is common to document only the important components, but the writer must have sufficient skills to determine which components are important.
[0005] A specific example is the selection of elements to be included in claims in a patent application. When deciding which elements to include in claims, it is necessary to select elements that express the characteristics of the invention without excess or deficiency.
[0006] Or, for example, imagine a case where you visit an exhibition and write a report on trends in the booths you visited. At this time, you need to choose which of the things you saw and heard should be included in the report. If you include too many general points, the report will become lengthy and redundant, but if you limit the information, you may miss important points.
[0007] For example, there is a survey to select a contractor for system development. In this case, it is desirable to research and compare the characteristics of each contractor and select based on rational criteria, but even if the information is collected, it is time-consuming to organize the characteristics of each contractor. If all the contractors have common characteristics, it becomes difficult to compare them, but if there are too few characteristics, the validity of the comparison is lost.
[0008] Thus, when creating a document, it is highly necessary to select exactly what should be written in the document, but the problem is that this cannot be supported by commonly used methods for checking grammar, etc.
[0009] Patent Document 1 discloses a method for recommending a style to be used based on questions and answers as a publicly known method for assisting in determining the outline of a document to be written. With this method, a user can identify a document template simply by answering questions presented by the system. Patent Document 2 also discloses a technology that makes it easier to organize answer data to questions by searching for question sentences using conceptual keywords. With this method, templates can be searched conceptually. [Prior art documents] [Patent documents]
[0010] [Patent Document 1] Japanese Patent Publication No. 2020-35165 [Patent Document 2] Japanese Patent Application Laid-Open No. 2004-213156 Summary of the Invention [Problem to be solved by the invention]
[0011] However, the above-mentioned conventional methods cannot be expected to have the effect of selecting those components that should be particularly emphasized among the multiple components. The present invention has been made in consideration of such circumstances, and when a user memorizes information on a case having multiple components, it becomes easy to consider what to include and what to omit in the process of incorporating that information into a document. [Means for solving the problem]
[0012] One aspect of the present invention is an apparatus for generating a document draft for a user, comprising one or more processors and one or more storage devices. The one or more storage devices store a case database containing multiple cases, each of which includes one or more components. The one or more processors repeatedly query the user about one or more components selected from the case database. The query processing includes presenting the one or more components selected from the case database to the user via an output device, presenting the user with a question about whether the one or more components correspond to the content of the document draft via the output device, obtaining the user's answer to the question via an input device, and including the answer in an answer history stored in the one or more storage devices. The one or more processors select one or more components to be next subjected to the query processing from one or more components in the multiple cases that have not yet been processed based on statistics of at least some of the components of the multiple cases, and generate the document draft based on the components indicated by the answer history as corresponding to the content of the user's document draft. [Effects of the Invention]
[0013] According to the present invention, important information that is only in the user's memory can be easily reflected in a document without much effort. [Brief explanation of the drawings]
[0014] [Figure 1] 1 shows a configuration example of a first embodiment. [Figure 2] 1 shows an example of hardware implementation of the first embodiment. [Figure 3] 1 shows an example of a processing flow of the first embodiment. [Figure 4] 1 shows an example of an initial screen in the first embodiment. [Figure 5] 10 shows an example of login information according to the first embodiment. [Figure 6] 10 shows an example of a user data table according to the first embodiment. [Figure 7] 10 shows an example of a case selection screen in the first embodiment. [Figure 8] 10 shows an example of a user case table according to the first embodiment. [Figure 9] 10 shows an example of a case selection record table according to the first embodiment. [Figure 10] 1 shows an example of a multiple-choice question screen according to the first embodiment. [Figure 11] 10 shows an example of a case table according to the first embodiment. [Figure 12] An example of the relevant / non-relevant information in Example 1 is shown below. [Figure 13] 1 shows an example of a descriptive question screen in the first embodiment. [Figure 14] 1 shows an example of a document draft display in the first embodiment. [Figure 15] 10 shows an example of a statistical calculation process according to the first embodiment. [Figure 16] 10 illustrates an example of a data structure of a user example component in the first embodiment. [Figure 17] 10 shows an example of a component table according to the first embodiment. [Figure 18] 10 shows a schematic diagram of a process using a user component according to the first embodiment. [Figure 19] 10 shows an example of unanswered components extracted in Example 1. [Figure 20] 1 shows a configuration example of a second embodiment. [Figure 21] 10 shows an example of a processing flow of the second embodiment. [Figure 22] 10 shows an example of a user data table according to the second embodiment. [Figure 23]10 illustrates an example of an external information management table according to the second embodiment. [Figure 24] 10 shows an example of a case selection screen in the second embodiment. [Figure 25] 10 shows an example of a multiple-choice question screen according to the second embodiment. [Figure 26] 10 shows an example of a descriptive question screen in the second embodiment. [Figure 27] 10 shows an example of a statistical calculation process according to the second embodiment. [Figure 28] 10 shows an example of a document draft display in the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0015] In the following, when necessary for convenience, the description will be divided into multiple sections or examples, but unless otherwise specified, they are not unrelated to each other, and one is related to the other as a partial or complete modification, detail, supplementary explanation, etc. Furthermore, in the following, when the number of elements, etc. (including the number, numerical value, amount, range, etc.) is mentioned, it is not limited to that specific number, and may be more or less than the specific number, unless otherwise specified or when it is clearly limited in principle to a specific number, etc.
[0016] The document creation support device described below infers the components that correspond to a proposed sentence by asking the user questions, and generates a proposed sentence based on the inference results. A sentence represents a single piece of content and is a string of symbols including words. A sentence can be composed of only words, for example, or can include symbols other than words, such as mathematical or chemical formulas. A sentence consists of one or more sentences. A document is information to be conveyed to a person that is written in some form, such as text or images, and can be composed of text, for example, or can include images.
[0017] When writing a document, it is often difficult to decide what to document. Generally, only important components are documented, but the writer needs sufficient skill to determine which components are important.
[0018] One example is the selection of elements to be included in the claims of a patent application. When deciding which elements to include in the claims, it is necessary to select elements that express the characteristics of the invention without excess or deficiency. Another example is when visiting an exhibition and writing a report on trends in the booths you visited. At this time, you need to select which of the cases you saw and heard should be included in the report. Another example is a research report to select a contractor for system development.
[0019] When creating a document, it is highly necessary to select just the right amount of information to be included in the document. According to one embodiment of this specification, it becomes easier to reflect important information that exists only in the user's memory in the document without too much effort. As a result, knowledge that exists only in the memory of intellectual workers and tends to be scattered can be preserved as a document and reused. [Example]
[0020] As an example of a document creation support device, FIG. 1 shows an example of the logical configuration of a device that supports the review of patent claims related to chemical formulas. Users can easily create claims related to chemical formulas. A claim is a single sentence and can constitute a document. The document creation support device 101 of this embodiment includes an input / output reception unit 110, a user information determination unit 112, a case selection question generation unit 114, a case statistics calculation unit 116, a multiple-choice question generation unit 117, a written question generation unit 118, and a document proposal generation unit 120. The document creation support device 101 further includes a user information database (DB) 111, a user case DB 113, a common case DB 115, and an answer history DB 119.
[0021] The input / output receiving unit 110 receives input from the user 102. The user information DB 111 stores information about the users. The user information determining unit 112 authenticates the users and identifies the users.
[0022] The user case DB 113 stores candidate cases that can be documented for each user. The case selection question generator 114 allows the user to select one of the cases stored in the user case DB 113 when using this device. The common case DB 115 stores information on cases that are common to all users.
[0023] The case statistics calculation unit 116 calculates statistical information based on information in the common case DB 115. The multiple-choice question generation unit 117 generates multiple-choice questions based on the results of the case statistics calculation unit 116. The descriptive question generation unit 118 generates descriptive questions for information that should be confirmed by the user 102 in a descriptive manner. The answer history DB 119 stores answers by the user 102 to multiple-choice questions and descriptive questions. The document proposal generation unit 120 is a document proposal generating unit that generates a document proposal based on the answer history.
[0024] 2 shows an example of the physical implementation of the document creation support device 101. The document creation support device 101 can be implemented using, for example, a general-purpose computer. The document creation support device 101 can be implemented using, for example, a general-purpose computer.
[0025] That is, the document creation support device 101 includes a processor 211 having high computing performance, and a DRAM 212, which is a main memory device that provides a volatile temporary storage area for storing programs and data executed by the processor 211. That is, the document creation support device 101 may further include an auxiliary storage device 213 that provides a permanent information storage area using an HDD (Hard Disk Drive) or flash memory, and an interface 216, such as a serial port, for communicating data with other devices. The document creation support device 101 also includes an input device 214, such as a mouse or keyboard, for performing operations, and a monitor 215 (an example of an output device) that presents the output results of each process to the user.
[0026] The programs executed by the processor 211 and the data to be processed are loaded from the auxiliary storage device 213 to the DRAM 212. The above-mentioned functions of the document creation support device 101 may be distributed among multiple computers. In this way, the document creation support device 101 includes one or more storage devices and one or more processors.
[0027] The input / output receiving unit 110, the user information determining unit 112, the case selection question generating unit 114, the case statistics calculating unit 116, the multiple-choice question generating unit 117, the written question generating unit 118, and the document proposal generating unit 120 can be realized by the processor 201 executing programs recorded in the auxiliary storage device 213. The user information DB 111, the user case DB 113, the common case DB 115, and the answer history DB 119 can be implemented by the processor 201 executing programs that store data in the auxiliary storage device 213.
[0028] The document creation support device 101 may be a physical computer system (one or more physical computers) as described above, or may be built on a group of computing resources (plural computing resources) such as a cloud platform. The computer system or the group of computing resources includes one or more interface devices (including, for example, a communication interface device and an input / output device), one or more storage devices (including, for example, a memory (main memory) and an auxiliary storage device), and one or more processors.
[0029] When a function is realized by a program being executed by a processor, the specified processing is performed using a storage device and / or an interface device, etc., as appropriate, and therefore the function may be considered to be at least a part of the processor. Processing described using a function as the subject may also be processing performed by a processor or a system having that processor. A program may be installed from a program source. The program source may be, for example, a program distribution computer or a computer-readable storage medium (e.g., a computer-readable non-transitory storage medium). The description of each function is an example, and multiple functions may be combined into one function, or one function may be divided into multiple functions.
[0030] 3 shows the operation procedure of the document creation support device 101 according to the first embodiment. When the document creation support device 101 is started up, the user information determination unit 112 generates an initial screen, presents it via the input / output reception unit 110, and waits for access from the user 102. An example 308 of this initial screen displayed on the monitor 205 is shown in FIG.
[0031] This initial screen 308 displays a text box 401 for entering the ID of the user 102 and a text box 402 for entering a password. When the user 102 enters these using the input device 204 and presses the start use button 403, the user information determination unit 112 generates login information 309 shown in Fig. 5 based on the entered information. Fig. 5 is a diagram for explaining an example of the configuration of the login information 309. Items of the login information 309 include a user ID 501 and a password 502.
[0032] In this example, the content of text box 401 is stored as user ID 501, and the content of text box 402 is stored as password 502. Thereafter, login information 309 is sent to user information determination unit 112. User information determination unit 112 collates the input information with the user data table in user information DB 111, and authenticates user 102 (S301).
[0033] 6 shows an example of the user data table 600. FIG. 6 is a diagram for explaining an example of the configuration of the user data table 600. The items of the user data table 600 include a user ID 601, a user name 602, and a password 603.
[0034] In this example, authentication is performed by comparing the value of user ID 501 in login information 309 with the value of user ID 601 in user data table 600, and further comparing the value of password 502 in login information 309 with the value of password 603 in user data table 600. Note that although a typical method using an ID and password is used as the method used to authenticate user 102 here, any method may be used as long as it can properly identify user 102.
[0035] Next, the case selection question generation unit 114 acquires cases related to the current user 102 from the user case DB 113. The case selection question generation unit 114 displays a case selection screen 310 via the input / output receiving unit 110 to allow the user 102 to select one case from among the cases acquired (S302).
[0036] 7 shows an example of the case example selection screen 310. In this example, an image (chemical structure formula) 701 based on information from the user case example DB 113 is displayed. The user 102 can select a case example by, for example, clicking a control 702 in the case example selection screen 310, and can confirm the selected case example by pressing a button 703.
[0037] FIG. 8 shows an example of a user case table 800 in the user case DB 113, which is used to generate the case selection screen 310. FIG. 8 is a diagram illustrating an example of the configuration of the user case table 800. The user case table 800 stores user-specific cases for each user. Common cases are referenced for all users, and user cases are referenced only for that user.
[0038] The items in the user case table 800 include a user ID 801, a case ID 802, and a case expression 803. The case expression 803 stores an expression of the case content. The user ID 801 corresponds to the user ID 601 in the user data table 600. By referencing the user ID 801, the case selection question generation unit 114 can select only records that correspond to the current user 102.
[0039] The case selection question generation unit 114 can display the image 701 of this case selection screen 310 by obtaining and drawing the SMILES character string representation of the chemical formula stored as the case expression 803. In this example, the case expression 803 is a SMILES character string representation, but an appropriate type of case expression 803 is selected depending on the type of sentence to be generated. This is also true for the common case database.
[0040] Any expression may be used as long as it indicates the case in question, and for example, image information may be stored as is. Also, by displaying the user name 602, it may be possible to check whether the user is using incorrect login information. Multiple user cases may be selectable, or the selection of a user case may be omitted. The selection of a user case facilitates the selection of a common case for identifying the scope of the claim.
[0041] When a user selects a case on the case selection screen 310, the value of the corresponding case ID 802 is sent from the input / output receiving unit 110 to the case selection question generation unit 114. The case selection question generation unit 114 receives it and stores it in a case selection record table 900 stored in the answer history DB 119 shown in Fig. 9. Fig. 9 is a diagram for explaining an example of the configuration of the case selection record table 900.
[0042] The items in the case selection record table 900 include case ID 902, answer 903, and case ID 904. The value of user ID 601 is stored in user ID 901. The value of the case selected in case ID 802 is stored in case ID 902. In this embodiment, answer 903 always stores a code meaning YES. In another example, case ID 902 may store all user cases including unselected cases, and answer 903 may store a code indicating the selected case and a code explicitly indicating that the case is not subject to claim consideration. Case ID 904 stores an identifier (case ID) assigned to a series of processes performed starting from case selection.
[0043] Next, the multiple-choice question generation unit 117 repeats the question processing (S303-S305). Based on at least some statistics of the components in the common case DB 115, the multiple-choice question generation unit 117 selects a common case to be next subjected to question processing from among the common cases for which question processing has not been performed.
[0044] Specifically, the multiple-choice question generation unit 117 executes a statistics calculation process (S303). This statistics calculation process calculates statistics of the components recorded in the answer history DB 119. The statistics calculation process calculates a component identification score and a case importance score based on the information recorded in the answer history DB 119. The component identification score indicates how well the scope of the claim has been identified from the answer history up to this point. The case importance score indicates how meaningful it is to ask a question about each case stored in the common case DB 115. Details of this process will be described later.
[0045] Thereafter, the multiple-choice question generation unit 117 performs threshold processing on the component identification score (S304). That is, the multiple-choice question generation unit 117 determines whether further questions are necessary (whether to stop question processing) by evaluating whether the component identification score is smaller than a predetermined threshold. If continuation is necessary (S304: continuation necessary), the multiple-choice question generation unit 117 displays a multiple-choice question screen for the case with the highest case importance score among the unprocessed (unquestioned) cases in the common case DB 115, and asks whether the case falls within the scope (content) of the claim (invention) currently being considered by the current user 102 (S305). A case is composed of one or more components, and this question asks whether the one or more components constituting the case, as a combination of components, fall within the content of the claim, which is a draft document.
[0046] FIG. 10 shows an example of this multiple-choice question screen 312. The chemical formula of the case with the highest case importance score is drawn in an area 1001 in the center of the screen. FIG. 11 is a diagram illustrating an example of the configuration of a case table 1100 in the common case DB 115 required for this drawing. Items in the case table 1100 include a case ID 1101 and a case expression 1102. The case ID 1101 stores the identifier (ID) of the common case. The case expression 1102 stores a SMILES string corresponding to the expression of the case identified by the case ID 1101. The information of the case expression 1102 is the object to be drawn in the drawing area 1001.
[0047] The user answers whether the presented case falls within the currently assumed scope, i.e., whether it contains all of the assumed components. The input / output receiving unit 110 returns pertinence information 313 indicating the answer of the user 102 to the multiple-choice question generating unit 117. Figure 12 is a diagram illustrating an example of the configuration of this pertinence information 313. The items of the pertinence information 313 include a case ID 1201, a pertinence flag 1202, a user ID 1203, and a case ID 1204.
[0048] The case ID 1201 stores the value of the case ID 1101 in the case table 1100. The relevant flag 1202 indicates YES / NO information. YES indicates that the case is a relevant case that falls within the currently assumed target range, and NO indicates that the case is a non-relevant case that does not fall within the target range.
[0049] User ID 1203 stores the value of user ID 501 in login information 309 to identify the current user 102. Case ID 1204 stores the value of case ID 904 in case selection record table 900. The non-pertinence information 313 is returned from input / output receiving unit 110 to multiple-choice question generation unit 117 and stored in the answer history table of answer history DB 119. The non-pertinence information 313 is stored as is in this table.
[0050] After the question and answer session S305 using the multiple-choice question screen 312 is completed, the statistical calculation process S303 is executed again, followed by the threshold process S304 for the component identification score. This process is repeated until it is determined in the threshold process S304 for the component identification score that continuation is not necessary.
[0051] If it is determined in the threshold process S304 for the component identification score that continuation is unnecessary, the descriptive question generation unit 118 presents a descriptive question screen 314 via the input / output receiving unit 110 and obtains an answer to the descriptive question (S306). Fig. 13 shows an example of the descriptive question screen 314 displayed on the monitor 205. An area 1301 in the center of the screen displays an image of the identified component. A field 1302 allows writing on the image in area 1301.
[0052] The procedure for generating an image of a component in area 1301 is related to the statistical calculation process S303 and will be described later. Furthermore, section 1303 allows the user 102 to add a component. For example, a representation of a component that was not included in the target range expected by the user 102 when selecting the relevant magnetism, or a representation of a component that is not found in the shared case DB 115, may be entered. After entering these, the user 102 presses the answer button 1304, whereby the entered information is passed from the input / output receiving unit 110 to the document draft generation unit 120, and a final document draft display is generated. Furthermore, on the descriptive question screen 314, the user 102 may be able to delete some of the presented components.
[0053] 14 shows an example of the document draft display 315 in the first embodiment. A document draft 1401 based on the components is displayed in the center of the screen. A button 1402 that allows the user to download the document file of this document draft may be provided. The generation of this document draft will be described later as it is related to the statistics calculation process S303.
[0054] 15 shows details of the statistics calculation process S303 in the first embodiment. First, the case statistics calculation unit 116 acquires a list of user case components corresponding to records in the case selection record table 900 from the user case DB 118 (S1501). The case selection record table 900 stores information on user cases selected by the user 102 as targets for claim review.
[0055] 16 is a diagram illustrating an example of the configuration of a user case component table 1600 that stores user case components. The user case component table 1600 is stored in the user case DB 113. The items in the user case component table 1600 include a user ID 1601, a case ID 1602, a component ID 1603, and a presence / absence flag 1604. The user ID 1601 indicates the ID of each user, and the case ID 1602 indicates the ID of each user case. The combination of the value of the user ID 1601 and the value of the case ID 1602 is an identifier that can identify each record in the case selection record table 900.
[0056] The component ID 1603 indicates the ID of each preset component. The presence / absence flag 1604 indicates whether each component indicated by the component ID 1603 is included in this example.
[0057] 16, the user case DB 113 stores component master data 1610. Items in the component master data 1610 include a component ID 1611 and a component representation 1612. The component ID 1611 indicates a plurality of pre-set component IDs that are common to all user cases in the user case DB 113. The component representation 1612 indicates the component representation of each component indicated by the component ID 1611.
[0058] The value of the component ID 1603 in the user case component table 1600 matches the value indicated by the component ID 1611 in the component master data 1610 , and the component representation corresponding to these values can be obtained from the component representation 1612 .
[0059] Next, the case statistics calculation unit 116 separates the cases (unanswered (unasked) cases) that are not recorded in the answer history table (list of non-relevant information 313) of the answer history DB 119 from the cases (answered (asked) cases) that are recorded in the answer history table from the common case DB 115, and acquires the components (S1502).
[0060] 17 is a diagram illustrating an example of a component table 1700 in the common case DB 115. The items in the component table 1700 include a case ID 1701, a component ID 1702, and a presence / absence flag 1703. The case ID 1701 stores a value that can identify each record in the case table 1100, and matches the value of the case ID 1101 in the case table 1100.
[0061] The component ID 1702 stores the ID of each predefined component. The presence / absence flag 1703 indicates whether each component indicated by the component ID 1702 is included in this case. The common case DB 115 also stores component master data that is the same as the component master data 1610 in the user case DB 113. The value of the component ID 1702 matches the value of the component ID in this component master data, and the component representation corresponding to these values can be obtained from the component representation in the component master data.
[0062] Next, the case statistics calculation unit 116 selects components whose essentiality has not been determined (S1503). This procedure will be explained using Fig. 18. Fig. 18 shows table 1801 listing the results of obtaining a list of user case components (S1501), and table 1802 listing the results of the component table 1701 of the common case DB 115 and the response history table in terms of applicability.
[0063] The records in table 1801 show information about the components of the selected user case. Each record in table 1802 shows information about the components of each answered common case. The numbers 1 to 8 in the top row of the component inclusion information are component IDs, and the rows below that are flags that indicate whether each component is included. If the flag is 1, it means that the component is included in the case.
[0064] Among the user case components shown in table 1801, components with a presence / absence flag of 0 can be inferred to be non-essential components in this example. On the other hand, components with a flag of 1 can be said to have uncertain necessity, in the sense that they may or may not be essential. Also, in table 1802, records with Y stored in the relevant / not relevant column are common cases that are the target of this example, and components with a presence / absence flag of 1 in all of these records have uncertain necessity (it is unclear whether they are essential or not).
[0065] In order to judge these together, the user cases and the relevant common cases (common cases where the relevant / non-relevant flag is YES) are extracted into one table 1803. Then, it is not possible to judge whether a component whose flag is 1 in all records is essential or not, and it is possible to judge whether a component whose presence / absence flag in any record is 0 is not essential. In this example, it is not possible to judge whether the components with component IDs 1, 6, and 8 are essential or not.
[0066] Next, the case statistics calculation unit 116 calculates a case importance score for each of the unanswered common cases (S1504). Fig. 19 shows a schematic diagram of the result 1901 of extracting only the unanswered common cases. The degree of information that each unanswered case provides for the components whose necessity was determined in the previous step S1503 can be determined by whether the presence / absence flag of each undetermined component is 0.
[0067] For example, the record in the third row stores 1 for the first component 1902 in Fig. 19. Therefore, even if this common example is presented to the user 102 and a question is asked as to whether it is the subject of this case, and a YES answer is received, it is still not possible to determine whether the first component 1902 is essential.
[0068] Therefore, the more 0-presence flags there are for components whose essentiality is uncertain, the more information there will be when a YES answer is obtained. However, in cases where there are many 0-presence flags for components, the answer is more likely to be NO. When a NO answer is obtained, the information obtained is that in that case, "all required components are not 1." Therefore, cases where there are many 0-presence flags for components provide almost no information.
[0069] Therefore, you should select a case where the probability of getting a NO answer is statistically low and where there are many components with the presence / absence flag set to 0. In other words, you should select a case that is dissimilar (far removed) from both cases where you got a YES answer and cases where you got a NO answer.
[0070] For example, the case statistics calculation unit 116 can calculate the case importance score of each unanswered common case and determine the next common case to ask a question about based on the value. One example of a case importance score calculation method is to calculate the case importance score based on the distribution of undetermined components of non-relevant cases in the answer history.
[0071] For example, from the cases in the common case DB 115, cases where the relevant flag 1202 of the relevant information 313 in the answer history table is set to NO are selected. The calculation method further selects from among these cases the case with the smallest number of components whose essentiality is not yet determined and whose presence or absence differs from that of the current unanswered case.
[0072] This calculation method calculates the case importance score as the sum of the products of the number of components whose presence / absence differs between the selected case and the current unanswered case and the number of 0 presence / absence flags in the current unanswered case. Instead of the number of 0 presence / absence flags for all the current unanswered cases, the number of 0 presence / absence flags for components whose essentiality is not yet determined may be used.
[0073] Another calculation method calculates the case importance score based on the distribution of components of relevant and non-relevant cases included in the answer history. For example, one method collects answered cases stored in the answer history table of the common case DB 115, regardless of whether the relevant / non-relevant flag is YES / NO, and calculates the proportion of presence / absence flags of 0 and 1 for each component. Furthermore, this calculation method selects, for each unanswered case, the proportion of 0 when the component presence / absence flag is 1 and the proportion of 1 when the component presence / absence flag is 0, and calculates the case importance score by accumulating (summing up) these. This makes it possible to select the next unanswered case to ask a question about, taking into account the characteristics of answered cases whose relevant / non-relevant flags are YES.
[0074] The second calculation method can further determine the weight to be used in the sum of products based on the distribution (statistics) of the components of unanswered (unprocessed) cases. Specifically, for unanswered cases, the proportion of presence / absence flags that are 0 and the proportion of presence / absence flags that are 1 are calculated for each component. The closer these proportions are to 50%, the larger the value is set as the weight for each component. For example, a value based on information entropy (proportion of log0 + proportion of log1) can be used. This can also reflect the importance of the statistics of unanswered cases in the common case DB 115, such as the fact that a component whose "presence / absence flag is 1 in all unanswered cases in the common case DB 115" has a low amount of information.
[0075] Another calculation method may be to calculate the sum of the products of the information entropies of the components of the unanswered cases as the case importance score of each unanswered case.
[0076] Next, the case statistics calculation unit 116 calculates and adds up the confidence levels for the undetermined components to obtain a component identification score (S1505). As mentioned above, this calculation is based on the fact that components whose presence / absence flag is 1 in all cases where the relevant / absent flag in the response history table indicates YES cannot at least be said to be unnecessary. In other words, an amount indicating the degree to which such a component is estimated to be essential is calculated, and this is treated as the confidence level.
[0077] The certainty of an undetermined component can be determined based on the distribution of the undetermined component in the answer history DB 119. For example, the certainty can be calculated based on statistics of the components of non-relevant cases. As an example, the case statistics calculation unit 116 can determine whether this component is essential based on the number of cases in which the component flag is 0 in cases in which the non-relevant flag is NO (outside the target range). The more cases in which the component flag is 0, the more likely the component is essential.
[0078] Alternatively, the case statistics calculation unit 116 can calculate the proportion of cases in the common case DB 115 for which the presence / absence flag is 1 for each undetermined component, evaluate the significance of the statistical bias of cases for which the presence / absence flag in the response history table is YES, and treat the index as the certainty of the essentiality of each component. For example, if the same number of samples as the number of cases for which the presence / absence flag is YES are randomly acquired from the common case DB 115, the probability that all the presence / absence flags of the component will be 1 can be treated as the certainty of the essentiality of the component.
[0079] By adding up these confidence levels, for example by taking the sum of them, it is possible to evaluate the validity of the judgment that all of the current ``components whose essentiality is not yet determined'' are necessary, and the case statistics calculation unit 116 uses this as the component identification score.
[0080] Each of the components that are deemed essential and estimated by this procedure is associated with its component expression by referencing the component master data in the common case DB 115. This can be used to generate an example 1301 on the descriptive question screen 314 in Fig. 13 and a document proposal display 1401 on the document proposal display 315 in Fig. 14.
[0081] With the above example, users can specify the essential components of a case consisting of multiple components simply by answering yes or no multiple-choice questions. For example, a list of substructures can be created as a draft of a patent claim for a chemical formula. This allows business personnel to draft patentability documents without wasting time and prevents the loss of empirical knowledge.
[0082] Furthermore, if claims are properly written, it will lead to a reduction in the number of unnecessary rejections during patent examination, enabling more efficient utilization of intellectual property. The features of this specification can be applied to the creation of claims in forms other than chemical formulas. Furthermore, the claim proposal is an example of a document proposal, and the features of this embodiment can be applied to the creation of other types of document proposals. [Example]
[0083] An example of the system configuration of the second embodiment is shown in Fig. 20. The second embodiment supports the user 102 in selecting elements to be written when writing a report on an event such as a conference that the user attended. One of the differences between the first and second embodiments is that the information of the user 102 is acquired from an external system 2001 or the like and used. For this reason, a user information acquisition unit 2002 is included in the document creation support device 101.
[0084] The external system 2001 may be another system that manages user information, such as a system that manages master data of employee information or work information, or another document writing support system. Information from multiple external systems may also be used simultaneously.
[0085] 21 shows the operation flow of the second embodiment. One of the differences from the first embodiment is that the user information acquisition unit 2002 acquires information about the user from the user information DB 111 or the external system 2001, and stores the information as one of the previously obtained answers in the answer history DB 119 (S2101). In addition, the statistical amount calculation process S2102 is also included in the main differences from the first embodiment.
[0086] 22 is a diagram illustrating an example of the configuration of a user data table 2201, which is information acquired by the user information acquisition unit 2002 from the user information DB 111. The user data table 2201 differs from the first embodiment in that user attributes 2202 are stored in association with the user. These user attributes are information that affects the type of document the user writes, such as the user's work responsibilities and work history, and by treating these as an answer to one question, the number of questions required can be reduced. The user attributes are included in the required components in step S2701, which will be described later.
[0087] 23 is a diagram illustrating an example of the configuration of an external information management table 2301 in the reply history DB 119, into which the information acquired by the user information acquisition unit 2002 is written. The external information management table 2301 includes an external information type 2304 and an external information content 2305 corresponding to a user ID 2302 and a case ID 2303, which can identify a user and a case.
[0088] The external information type 2304 stores a code value indicating the origin of the external information. For example, a code value indicating the data acquired from the user data table 2201 is used. For information on the external system 2001, an identifier identifying the external system 2001 is used.
[0089] FIG. 24 shows an example of a case selection screen 310 for user cases in the second embodiment. Unlike in the first embodiment, each case shows information about a meeting rather than a chemical formula. In the example of FIG. 24, the case selection screen 310 shows a list 2401 of multiple meetings. Each record shows information about one meeting, specifically, the date and time, name, and participants. The case selection screen 310 can be generated in the same way as in the first embodiment. That is, the case expression 803 stored in the user case table 800 in FIG. 8 only needs to store information about the meeting. The user selects one meeting from the list 2401.
[0090] In the first embodiment, cases contained in the common case DB 115 are presented to the user to ask whether they apply to the content of the idea the user is considering. In contrast to this, in the second embodiment, components are directly used as multiple-choice questions, rather than cases. The document creation support device 101 asks the user whether the presented components apply to the content of the idea the user is currently considering.
[0091] 25 shows an example of a multiple-choice question screen 312 according to the second embodiment. The components of a case are, for example, keywords in reports. In the example of FIG. 25, the keywords 2501 are presented to the user, and the user is asked whether they are related to the user's thoughts.
[0092] The component table 1700 in Fig. 17 can be constructed using data that identifies the components of past report examples based on the results of examining past report examples according to a dictionary created in advance and examining whether each keyword appears. In the second embodiment, the user case component table 1600 in Fig. 16 is constructed so that the presence / absence flag 1604 is always 0. Furthermore, the component expression 1612 in the component master data 1610 stores expressions as keywords and expressions used when creating a document draft. This can be changed depending on the type of document draft to be created.
[0093] For example, if the document to be created is not a document consisting of text only, but a presentation slide containing drawings, the component representation 1612 stores information about what kind of drawings and characters should be inserted in what position in the slide, and what kind of design. In this way, the contents of the component representation 1612 can correspond to any form.
[0094] FIG. 26 shows an example of the descriptive question screen 314 in the second embodiment. In the second embodiment, the generated draft document (draft document) takes the form of a report. Therefore, the descriptive question generation unit 118 asks for further specificity in the descriptive question for each of the keywords (components) identified in the multiple-choice questions. In the example of FIG. 26, each component 2601 and a text box 2602 in which the user 102 writes its details are listed. The user can change the order of the descriptions using the controls on the left.
[0095] Furthermore, the descriptive question generation unit 118 selects components for which no explicit answer was obtained from the user and for which the certainty of the essentiality of the component calculated in the statistics calculation process S2102 exceeds a predetermined threshold and is sufficiently high, and includes the components in the descriptive question screen 314. The upper limit of the number of components to be selected may be set in advance. In the example of FIG. 26, the component is underlined and displayed in a manner that makes it distinguishable from the component selected by the user. If the user 102 makes an incorrect guess and the component is unnecessary, the user 102 can delete the component by pressing the button marked with an X. By showing the user 102 unselected components with a high certainty of essentiality, it becomes possible to create a more appropriate document.
[0096] To perform the display as described above, a statistical quantity calculation process S2102 different from that of the first embodiment is executed. A flowchart of an example of the statistical quantity calculation process S2102 in the second embodiment is shown in FIG. 27. One of the differences between the statistical quantity calculation process S2102 of the second embodiment and the statistical calculation process of the first embodiment is as follows. In the first embodiment, a case importance score is calculated for each case to select a question. On the other hand, in the second embodiment, a component importance score, which is a score for each component, is calculated (S2702).
[0097] In the flowchart of Fig. 27, step S1501 is the same as the statistical calculation process of the first embodiment shown in Fig. 15. The list of user case components generated in step S1501 indicates the presence / absence flag of 0 for all components. Step S1501 may be omitted.
[0098] Next, based on the information in the answer history DB 119, the case statistics calculation unit 116 extracts from the common case DB 115 cases (target cases) that include components previously determined to be essential and cases (non-target cases) that do not include those components (S2701). In the first loop, all cases are selected. Next, the case statistics calculation unit 116 selects components whose essentiality has not been determined, that is, components for which no answer has been given, from the target cases (S1503).
[0099] Next, the case statistics calculation unit 116 calculates the component importance score of the selected unanswered component. This component importance score can be calculated based on the distribution of unasked components. For example, information entropy calculated from the proportion of 0s and 1s for each component in the target case, such as "proportion of log0 + proportion of log1", can be used. This makes it possible to determine the components for which asking questions is effective.
[0100] 21, the case statistics calculation unit 116 asks the user 102 whether the unanswered components are to be included in the document proposal, starting with the component with the highest component importance score. The components selected as the components to be included in the document proposal are the components that were previously determined to be essential in step S2701.
[0101] 27, the case statistics calculation unit 116 calculates the certainty of the undetermined components and calculates the component identification score from these values. The undetermined components are unanswered, i.e., unquestioned, components in the common case DB 115. The certainty of the undetermined components can be calculated based on statistics of the undetermined components in the common cases.
[0102] For example, the case statistics calculation unit 116 selects all cases that include any of the components determined to be essential. As described in the first embodiment, the probability that the unanswered components in these cases will be 1 is compared with the probability that the unanswered components in the common case DB 115 will be 1, thereby determining the confidence level of each unanswered component.
[0103] In another example, the case statistics calculation unit 116 forms a group of cases including each component determined to be essential. Each case in the group includes at least one essential component corresponding to the group. The case statistics calculation unit 116 compares the maximum probability of the unanswered component in each group being 1 with the probability of the unanswered component being 1 in the common case DB 115. The case statistics calculation unit 116 can determine the certainty of each unanswered component based on the comparison result.
[0104] For example, if the sum of the confidence levels of the unanswered components is less than a threshold, it is determined that there is no need to continue the question. If there are any unanswered components whose confidence levels exceed the threshold, they are presented to the user 102 as recommended components.
[0105] By adopting the embodiment of Example 2, a draft report document can be created, such as the example document shown in Figure 28. This draft document has some blank spaces and insufficient information, but the user can press the download button to obtain a document file that can be edited using known software, and then edit the file to complete the report. Compared to writing a document from scratch with a blank slate, it is relatively easy to fill in blank spaces in a document, so the document can be completed in a shorter time.
[0106] The features of the embodiments of this specification can be applied to assist in the creation of documents of different types than claims in patent applications or conference reports, such as reports on booths visited at an exhibition or research reports for selecting contractors for system development.
[0107] The present invention is not limited to the above-described embodiments and includes various modifications. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those including all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, or to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is possible to add, delete, or replace part of the configuration of each embodiment with other configurations.
[0108] Furthermore, the above-mentioned components, functions, processing units, etc. may be realized in part or in whole by hardware, for example, by designing them as integrated circuits. Furthermore, the above-mentioned components, functions, etc. may be realized in software by a processor interpreting and executing a program that realizes each function. Information such as the programs, tables, and files that realize each function can be stored in a memory, a recording device such as a hard disk or SSD (Solid State Drive), or a recording medium such as an IC card or SD card.
[0109] In addition, the control lines and information lines shown are those that are considered necessary for the explanation, and do not necessarily show all the control lines and information lines in the product. In reality, it can be assumed that almost all components are interconnected. [Explanation of symbols]
[0110] 101 Document creation support device 115 Common Case DB 116 Case Statistics Calculation Department 117 Multiple Choice Question Generation Unit 1202 Applicable / inapplicable flag 1703 Presence / absence flag
Claims
1. A document creation support device, one or more processors; one or more storage devices; the one or more storage devices store a case database that stores a plurality of cases, each of which includes one or more components; the one or more processors repeatedly ask a user questions about cases selected from the case database; The question processing includes: presenting the case selected from the case database to the user by an output device; a question is presented to the user by the output device as to whether the example corresponds to the content of the document draft; obtaining an answer from the user to the question via an input device; including the answer in an answer history stored in the one or more storage devices; The one or more processors: calculating an importance score indicating the degree of significance of asking a question for each of the cases in the case database for which the question processing has not been performed, and selecting a case for which the question processing will be performed next from among the cases for which the question processing has not been performed in accordance with the importance score; Excluding components that are not included in the cases that the answer history indicates correspond to the content of the document draft of the user, and creating the document draft that includes components identified based on the confidence level among the components of the cases that the answer history indicates correspond to the content of the document draft of the user; the importance score is calculated based on a predetermined similarity between the case for which the question processing has not been performed and the case for which the question processing has been performed, or based on an amount of information of the case for which the question processing has not been performed, The similarity is based on a similarity / difference relationship between the components of the case not processed by the query processing and the components of the case not corresponding to the query processing, and a lower similarity indicates a higher importance; the amount of information is expressed by an index value whose value increases as the ratio of each of the components of the case whose query processing has not been processed to the ratio of each of the components of the case whose query processing has not been processed to the cases in the case database whose query processing has not been processed approaches 50%, A document creation support device in which the certainty of the component is based on the number of cases in which the response history indicates that the component does not match the content of the user's document proposal, or the percentage of cases in the case database that include the component.
2. A document creation support device, one or more processors; one or more storage devices; the one or more storage devices store a case database that stores a plurality of cases, each of which includes one or more components; the one or more processors repeatedly query the user about the components selected from the case database; The question processing includes: presenting the component selected from the case database to the user by an output device; presenting a question to the user via the output device as to whether the constituent element corresponds to the content of the document draft; obtaining an answer from the user to the question via an input device; including the answer in an answer history stored in the one or more storage devices; The one or more processors: calculating an importance score indicating the degree of significance of asking a question for each of the components that have not been processed by the question processing in the case database based on the amount of information of the components that have not been processed by the question processing, and selecting a component to be subjected to the question processing next from the components that have not been processed by the question processing in accordance with the importance score; creating a document draft that includes, in its content, a component that the response history indicates corresponds to the content of the document draft of the user; A document creation support device in which the amount of information is represented by an index value that becomes larger as the ratio of components that have not been processed by the question processing to those that are included in cases in the case database that contain components that correspond to the content of the document proposal approaches 50%.
3. 2. The document creation support device according to claim 1, the case database stores information on components of each of the plurality of cases; The information on the components of the case includes information on whether each predefined component is included in the case; The answer history includes cases that correspond to the content of the document draft and cases that do not correspond to the content of the document draft.
4. A document creation support device according to claim 1, Each of the plurality of instances represents a chemical formula: A document creation support device, wherein each component of each of the plurality of cases is a partial structure of a chemical formula.
5. A document creation support device according to claim 1, The response history includes cases that correspond to the content of the document proposal and cases that do not correspond to the content of the document proposal, The answer history includes undetermined components that are components included in all of the corresponding cases, The one or more processors: determining a certainty level for each undetermined component based on the number of cases in which each undetermined component is not included in the non-applicable cases in the response history; The document creation support device determines whether or not to stop the query processing based on the certainty of the undetermined constituent.
6. A document creation support method, comprising: The apparatus includes a case database storing a plurality of cases, each of the cases including one or more components; The document creation support method includes: the device repeatedly asks the user questions about the cases selected from the case database; The question processing includes: presenting the case selected from the case database to the user by an output device; a question is presented to the user by the output device as to whether the example corresponds to the content of the document draft; obtaining an answer from the user to the question via an input device; including the answer in an answer history; The device, calculating an importance score indicating the degree of significance of asking a question for each of the cases in the case database for which the question processing has not been performed, and selecting a case for which the question processing will be performed next from among the cases for which the question processing has not been performed in accordance with the importance score; excluding components that are not included in the case that the answer history indicates corresponds to the content of the document draft of the user, and creating and determining the document draft that includes components identified based on the confidence level among the components of the case that the answer history indicates corresponds to the content of the document draft of the user; the importance score is calculated based on a predetermined similarity between the case for which the question processing has not been performed and the case for which the question processing has been performed, or based on an amount of information of the case for which the question processing has not been performed, The similarity is based on a similarity / difference relationship between the components of the case not processed by the query processing and the components of the case not corresponding to the query processing, and a lower similarity indicates a higher importance; the amount of information is expressed by an index value whose value increases as the ratio of each of the components of the case whose query processing has not been processed to the ratio of each of the components of the case whose query processing has not been processed to the cases in the case database whose query processing has not been processed approaches 50%, A document creation support method in which the certainty of the component is based on the number of cases in which the response history indicates that the component does not match the content of the user's document proposal, or the percentage of cases in the case database that include the component.
7. A document creation support method, comprising: The apparatus includes a case database storing a plurality of cases, each of the cases including one or more components; The document creation support method includes: the device repeatedly asks the user questions about the components selected from the case database; The question processing includes: presenting the component selected from the case database to the user by an output device; presenting a question to the user via the output device as to whether the constituent element corresponds to the content of the document draft; obtaining an answer from the user to the question via an input device; including the answer in an answer history; The device, calculating an importance score indicating the degree of significance of asking a question for each of the components that have not been processed by the question processing in the case database based on the amount of information of the components that have not been processed by the question processing, and selecting a component to be subjected to the question processing next from the components that have not been processed by the question processing in accordance with the importance score; creating a document draft that includes, in its content, a component that the response history indicates corresponds to the content of the document draft of the user; A document creation support method in which the amount of information is expressed as an index value that becomes larger as the ratio of components that have not been processed by the question processing to those that are included in cases in the case database that contain components that correspond to the content of the document proposal approaches 50%.
Citation Information
Patent Citations
Interactive diary generating device
JP2001134559A
Electronic document producing device, electronic document producing method, and program for making computer execute its method
JP2004213156A
Text generation device and text generation method
JP2012141832A
Legal document creation support system, legal document creation support method and program
JP2020035165A
Document generation support system, document generation support method, and computer program
JP2020140369A