Document instruction extraction system, document instruction extraction method, and document instruction extraction program
The intra-document instruction extraction system addresses the challenge of varying terminologies and structures in industrial documents by using meta information and representative word dictionaries to accurately extract and classify instruction meanings, enhancing the efficiency of design change analysis.
Patent Information
- Application Number
- JP2022093076
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-06-08
- Publication Date
- 2025-10-02
- Estimated Expiration
- 2042-06-08
AI Technical Summary
Existing technologies struggle to efficiently extract and interpret instructions in industrial product design documents due to varying terminologies, organizational structures, and rapid change cycles, which complicates the identification of representative words and trends in design changes.
An intra-document instruction extraction system that utilizes a processor and memory unit to identify instruction content and representative words by employing meta information and representative word dictionaries, enabling accurate extraction and classification of instruction meanings.
The system effectively extracts and identifies representative words from industrial product design documents, improving the efficiency of analyzing design changes and reducing the time required for manual interpretation.
Smart Images

Figure 0007748337000001 
Figure 0007748337000002 
Figure 0007748337000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a technique for extracting and processing instructions written in documents such as specifications for designing industrial products. [Background technology]
[0002] To improve quality and reduce costs in the design and manufacturing of industrial products, design teams analyze past specifications to identify product parts that have undergone repeated design changes, shape features, types of changes, and other trends. This analysis requires someone with product knowledge to read through all of the instructions in the specifications, interpret their meaning, and classify them, which takes a significant amount of time.
[0003] Regarding the analysis of documents in the medical field, there is a known technology that targets medical observation documents such as electronic medical records and report systems, and identifies problems from free-text text using rules that define categories and keywords (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2009-245232 Summary of the Invention [Problem to be solved by the invention]
[0005] For example, the method in Patent Document 1 targets the medical field, constructing judgment rules consisting of categories and keywords for problems, and then sequentially editing and expanding them using an editor. In the medical field, medical terminology is relatively standardized even across different medical settings, so this method makes it possible to expand a universal model.
[0006] On the other hand, in the field of industrial product design and manufacturing, the shapes and component configurations of objects vary enormously, there are many proper nouns such as product models and component names, the design and manufacturing processes span multiple organizations, and product change cycles tend to be rapid. For this reason, the meaning of instructions described in documents such as specifications can vary depending on the design site, product, etc., and the same term may be used with completely different meanings. Furthermore, because many organizations are involved, the structure of specifications and other documents can differ depending on the process, and the location and method of instructions within the specifications can also vary widely. This can also be seen in documents outside the field of industrial product design and manufacturing.
[0007] The present invention has been developed in consideration of the above circumstances, and its purpose is to provide a technology that can appropriately extract the instruction content contained in a document and appropriately identify representative words that are representative words corresponding to the words contained in the instruction content. [Means for solving the problem]
[0008] In order to achieve the above-mentioned object, one aspect of an intra-document instruction extraction system is an intra-document instruction extraction system that extracts instruction content from a document containing the instruction content and identifies a representative word that corresponds to a phrase contained in the instruction content, the intra-document instruction extraction system having a processor and a memory unit, the memory unit storing one or more of the documents, instruction content identification information for identifying the instruction content from the document, and representative word identification information for identifying the representative word from the instruction content, and the processor identifies the instruction content from the document based on the instruction content identification information, and identifies a representative word from the instruction content based on the representative word identification information. [Effects of the Invention]
[0009] According to the present invention, it is possible to properly extract instructions contained in a document and properly identify representative words corresponding to words contained in the instructions. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a configuration diagram of an example of a product according to the first embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of a specification according to the first embodiment. [Figure 3] FIG. 3 is a diagram showing the overall configuration of the document instruction extraction system according to the first embodiment. [Figure 4] FIG. 4 is a functional configuration diagram of the document instruction extraction system according to the first embodiment. [Figure 5] FIG. 5 is a diagram showing the configuration of the meta information cue word dictionary according to the first embodiment. [Figure 6] FIG. 6 is a diagram showing the configuration of the instruction sentence cue word dictionary according to the first embodiment. [Figure 7] FIG. 7 is a diagram showing the configuration of a representative word dictionary relating to a design object according to the first embodiment. [Figure 8] FIG. 8 is a diagram showing the configuration of a representative word dictionary relating to CAD shapes according to the first embodiment. [Figure 9] FIG. 9 is a diagram showing the configuration of a representative word dictionary relating to types of changes according to the first embodiment. [Figure 10] FIG. 10 is a flowchart of the reference sentence extraction and group word identification process according to the first embodiment. [Figure 11] FIG. 11 is a flowchart of the meta information / instruction sentence extraction process according to the first embodiment. [Figure 12] FIG. 12 is a flowchart of the directive statement division process according to the first embodiment. [Figure 13] FIG. 13 is a flowchart of the group word identification process according to the first embodiment. [Figure 14] FIG. 14 is a flowchart of the determination process according to the first embodiment. [Figure 15] FIG. 15 is a flowchart of an example of a determination subprogram process according to the first embodiment. [Figure 16] FIG. 16 is a flowchart of another example of the determination subprogram processing according to the first embodiment. [Figure 17]FIG. 17 is a diagram showing an example of a determination result of a representative word related to a design object according to the first embodiment. [Figure 18] FIG. 18 is a diagram showing an example of a determination result of a representative word related to a CAD shape according to the first embodiment. [Figure 19] FIG. 19 is a diagram showing the configuration of the instruction content list according to the first embodiment. [Figure 20] FIG. 20 is a diagram showing the layout of a summary table according to the first embodiment. [Figure 21] FIG. 21 is a configuration diagram of the tally screen according to the first embodiment. [Figure 22] FIG. 22 is a functional configuration diagram of the document instruction extraction system according to the second embodiment. [Figure 23] FIG. 23 is a diagram showing the configuration of a CAD model display screen according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] The following description of the embodiments will be given with reference to the drawings. Note that the embodiments described below do not limit the scope of the invention as claimed, and not all of the elements and combinations thereof described in the embodiments are necessarily essential to the solution of the invention.
[0012] In this embodiment, items in dictionaries and tables are enclosed in <>, and specific examples of data values are enclosed in " " and written in italics in the figures.
[0013] (First embodiment) Before describing the intra-document instruction extraction system 101, an example of a product that is the target of the document (specification) used in the intra-document instruction extraction system 101, and a specification as an example of a document related to the product will be described.
[0014] FIG. 1 is a configuration diagram of an example of a product according to the first embodiment.
[0015] Product 300 is assembled from product main body 301 and component 302 (also referred to as component A). Component 302 is attached with harness 303. In product 300, rib 307 is formed on the top of product main body 301. Component mounting portion 305 to which component 302 is attached is formed on front surface 304 of product main body 301. In addition, boss 306 is provided on the top of product main body 301 for fixing harness 303 of component 302.
[0016] Fig. 2 is a diagram illustrating an example of a specification according to the first embodiment. The specification in Fig. 2 illustrates an example of a specification in which instructions (instruction contents) for design changes to the product 300 shown in Fig. 1 are written.
[0017] The specification 201 includes a main title display area 202, an issue date display area 203, a specification number display area 204, a product model display area 205, a change title display area 206, an instruction statement description range 207, a summary display area 208, and a page display area 209.
[0018] The main title display area 202 contains the main title "Specifications." The issue date display area 203 contains the subtitle "Issue Date" and the issue date of the specifications 201, "2022 / 04 / 01." The specification number display area 204 contains the subtitle "Specifications Number" and the specification number, "K1001." The product model display area 205 contains the subtitle "Product Model" and the model of the product 300, "KA1." The change title display area 206 contains the subtitle "Changes." The instruction statement field 207 contains one or more instruction statements corresponding to the changes to the design of the product 300. In the example of FIG. 2, the instruction statement field 207 contains instruction statements indicating each instruction, in bulleted order, with a predetermined first character at the beginning of each line. The deadline display area 208 displays the subtitle "Date" and the deadline for the product 300, "2022 / 05 / 01." The page display area 209 displays the word "Page" and the page number, "1 / 2."
[0019] When it is desired to grasp the trend of design changes based on multiple specifications such as those shown in Figure 2, the document instruction extraction system 101 performs a process of identifying instruction sentences in the specifications, and identifying and classifying representative words that express the original meaning of the instructions from the instruction sentences.
[0020] In this process, simply extracting words (characteristic words) contained in the instruction sentence and comparing them with a dictionary may not be enough to identify and classify the representative words that originally meant the sentence.
[0021] For example, if the instruction for a design change is "Expand the R of the upper part of the fixing part of Part A," even if you try to identify the representative word for the part of the product that this instruction applies to, the words in the instruction do not contain information that the part is the front of the product body that the instruction is intended to refer to. Similarly, if the instruction is "Expand the R of the base of the harness boss of Part A," even if you try to identify the representative word for the part of the product that this instruction applies to, the words in the instruction do not contain information that the part is the upper part of the product body that the instruction is intended to refer to. Furthermore, although these two instructions contain the same word, "Part A," the parts of the product 300 that these instructions apply to are different.
[0022] Furthermore, although the instruction "Enlarge upper R of fixed part of part A" contains the word "upper part," this upper part does not refer to the upper part of the main body of product 300. Therefore, it is not possible to identify a representative word that indicates the target part of product 300 from the words and phrases in the instruction alone.
[0023] As another example, when trying to identify a representative word that represents the type of design change from the instruction for a design change, in the instruction "Enlarge the R of the upper part of the fixed part of part A," "enlarge R" means, for example, enlarging the convex R of the fixed part of part A, whereas in the instruction "Enlarge the R of the base of the boss for part A harness," it means enlarging the concave R of the base of the boss.Even though the same phrase "enlarge R" is used, these are different types of design changes, and it is not possible to directly identify a representative word that corresponds to the type of design change from the phrase in the instruction.
[0024] Furthermore, in the design of industrial products, the relative positions of the shapes of parts or portions of the product may change significantly depending on minor changes to the model or the manufacturing period. For example, boss 306 for securing harness 303 of part A of product 300 may be positioned on the bottom surface of product body 301 at a certain point in time. If the position of boss 306 is changed in this way, the intended content of the instruction "Expand the radius of the boss base for harness A" must be recognized as a design change related to the bottom surface of the body, and a representative term must be specified.
[0025] In this way, in order to identify and classify the appropriate representative word from a directive sentence, it is necessary to recognize the original intended content according to the combination of words in the directive sentence, their order, etc. It is also necessary to recognize that even if the same words are used, different intentions may be intended depending on other conditions.
[0026] The document instruction extraction system 101 takes into consideration the above-mentioned circumstances and performs processing to appropriately identify and classify representative words that express the original meaning of instructions from instruction sentences.
[0027] Next, the document instruction extraction system 101 will be described in detail.
[0028] (Hardware configuration) FIG. 3 is a diagram showing the overall configuration of the document instruction extraction system according to the first embodiment.
[0029] The document instruction extraction system 101 is configured by a computer such as a PC (Personal Computer) or a general-purpose server. The document instruction extraction system 101 includes a CPU (Central Processing Unit) 11 as an example of a processor, a RAM (Random Access Memory) 12, a ROM (Read Only Memory) 13, a storage device 14, a communication interface (I / F) 15, a recording medium reading device 17, an input device 18, and a display device 19. The CPU 11, RAM 12, ROM 13, storage device 14, communication I / F 15, recording medium reading device 17, input device 18, and display device 19 are connected to each other so as to be able to communicate with each other via a bus, for example.
[0030] The CPU 11 executes programs stored in the RAM 12 and ROM 13 to perform various calculations and centrally control each part of the document instruction extraction system 101. The RAM 12 functions as a work area for storing the programs executed by the CPU 11 and necessary information. The ROM 13 stores the BIOS (Basic Input Output System).
[0031] The storage device 14 is an example of a storage unit, and is a non-volatile storage device such as a hard disk drive (HDD) or a solid state drive (SSD), and stores various types of information. In this embodiment, the storage device 14 stores an intra-document instruction extraction program 20 and one or more document files 21. The document file 21 is, for example, a file of a document (e.g., a specification) related to the design or manufacture of a product, and may include text data and images created using document editing software or the like, or may be obtained by performing optical character recognition (OCR) on image data of the specification to convert characters in the image data into text data.
[0032] The communication I / F 15 is an interface for communicating with other devices via a network such as the Internet (not shown). The recording medium reader 17 is a device that reads data from a recording medium 16 such as a Blu-ray (registered trademark) Disc, a DVD (Digital Versatile Disc), or a CD (Compact Disc).
[0033] The input device 18 is, for example, a keyboard, a mouse, etc., and accepts information input by a user. The display device 19 is, for example, a liquid crystal display, an organic EL display, etc., and displays and outputs a user interface including various types of information.
[0034] In the document instruction extraction system 101, the document instruction extraction program 20 set up in the storage device 14 may be obtained by downloading it from a device on the Internet via the communication I / F 15, or it may be read by the recording medium reading device 17 from the recording medium 16 on which the document instruction extraction program 20 is recorded.
[0035] 3 shows an example in which the intra-document instruction extraction system 101 is configured by one computer, but the present invention is not limited to this, and the intra-document instruction extraction system 101 may be configured by multiple server devices. When the intra-document instruction extraction system 101 is configured by multiple server devices in this way, the intra-document instruction extraction program 20 may be executed in cooperation with the multiple server devices. In this case, the intra-document instruction extraction program 20 corresponds to a collection of programs that are distributed and set up on multiple server devices.
[0036] Next, the functional configuration of the document instruction extraction system 101 will be described.
[0037] (Functional configuration) FIG. 4 is a functional configuration diagram of the document instruction extraction system according to the first embodiment.
[0038] The intra-document instruction extraction system 101 includes a dictionary data storage unit 107, a determination subprogram storage unit 109, a file input unit 102, a text data extraction unit 103, a directive sentence extraction unit 104, a target word identification unit 105, a counting unit 110, and an output unit 111. The dictionary data storage unit 107 and the determination subprogram storage unit 109 are mainly configured by the RAM 12 and / or the storage device 14, and the file input unit 102, the text data extraction unit 103, the directive sentence extraction unit 104, the target word identification unit 105, the counting unit 110, and the output unit 111 are mainly configured by the CPU 11 executing the intra-document instruction extraction program 20.
[0039] The dictionary data storage unit 107 stores a meta information cue word dictionary D1, a demonstrative cue word dictionary D2, and one or more representative word dictionaries D3. Details of these dictionaries will be described later. The determination subprogram storage unit 109 stores one or more determination subprograms (an example of a determination program) for determining representative words. The determination subprograms are called and executed by the second determination unit 108.
[0040] The file input unit 102 reads the document file 21 to be processed from the storage device 14 and passes the data of the document file to the text data extraction unit 103. The document file 21 may also be read from an external device via a network. The text data extraction unit 103 extracts text data from the data of the document file 21 and passes it to the instruction sentence extraction unit 104.
[0041] The instruction sentence extraction unit 104 extracts instruction sentences of the instruction content in the document file based on the text data of the document file, the meta information clue word dictionary D1, and the instruction sentence clue word dictionary D2, generates a data set including a list of instruction sentences (instruction sentence list) and meta information data, and passes it to the reference word identification unit 105.
[0042] The target word identification unit 105 has a first determination unit 106 and a second determination unit 108. The first determination unit 106 searches for characteristic words in instruction sentences based on the instruction sentence list and the target word dictionary D3, and performs processing to identify target word candidates. When it is necessary to execute a determination subprogram to determine whether a target word candidate is a target word, the second determination unit 108 executes the determination subprogram to determine whether the candidate is a target word, and based on the result, generates a list of target words (target word list), and passes the target word list, instruction sentence list, and meta information data to the aggregation unit 110.
[0043] The counting unit 110 passes a data set that compiles the meta information data, the instruction sentence list, and the representative word list, and data counted based on this data set (counted data: see FIG. 20) to the output unit 111. The output unit 111 stores the passed data set and counted data in the storage device 14, and outputs to the display device 19 a screen of an instruction content list (see FIG. 19) including information on the data set, and a counting screen (see FIG. 21) based on the counted data.
[0044] (dictionary) Next, the meta information cue word dictionary D1, the instruction cue word dictionary D2, and the representative word dictionary D3 will be described in detail.
[0045] FIG. 5 is a diagram showing the configuration of the meta information cue word dictionary according to the first embodiment.
[0046] The meta information clue word dictionary D1 includes the items <specification number start word>, <specification number format>, <specification date start word>, and <specification date format>.
[0047] <Specification number start word> stores a word that indicates the start of the specification number in the specification. In the example of Figure 5, <Specification number start word> stores "Specification number", "Specification #", etc. <Specification number format> stores the format of the specification number. In the example of Figure 5, "A0001" is stored as the specification format. <Specification date start word> stores a word that indicates the start of the specification date in the specification. In the example of Figure 5, <Specification date start word> stores "Creation date", "Registration date", etc. <Specification date format> stores the format of the specification date. In the example of Figure 5, <Specification date format> stores "0000 / 00 / 00".
[0048] FIG. 6 is a diagram showing the configuration of the instruction sentence cue word dictionary according to the first embodiment.
[0049] The instruction sentence clue word dictionary D2 is an example of instruction content identification information and individual instruction identification information, and includes the items of <specification start word>, <specification end word>, <instruction sentence description range start word>, <instruction sentence description range end word>, <line start character type>, <line number>, and <line start character>.
[0050] <Specification start word> stores a word that indicates the start position of the range in the specification to be processed. In the example in Figure 6, <specification start word> stores "specification", "specification", etc. <specification end word> stores a word that indicates the end position of the range in the specification to be processed. In the example in Figure 6, <specification end word> stores "page", etc.
[0051] <Directive statement range start word> stores a word that indicates the start position of the range in the specification where the directive is written (directive statement range). In the example in Figure 6, <Directive statement range start word> stores "Change content". <Directive statement range end word> stores a word that indicates the end position of the directive statement range in the specification. In the example in Figure 6, <Directive statement range end word> stores "Important period".
[0052] <Line start character type> stores the character type that indicates the start of a line of an itemized directive in the directive description range. In the example of Figure 6, <Line start character type> stores "Num Type1" indicating that the line start character is type [n], "Num Type2" indicating that the line start character is type (n), etc. <Line number> stores the line number of the directive in the directive description range. In the example of Figure 6, <Line number> stores numbers from 1 to a predetermined number in order for each line start character type. <Line start character> stores the line start character that is displayed in the line number of the character line start type.
[0053] For example, in the example in Figure 6, the following are stored as "〈Line start character type〉,〈Line number〉,〈Line start character〉": "'Num Type1', '1', '[1]'", "'Num Type1', '2', '[2]'", ..., "'Num Type2', '1', '(1)'", "'Num Type2', '2', '(2)'", ...
[0054] In this embodiment, the target word dictionary D3 (an example of target word identification information) includes multiple target word dictionaries for target words of different categories, including a target word dictionary D3-1 related to the design object, a target word dictionary D3-2 related to CAD shapes, and a target word dictionary D3-3 related to types of changes.
[0055] FIG. 7 is a diagram showing the configuration of a representative word dictionary relating to a design object according to the first embodiment.
[0056] The representative word dictionary D3-1 is a dictionary for identifying representative words related to the design object (an example of a category), and includes items such as <process number>, <representative word>, <judgment code>, <characteristic word (1)>, <characteristic word (2)>, etc. In the example of Fig. 7, there are two items for storing characteristic words, <characteristic word (1)> and <characteristic word (2)>, but if there are many characteristic words, there may be three or more items.
[0057] The <Processing Number> field stores the serial number (Processing Number) for the record in the target word dictionary D3-1. The <Target Word> field stores a target word candidate (target word). The <Determination Code> field stores determination process information related to the determination process for determining whether the target word candidate for the <Target Word> corresponding to the record is the target word. The determination process information may include program-specific information that identifies the determination subprogram used in the determination process and information on the arguments to be passed to the determination subprogram. For example, the "no-" in "No-Part A, Part B" in the <Determination Code> field of a record with a processing number of 2 is program-specific information that indicates the determination subprogram "Sub-n," and "Part A, Part B" are the arguments. The determination subprogram "Sub-n" is a program that executes a determination process to determine whether the directive does not contain a string corresponding to the argument. In addition, the "date-after" in "date-after-20210501" in the <Decision Code> of the record with processing number 4 is program specific information that indicates the judgment subprogram "date-after," and "20210501" is the argument. The judgment subprogram "date-after" is a program that executes a judgment process to determine whether or not a directive statement conforms to the date specified in the argument. If no data is set in the <Decision Code>, the word in the <Representative Word> becomes the representative word without performing the judgment process.
[0058] <Feature Word (1)> and <Feature Word (2)> store one or more feature words that are the conditions for determining a record as a representative word. In this embodiment, one of the conditions for a word in <Representative Word> of a record to become a representative word is that the instruction sentence contains a combination of one of the feature words in <Feature Word (1)> and one of the feature words in <Feature Word (2)>.
[0059] For example, in the fourth record in Figure 7, i.e., the record with a <processing number> of "4," the <representative word> stores "bottom of main body," the <determination code> stores "data-after-20210501," the <characteristic word (1)> stores "Part A," and the <characteristic word (2)> stores "harness, code." This record indicates that if the combination of Part A and harness, or Part A and code, is included in the directive, and the determination subprogram "date-after" indicated by "data-after" performs a determination process with "20210501" as an argument and determines that the combination is compatible, then the representative word is identified as "bottom of main body."
[0060] Here, a specific product orientation can be expressed in a variety of ways, such as by expressions based on the orientation of the product in its normal state, by expressions based on the orientation of the product on design drawings, or by expressions based on the orientation during the manufacturing process. Therefore, even if a phrase indicating the same direction is used, the actual meaning of the direction may be completely different. Therefore, the target word dictionary may include multiple correspondences for the same target word, including a correspondence relationship with a feature word (first feature word) related to the orientation of the product in its normal state, a correspondence relationship with a feature word (second feature word) related to the orientation on the design drawings, and a correspondence relationship with a feature word (third feature word) related to the orientation during the manufacturing process. For example, the target word dictionary may store at least multiple records, including a record including a feature word related to the orientation of the target word for the upper part of the main body in its normal state, a record including a feature word related to the orientation of the upper part of the main body on the design drawings, and a record including a feature word related to the orientation of the upper part of the main body during the manufacturing process. This makes it possible to identify the same target word regardless of the expression used to express the direction.
[0061] FIG. 8 is a diagram showing the configuration of a representative word dictionary relating to CAD shapes according to the first embodiment.
[0062] The representative word dictionary D3-2 is a dictionary for identifying representative words related to CAD shapes (an example of a category), and differs from the representative word dictionary D3-1 only in that the data registered therein relates to representative words related to CAD shapes; the item structure is the same.
[0063] FIG. 9 is a diagram showing the configuration of a representative word dictionary relating to types of change according to the first embodiment.
[0064] The representative word dictionary D3-3 is a dictionary for identifying representative words related to change types (an example of a category), and differs from the representative word dictionary D3-1 only in that the data registered therein relates to representative words related to change types in the design content of the product; the item structure is the same.
[0065] (Processing operation) Next, the processing operation of the document instruction extraction system 101 will be described.
[0066] (Instruction extraction and representative word identification processing) First, the process of extracting directives and identifying representative words will be explained.
[0067] FIG. 10 is a flowchart of the reference sentence extraction and group word identification process according to the first embodiment.
[0068] When the intra-document instruction extraction system 101 is in a standby state and receives a trigger to start processing (for example, a processing start instruction from the user via the input device 18), the file input unit 102 accepts a selection from the user of one or more document files to be processed, reads the selected document files from the storage device 14, and passes the document files as data to the text data extraction unit 103 (S101).
[0069] Next, the text data extraction unit 103 extracts text data from the data of the document file and passes it to the instruction statement extraction unit 104 (S102).
[0070] Next, the instruction sentence extraction unit 104 executes a meta information / instruction sentence extraction process (see FIG. 11) to extract meta information of the document file and instruction sentences in the document file based on the text data of the document file, the meta information clue word dictionary D1, and the instruction sentence clue word dictionary D2, and passes a data set including a list of the extracted instruction sentences (instruction sentence list) and meta information data to the target word identification unit 105 (S103).
[0071] Next, the target word identification unit 105 searches for characteristic words based on the instruction sentence list, meta information data, and target word dictionary D3, executes a target word identification process (see FIG. 13) to identify target words, generates a target word list, and passes the target word list, instruction sentence list, and meta information data to the aggregation unit 110 (S104). In this embodiment, the target word identification unit 105 executes the process of step S104 for all document files to be processed, for each category of target words for each document file. For example, the target word identification unit 105 outputs the target word list, instruction sentence list, and meta information data obtained for all document files to be processed as a table-format file corresponding to the instruction content list R1 (see FIG. 19).
[0072] Next, the counting unit 110 reads the tabular file output from the target word identification unit 105, counts the number of target words identified for each target word category, creates a count table R2 (see Figure 20), stores it in the storage device 14, and outputs it as a tabular file (S105).
[0073] Next, the output unit 111 draws a graph showing the identified number of representative words for each representative word category based on the summary table R2, outputs the graph to the storage device 14 as an image file, and displays the summary screen 401 (see Figure 21) including the graph on the display device 19 (S106), and terminates the processing.
[0074] (Meta-information and instruction sentence extraction processing) Next, the meta-information / instruction statement extraction process in step S103 will be described.
[0075] FIG. 11 is a flowchart of the meta information / instruction sentence extraction process according to the first embodiment.
[0076] First, the instruction extraction unit 104 receives the text data of the document file from the text data extraction unit 103 and stores the text data (S201).
[0077] Next, the instruction sentence extraction unit 104 extracts and stores the specification number from the text data based on the data of <specification number start word> and <specification number format> in the meta information clue word dictionary D1 (S202).
[0078] Next, the instruction sentence extraction unit 104 extracts and stores the specification date from the text data based on the data of <specification date start word> and <specification date format> in the meta information clue word dictionary D1 (S203).
[0079] Next, the instruction sentence extraction unit 104 extracts and stores the instruction sentence description range from the text data based on the data of the instruction sentence description range start word and the instruction sentence description range end word in the instruction sentence clue word dictionary D2 (S204).
[0080] Next, the instruction sentence extraction unit 104 performs an instruction sentence division process (see Figure 12) to divide the contents of the instruction sentence description range into individual instruction sentences based on the data of the <line start character type>, <line number>, and <line start character> in the instruction sentence clue word dictionary D2 (S205).
[0081] Next, the instruction sentence extraction unit 104 stores the stored meta-information data such as the specification number and specification date, and the data list of the instruction sentences, as data available to the group word identification unit 105 (S206), and ends the process.
[0082] (Directive division processing) Next, the instruction statement division process in step S205 will be described.
[0083] FIG. 12 is a flowchart of the directive statement division process according to the first embodiment.
[0084] When the instruction statement extraction unit 104 receives a trigger to start processing, it resets the value of the extracted flag to "False" and prepares an instruction statement list as an empty array variable (S211). Here, the instruction statement list has, for example, a pair of a number and a corresponding instruction statement as one element.
[0085] Next, the instruction sentence extraction unit 104 assigns the instruction sentence description range stored in step S204 to variable t1, and assigns the data of <type of first character> of the instruction sentence cue word dictionary D2 to array variable L1 so as not to overlap (S212). As a result, all types of first character types registered in <type of first character> of the instruction sentence cue word dictionary D2 are stored in array variable L1.
[0086] Next, the directive extraction unit 104 repeats the processing of loop A (S213 to S220) by substituting the values of the array variable L1 one by one into the variable n.
[0087] In the processing of loop A, the instruction extraction unit 104 checks the value of the extracted flag (S213), and if the extracted flag is "True" (S213: Yes), it indicates that the instruction has been divided, so it exits loop A and ends the processing.
[0088] On the other hand, if the extracted flag is “False” (S213: No), the instruction sentence extraction unit 104 assigns all lists of pairs of line numbers in the <line number> and line start characters in the <line start character> corresponding to the <line start character type> = variable n in the instruction sentence clue word dictionary D2 to the array variable L2 (S214).
[0089] Next, the instruction extraction unit 104 repeats the processing of loop B (S215 to S220) by assigning the line number of <line number> to variable m and the first character of <line first character> to variable s, starting from <line number> = 1 for the array variable L2.
[0090] In the processing of loop B, the directive extraction unit 104 searches for variable s in variable t1 and determines whether variable s is present in variable t1 (S215). As a result, if variable s is not present in variable t1 (S215: No), the directive extraction unit 104 determines whether the value of variable m is 1 (S221). If it is determined that variable m is 1 (S221: Yes), that is, if the processing is the first iteration of loop B, the directive extraction unit 104 exits from the processing of loop B. On the other hand, if it is determined that variable m is not 1 (S221: No), that is, if the processing is the second or subsequent iteration of loop B, the directive extraction unit 104 adds variable t1 to the directive list as the (m-1)th directive (S220) and exits from the processing of loop B.
[0091] On the other hand, if variable s is found in variable t1 (S215: Yes), the instruction extraction unit 104 sets the position where variable s is found in variable t1 as p, assigns "True" to the extracted flag, and assigns everything from p to the end of variable t1 to variable t2 (S216).
[0092] Next, the instruction extraction unit 104 determines whether the value of variable m is 1 or not (S217), and if it determines that variable m is 1 (S217: Yes), that is, if this is the first processing iteration of loop B, it proceeds to step S219, whereas if it determines that variable m is not 1 (S217: No), that is, if this is the second or subsequent processing iteration of loop B, it proceeds to step S218.
[0093] In step S218, the instruction extraction unit 104 adds the (m-1)th instruction statement from the beginning of the variable t1 to p to the instruction statement list, and proceeds to step S219. In step S219, the instruction extraction unit 104 assigns the variable t2 to the variable t1 (S219).
[0094] When the processing of step S219 is completed, the instruction extraction unit 104 performs processing of loop B with the next line number as the processing target, and when processing of loop B has been performed for all line numbers of the array variable, it exits loop B.
[0095] Next, the instruction extraction unit 104 processes loop A, targeting the next value of array variable L1, and when loop A has been processed for all values of array variable L1, it exits loop A and terminates the instruction splitting process.
[0096] (Representative word identification processing) Next, the target word specification process in step S104 will be described.
[0097] FIG. 13 is a flowchart of the group word identification process according to the first embodiment.
[0098] First, when the first determination unit 106 of the target word identification unit 105 receives a trigger to start processing, it stores the instruction statement list generated by the instruction statement extraction unit 104 (S301).
[0099] Next, the first determination unit 106 repeats the processing of loop C (S302 to S309) by substituting the values of the directive statement list one by one into the variable t3.
[0100] In the processing of loop C, the first judgment unit 106 repeats the processing of loop D (S302 to S309) by obtaining the data of the <representative word>, <second judgment code>, <characteristic word (1)>, and <characteristic word (2)> of one record in the representative word dictionary D3, starting from the top.
[0101] In the processing of Loop D, the first judgment unit 106 divides the data of each of <Feature Word (1)> and <Feature Word (2)> by a comma ",", and creates keywords by combining and concatenating the character strings obtained by dividing the data of <Feature Word (1)> with the character strings obtained by dividing the data of <Feature Word (2)> in a round-robin manner, and generates a list of these keywords (keyword list) (S302). For example, when the target record is a record with a <Processing Number> of 4 in Figure 7, a list including the keywords "Part A Harness" and "Part A Code" is generated.
[0102] Next, the first determination unit 106 searches for a keyword in the keyword list in the variable t3, and determines whether or not the keyword is found (S303).
[0103] As a result, if the keyword is found in variable t3 (S303: Yes), the first judgment unit 106 proceeds to step S304, whereas if the keyword is not found in variable t3 (S303: No), the first judgment unit 106 proceeds to step S308.
[0104] In step S304, the first determination unit 106 determines whether data is stored in the <determination code> of the record. As a result, if data is stored in the <determination code> (S304: Yes), the first determination unit 106 causes the second determination unit 108 to execute a determination process (see FIG. 14) based on the data included in the <determination code> (S305).
[0105] Next, the first determination unit 106 receives the determination result from the determination process and determines whether the determination result is suitable as a target word (S306). As a result, if the determination result is suitable as a target word (S306: Yes), the first determination unit 106 proceeds to step S307, whereas if the determination result is not suitable as a target word (S306: No), the first determination unit 106 proceeds to step S308.
[0106] If no data is stored in the <determination code> in step S304 (S304: No), or if the determination result in step S306 is that the record is suitable as a representative word (S306: Yes), the first determination unit 106 determines the character string in the <representative word> of the record as the representative word, associates it with variable t3 (instruction statement), and adds it to the representative word list (S307).
[0107] In step S308, the first judgment unit 106 judges whether there is a next record in the target word dictionary. If there is no next record (S308: No), the target word is added to the target word list as "unconfirmed," while if there is a next record (S308: Yes), processing for this record is terminated.
[0108] When the processing of loop D for one record is completed, the first judgment unit 106 performs processing of loop D for the next record, and when the processing of loop D has been performed for all records, the first judgment unit 106 exits loop D.
[0109] When loop D is exited, the first judgment unit 106 processes loop C with the next value in the instruction sentence list as the processing target, and when loop C has been processed with all values in the instruction sentence list as the processing target, it exits loop C and terminates the target word identification process.
[0110] (Determination process) Next, the determination process in step S305 will be described.
[0111] FIG. 14 is a flowchart of the determination process according to the first embodiment.
[0112] The second judgment unit 108 determines the judgment subprogram to be used in the judgment process and the arguments in the judgment subprogram based on the judgment code (S401). For example, if the judgment code is the judgment code of a record in which the <processing number> in FIG. 7 is 2, the second judgment unit 108 determines that the judgment subprogram to be used is the judgment subprogram "Sub no." corresponding to "no-," and determines the arguments to be "Part A, Part B." Next, the second judgment unit 108 executes the determined judgment subprogram using the determined arguments, thereby performing the judgment subprogram process (see FIG. 15) (S402). Next, the second judgment unit 108 returns the judgment result from the judgment subprogram process (S403), and ends the judgment process.
[0113] (Decision subprogram processing) Next, an example of the determination subprogram processing in step S402 will be described.
[0114] Fig. 15 is a flowchart of an example of the judgment subprogram processing according to the first embodiment. The judgment subprogram processing in Fig. 15 is realized by the second judgment unit 108 executing the judgment subprogram "Sub-n".
[0115] The second judgment unit 108, which executes the judgment subprogram "Sub-n", generates a list of keywords from arguments (S501).
[0116] Next, the second determination unit 108 searches for a keyword in the keyword list in the variable t3 (instruction statement) and determines whether or not the keyword is found (S502).
[0117] As a result, if a keyword is found (S502: Yes), it means that the character string of the target word in the record being processed is not a target word, so the second judgment unit 108 determines the judgment result to be incompatible (S503) and proceeds to step S505.
[0118] On the other hand, if the keyword is not found (S502: No), this means that the character string in the target record is the target word, so the second judgment unit 108 determines the judgment result to be a match (S504) and proceeds to step S505.
[0119] In step S505, the second determination unit 108 returns the determination result and ends the determination subprogram processing.
[0120] Next, another example of the determination subprogram processing in step S402 will be described.
[0121] Fig. 16 is a flowchart of another example of the judgment subprogram processing according to the first embodiment. The judgment subprogram processing in Fig. 16 is realized by the second judgment unit 108 executing the judgment subprogram "date-after".
[0122] The second determination unit 108, which executes the determination subprogram "date-after," acquires an argument as a reference date (S601).
[0123] Next, the second determination unit 108 compares the date of the specification including the directive of variable t3 (specification date) with the reference date (S602), and determines whether the specification date is later than the reference date (S603).
[0124] As a result, if the specification date is later than the reference date (S603: Yes), this means that the character string in the <target word> of the record being processed is a target word, so the second judgment unit 108 determines the judgment result to be conforming (S604) and proceeds to step S606.
[0125] On the other hand, if the specification date is not later than the reference date (S603: No), this means that the character string of the target word in the record being processed is not a target word, so the second judgment unit 108 determines the judgment result to be incompatible (S605) and proceeds to step S606.
[0126] In step S606, the second determination unit 108 returns the determination result and ends the determination subprogram processing.
[0127] Although the above description uses two judgment subprograms, "Sub-n" and "date-after," as examples, more judgment subprograms may be used. In this case, each judgment subprogram is stored in the judgment subprogram storage unit 109, and program identification information that enables the second judgment unit 108 to identify the judgment subprogram is stored in the <judgment code>, and the second judgment unit 108 identifies and executes the judgment subprogram corresponding to the program identification information. Furthermore, when a new judgment subprogram is added, the judgment subprogram is stored in the judgment subprogram storage unit 109, and the second judgment unit 108 is configured to be able to identify the new judgment subprogram from the judgment code. This allows the judgment subprogram to easily perform appropriate judgment processing to determine whether or not a word is a representative word.
[0128] (Results of the target word identification process) Next, the determination result of the target word identification process for target words related to the design object will be described.
[0129] Fig. 17 is a diagram showing an example of the determination result of the target word related to the design object according to the first embodiment. Fig. 17 shows the determination result when the target word identification process is performed using the target word dictionary D3 of the target words related to the design object shown in Fig. 7, with the specification 201 shown in Fig. 2 as the processing object.
[0130] The determination results shown in Figure 17 show each instruction statement, the matching results for each record in the target word dictionary for each instruction statement, and the target word identified as a result. Specifically, the instruction statement "Expand R of upper part of fixed part of part A" matches the record with process number 1 in the target word dictionary, thereby identifying "front surface of main body" as the target word. Furthermore, the instruction statement "Expand R of base of upper rib" matches the record with process number 2 in the target word dictionary, thereby identifying "top of main body" as the target word. Furthermore, the instruction statement "Expand R of base of harness boss of part A" matches the record with process number 3 in the target word dictionary, thereby identifying "top of main body" as the target word.
[0131] Next, the determination result of the target word specification process for target words related to CAD shapes will be described.
[0132] Fig. 18 is a diagram showing an example of a determination result of a target word related to a CAD shape according to the first embodiment. Fig. 18 shows a determination result when a target word identification process is performed using the target word dictionary D3 of target words related to a CAD shape shown in Fig. 8, with the specification 201 shown in Fig. 2 as the processing target.
[0133] The determination results shown in Figure 18 show each instruction statement, the matching results for each record in the target word dictionary for each instruction statement, and the target word identified as a result. Specifically, the instruction statement "Enlarge R of upper fixed part of part A" matches the record with process number 3 in the target word dictionary, thereby identifying "other" as the target word. Also, the instruction statement "Enlarge R of base of upper rib" matches the record with process number 1 in the target word dictionary, thereby identifying "rib" as the target word. Also, the instruction statement "Enlarge R of base of boss for part A harness" matches the record with process number 2 in the target word dictionary, thereby identifying "boss" as the target word.
[0134] (Instruction List)
[0135] Next, the instruction list R1 will be explained.
[0136] FIG. 19 is a diagram showing the configuration of the instruction content list according to the first embodiment.
[0137] The instruction content list R1 has the following items: <specification number>, <creation date>, <line number>, <instruction statement>, <design object>, <CAD shape>, and <change type>.
[0138] <Specification number> stores the specification number. <Creation date> stores the date the specification was created. <Line number> stores the line number within the directive range of the directive corresponding to the line. <Directive> stores the directive corresponding to the line. <Design object> stores the representative word related to the design object in the directive corresponding to the line. <CAD shape> stores the representative word related to the CAD shape in the directive corresponding to the line. <Change type> stores the representative word related to the change type in the directive corresponding to the line.
[0139] (Summary table) Next, summary table R2 will be explained.
[0140] FIG. 20 is a diagram showing the layout of a summary table according to the first embodiment.
[0141] Summary table R2 has the items of <Instruction Category>, <Representative Word>, and <Number of Instructions>. <Instruction Category> stores the category of the representative word, such as the design object, CAD shape, or change type. <Representative Word> stores the identified representative word. <Number of Instructions> stores the number of instructions for which the representative word is identified.
[0142] Next, the tally screen 401 will be described.
[0143] FIG. 21 is a configuration diagram of the tally screen according to the first embodiment.
[0144] The summary screen 401 includes a design object graph display area 402 in which a graph showing the identified number of representative words for the design object is displayed, a CAD shape graph display area 403 in which a graph showing the identified number of representative words for the CAD shape is displayed, and a change type graph display area 404 in which a graph showing the identified number of representative words for the change type is displayed.
[0145] The summary screen 401 allows the user to easily grasp the identified target words and the ratio of the target words for each category, thereby making it easy to grasp which target words are attracting attention in a document.
[0146] (Second embodiment) Next, a document instruction extraction system 101A according to a second embodiment will be described. Note that the same components as those in the document instruction extraction system 101 according to the first embodiment will be denoted by the same reference numerals.
[0147] (Functional configuration) FIG. 22 is a functional configuration diagram of the document instruction extraction system according to the second embodiment.
[0148] The document instruction extraction system 101A further includes a product CAD model storage unit 112 and a CAD display unit 113 in addition to the components of the document instruction extraction system 101.
[0149] The product CAD model storage unit 112 stores product CAD model information in which shape information (graphic data) of each object (component) constituting the product is associated with a representative word that indicates each object. The CAD display unit 113 acquires, for example, an instruction content list R1 or a summary table R2 from the output unit 111, identifies objects associated with representative words that are frequently identified, and displays a CAD model display screen 505 (see FIG. 23) including a CAD model of the product in which the identified objects are emphasized (e.g., highlighted). For example, the representative word that is frequently identified may be the representative word that is identified the most frequently, or may be one or more representative words that are identified a number greater than a predetermined number.
[0150] Next, the CAD model display screen 505 will be described.
[0151] Fig. 23 is a configuration diagram of a CAD model display screen according to the second embodiment. Note that the CAD model display screen 505 in Fig. 23 is a screen that is displayed when the boss 503 on the side of the main body is most frequently identified as the representative word.
[0152] A CAD model 502 of the product is displayed on a CAD model display screen 505, and a boss 503 on the side of the main body, which is most frequently identified as the representative word, is highlighted in the CAD model 502. This CAD model display screen 505 makes it easy to grasp the part most frequently identified as the representative word, i.e., the part most frequently specified in instruction sentences.
[0153] The present invention is not limited to the above-described embodiment, and can be modified appropriately without departing from the spirit of the present invention.
[0154] For example, in the above-described embodiments, some or all of the processing performed by the CPU may be performed by a hardware circuit. Also, the programs in the above-described embodiments may be installed from a program source. The program source may be a program distribution server or a storage medium (e.g., a portable storage medium). [Explanation of symbols]
[0155] 11...CPU, 14...storage device, 20...intra-document instruction extraction program, 21...document file, 101, 101A...intra-document instruction extraction system, D1...meta-information cue word dictionary, D2...instruction sentence cue word dictionary, D3...representative word dictionary
Claims
1. A document instruction extraction system that extracts instruction content from a document containing the instruction content and identifies a representative word that is a representative word corresponding to a phrase included in the instruction content, The document instruction extraction system includes a processor and a storage unit. The storage unit one or more of said documents; instruction content identification information for identifying the instruction content from the document; and target word specification information for specifying the target word from the instruction content, The processor: Identifying instruction content from the document based on the instruction content identification information; Identifying a target word from the instruction content based on the target word identification information; The representative word identification information corresponds to one or more characteristic words included in the instruction content, candidate representative words corresponding to the one or more characteristic words, and determination process information related to a determination process for determining that the candidate representative words are the representative words corresponding to the one or more characteristic words, The processor: determining whether the candidate for the target word is a target word based on the determination process information; The storage unit storing one or more determination programs for executing a determination process; the determination process information includes information for identifying a determination program to be used in the determination process and information on arguments to be passed to the determination program; The processor: The determination program is executed using the arguments to determine whether the candidate for the target word is the target word. A system for extracting instructions from documents.
2. A document instruction extraction system that extracts instruction content from a document containing the instruction content and identifies a representative word that is a representative word corresponding to a phrase included in the instruction content, The document instruction extraction system includes a processor and a storage unit. The storage unit one or more of said documents; instruction content identification information for identifying the instruction content from the document; and target word specification information for specifying the target word from the instruction content, The processor: Identifying instruction content from the document based on the instruction content identification information; Identifying a target word from the instruction content based on the target word identification information; For each target term identified from the plurality of documents, tallying the number of instructions in which the target term was identified; Based on the aggregated results, create and display a graph showing the number of instructions for multiple representative words. A system for extracting instructions from documents.
3. A document instruction extraction system that extracts instruction content from a document containing the instruction content and identifies a representative word that is a representative word corresponding to a phrase included in the instruction content, The document instruction extraction system includes a processor and a storage unit. The storage unit one or more of said documents; instruction content identification information for identifying the instruction content from the document; and target word specification information for specifying the target word from the instruction content, The processor: Identifying instruction content from the document based on the instruction content identification information; Identifying a target word from the instruction content based on the target word identification information; The storage unit storing graphic data in which representative words are associated with components of the product; The processor: When displaying an image of a product based on the graphic data, a component of the product corresponding to a representative word identified from a document relating to the product is highlighted. A system for extracting instructions from documents.
4. A document instruction extraction system that extracts instruction content from a document containing the instruction content and identifies a representative word that is a representative word corresponding to a phrase included in the instruction content, The document instruction extraction system includes a processor and a storage unit. The storage unit one or more of said documents; instruction content identification information for identifying the instruction content from the document; and target word specification information for specifying the target word from the instruction content, The processor: Identifying instruction content from the document based on the instruction content identification information; Identifying a target word from the instruction content based on the target word identification information; The document is a document related to the design and manufacturing of a specific product, The target word specifying information is For the same representative word, a plurality of correspondence relationships are included between at least a plurality of characteristic words among a first characteristic word relating to the direction of the product in normal operation, a second characteristic word relating to the direction in the design drawing of the product, and a third characteristic word relating to the direction of the product during manufacturing. A system for extracting instructions from documents.
5. A method for extracting instructions in a document using an instruction extraction system that extracts instructions from a document containing the instructions and extracts representative words that correspond to words included in the instructions, comprising: The document instruction extraction system includes: Identifying the instruction content from the document based on instruction content identification information for identifying the instruction content from the document; Identifying a target word from the instruction content based on target word identification information for identifying the target word from the instruction content; The representative word identification information corresponds to one or more characteristic words included in the instruction content, candidate representative words corresponding to the one or more characteristic words, and determination process information related to a determination process for determining that the candidate representative words are the representative words corresponding to the one or more characteristic words, The document instruction extraction system includes: determining whether the candidate for the target word is a target word based on the determination process information; The document instruction extraction system includes: storing one or more determination programs for executing a determination process; the determination process information includes information for identifying a determination program to be used in the determination process and information on arguments to be passed to the determination program; The document instruction extraction system includes: The determination program is executed using the arguments to determine whether the candidate for the target word is the target word. Intra-document instruction extraction method.
6. A method for extracting instructions in a document using an instruction extraction system that extracts instructions from a document containing the instructions and extracts representative words that correspond to words included in the instructions, The document instruction extraction system includes: Identifying the instruction content from the document based on instruction content identification information for identifying the instruction content from the document; Identifying a target word from the instruction content based on target word identification information for identifying the target word from the instruction content; For each target term identified from the plurality of documents, tallying the number of instructions in which the target term was identified; Based on the aggregated results, create and display a graph showing the number of instructions for multiple representative words. Intra-document instruction extraction method.
7. A method for extracting instructions in a document using an instruction extraction system that extracts instructions from a document containing the instructions and extracts representative words that correspond to words included in the instructions, comprising: The document instruction extraction system includes: Identifying the instruction content from the document based on instruction content identification information for identifying the instruction content from the document; Identifying a target word from the instruction content based on target word identification information for identifying the target word from the instruction content; When displaying an image of a product based on graphic data in which a representative word is associated with a component of the product, the component of the product corresponding to the representative word identified from a document relating to the product is highlighted. Intra-document instruction extraction method.
8. A method for extracting instructions in a document using an instruction extraction system that extracts instructions from a document containing the instructions and extracts representative words that correspond to words included in the instructions, comprising: The document instruction extraction system includes: Identifying the instruction content from the document based on instruction content identification information for identifying the instruction content from the document; Identifying a target word from the instruction content based on target word identification information for identifying the target word from the instruction content; The document is a document related to the design and manufacturing of a specific product, The target word specifying information is For the same representative word, a plurality of correspondence relationships are included between at least a plurality of characteristic words among a first characteristic word relating to the direction of the product in normal operation, a second characteristic word relating to the direction in the design drawing of the product, and a third characteristic word relating to the direction of the product during manufacturing. Intra-document instruction extraction method.
9. A document instruction extraction program executed by a computer to extract instruction contents from a document containing the instruction contents and extract representative words that are representative words corresponding to words included in the instruction contents, The computer, Identifying the instruction content from the document based on instruction content identification information for identifying the instruction content from the document; Identifying a target word from the instruction content based on target word identification information for identifying the target word from the instruction content; The representative word identification information corresponds to one or more characteristic words included in the instruction content, candidate representative words corresponding to the one or more characteristic words, and determination process information related to a determination process for determining that the candidate representative words are the representative words corresponding to the one or more characteristic words, The computer, determining whether the candidate representative word is a representative word based on the determination processing information; the determination process information includes information for identifying a determination program to be used in the determination process and information on arguments to be passed to the determination program; The computer, The determination program is executed using the arguments to determine whether the candidate for the representative word is the representative word. Document instruction extractor.
10. A document instruction extraction program executed by a computer to extract instruction contents from a document containing the instruction contents and extract representative words that are representative words corresponding to words included in the instruction contents, The computer, Identifying the instruction content from the document based on instruction content identification information for identifying the instruction content from the document; Identifying a target word from the instruction content based on target word identification information for identifying the target word from the instruction content; For each target word identified from the plurality of documents, counting the number of instructions in which the target word was identified; Based on the aggregated results, create and display a graph showing the number of instructions for multiple representative words. Document instruction extractor.
11. A document instruction extraction program executed by a computer to extract instruction contents from a document containing the instruction contents and extract representative words that are representative words corresponding to words included in the instruction contents, The computer, Identifying the instruction content from the document based on instruction content identification information for identifying the instruction content from the document; Identifying a target word from the instruction content based on target word identification information for identifying the target word from the instruction content; When displaying an image of a product based on graphic data in which a representative word is associated with a component part of the product, the component part of the product corresponding to the representative word identified from a document relating to the product is highlighted. Document instruction extractor.
12. A document instruction extraction program executed by a computer to extract instruction contents from a document containing the instruction contents and extract representative words that are representative words corresponding to words included in the instruction contents, The computer, Identifying the instruction content from the document based on instruction content identification information for identifying the instruction content from the document; Identifying a target word from the instruction content based on target word identification information for identifying the target word from the instruction content; The document is a document related to the design and manufacturing of a specific product, The target word specifying information is For the same representative word, a plurality of correspondence relationships are included between at least a plurality of characteristic words among a first characteristic word relating to the direction of the product in normal operation, a second characteristic word relating to the direction in the design drawing of the product, and a third characteristic word relating to the direction of the product during manufacturing. Document instruction extractor.
Citation Information
Patent Citations
Software development support apparatus
JP2005250946A
Dedicated rule editor for generating rule definition of problem extraction from free description sentence of medical observation document
JP2009245232A
Consistency determining system, method and program
JP2013125442A