Keyword dictionary creation method, work support method, keyword dictionary creation system, and work support system

The keyword dictionary creation method automates the registration of keywords from business utterances, addressing the inefficiencies of manual dictionary creation and enhancing the quality and speed of keyword dictionary development.

JP7680380B2Active Publication Date: 2025-05-20HITACHI GE NUCLEAR ENERGY LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022004540
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-01-14
Publication Date
2025-05-20
Estimated Expiration
2042-01-14

AI Technical Summary

Technical Problem

Conventional work support systems require manual creation of keyword dictionaries, which is time-consuming and labor-intensive, especially when there is a shortage of skilled workers or when dealing with new types of work.

Method used

A keyword dictionary creation method that automatically registers keywords detected from utterances made during business activities, using a voice recognition step, word extraction step, and registration step based on predetermined conditions.

Benefits of technology

Enables the creation of high-quality keyword dictionaries efficiently, reducing the need for manual creation and overcoming challenges related to skilled labor and new work types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007680380000001
    Figure 0007680380000001
  • Figure 0007680380000002
    Figure 0007680380000002
  • Figure 0007680380000003
    Figure 0007680380000003
Patent Text Reader

Abstract

To prepare a high-quality keyword dictionary.SOLUTION: In a keyword dictionary preparing method, a keyword dictionary preparing system prepares a keyword dictionary in which a keyword to be detected from an utterance made in a business is registered. The method comprises: a voice recognition step of performing voice recognition on a first utterance made in the business and a second utterance made following the first utterance; a word extraction step of extracting words from a first voice recognition result of the first utterance and a second voice recognition result of the second utterance obtained by the voice recognition step respectively; and a registration step of determining whether or not a first word extracted from the first voice recognition result and a second word extracted from the second voice recognition result by the word extraction step satisfies a predetermined condition corresponding to characteristics of the business, and registering the first word or the second word as a keyword in a keyword dictionary when the predetermined condition is satisfied.SELECTED DRAWING: Figure 20
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a keyword dictionary creating method, a task support method, a keyword dictionary creating system, and a task support system. [Background technology]

[0002] With the remarkable progress of artificial intelligence and IoT (Internet of Things) technology, momentum is building for using media recognition technology to support work, even in maintenance and inspection sites where the introduction of digital technology has not progressed much. For example, at maintenance and inspection sites, work efficiency can be greatly improved by recognizing the speech of a worker wearing a microphone and automatically recording the inspection contents by estimating them through text processing of the recognition results. In addition, human error by workers can be prevented by automatically checking the estimated inspection contents and issuing a warning to notify the worker or supervisor if the procedure is incorrect. In such systems, the method of estimating the inspection contents by extracting keywords from the speech is common.

[0003] For example, Patent Document 1 discloses an example of a work support system that assists in the creation of work reports by extracting words contained in a pre-created keyword dictionary from the speaker's voice and automatically creating a file that associates the keywords with the recorded voice. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Patent Publication No. 2021-2391 Summary of the Invention [Problem to be solved by the invention]

[0005] However, in conventional work support systems, the keyword dictionary used to estimate the inspection contents must be manually created in advance, which means that building the system is time-consuming and labor-intensive. In addition, creating keywords requires deep knowledge and experience of the relevant work, but in cases where there is a shortage of skilled workers or the relevant work is a new type of work that no one has performed before, it is difficult to create a high-quality keyword dictionary.

[0006] The present invention has been made in view of the above, and has as its object to create a high-quality keyword dictionary. [Means for solving the problem]

[0007] In order to solve the above-mentioned problems, one aspect of the present invention provides a keyword dictionary creation method in which a keyword dictionary creation system creates a keyword dictionary in which keywords detected from utterances made in business are registered, the method comprising: a voice recognition step of performing voice recognition on a first utterance made in the business and a second utterance made following the first utterance; a word extraction step of extracting words from a first voice recognition result of the first utterance by the voice recognition step and a second voice recognition result of the second utterance by the word extraction step; and a registration step of determining whether or not a first word extracted from the first voice recognition result by the word extraction step and a second word extracted from the second voice recognition result satisfy a predetermined condition according to characteristics of the business, and if the predetermined condition is satisfied, registering the first word or the second word in the keyword dictionary as the keyword. Effect of the Invention

[0008] According to one aspect of the present invention, for example, a high-quality keyword dictionary can be created. [Brief description of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram for explaining maintenance and inspection work performed by two people. [Diagram 2] FIG. 1 is a diagram showing an example of the configuration of a work support system according to a first embodiment. [Diagram 3] FIG. 4 is a diagram showing another example of the configuration of the work support system according to the first embodiment. [Figure 4] FIG. 4 is a diagram showing another example of the configuration of the work support system according to the first embodiment. [Diagram 5] FIG. 2 is a diagram showing an example of the configuration of a keyword detector in the work support system according to the first embodiment. [Figure 6] FIG. 4 is a diagram showing another example of the configuration of a keyword detector in the work support system according to the first embodiment. [Figure 7] FIG. 4 is a diagram showing an example of visualization information generated by a result visualization unit according to the first embodiment. [Figure 8] FIG. 4 is a diagram showing an example of a voice generated by a voice synthesis unit according to the first embodiment. [Figure 9] 5 is a flowchart showing an example of a keyword detection process performed by the keyword detector according to the first embodiment. [Figure 10] FIG. 11 is a diagram showing an example of the configuration of a keyword dictionary creating system according to a second embodiment. [Figure 11] FIG. 11 is a diagram showing another example of the configuration of a keyword dictionary creating system according to the second embodiment. [Figure 12] FIG. 11 is a diagram showing another example of the configuration of a keyword dictionary creating system according to the second embodiment. [Figure 13] FIG. 11 is a diagram showing an example of the configuration of a keyword generator of the keyword dictionary creation system according to the second embodiment. [Figure 14] FIG. 11 is a diagram showing an example of visualization information generated by a result visualization unit according to the second embodiment. [Figure 15] FIG. 11 is a diagram showing an example of visualization information generated by a result visualization unit according to the second embodiment. [Figure 16] FIG. 11 is a diagram showing an example of a voice generated by a voice synthesis unit according to the second embodiment. [Figure 17] FIG. 11 is a diagram for explaining a specific example of a method for automatically creating a keyword dictionary by a keyword generator according to the second embodiment. [Figure 18] FIG. 11 is a diagram for explaining a specific example of a method for automatically creating a keyword dictionary by a keyword generator according to the second embodiment. [Figure 19] FIG. 11 is a diagram for explaining a specific example of a method for automatically creating a keyword dictionary by a keyword generator according to the second embodiment. [Figure 20] 11 is a flowchart showing an example of a keyword generation process performed by a keyword generator according to the second embodiment. [Figure 21] FIG. 13 is a diagram showing an example of the configuration of a keyword generator according to the third embodiment. [Figure 22] FIG. 13 is a diagram showing an example of statistical information collected by a statistical information collecting unit according to the third embodiment. [Diagram 23] FIG. 13 is a diagram showing an example of a keyword dictionary created by a keyword generator according to the third embodiment. [Figure 24] FIG. 13 is a diagram showing another example of statistical information collected by the statistical information collecting unit according to the third embodiment. [Diagram 25] FIG. 13 is a diagram showing another example of a keyword dictionary created by the keyword generator according to the third embodiment. [Figure 26] 13 is a flowchart showing an example of a keyword generation process performed by a keyword generator according to a third embodiment. [Figure 27] FIG. 1 is a diagram for explaining maintenance and inspection work performed by one person. [Figure 28] FIG. 13 is a diagram showing an example of the configuration of a work support system according to a fourth embodiment. [Figure 29] FIG. 13 is a diagram showing an example of the configuration of a keyword detector in the work support system according to the fourth embodiment. [Diagram 30] FIG. 13 is a diagram showing an example of the configuration of a keyword dictionary creating system according to a fifth embodiment. [Diagram 31] FIG. 13 is a diagram showing an example of the configuration of a keyword generator of the keyword dictionary creating system according to the fifth embodiment. [Diagram 32] FIG. 13 is a diagram for explaining a specific example of a method for automatically creating a keyword dictionary by a keyword generator according to the fifth embodiment. [Diagram 33] FIG. 13 is a diagram for explaining a specific example of a method for automatically creating a keyword dictionary by a keyword generator according to the fifth embodiment. [Diagram 34] FIG. 13 is a diagram for explaining a specific example of a method for automatically creating a keyword dictionary by a keyword generator according to the fifth embodiment. [Diagram 35] 13 is a flowchart showing an example of a keyword generation process performed by a keyword generator according to a fifth embodiment. [Diagram 36] FIG. 2 is a hardware diagram showing an example of the configuration of a computer. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. Note that the embodiment described below, including the drawings, is merely an example and does not limit the technology disclosed in this application. In addition, not all of the elements and combinations thereof described in the embodiments are necessarily essential to the solution of the invention. In addition, illustration and explanation of a configuration that is essential to the configuration of the invention but is well known may be omitted. In addition, the number of each element shown in each figure is an example and is not limited to the number shown in the figure.

[0011] In the following description, part or all of a certain embodiment may be combined with part or all of another embodiment within the scope and consistency of the spirit of this application, and such combinations are also included in the embodiments disclosed in this application.

[0012] In the following description, the following embodiments will be described with a focus on the differences from the previous embodiments, and duplicated descriptions may be omitted. In addition, configurations and processes with the same names but different reference numerals include similar functions or processes even if there are some differences.

[0013] In the following description, the program may be installed in a device such as a computer, or may be, for example, in a program distribution server or a computer-readable (for example, non-transitory) recording medium. Also, in the following description, two or more programs may be realized as one program, or one program may be realized as two or more programs.

[0014] In the following description, the term "CPU (Central Processing Unit)" refers to one or more. The term is not limited to a microprocessor as typified by a CPU, but may refer to other types of processors such as a GPU (Graphics Processing Unit). The CPU may be single-core or multi-core. The CPU may also be replaced by a broader processor such as a hardware circuit (for example, an FPGA (Field-Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit)) that performs part or all of the processing. In the following description, the configuration referred to as "xxx unit" refers to a functional unit that is realized by program execution in cooperation with the CPU and memory.

[0015] In the following description, information described as "xxx information" may be data of any structure. In the following description, the items and structures of each piece of information are merely examples, and one piece of information may be divided into two or more pieces of information, or two or more pieces of information may be combined into one piece of information in whole or in part.

[0016] [Embodiment 1] (Maintenance and inspection work done by two people) Prior to describing the first embodiment, a maintenance and inspection work performed by multiple people will be described. Fig. 1 is a diagram for explaining maintenance and inspection work performed by a two-person team. As shown in Fig. 1, at the site of maintenance and inspection work, it is common for multiple workers to share work roles, such as an instructor 101 who gives work instructions and an operator 102 who performs work according to the instructions given by the instructor 101.

[0017] First, the instructor 101 "1. Confirm the next task using the checklist 103". Next, the instructor 101 gives "2. Instructions to the worker 102". The worker 102 then operates the equipment of the maintenance target 104 according to the instructions and "3. Repeats the instructions". For example, in FIG. 1, in response to the instruction from the instructor 101 (confirmation of the equipment number) of "The equipment number is 53AC1", the worker 102 repeats the equipment number "53AC1" as "53AC1, OK". In this way, the worker 102 repeats the important parts of the instruction from the instructor 101, thereby preventing human error due to mishearing the instructions. Then, the instructor 101 listens to the repetition of the instruction from the worker 102, confirms that the instructions have been correctly conveyed, and "4. Puts a check mark on the checklist to indicate that the task has been completed".

[0018] However, at a maintenance and inspection site, for example, when work is performed under noisy conditions, there is a problem that the voices of the instructor 101 and the worker 102 do not easily reach each other, and the work cannot be performed smoothly. Also, when the instructor 101 gives instructions to the worker 102 remotely, if wireless devices are used, the influence of noise is easily felt, making it difficult to hear the other person's voice. Therefore, in order to perform work efficiently even under such conditions, a work support system that recognizes the other person's voice and presents the results of an estimate of the work content is useful.

[0019] (Configuration of work support system 1S according to embodiment 1) Fig. 2 is a diagram showing a configuration example of a work support system 1S according to embodiment 1. Fig. 2 shows an example in which the work support system 1S recognizes the speech of a worker 102, estimates the work content, and presents the result to an instructor 101. The work support system 1S has a configuration in which three devices 210 (device 3), 211 (device 1), and 212 (device 2), each having a processor such as a personal computer, tablet, or smartphone, are connected via a network 203.

[0020] First, the device 211 (device 1) transmits the speech of the worker 102 recorded by the microphone 201 to the device 212 (device 2) via the network 303 by the transmitter 202. The device 212 (device 2) refers to the keyword dictionary 206 stored in a storage (not shown), and extracts keywords for estimating the work content from the voice of the worker received by the transceiver 204 by the keyword detector 205. Then, the device 212 (device 2) transmits the extracted keywords to the device 210 (device 3) via the transceiver 204 and the network 203. The device 210 (device 3) presents information related to the keywords received by the receiver 207 to the instructor 101 by the speaker 208 and the display terminal 209.

[0021] The speaker 208 may be any device capable of reproducing sound, such as earphones, and the display terminal 209 may be any device capable of visualizing information, such as a display device such as a tablet or display, or a wearable device such as a head-mounted display or AR glasses. Hereinafter, the same applies when referring to a speaker and a display terminal.

[0022] Alternatively, the work support system 1S may have a configuration as shown in Fig. 3 or Fig. 4. Fig. 3 and Fig. 4 are diagrams showing another example of the configuration of the work support system 1S according to the first embodiment.

[0023] In the case of FIG. 3, the work support system 1S has a configuration in which two devices 309 (device 1) and 310 (device 2) each having a processor are connected via a network 303. First, the device 309 (device 1) transmits the speech of the worker 102 recorded by a microphone 301 to the device 310 (device 2) via the network 303 by a transmitter 302. The device 310 (device 2) refers to a keyword dictionary 306 stored in a storage (not shown), and extracts keywords for estimating the work content from the speech of the worker 102 received by a receiver 304 by a keyword detector 305. Then, the device 310 (device 2) presents information related to the extracted keywords to the instructor 101 via a speaker 307 and a display terminal 308.

[0024] 4, the work support system 1S is realized by one device 406 having a processor. First, the device 406 refers to a keyword dictionary 403 stored in a storage (not shown), and extracts keywords for estimating work content from the speech of the worker 102 recorded by a microphone 401 using a keyword detector 402. Then, the work support system 1S presents information related to the extracted keywords to the instructor 101 via a speaker 404 and a display terminal 405.

[0025] (Configuration of the keyword detector according to the first embodiment) 5 is a diagram showing an example of the configuration of the keyword detectors 205, 305, 402 of the task support system 1S according to embodiment 1. The keyword detectors 205, 305, 402 each include a voice recognition unit 501, a morphological analysis unit 502, a keyword candidate word extraction unit 503, a keyword extraction unit 504, a result visualization unit 506, and a voice synthesis unit 507.

[0026] First, the voice recognition unit 501 recognizes the speech of the worker 102 acquired via a microphone and converts it into a text sentence. The morphological analysis unit 502 breaks down the text sentence converted by the voice recognition unit 501 into parts of speech and breaks it down into words.

[0027] The keyword candidate word extraction unit 503 extracts words that can become keywords from the word string obtained by the word decomposition by the morphological analysis unit 502. For example, the keyword candidate word extraction unit 503 extracts only words of specific parts of speech (nouns, verbs, adjectives) and excludes others (particles, conjunctions, interjections, etc.). The keyword extraction unit 504 then checks whether or not a word included in the keyword dictionary 505 is present among the words extracted as keyword candidates by the keyword candidate word extraction unit 503, and if so, transmits the word to the result visualization unit 506 and the speech synthesis unit 507.

[0028] In the example of FIG. 5, the voice recognition unit 501 outputs the text sentence "53AC1 is understood" from the recognition result of the microphone input. The morphological analysis unit 502 breaks down "53AC1 is understood" into parts of speech by morphological analysis to obtain " / 53AC1 / understanding / is". The keyword candidate word extraction unit 503 extracts the word "53AC1" and the word "understanding" as keyword candidates. The keyword dictionary 505 stores "53AC1, No. 5, cover, disconnection, none". Therefore, the keyword extraction unit 504 recognizes and extracts the word "53AC1" as matching a word in the keyword dictionary 505. The result visualization unit 506 and the voice synthesis unit 507 convert the keyword "53AC1" extracted by the keyword extraction unit 504 into visualization information and voice information, respectively, and present it to the instructor 101 using a display terminal and a speaker.

[0029] The process of extracting parts of speech from a text sentence recognized and converted by the speech recognition unit 501 is not limited to morphological analysis, and other part of speech extraction methods may be used.

[0030] 2 to 5, when it is desired to recognize the speech of the instructor 101 and present the result to the worker 102, the instructor 101 and the worker 102 in each figure may be interchanged.

[0031] Alternatively, as shown in Fig. 6, the spoken voices of both the instructor 101 and the worker 102 may be recognized to extract keywords. Fig. 6 is a diagram showing another example of the configuration of a keyword detector in the work support system according to the first embodiment. In this case, a microphone is also provided on the instructor 101 side, and another set of a voice recognition unit 3601, a morphological analysis unit 3602, and a keyword candidate word extraction unit 3603 is provided in addition to the worker 102 side. The keyword extraction unit 3604 outputs the word "53AC1" which is registered in the keyword dictionary 3505 and is common to the two output results from the keyword candidate word output units 503 and 3603.

[0032] (Visualization information according to the first embodiment) 7 is a diagram showing an example of visualization information generated by the result visualization unit 506 according to the first embodiment. The visualization information is displayed on a display terminal as a checklist in which the procedure of each task is represented by a keyword string, and a series of tasks is represented by an enumeration of the keyword string, and a check mark (an "X" mark in FIG. 7) is automatically added to a completed task. For example, if both (or either one) of the keywords 601 "equipment number" and "53AC1" are spoken, the number 1 check mark is added. The result visualization unit 506 is an example of an output unit that outputs visualization information.

[0033] (Sound according to embodiment 1) 8 is a diagram showing an example of a voice generated by the voice synthesis unit 507 according to the first embodiment. A voice 701 that reads out a completed task together with the task number, and a voice 702 that reads out a keyword for the next task together with the number keyword are output from a speaker. The voice synthesis unit 507 is an example of an output unit that outputs a voice.

[0034] 9 is a flowchart showing an example of a keyword detection process by the keyword detector 205 (305, 402) according to embodiment 1. First, in step S801, the speech recognition unit 501 generates a text sentence by performing speech recognition on the speech of the worker 102. Next, in step S802, the morphological analysis unit 502 performs a morphological analysis on the text sentence generated by the speech recognition unit 501, and breaks it down into words.

[0035] Next, in step S803, the keyword candidate word extraction unit 503 checks whether there is a word that can be a keyword candidate among the words resolved by the morphological analysis unit 502. Next, in step S804, the keyword extraction unit 504 refers to the keyword dictionary 505, and if a keyword exists (step S804 YES), the process proceeds to step S805, and if not (step S804 NO), the process proceeds to step S806.

[0036] In step S805, the result visualization unit 506 and / or the voice synthesis unit 507 create information to be presented to the instructor 101. In step S806, the result visualization unit 506 and / or the voice synthesis unit 507 present the created information from a speaker and / or a display terminal. In step S807, the keyword detector 205 (305, 402) determines whether all tasks have been completed, and ends the keyword detection process if all tasks have been completed (step S807YES), and returns to step S801 if all tasks have not been completed (step S807NO), and repeats the processes of steps S801 to S806 until all tasks are completed.

[0037] [Embodiment 2] The keyword dictionary 206 (306, 403, 505) needs to be prepared in advance. However, creating a keyword dictionary manually takes time and effort. Furthermore, creating a high-quality keyword dictionary manually requires in-depth knowledge and experience of the work, and it is therefore desirable for the keyword dictionary to be created by someone with the necessary knowledge.

[0038] Therefore, in the second embodiment, a means for automatically creating a high-quality keyword dictionary is provided by automatically extracting keywords from utterances made by the instructor 101 and the worker 102 when they actually perform the work. A keyword dictionary creating system 2S according to the second embodiment will be described below.

[0039] (Configuration of keyword dictionary creation system 2S according to embodiment 1 and embodiment 2) 10 is a diagram showing an example of the configuration of a keyword dictionary creation system 2S according to embodiment 2. The keyword dictionary creation system 2S has a configuration in which three devices 913, 914, and 915, each having a processor such as a personal computer, a tablet, or a smartphone, are connected via a network 909.

[0040] First, the device 913 (device 1) and the device 914 (device 2) record the speech of the worker 102 and the instructor 101 with the microphones 901 and 905, respectively, and transmit the speech to the device 915 (device 3) via the network 909 with the transceivers 904 and 908. The device 915 (device 3) extracts keywords from the speech of the instructor 101 and the worker 102 with the keyword generator 911 and registers the keywords in the keyword dictionary 912, and transmits the extracted keywords to the device 913 (device 1) and the device 914 (device 2) via the transceiver 910 and the network 909, and notifies the worker 102 and the instructor 101 from the display terminals 902 and 906 and the speakers 903 and 907, respectively, that the keywords have been added to the keyword dictionary 912.

[0041] Alternatively, the keyword dictionary creating system 2S may have a configuration as shown in Fig. 11 or Fig. 12. Fig. 11 and Fig. 12 are diagrams showing another example of the configuration of the keyword dictionary creating system 2S according to the second embodiment.

[0042] 11, the keyword dictionary creation system 2S has a configuration in which two devices 1012 (device 1) and 1013 (device 2) each having a processor are connected via a network 1008. First, the device 1012 (device 1) records the speech of the worker 102 with a microphone 1001, and transmits it to the device 1013 (device 2) via the network 1008 by the transceiver 1004. The device 1013 (device 2) records the speech of the instructor 101 with a microphone 1005.

[0043] Then, the device 1013 (device 2) extracts keywords from the voices of the instructor 101 and the worker 102 using a keyword generator 1010, registers the keywords in a keyword dictionary 1011, and transmits the extracted keywords to the device 1012 (device 1) via the transceiver 1009 and the network 1008. The devices 1012 (device 1) and 1013 (device 2) notify the worker 102 and the instructor 101 that the keywords have been added to the keyword dictionary 1011 via the display terminals 1002 and 1006 and the speakers 1003 and 1007, respectively.

[0044] 12, the keyword dictionary creation system 2S is realized by one device 1109 having a processor. The device 1109 (device 1) records the speech of the worker 102 and the instructor 101 with microphones 1101 and 1104, respectively, extracts keywords from the speech of the worker 102 and the instructor 101 with a keyword generator 1107, and registers them in the keyword dictionary 1108, and notifies the worker 102 and the instructor 101 that the keywords have been added to the keyword dictionary 1108 with display terminals 1102 and 1105 and speakers 1103 and 1106.

[0045] (Configuration of the keyword generators 911, 1010, 1107 according to the second embodiment) 13 is a diagram showing an example of the configuration of keyword generators 911, 1010, and 1107 of a keyword dictionary creation system 2S according to the second embodiment. The keyword generators 911, 1010, and 1107 include speech recognition units 1201 and 1204, morphological analysis units 1202 and 1205, keyword candidate word extraction units 1203 and 1206, overlapping word extraction unit 1207, result visualization unit 1209, and speech synthesis unit 1210. The speech recognition units 1201 and 1204, the morphological analysis units 1202 and 1205, the keyword candidate word extraction units 1203 and 1206, the result visualization unit 1209, and the speech synthesis unit 1210 have the same functions as the speech recognition unit 501, the morphological analysis unit 502, the keyword candidate word extraction unit 503, the result visualization unit 506, and the speech synthesis unit 507 of the first embodiment, respectively.

[0046] First, the voice recognition unit 1201 recognizes the voice of the worker 102 recorded by the microphone and converts it into a text sentence. The voice recognition unit 1204 recognizes the voice of the instructor 101 recorded by the microphone and converts it into a text sentence.

[0047] The morphological analysis unit 1202 breaks down the text sentence converted by the speech recognition unit 1201 into words by breaking down the parts of speech. The morphological analysis unit 1205 breaks down the text sentence converted by the speech recognition unit 1204 into words by breaking down the parts of speech. The keyword candidate word extraction unit 1203 extracts keyword candidate words that can become keywords from the word string acquired by the morphological analysis unit 1202. The keyword candidate word extraction unit 1206 extracts keyword candidate words that can become keywords from the word string acquired by the morphological analysis unit 1205.

[0048] The overlapped word extraction unit 1207 extracts overlapping words (i.e., repeated words) from among the keyword candidate words extracted from the speech of each of the instructor 101 and the worker 102 by the keyword candidate word extraction units 1203 and 1206. The overlapped word extraction unit 1207 adds the extracted overlapped words to the keyword dictionary 1208 and transmits them to the result visualization unit 1209 and the voice synthesis unit 1210. The result visualization unit 1209 and the voice synthesis unit 1210 convert the overlapped words received from the overlapped word extraction unit 1207 into visualization information and voice information, respectively, and present them to the worker 102 and the instructor 101 using a display terminal and a speaker. The overlapped word extraction unit 1207 is an example of a registration unit that registers keywords in the keyword dictionary 1208 and deletes keywords from the keyword dictionary 1208.

[0049] In the example of Figure 13, the word "53AC1" is duplicated (repeated) between the spoken content of instructor 101 "The device number is 53AC1" and the spoken content of worker 102 "53AC1 understood", so the word "53AC1" is extracted as a duplicated word and registered in keyword dictionary 1208.

[0050] (Visualization information according to the second embodiment) 14 and 15 are diagrams showing an example of visualization information generated by the result visualization unit 1209 according to embodiment 2. As shown in Fig. 14, the visualization information is displayed on a display terminal as a registered keyword string 1301 that has been registered in the keyword dictionary 1208 and a newly detected keyword string 1302 that has been newly detected.

[0051] 14 is pressed, the keyword generator 911 (1010, 1107) registers the words included in the newly detected keyword column 1302 in the keyword dictionary 1208, and moves the newly detected keyword column 1302 to the field of the registered keyword column 1301 as shown in Fig. 15. Furthermore, when the delete button 1304 shown in Fig. 14 is pressed, the keyword generator 911 (1010, 1107) deletes the newly detected keyword column 1302 from the screen display shown in Fig. 14.

[0052] (Sound according to embodiment 2) Fig. 16 is a diagram showing an example of a voice generated by the voice synthesis unit 1210 according to the second embodiment. As shown in Fig. 16, a speaker outputs a voice 1501 for reading out a newly detected newly detected keyword string, and a voice 1502 for confirming whether or not to register the newly detected keyword string in the keyword dictionary 1208. When the keyword generator 911 (1010, 1107) recognizes that the worker 102 or the instructor 101 has answered "Yes" to the voice 1502, the keyword generator 911 (1010, 1107) registers the newly detected keyword string in the keyword dictionary 1208, and when the keyword generator 911 recognizes that the worker 102 or the instructor 101 has answered "No" or has not answered "Yes", the keyword generator 911 does not register the newly detected keyword string in the keyword dictionary 1208 and erases it from the screen display shown in Fig. 14, for example.

[0053] (Specific example of the automatic keyword dictionary creation method according to the second embodiment) 17 to 19 are diagrams for explaining a specific example of a method for automatically creating a keyword dictionary by the keyword generator 911 (1010, 1107) according to the second embodiment.

[0054] Fig. 17 shows the results of speech recognition of the speech of the worker 102 and the instructor 101. Fig. 18 shows spoken sentences stored in memory during speech recognition. For ease of understanding, the recognized spoken sentences are written with punctuation marks added in Fig. 17 and Fig. 18. In Fig. 17 and Fig. 18, one memory each for storing the spoken sentences of the worker 102 and the instructor 101 (two in total) is provided, and the contents of these memories are overwritten and updated as appropriate.

[0055] 19 shows keyword strings extracted from the utterances of the worker 102 and instructor 101. By recording these keyword strings in chronological order in units of utterances in the order in which they were extracted, the order of tasks can be memorized, and it becomes easy to use them as the checklist shown in FIG.

[0056] As shown in Fig. 17, first, in step S1601, the keyword generator 911 (1010, 1107) recognizes the utterance of the instructor 101, "Start checking for cable breaks." Next, in step S1602, the keyword generator 911 (1010, 1107) recognizes the response of the worker 102, "Yes," to step S1601. Then, the keyword generator 911 (1010, 1107) determines that the speaking right has been transferred from the instructor 101 to the worker 102, and writes the utterance of the instructor 101 recognized in step S1601 into the memory for the instructor, as shown in step S1701 (Fig. 18).

[0057] Next, in step S1603, when the keyword generator 911 (1010, 1107) recognizes the speech of the instructor 101, it determines that the speaking right has been transferred from the worker 102 to the instructor 101, and writes the spoken sentence of step S1602 to the worker memory as shown in step S1702 (Fig. 18). Then, the keyword generator 911 (1010, 1107) compares the spoken sentence written in the instructor memory in step S1701 with the spoken sentence written in the worker memory in step S1702, and if there is a duplicated word (repeated keyword), it adds it to the keyword dictionary in Fig. 19. In this way, the process of overwriting and updating the data in the memory every time the speaking right is transferred and the process of comparing the spoken sentences in the memory are repeated until the inspection work is completed (no more speaking is performed).

[0058] Here, when the keyword generator 911 (1010, 1107) recognizes, for example, the repetition of the device number "53AC1" by the worker 102 (step S1604 in FIG. 17), since the spoken sentence written in the memory for the instructor (step S1703 in FIG. 18) and the spoken sentence written in the memory for the worker (step S1704 in FIG. 18) contain the same word "53AC1", the keyword generator 911 adds this word to the dictionary as a keyword (step S1801 in FIG. 19).

[0059] In addition, in cases where the instructor 101 makes successive utterances (steps S1605 and S1606, and steps S1609 and S1610 in FIG. 17), these utterances are recorded concatenated in the instructor memory as if there had been no transition of speaking rights in between (step S1705 in FIG. 18).

[0060] Furthermore, when multiple words are repeated at the same time (step S1607 in FIG. 17) (in this example, the words "number 5" and "miss" are repeated at the same time), these words are registered in the dictionary as one keyword string (step S1802 in FIG. 19). Note that it is not always the case that the worker 102 repeats what the instructor 101 says, and the reverse is also possible. That is, there is also the case where the instructor repeats (step S1609) what the worker 102 says (step S1608). In this case as well, the repeated words are registered in the keyword dictionary (step S1803).

[0061] (Keyword generation process according to the second embodiment) FIG. 20 is a flowchart showing an example of a keyword generation process performed by the keyword generator 911 (1010, 1107) according to the second embodiment.

[0062] First, in step S1901, the voice recognition units 1201 and 1204 recognize the speech of the worker 102 or instructor 101. Next, in step S1902, the voice recognition units 1201 and 1204 write the recognition result of step S1902 into memory. The voice recognition units 1201 and 1204 repeat steps S1901 and S1902 until the right to speak is transferred to the other party (step S1903 YES). If the right to speak is transferred to the other party (step S1903 YES), the voice recognition units 1201 and 1204 proceed to step S1904, and if the right to speak has not been transferred to the other party (step S1903 NO), they return to step S1901.

[0063] In step S1904, the morphological analysis units 1202 and 1205 perform morphological analysis on the spoken sentences in the memories for the worker and the instructor, and the keyword candidate word extraction units 1203 and 1206 extract keyword candidates. Next, in step S1905, the morphological analysis units 1202 and 1205 check whether there are any overlapping words between the worker 102 and the instructor 101 among the keyword candidates extracted in step S1904, and if there are any (step S1905 YES), they register the corresponding words as keywords in the keyword dictionary 1208 (step S1906), and if there are no overlapping words (step S1905 NO), the process proceeds to step S1907. The speech recognition units 1201 and 1204, the morphological analysis units 1202 and 1205, and the duplicate word extraction unit 1207 repeat the processes of steps S1901 to S1906 until the work is completed (YES in step S1907).

[0064] The keyword dictionary 1208 created by the keyword dictionary creation system 2S in Fig. 10 to Fig. 12 can be used in the work support system 1S of the first embodiment in Fig. 2 to Fig. 4. In addition to the work support system 1S that supports maintenance and inspection work, the keyword dictionary 1208 can be used in various work support systems, such as a work history summarization system, a speech history recording system in a call center, a meeting minutes creation system, and an automatic creation system for a recognition word dictionary used in voice recognition.

[0065] Furthermore, while the keyword dictionary creation system 2S assumes that spoken voice is input through a microphone, any input format is acceptable, for example, text sentences can be input through a keyboard. Furthermore, the keyword dictionary creation system 2S can be expanded to cases where three or more people work together. That is, if all the people working use microphones and two or more people utter the same word within a specified time, the word is regarded as having been repeated and registered in the dictionary.

[0066] The keyword generator 911 (1010, 1107) may generate a task list of the work in which the keywords registered in the keyword dictionary 1208 are listed in chronological order of the task sequence. The work support system 1S may recognize the contents of the speech of the instructor 101 and the worker 102 by voice recognition, and use this task list as a checklist to mark completed tasks to indicate that they are completed, thereby managing the progress of the tasks.

[0067] [Embodiment 3] In the second embodiment, an example was shown in which a pair (a worker and an instructor) actually performed a task, and the repeated important words were registered in a keyword dictionary in real time. However, in this case, problems occur such as an important word not being registered as a keyword in the dictionary if it is accidentally forgotten to be repeated, and an unimportant word or an omissible word being uttered by the two people by chance being mistakenly registered in the dictionary. Therefore, in this embodiment, repeated word information is collected from multiple pairs, and keywords are determined based on the statistical information. This makes it possible to assign importance to repeated word strings and extract only important ones to create a high-quality keyword dictionary.

[0068] (Configuration of the keyword generator 200 according to the third embodiment) Fig. 21 is a diagram showing an example of the configuration of a keyword generator 200 according to embodiment 3. In this embodiment, the keyword generator 200 in Fig. 21 is used instead of the keyword generators 911, 1010, and 1107 in Fig. 13. Compared with the keyword generators 911, 1010, and 1107, the keyword generator 200 has a configuration in which a statistical information collecting unit 2008, statistical information 2009, and a keyword determining unit 2010 are added.

[0069] The speech recognition units 2001 and 2004, the morphological analysis units 2002 and 2005, and the keyword candidate word extraction units 2003 and 2006 have the same functions as the speech recognition units 1201 and 1204, the morphological analysis units 1202 and 1205, the keyword candidate word extraction units 1203 and 1206, and the duplicate word extraction unit 1207 of the second embodiment, respectively.

[0070] Furthermore, when deciding keywords based on the statistical information 2009, it is not necessary to create the keyword dictionary 2011 in real time. Therefore, in the keyword generator 200, a display terminal, a result visualization unit, a speaker, and a voice synthesis unit for presenting the newly added keywords to the worker 102 and the instructor 101 are omitted.

[0071] 21, words spoken by the instructor 101 and repeated by the worker 102 are collected by a statistical information collecting unit 2008, and the collected statistical information 2009 is stored in a storage. This collection process is performed for a plurality of pairs of the instructor 101 and the worker 102, and a keyword determining unit 2010 determines words to be keywords based on the statistical information 2009, and adds the words to a keyword dictionary 2011. The keyword determining unit 2010 is an example of a registration unit that registers keywords in the keyword dictionary 2011 and deletes keywords from the keyword dictionary 2011.

[0072] (Statistical information according to the third embodiment) Fig. 22 is a diagram showing an example of statistical information 2009 collected by the statistical information collecting unit 2008 according to the third embodiment. Fig. 23 is a diagram showing an example of a keyword dictionary created by the keyword generator 2000 according to the third embodiment.

[0073] In this example, the statistical information 2009 is given in the form of a table in which each row represents an extracted (i.e., repeated) word string (e.g., "53AC1" (row 2101)) and each column represents a pair number (e.g., "Pair 1" (column 2104)). For each pair, the statistical information 2009 indicates whether the corresponding word string was extracted (indicated by a "○" in the table) or not (indicated by a "-" in the table). In addition, the last column lists the number of pairs in which the corresponding word string was detected in each row ("Total" (column 2105)), i.e., the number of "○" marks in each row, which is tallied.

[0074] At this time, since a word string with a large "total value" (column 2105) is repeated by more pairs and can be said to be more important, the keyword determination unit 2010 adds word strings with a total value (column 2105) equal to or greater than a threshold to the keyword dictionary 2011. In this example, if the threshold is set to 5, for example, a keyword dictionary 2011 like that shown in FIG. 23 is created.

[0075] In the example of Figure 23, the word string "No. 5, Remove" (row 2102) and the word string "No. 5, Cover, Remove" (row 2103) both have a "Total Value" (column 2105) that is greater than the threshold (9 and 5, respectively), but these word strings contain the same word string "No. 5, Remove." It is possible to exclude such duplicate word strings and register only one of them. For example, in Figure 22, of these word strings, the word string "No. 5, Remove" (row 2102) with the larger "Total Value" (column 2105) is registered in the keyword dictionary (row 2201 in Figure 23).

[0076] Fig. 24 is a diagram showing another example of the statistical information 2009 collected by the statistical information collecting unit 2008 according to the third embodiment. Fig. 25 is a diagram showing another example of the creation of a keyword dictionary by the keyword generator 2000 according to the third embodiment.

[0077] In this example, extracted (i.e., repeated) words (e.g., "53AC1") are given in a table format with rows 2301 and columns 2302 set to each other, and the number of pairs in which two corresponding words are extracted simultaneously (i.e., included in the same word string) is tallied. However, on the diagonal of the table, the number of pairs in which one corresponding word is extracted once is tallied.

[0078] For example, the number of pairs in which the word strings "53AC1" and "No. 5" are extracted (i.e., repeated) is 1 (total value column 2303), and the number of pairs in which the word strings "No. 5" and "Omit" are extracted is 8 (total value column 2304). At this time, words with a large total value have a high co-occurrence and are highly likely to be related to each other. Therefore, the keyword determination unit 2010 registers a combination of words having a total value equal to or greater than a threshold as a keyword for the same word string. In this example, if the threshold is set to 5, for example, a keyword dictionary 2011 as shown in FIG. 25 is created from the statistical information 2009 in FIG. 24. The total value is an example of a weight index that measures the co-occurrence of two words. A combination of words with a large "total value" is repeated by more pairs and is therefore considered to be important.

[0079] 24, keywords already registered in the keyword dictionary 2011 may be set in each row 2301, extracted (i.e., repeated) words may be set in each column 2302, and the number of pairs in which the two corresponding words are extracted simultaneously (i.e., included in the same word string) may be tallied. In this way, keywords highly related to keywords already registered in the keyword dictionary 2011 may be newly registered in the keyword dictionary 2011.

[0080] (Keyword generation process according to the third embodiment) FIG. 26 is a flowchart showing an example of a keyword generation process performed by the keyword generator 2000 according to the third embodiment.

[0081] First, in step S2501, the keyword generator 2000 recognizes the change of pair between the instructor 101 and the worker 102. Next, in step S2502, the voice recognition units 2501 and 2504 recognize the speech of the worker 102 or the instructor 101. Next, in step S2503, the voice recognition units 2501 and 2504 write the recognition result of step S2502 to memory. Next, in step S2504, the voice recognition units 2501 and 2504 determine whether the right to speak has been transferred from one to the other. If the right to speak has been transferred from one to the other (step S2504 YES), the voice recognition units 2501 and 2504 proceed to step S2505, and if the right to speak has not been transferred (step S2504 NO), they return to step S2502.

[0082] In step S2505, the morphological analyzers 2502 and 2505 perform morphological analysis on the spoken sentences in the worker and instructor memories, and the keyword candidate word extractors 1203 and 1206 extract keyword candidates from the words extracted by the morphological analysis.

[0083] Next, in step S2506, the duplicated word extraction unit 1207 determines whether there are any duplicated words among the keyword candidates spoken by the instructor 101 and the worker 102. If there are any duplicated words (step S2506 YES), the process proceeds to step S2507. If there are no duplicated words (step S2506 NO), the process proceeds to step S2508.

[0084] In step S2507, the statistical information collecting unit 2008 updates the statistical information 2009 using the overlapping words in step S2508. In step S2508, the statistical information collecting unit 2008 judges whether sufficient statistical information 2009 has been collected. If sufficient statistical information 2009 has been collected (step S2509 YES), the statistical information collecting unit 2008 moves the process to step S2510, and if sufficient statistical information 2009 has not been collected (step S2509 NO), the statistical information collecting unit 2008 returns the process to step S2501, and repeats the processes of steps S2502 to S2509 while alternating pairs of instructor 101 and worker 102. Then, if a sufficient amount of statistical information has been collected (step S2510 YES), the keyword determining unit 2010 determines a word string that will be a keyword based on the collected statistical information 2009, and registers the determined keyword in the keyword dictionary 2011 (step S2510).

[0085] In the third embodiment, a word whose total number of occurrences, which is a weight for measuring the co-occurrence of words, is equal to or greater than a threshold value is registered as a keyword in the keyword dictionary 2011, but this is not limiting. For example, instead of being limited to the total number of occurrences being equal to or greater than a threshold value, a word whose total number of occurrences accounts for a percentage of the total number of occurrences relative to the total number of occurrences may be registered as a keyword in the keyword dictionary 2011.

[0086] In parallel with the work support of the work support system of the second embodiment, as in the third embodiment, statistics are taken of the frequency of occurrence of words in the utterances of the instructor 101 and the worker 102 uttered in this work. Then, words whose total number of occurrences is equal to or greater than a threshold value are registered in the keyword dictionary as meeting the registration requirements, and words whose total number of occurrences is less than the threshold value are deleted from the keyword dictionary 2011 as meeting the non-registration requirements, thereby increasing or decreasing the number of keywords registered in the keyword dictionary. That is, as the utterances of the instructor 101 and the worker 102 are accumulated, keywords are added to the keyword dictionary and keywords are deleted from the keyword dictionary. In this way, appropriate keywords reflecting the latest current state of utterances can be registered in the keyword dictionary while preventing the keyword dictionary from taking up too much storage resources.

[0087] [Embodiment 4] (Maintenance and inspection work done by one person) In the above embodiments, examples have been described in which a plurality of people work together. In the fourth embodiment, an example in which one person works will be described. Before describing the fourth embodiment, maintenance and inspection work performed by one person will be described. FIG. 27 is a diagram for explaining maintenance and inspection work performed by one person.

[0088] First, the worker 2601 checks the next task using the checklist 2602, and while operating the equipment of the maintenance target 2603, he / she confirms the task content out loud. At this time, it is common to perform a pointing and calling confirmation by pointing at the maintenance target 2603 and speaking to confirm. At this time, the same words are often spoken before and after the task. For example, in FIG. 27, after confirming the task content (confirming the model number) by saying "The model number is 53AC1," the worker confirms by saying "53AC1, OK" that the task is complete. In this way, the worker 2601 repeats the important word "53AC1" by himself / herself to prevent human error.

[0089] When working alone, problems such as difficulty in hearing the other person's voice, which occur when working with multiple people, do not occur. However, even in this case, in order to prevent human errors such as missing tasks, it is useful to provide a work support system that has a function of recognizing the voice of the worker 2601 himself and presenting the results of an estimation of the work content (for example, a function of automatically generating a checklist and automatically marking completed tasks as completed).

[0090] (Configuration of the work support system 4S according to the fourth embodiment) Fig. 28 is a diagram showing a configuration example of a work support system 4S according to embodiment 4. The work support system 4S for maintenance and inspection work by one person corresponds to the work support system 1S (see Fig. 5) for maintenance and inspection work by two people in embodiment 1. As shown in Fig. 28, the work support system 4S is realized using one device 2706 having a processor.

[0091] First, the work support system 4S refers to a keyword dictionary 2705 stored in storage from the speech of the worker 2601 recorded by a microphone 2701, and extracts keywords for estimating the work content by a keyword detector 2704. Then, information related to the extracted keywords is presented to the worker 2601 by a speaker 2702 and a display terminal 2703.

[0092] (Configuration of the keyword detector 2704 according to the fourth embodiment) Fig. 29 is a diagram showing an example of the configuration of the keyword detector 2704 of the task support system 4S according to the fourth embodiment. The keyword detector 2704 has a voice recognition unit 2801, a morphological analysis unit 2802, a keyword candidate word extraction unit 2803, a keyword extraction unit 2804, a result visualization unit 2806, and a voice synthesis unit 2807. The voice recognition unit 2801, the morphological analysis unit 2802, the keyword candidate word extraction unit 2803, the keyword extraction unit 2804, the result visualization unit 2806, and the voice synthesis unit 2807 have the same functions as the voice recognition unit 501, the morphological analysis unit 502, the keyword candidate word extraction unit 503, the keyword extraction unit 504, the result visualization unit 506, and the voice synthesis unit 507 of the first embodiment (see Fig. 5). The visualization information and the voice information generated by the result visualization unit 2806 and the voice synthesis unit 2807 are the same as those of the first embodiment (see Figs. 7 and 8).

[0093] [Embodiment 5] 30 is a diagram showing a configuration example of a keyword dictionary creation system 5S according to embodiment 5. The keyword dictionary creation system 5S for maintenance and inspection work by one person corresponds to the keyword dictionary creation system 2S for maintenance and inspection work by two people according to embodiment 2 (see FIG. 12).

[0094] The keyword dictionary creation system 5S is realized by using one device 2906 having a processor, similar to the keyword dictionary creation system 2S shown in Fig. 12. First, the device 2908 records the speech of the worker 2601 with a microphone 2901, and extracts important words from the speech with a keyword generator 2904 and registers them in the keyword dictionary 2905. At that time, the device 2908 notifies the worker 2601 by a display terminal 2902 and a speaker 2903 that the keyword has been added to the keyword dictionary 2905.

[0095] (Configuration of the keyword generator 2904 according to the fifth embodiment) 31 is a diagram showing an example of the configuration of a keyword generator 2904 of a keyword dictionary creation system 5S according to the fifth embodiment. The keyword generator 2904 includes a speech recognition unit 3001, a morphological analysis unit 3002, a keyword candidate word extraction unit 3003, a duplicated word extraction unit 3004, a result visualization unit 3006, and a speech synthesis unit 3007. The speech recognition unit 3001, the morphological analysis unit 3002, the keyword candidate word extraction unit 3003, the duplicated word extraction unit 3004, the result visualization unit 3006, and the speech synthesis unit 3007 have the same functions as the speech recognition unit 1201, the morphological analysis unit 1202, the keyword candidate word extraction unit 1203, the duplicated word extraction unit 1207, the result visualization unit 1209, and the speech synthesis unit 1210 of the second embodiment (see FIG. 13). The visualization information and voice information generated by the result visualization unit 3006 and the voice synthesis unit 3007 are the same as those in the second embodiment (see FIGS. 14 and 16).

[0096] (Specific example of the automatic keyword dictionary creation method according to the fifth embodiment) 32 to 34 are diagrams for explaining a specific example of a method for automatically creating a keyword dictionary by the keyword generator 2904 according to the fifth embodiment.

[0097] Fig. 32 shows the result of speech recognition of the speech of the worker 2601. Fig. 33 shows the spoken sentences stored in the memory during the speech recognition. For ease of understanding, the recognized spoken sentences are written with punctuation marks in Fig. 32 and Fig. 33. In Fig. 32 and Fig. 33, the contents of the memory for storing the spoken sentences of the worker 2601 are overwritten and updated as appropriate.

[0098] 34 shows keyword strings extracted from the utterances of the worker 2601. By recording these keyword strings in chronological order in units of utterances in the order in which they were extracted, the order of tasks can be memorized, and it becomes easy to use them as the checklist shown in FIG.

[0099] As shown in FIG. 32, first, in step S3101, the keyword generator 2904 recognizes the utterance of the worker 2601, "Start inspection for cable breakage." Next, in step S3102, the keyword generator 2904 recognizes the next utterance of the worker 2601, "The model number is 53AC1." At this time, it is assumed that there is a time interval of 45 seconds between the utterance of step S3101 and the utterance of step S3102. In this embodiment, if this time interval is less than a threshold, the previous and following utterances are combined into one utterance and stored in memory. Here, if the threshold is set to 30 seconds, the keyword generator 2904 does not combine with the utterance sentence of step S3102, and stores the utterance sentence of step S3101 alone in memory, as shown in step S3201 (FIG. 33).

[0100] Then, the keyword generator 2904 checks whether there are any duplicated words (i.e., repeated words) in the utterances stored in the memory in step S3201. If there are any duplicated words, the keyword generator 2904 registers the duplicated words in the keyword dictionary 3005. In this case, there are no duplicated words (there are no words that appear multiple times in the utterances stored in the memory), so nothing is registered in the keyword dictionary 3005.

[0101] Next, in step S3103, the keyword generator 2904 recognizes the next utterance of the worker 2601. At this time, it is assumed that there was a time interval of 20 seconds between the utterance in step S3102 and the utterance in step S3103. Since the time interval of this utterance is less than the threshold, the keyword generator 2904 concatenates the sentence uttered in step S3102 and the sentence uttered in step S3103 and stores them in memory, as shown in step S3202 (FIG. 33).

[0102] Then, the keyword generator 2904 checks the spoken sentences stored in the memory, and checks whether there are any duplicated words (i.e., words being repeated) in the spoken sentences stored in the memory in step S3202. If there are any duplicated words, the keyword generator 2904 registers the duplicated words in the keyword dictionary 3005. In this case, since there is a duplicated word "53AC1", "53AC1" is registered in the keyword dictionary 3005 as shown in step S3301 (FIG. 34).

[0103] When the same process is performed on the remaining utterances in steps S3104 and S3105, these utterances are concatenated and overwritten in the memory, and as shown in step S3302 (FIG. 34), the overlapping word strings "No. 5" and "Cover" are registered in keyword dictionary 3005. In this way, each time an utterance is recognized, a determination is made as to whether or not the recognized utterance should be concatenated with the previous utterance, and the process of overwriting and updating the data in memory according to the determination result, a comparison process of the utterances in memory, and a registration process in keyword dictionary 3005 are repeated until the inspection work is completed (when no more utterances are being made).

[0104] In addition, "Yoshi" in steps S3103 and S3105 is a specific word specific to the task that is spoken when the keyword is self-repeated. Therefore, the word "53AC1" uttered immediately before "Yoshi" in the same utterance, and the word "53AC1" uttered in the utterance immediately before that, can be considered to be the keyword. In this case, it may not be necessary to divide the utterance sentence in consideration of the time interval of the utterance.

[0105] (Keyword generation process according to the fifth embodiment) FIG. 34 is a flowchart showing an example of a keyword generation process performed by the keyword generator 2904 according to the fifth embodiment.

[0106] First, in step S3401, the voice recognition unit 3001 recognizes the voice uttered by the worker 2601. Next, in step S3402, the voice recognition unit 3001 writes the text sentence obtained by converting the recognition result of step S3401 into memory. Next, in step S3403, the voice recognition unit 3001 determines whether or not a time interval equal to or greater than a threshold has elapsed since the utterance that caused the text sentence to be written into memory in step S3402. If a time interval equal to or greater than the threshold has elapsed (step S3403 YES), the voice recognition unit 3001 proceeds to step S3404, and if not (step S3403 NO), the voice recognition unit 3001 returns to step S3401.

[0107] In step S3404, the morphological analysis unit 3002 performs morphological analysis on the utterance in the memory, and the keyword candidate word extraction unit 3003 extracts keyword candidates from the morphological analysis results. Next, in step S3405, the duplicated word extraction unit 3004 determines whether or not a word that overlaps with the keyword candidate exists in another keyword candidate. If an overlapping word exists (step S3405 YES), the duplicated word extraction unit 3004 proceeds to step S3406, and if not (step S3405 NO), the duplicated word extraction unit 3004 proceeds to step S3407.

[0108] In step S3406, the redundant word extraction unit 3004 registers the relevant word determined to exist in step S3405 as a keyword in the keyword dictionary 3005. In step S3407, the redundant word extraction unit 3004 determines whether the series of tasks has been completed. If the series of tasks has been completed (step S3407YES), the redundant word extraction unit 3004 ends the keyword generation process, and if the series of tasks has not been completed (step S3407NO), the process returns to step S3401. The redundant word extraction unit 3004 repeats the processes of steps S3401 to S3406 until the series of tasks is completed (step S3407YES).

[0109] In this embodiment, important words are extracted from the work of one worker to generate a keyword dictionary, but in the same manner as in embodiment 3, statistical information (e.g., Fig. 22 or Fig. 24) may be created from work data acquired by multiple workers in shifts, and a keyword dictionary may be created based on weights such as the number of occurrences of extracted word strings. In this case, in order to obtain word occurrence statistics for each worker, the pair numbers (e.g., Pair 1, Pair 2, ...) in each column of Fig. 22 are replaced with worker numbers (e.g., Worker 1, Worker 2, ...).

[0110] In the above embodiment, keywords are automatically extracted from words extracted from the speech of a person performing a task according to the characteristics of the task (if the task is performed alone, the worker will immediately respond to a keyword uttered by the person by saying "good" or "OK", etc., and if the task is performed by two or more people, a keyword uttered by one worker will be repeated by other workers, etc.) and registered in a keyword dictionary. This eliminates the need to manually create a keyword dictionary, making it possible to build a task support system in a short time. In addition, since it is possible to register keywords that reflect the actual task content, it has the effect of enabling the creation of a high-quality keyword dictionary.

[0111] In the above embodiment, statistical information on words extracted from the speech of a person performing a task is obtained, and words whose importance based on the statistical information is equal to or greater than a threshold are registered as keywords in a keyword dictionary. This allows prioritization of keywords, and important keywords with high priorities can be registered in the keyword dictionary.

[0112] Furthermore, according to the above embodiment, it is possible to register highly important words related to keywords already registered in the keyword dictionary as keywords in the keyword dictionary.

[0113] Furthermore, according to the above embodiment, repeated words, two words with high co-occurrence, or words that are immediately followed by specific words such as "good" or "OK" can be registered as keywords in the keyword dictionary.

[0114] According to the above embodiment, if a word satisfies the requirements for registration in the keyword dictionary based on statistical information, the word is registered in the keyword dictionary as a keyword, and if the requirements for non-registration are met, the word is deleted from the keyword dictionary as a keyword. This makes it possible to appropriately select important keywords that change from moment to moment in accordance with the latest situation and register them in the keyword dictionary.

[0115] Furthermore, according to the above embodiment, whether the task is performed by one person or by two or more people, keywords can be extracted from speech during the task and registered in the keyword dictionary.

[0116] In addition, in the above embodiments, a task list is generated in which the keywords registered in the keyword dictionary are listed in chronological order of the work sequence, so that workers can check the work procedures based on the task list.

[0117] Furthermore, by simultaneously running the keyword dictionary creation system of each of the above-described embodiments and the work support system and extracting keywords from actual speech during work to update the keyword dictionary, the keyword dictionary can be made more up-to-date and accurate in line with the progress of work.

[0118] (Computer 1 Hardware) 36 is a hardware diagram showing a configuration example of the computer 1. For example, the computer 1 realizes task support systems 1S, 4S including keyword detectors 205, 305, 402, 2704, 2904, keyword dictionary creation systems 2S, 5S including keyword generators 911, 1010, 1107, 2000, or devices that appropriately integrate these.

[0119] The computer 1 is a computer equipped with a processor 11 including a CPU, a main memory device 12, an auxiliary memory device 13, a network interface 14, an input device 15, and an output device 16, all of which are interconnected via an internal communication line 19 such as a bus.

[0120] The processor 11 controls the overall operation of the computer 1. The main storage device 12 is composed of, for example, a volatile semiconductor memory, and is used as a work memory for the processor 11. The auxiliary storage device 13 is composed of a large-capacity nonvolatile storage device such as a hard disk device, an SSD (Solid State Drive), or a flash memory, and is used to store various programs and data for a long period of time.

[0121] The executable program 13P stored in the auxiliary memory device 13 is loaded into the main memory device 12 when the computer 1 is started up or when needed, and the processor 11 executes the executable program 13P loaded into the main memory device 12 to realize each of the aforementioned devices that perform various processes.

[0122] The executable program 13P may be recorded on a non-transitory recording medium, read from the non-transitory recording medium by a medium reading device, and loaded into the main storage device 12. Alternatively, the executable program 13P may be obtained from an external computer via a network and loaded into the main storage device 12.

[0123] The network interface 14 is an interface device for connecting the computer 1 to each network in the system or for communicating with other computers. The network interface 14 is, for example, configured with a NIC (Network Interface Card) for a wired LAN (Local Area Network) or a wireless LAN.

[0124] The input device 15 is composed of a keyboard, a pointing device such as a mouse, and the like, and is used by the user to input various instructions and information to the computer 1. The output device 16 is composed of a display device such as a liquid crystal display or an organic EL (Electro Luminescence) display, and an audio output device such as a speaker, and is used to present necessary information to the user when necessary.

[0125] The present invention is not limited to the above-described embodiments, and includes various modified examples. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and are not necessarily limited to those having all of the configurations described. In addition, it is also possible to replace a part of the configuration of one embodiment with the configuration of another embodiment, and to add the configuration of another embodiment to the configuration of one embodiment, as long as there is no contradiction. In addition, it is possible to add, delete, replace, integrate, or distribute a part of the configuration of each embodiment. In addition, the configurations and processes shown in the embodiments can be appropriately distributed, integrated, or replaced based on processing efficiency or implementation efficiency. [Explanation of symbols]

[0126] 1: Computer, 1S, 4S: Work support system, 2S, 5S: Keyword dictionary creation system, 205, 305, 402, 2704, 2904: Keyword detector, 911, 1010, 1107, 2000: Keyword generator.

Claims

1. A keyword dictionary creation method in which a keyword dictionary creation system creates a keyword dictionary in which keywords detected from utterances made in business are registered, comprising the steps of: a speech recognition step of performing speech recognition on a first utterance uttered in the business and a second utterance uttered following the first utterance; a word extraction step of extracting words from each of a first speech recognition result of the first utterance and a second speech recognition result of the second utterance by the speech recognition step; a registration step of determining whether or not a first word extracted from the first speech recognition result by the word extraction step and a second word extracted from the second speech recognition result satisfy a predetermined condition according to characteristics of the business, and if the predetermined condition is satisfied, registering the first word or the second word in the keyword dictionary as the keyword; A keyword dictionary creating method comprising the steps of:

2. 2. The keyword dictionary creation method according to claim 1, a statistical information collection step of collecting statistical information on the first words and the second words that satisfy the predetermined condition and are extracted by the word extraction step, In the registration step, the first word or the second word whose importance based on the statistical information is equal to or greater than a threshold is registered in the keyword dictionary as the keyword, and the first word or the second word whose importance is less than the threshold is not registered in the keyword dictionary. A keyword dictionary creating method comprising:

3. 3. The keyword dictionary creation method according to claim 2, the statistical information is a statistic of the number of times a keyword registered in the keyword dictionary and a related word related to the keyword are spoken; In the registration step, if the number of times is equal to or greater than a threshold, the related word is registered in the keyword dictionary. A keyword dictionary creating method comprising:

4. 3. The keyword dictionary creation method according to claim 2, the predetermined condition is that the first word and the second word overlap, or that the first word and the second word have co-occurrence, The statistical information is the number of occurrences of the word that satisfies the predetermined condition. A keyword dictionary creating method comprising:

5. 3. The keyword dictionary creation method according to claim 2, In the registration step, if the statistical information satisfies a registration requirement for the first word or the second word, registering the first word or the second word as the keyword in the keyword dictionary; When the statistical information satisfies a non-registration requirement for the first word or the second word, the first word or the second word is deleted from the keyword dictionary. A keyword dictionary creating method comprising:

6. 2. The keyword dictionary creation method according to claim 1, The first utterance is an utterance made by a first person, and the second utterance is an utterance made by a second person different from the first person. A keyword dictionary creating method comprising:

7. 2. The keyword dictionary creation method according to claim 1, The first utterance and the second utterance are utterances made by the same person. A keyword dictionary creating method comprising:

8. A keyword dictionary creation method according to claim 7, comprising: the predetermined condition is that the second word corresponds to a specific word; In the registration step, the first word is registered in the keyword dictionary as the keyword if the predetermined condition is met. A keyword dictionary creating method comprising:

9. 2. The keyword dictionary creation method according to claim 1, A task list for the business is generated by listing the keywords registered in the keyword dictionary in the registration step in chronological order of the tasks. A keyword dictionary creating method comprising:

10. A method according to any one of claims 1 to 9, a detection step of determining whether or not the first word extracted from the first speech recognition result by the word extraction step and the second word extracted from the second speech recognition result satisfy the predetermined condition according to the characteristics of the business, and if the predetermined condition is satisfied, checking whether the first word or the second word is registered in the keyword dictionary, and if registered, detecting the first word or the second word as the keyword; an output step of outputting the keyword detected by the detection step to a user; A work support method comprising the steps of:

11. A keyword dictionary creation system that creates a keyword dictionary in which keywords detected from utterances made in business are registered, a speech recognition unit that recognizes a first utterance uttered in the business and a second utterance uttered following the first utterance; a word extraction unit that extracts words from each of a first speech recognition result of the first utterance and a second speech recognition result of the second utterance by the speech recognition unit; a registration unit that determines whether a first word extracted from the first speech recognition result by the word extraction unit and a second word extracted from the second speech recognition result satisfy a predetermined condition according to a characteristic of the business, and registers the first word or the second word as the keyword in the keyword dictionary when the predetermined condition is satisfied; and A keyword dictionary creating system comprising:

12. The keyword dictionary creation system according to claim 11 ; a detection unit that determines whether the first word extracted from the first speech recognition result by the word extraction unit and the second word extracted from the second speech recognition result satisfy the predetermined condition according to the characteristics of the business, and if the predetermined condition is satisfied, checks whether the first word or the second word is registered in the keyword dictionary, and if registered, detects the first word or the second word as the keyword; an output unit that outputs the keyword detected by the detection unit to a user; A work support system comprising:

Citation Information

Patent Citations

  • In-speech important word extraction device and in-speech important word extraction using the device, and method and program thereof

    JP2015099289A

  • Dictionary creation device, dictionary creation program, speech recognition device, speech recognition program and recording medium

    JP2018049230A

  • Crane

    JP2019214466A

  • Automatic report creation system

    JP2021002391A

  • Work support management system

    JP2021051664A