Control device

JP2025087909A5Active Publication Date: 2025-06-30MITSUBISHI HEAVY IND LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025041376
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-30
Estimated Expiration
2041-05-31

AI Technical Summary

Technical Problem

The efficiency of word registration in the pronunciation dictionary of a speech recognition engine used for steering a moving body is low due to the time-consuming process of trial and error required to improve recognition accuracy.

Method used

A method and device for registering words in the pronunciation dictionary that involves extracting words from phrases, combining words with matching vowel arrangement patterns, and registering these combined words to improve recognition efficiency.

Benefits of technology

This approach enhances the efficiency of word registration by reducing misrecognition errors and streamlining the process, thereby improving the overall accuracy of speech recognition for steering commands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To improve efficiency of registering words to a pronunciation dictionary of a voice recognition engine used for maneuvering a moving object.SOLUTION: A word registration method for registering words in a pronunciation dictionary of a voice recognition engine used to input dialogues uttered during maneuvering of a moving object, comprises: a step of extracting words included in the dialogues uttered during the maneuvering; a first gathering step of gathering one of the words with the same vowel arrangement pattern into one word by combining it with preceding and following words in the dialogues, when multiple words with the same vowel arrangement pattern among the extracted words exist; and a step of registering the extracted words and the words gathered into the one word to the pronunciation dictionary.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a word registration method and a word registration device.

Background Art

[0002] Normally, in a moving body such as a ship, a helmsman listens to a steering order uttered by an issuer (such as a captain), recognizes its meaning, and performs steering. In recent years, in order to promote labor saving due to a decrease in the number of workers, it has been desired to use a speech recognition engine to perform steering without going through a helmsman by the utterance of a steering order by the issuer. For this purpose, it is necessary to improve the accuracy of speech recognition of the steering order.

[0003] As a technique related to the present disclosure, Patent Document 1 discloses specifying instruction operation information by performing speech recognition processing on instruction speech data.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] In a general speech recognition engine, a word to be recognized as a single word can be customized by registering a combination of a word (text information) and a pronunciation (acoustic model) in a pronunciation dictionary. In order to improve the recognition rate of the speech recognition engine, it is necessary to create this pronunciation dictionary, but it takes time because it requires trial and error while grasping the tendency of misrecognition of the speech recognition engine.

[0006] An object of the present disclosure is to improve the efficiency of word registration work in a pronunciation dictionary of a speech recognition engine used for steering (helm) of a moving body.

Means for Solving the Problems

[0007] According to one aspect of the present disclosure, a word registration method is a method for registering a word in a pronunciation dictionary of a speech recognition engine used for controlling a moving body, the method including: extracting a word included in a phrase uttered during the control; and when there are a plurality of words having the same vowel arrangement pattern among the extracted words, performing a first combining process of combining one of the words having the same vowel arrangement pattern with the words before and after in the phrase into one word; and registering the extracted words and the words combined into one word in the pronunciation dictionary.

[0008] According to one aspect of the present disclosure, a word registration device is a device for registering a word in a pronunciation dictionary of a speech recognition engine used for controlling a moving body, the device including: an extraction unit configured to extract a word included in a phrase uttered during the control; a first combining unit configured to, when there are a plurality of words having the same vowel arrangement pattern among the extracted words, combine one of the words having the same vowel arrangement pattern with the words before and after in the phrase into one word; and a registration unit configured to register the extracted words and the words combined into one word in the pronunciation dictionary.

[0009] According to one aspect of the present disclosure, a word registration method is a method for registering a word in a pronunciation dictionary of a speech recognition engine used for controlling a moving body, the method including: extracting a word included in a phrase uttered during the control; and when the last vowel of the first word and the first vowel of the second word match among pairs of two words used in combination in the phrase, combining the two words belonging to the pair into one word; and registering the extracted words and the words combined into one word in the pronunciation dictionary.

Effect of the Invention

[0010] According to each of the above aspects, it is possible to improve the efficiency of the word registration work in the pronunciation dictionary of the speech recognition engine used for controlling the moving body.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Mode for Carrying Out the Invention

[0012] <First Embodiment> Hereinafter, a method for registering a word in the pronunciation dictionary according to the first embodiment will be described in detail with reference to FIGS. 1 to 5.

[0013] (Overall Configuration of the Steering Control Device) FIG. 1 is a diagram showing the overall configuration of the steering control device according to the first embodiment. The steering control device 9 shown in FIG. 1 is used for steering (rudder control) of a ship, which is an example of a moving body. This steering control device automatically controls the steering (rudder control) of the ship according to the voice of the steering command issued by the issuer (without going through the operator).

[0014] The steering control device 9 includes a voice recognition engine 1 for recognizing the issued steering command. The voice recognition engine 1 inputs the voice of the steering command and converts it into text information. The steering control device 9 interprets the meaning of the uttered content (steering command) based on the text information converted by the voice recognition engine 1 and performs steering control. Note that the voice recognition engine 1 according to this embodiment may utilize existing technologies.

[0015] The speech recognition engine 1 includes a pronunciation dictionary 10. The pronunciation dictionary 10 is an information table in which words (text information) to be recognized and their pronunciations (acoustic models) are registered in association with each other. By referring to this pronunciation dictionary 10, the speech recognition engine 1 identifies a word whose input speech matches the acoustic model and converts it into text information.

[0016] (Example of pronunciation dictionary) FIG. 2 is a diagram showing an example of the data configuration of the pronunciation dictionary according to the first embodiment. As shown in FIG. 2, in the pronunciation dictionary 10, a word (text) and a pronunciation (acoustic model) are associated with each other. The "word" in FIG. 2 is text information, and in addition to arithmetic numbers such as "2" and "3", words such as "menjo" and "ippai" are registered. The "pronunciation" in FIG. 2 is the pronunciation corresponding to each "word". When there are multiple pronunciations for one word, all of them are associated. For example, for the word "2" as a number, multiple pronunciations such as "ni" (NI), "nii" (NII), and "huta" (HUTA) are associated.

[0017] (Flow of pronunciation dictionary registration work) FIG. 3 is a diagram showing the flow of the work of registering words in the pronunciation dictionary according to the first embodiment. FIGS. 4 and 5 are diagrams showing specific examples of each process in the registration work according to the first embodiment. Hereinafter, with reference to FIGS. 3 to 5, the flow of the word registration method according to the present embodiment will be described in detail.

[0018] The processing flow shown in FIG. 3 shows the flow of the work of registering words to be recognized by the speech recognition engine 1 in the pronunciation dictionary 10 at the stage before the operation start of the control device 9.

[0019] As shown in FIG. 3, first, the operator views all the phrases (for example, about 100 phrases) used as the ship's steering commands from above and extracts the words to be registered (the words to be candidates for registration) (step S01). Here, for example, frequently used phrases such as "fully steering" (とりかじいっぱい) and "fully hard steering" (おもかじいっぱい) are added to the registration candidates with the phrase itself as one word. Also, since "いっぱい" (ippai) is used in other phrases as well, it is added to the registration candidates as a word different from "とりかじいっぱい" and "おもかじいっぱい". Also, as described above, for a word that has multiple pronunciations, such as "2" (に, にー, ふた), all of its pronunciations are associated and added to the registration candidates. Hereinafter, in step S01, the list of all combinations of words and pronunciations listed as registration candidates is also referred to as the "registration candidate word group".

[0020] Next, the operator determines whether there are multiple words with matching vowel arrangement patterns among the registration candidate word group listed in step S01 (step S02). If there are multiple words with matching vowel arrangement patterns (step S02; YES), the operator replaces one of the words with another single word combined with the words before and after in the phrase (step S03).

[0021] Here, the "vowel arrangement pattern" refers to the type of vowels ("A", "I", "U", "E", "O", 5 types), the number, and the order in the pronunciation of each word. "The vowel arrangement patterns match" means that in the pronunciation of each word, the type, number, and order of vowels completely match.

[0022] The processing of steps S02 to S03 described above will be explained in more detail with reference to FIG. 4.

[0023] As shown in FIG. 4, assume that among the registered candidate word group, there are words "8" (hachi: HACHI) and "rudder" (kaji: KAJI). In this case, these two words "8" and "rudder" are words whose vowel arrangement patterns match from "A" → "I".

[0024] In this case, the operator selects one of the target words (either "8" or "rudder") and replaces it with a single word combined with the words before and after the word appears in the dialogue. Here, for example, if the operator selects "rudder", the operator checks the words before and after "rudder" in all the dialogues used for the steering command.

[0025] As a result, assume that the word "rudder" is used in two ways, "rudder break" and "rudder take", in all the dialogues of the steering command. In this case, the operator excludes the word "rudder" (kaji: KAJI) from the registered candidate word group, and instead adds two words, "rudder break" (kajikire: KAJIKIRE) and "rudder take" (kajitore: kajitore), to the registered candidates. Here, the word "rudder break" is a word formed by combining "break" that is in a context relationship with "rudder" in the dialogue into a single word, and the word "rudder take" is a word formed by combining "take" that is in a context relationship with "rudder" in the dialogue into a single word. In the above manner, the word "rudder" will be replaced with two words, "rudder break" and "rudder take". Note that the two newly added words, "rudder break" and "rudder take", are not words whose vowel arrangement patterns match.

[0026] Returning to FIG. 3, next, the operator determines whether there is a pair of two words that are used connected in the dialogue among the registered candidate word group obtained through steps S01 to S03, and whether the last vowel of the first word and the first vowel of the second word match (step S04). If there is a pair of words where the last vowel of the first word and the first vowel of the second word match (step S04; YES), the operator adds a single word formed by combining these two words to the registered candidate word group (step S05).

[0027] The processing of steps S04 to S05 described above will be described in more detail with reference to FIG. 5.

[0028] As shown in FIG. 5, assume that among the registration candidate word group, there are words "speed" (hayasa: HAYASA) and "3" (san: SAN). These two words "speed" and "3" may be used by connecting them before and after in the dialogue. For example, the dialogue "speed 30" is composed of connecting the word "speed" (hayasa: HAYASA) and the word "3" (san: SAN) in this order. Furthermore, the last vowel "A" of the first word "speed" (hayasa: HAYASA) matches the first vowel "A" of the second word "3" (san: SAN).

[0029] In this case, the operator combines the word pair of "speed" and "3" into one word "speed 3" (hayasasan: HAYASASAN), and adds this to the registration candidate word group.

[0030] Similarly, assume that among the registration candidate word group, there are words "5" (go: GO) and "degree" (do: DO). These two words "5" and "degree" may be used by connecting them before and after in the dialogue. For example, the dialogue "15 degrees" is composed of connecting the word "5" (5: GO) and the word "degree" (do: DO) in this order. Furthermore, the last vowel "O" of the first word "5" (go: GO) matches the first vowel "O" of the second word "degree" (do: DO).

[0031] In this case, the operator combines the word pair of "5" and "degree" into one word "5 degrees" (godo: GODO), and adds this to the registration candidate word group.

[0032] Returning to FIG. 3, next, the operator registers the words included in the registration candidate word group obtained through the processing of steps S02 to S03 and steps S04 to S05 in the pronunciation dictionary. And confirm the actual recognition rate for each word (step S06). Determine whether the recognition rate for each word is less than 100%. For words with a recognition rate less than 100% (step S07; NO), customize the pronunciation dictionary individually for each word (step S08). Examples of individual customization for each word are as follows. (Example) For misrecognition where a number that does not exist in the dictionary's dialogue (steering command) appears (for example, the dictionary dialogue "Full rudder" is misrecognized as "Full rudder 5"), perform logic processing to limit the use of the number so that it is only used together with units such as "degree".

[0033] If the recognition rate of 100% is obtained for all words, end the registration work (step S07; YES).

[0034] (Function, effect) As described above, the method for registering a pronunciation dictionary according to the present embodiment is characterized in that, in the registration candidate word group, when there are a plurality of words having the same vowel arrangement pattern, one of the words having the same vowel arrangement pattern is combined with the words before and after in the dialogue to form one word (first combining process) (steps S02 to S03 in FIG. 3).

[0035] Here, as a result of investigating the tendency of misrecognition with respect to the speech recognition engine 1 (speech recognition engine of the prior art), it was found that words having the same vowel arrangement pattern tend to be misrecognized with each other. For example, it was found that the word "8" (hachi: HACHI) and the word "rudder" (kaji: KAJI) are likely to be misrecognized by the speech recognition engine with each other. Therefore, by doing as described above, it is possible to avoid registering words that are likely to be misrecognized with each other in the pronunciation dictionary. For example, although the word "8" (hachi: HACHI) and the word "rudder" (kaji: KAJI) are likely to be misrecognized, if it is a relationship between "8" (hachi: HACHI) and "rudder break" (rudder break: KAJIKIRE) or "rudder removal" (kajitore: KAJITORE), since the vowel arrangement patterns are different from each other, it does not become a relationship that is likely to be misrecognized by the speech recognition engine 1.

[0036] In addition, the method for registering a pronunciation dictionary according to the present embodiment is characterized in that when the last vowel of the first word and the first vowel of the second word in a pair of two words used consecutively in a line of dialogue match, a process of combining the two words belonging to the pair into one word (second combining process) is performed (steps S04 to S05 in FIG. 3).

[0037] Here, as a result of investigating the tendency of misrecognition with respect to the speech recognition engine 1, it was found that misrecognition is likely to occur in the speech recognition engine 1 when the last vowel of the first word and the first vowel of the second word in a pair of two words used consecutively match. For example, when "speed 3" is included in a line of dialogue, the speech recognition engine 1 tends to have difficulty recognizing this as a combination of the words "speed" and "3". Similarly, when "5 degrees" is included in a line of dialogue, it was found that the speech recognition engine 1 tends to have difficulty recognizing this as a combination of the words "5" and "degrees". Therefore, by doing the above, "speed 3" that appears in a line of dialogue is recognized as one word, "speed 3" (nagasa-san: HAYSASAN). As a result, the speech recognition engine 1 will recognize the line of dialogue "speed 3" as one word, "speed 3", rather than as a combination of two words, "speed" and "3", so misrecognition is suppressed.

[0038] In addition, the method for registering a pronunciation dictionary according to the present embodiment executes steps S04 to S05 (second combining process) after executing steps S02 to S03 (first combining process) in FIG. 3. In this way, by first executing the processes of steps S02 to S03, the number of words targeted for the processes of steps S04 to S05 can be reduced. Therefore, the processing load of the entire processing flow in FIG. 3 can be reduced.

[0039] (Modification example of the first embodiment) In the above-described first embodiment, it has been described as including both steps S02 to S03 (first combining process) and steps S04 to S05 (second combining process) in FIG. 3. However, in other embodiments, it is not limited to this mode. For example, in the word registration method according to other embodiments, it may be a mode in which only one of steps S02 to S03 (first combining process) and steps S04 to S05 (second combining process) is performed.

[0040] Also, in the above-described first embodiment, it has been described that words are replaced in step S03 (first combining process) in FIG. 3. However, in other embodiments, it is not limited to this mode. For example, in other embodiments, in step S03 of FIG. 3, the words grouped into one word may be added to the registration candidate word group.

[0041] <Second Embodiment> In the first embodiment, it has been described that the operator executes the processing flow shown in FIG. 3 and performs the registration work manually. However, in other embodiments, it is not limited to this mode. That is, the processing flow shown in FIG. 3 may be automatically executed by the device. The word registration device according to the second embodiment will be described with reference to FIG. 6.

[0042] (Configuration of Word Registration Device) FIG. 6 is a diagram showing the configuration of the word registration device according to the second embodiment. As shown in FIG. 6, the word registration device 2 is a device for registering words in the pronunciation dictionary 10 of the speech recognition engine 1 used for steering (control) of a ship (mobile body). The word registration device 2 has a configuration of a general computer and includes a CPU 20, a memory 21, an input / output interface 22, and a recording medium 23.

[0043] The CPU 20 is a processor that controls the overall operation of the word registration device 2 and exhibits various functions by operating according to a predetermined program. Each function of the CPU 20 will be described later. The memory 21 is a so-called main memory device, and programs and the like for operating the CPU 20 are loaded into it. The input / output interface 22 is a communication interface for exchanging information with external devices. In the present embodiment, the input / output interface 22 receives an input of information indicating the content of all the lines used as steering commands. The recording medium 23 is a large-capacity storage device such as an HDD or an SSD.

[0044] The CPU 100 functions as an extraction unit 201, a first combination processing unit 202, a second combination processing unit 203, and a registration unit 204 by operating according to a program.

[0045] The extraction unit 201 extracts words included in the lines uttered during steering (operation) (step S01 in FIG. 3).

[0046] When there are a plurality of words having the same vowel arrangement pattern among the extracted words, the first combination processing unit 202 combines one of the words having the same vowel arrangement pattern with the preceding and succeeding words in the line into one word (steps S02 to S03 in FIG. 3).

[0047] When the last vowel of the first word and the first vowel of the second word match among the pairs of two words used connected in the line, the second combination processing unit 203 combines the two words belonging to the pair into one word (steps S04 to S05 in FIG. 3).

[0048] The registration unit 204 registers the extracted words and the words combined into one word in the pronunciation dictionary 10 (step S06 in FIG. 3).

[0049] In the above-described embodiments, the processes of various processes executed by the word registration device 2 are stored in a computer-readable recording medium in the form of a program, and the computer reads and executes this program to perform the above-described various processes. Further, the computer-readable recording medium refers to a magnetic disk, a magneto-optical disk, a CD-ROM, a DVD-ROM, a semiconductor memory, or the like. Further, this computer program may be distributed to a computer via a communication line, and the computer that has received this distribution may execute the program.

[0050] The above program may be for realizing a part of the functions described above. Further, it may be a so-called difference file (difference program) that can realize the above-described functions in combination with a program already recorded in a computer system.

[0051] As described above, some embodiments according to the present disclosure have been described, but all of these embodiments are presented as examples and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, replacements, and changes can be made without departing from the gist of the invention. These embodiments and their modifications are included in the scope and gist of the invention, and are also included in the invention described in the claims and the equivalent scope thereof.

[0052] <Supplementary Note> The word registration method and the word registration device described in each embodiment are understood as follows, for example.

[0053] (1) In a first aspect, a word registration method is a method for registering a word in a pronunciation dictionary 10 of a speech recognition engine 1 used for steering (operation) of a ship (mobile body). The method includes steps of extracting a word included in a phrase pronounced during steering, and when there are a plurality of words having the same vowel arrangement pattern among the extracted words, performing a first combining process of combining one of the words having the same vowel arrangement pattern with the words before and after in the phrase into one word, and registering the extracted words and the words combined into one word in the pronunciation dictionary.

[0054] (2) In a second aspect, the word registration method further includes a step of performing a second combining process of combining two words belonging to a set into one word when the last vowel of the first word and the first vowel of the second word match among two-word sets used connectedly in a phrase.

[0055] (3) In a third aspect, the word registration method executes the second combining process after executing the first combining process.

[0056] (4) In a fourth aspect, a word registration device 2 is a device for registering a word in a pronunciation dictionary of a speech recognition engine used for operation of a mobile body. The device includes an extraction unit 201 that extracts a word included in a phrase pronounced during operation, a first combining process unit 202 that, when there are a plurality of words having the same vowel arrangement pattern among the extracted words, combines one of the words having the same vowel arrangement pattern with the words before and after in the phrase into one word, and a registration unit 204 that registers the extracted words and the words combined into one word in the pronunciation dictionary.

[0057] (5) In a fifth aspect, the word registration method is a method for registering words in a pronunciation dictionary of a speech recognition engine used for controlling a moving body, the method including: extracting words included in a script uttered during the control; combining, into one word, two words belonging to a pair of two words used consecutively in the script when the last vowel of the first word and the first vowel of the second word match; and registering the extracted words and the words combined into one word in the pronunciation dictionary.

Explanation of Signs

[0058] 9 Steering control device 1 Speech recognition engine 10 Pronunciation dictionary 2 Word registration device

Claims

1. A control device for controlling a moving object, comprising: An extraction unit that extracts words included in lines uttered during the operation; a first combining unit that combines, when there are a plurality of words having the same vowel arrangement pattern among the extracted words, one of the words having the same vowel arrangement pattern with a word before and after the word in the dialogue to form a single word; a registration unit that registers the extracted words and words grouped into one word in a pronunciation dictionary; a speech recognition engine that receives an input of a speech spoken during flight, identifies words registered in the pronunciation dictionary from the speech, and converts the words into text information; A means for interpreting the meaning of the spoken content based on the text information converted by the voice recognition engine and performing steering control; A steering control device comprising:

2. The system further includes a second combining unit that combines two words in a pair of words used in conjunction in the dialogue into one word when the final vowel of the first word matches the initial vowel of the second word. The steering control device according to claim 1 .

3. After the first combining processing unit executes the process of combining into one word, the second combining processing unit executes the process of combining into one word. The steering control device according to claim 2.