Control system
By combining words with the same vowel arrangement or used together into single entries, the efficiency and accuracy of voice recognition for steering commands are improved, addressing the inefficiencies in existing systems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2026-04-10
AI Technical Summary
Existing voice recognition systems for steering commands in moving bodies require time-consuming trial and error to create pronunciation dictionaries due to misrecognition tendencies, necessitating an improvement in the efficiency of word registration.
A method and device for registering words in a pronunciation dictionary by combining words with the same vowel arrangement pattern or used together in dialogue into single words, reducing misrecognition and processing load.
Enhances the efficiency of registering words in the pronunciation dictionary, improving recognition accuracy and reducing misrecognition in voice recognition engines for steering systems.
Smart Images

Figure 0007843884000001 
Figure 0007843884000002 
Figure 0007843884000003
Abstract
Description
Technical Field
[0001] This disclosure relates to Control system .
Background Art
[0002] Normally, in a moving body such as a ship, a helmsman listens to a steering order uttered by an issuer (such as a captain), recognizes its meaning, and steers the ship. In recent years, in order to promote labor savings due to a decrease in the number of workers, it has been desired to use a voice recognition engine to steer without going through a helmsman by the utterance of a steering order by the issuer. For this purpose, it is necessary to improve the accuracy of voice recognition of the steering order.
[0003] As a technology related to this disclosure, Patent Document 1 discloses specifying instruction operation information by performing voice recognition processing on instruction voice data.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] In a general voice recognition engine, it is possible to customize a word to be recognized as a single word by registering a combination of a word (text information) and a pronunciation (acoustic model) in a pronunciation dictionary. In order to improve the recognition rate of the voice recognition engine, it is necessary to create this pronunciation dictionary, but it takes time because it requires trial and error while grasping the tendency of misrecognition of the voice recognition engine.
[0006] An object of this disclosure is to improve the efficiency of registering words in the pronunciation dictionary of a voice recognition engine used for steering (helm) of a moving body.
Means for Solving the Problems
[0007] According to one aspect of the present disclosure, a word registration method is a method for registering words in a pronunciation dictionary of a speech recognition engine used for controlling a mobile body, comprising the steps of: extracting words contained in lines spoken during the control; performing a first combining process in which, if there are multiple words among the extracted words that have the same vowel arrangement pattern, one of the words with the same vowel arrangement pattern is combined with the preceding and succeeding words in the line to form a single word; and registering the extracted words and the combined single word in the pronunciation dictionary.
[0008] According to one aspect of the present disclosure, a word registration device is a device for registering words in a pronunciation dictionary of a speech recognition engine used for controlling a mobile body, and comprises: an extraction unit for extracting words contained in lines spoken during the control; a first combining unit for combining one of the words with the same vowel arrangement pattern with the preceding and succeeding words in the line of speech to form a single word when there are multiple words with the same vowel arrangement pattern among the extracted words; and a registration unit for registering the extracted words and the combined word in the pronunciation dictionary.
[0009] According to one aspect of the present disclosure, a word registration method is a method for registering words in a pronunciation dictionary of a speech recognition engine used for controlling a mobile body, comprising the steps of: extracting words contained in lines spoken during the control; combining two words belonging to a pair of words used together in the lines into a single word when the last vowel of the first word matches the first vowel of the second word; and registering the extracted words and the combined word into the pronunciation dictionary. [Effects of the Invention]
[0010] According to the embodiments described above, the process of registering words into the pronunciation dictionary of a speech recognition engine used for controlling a mobile device can be made more efficient. [Brief explanation of the drawing]
[0011] [Figure 1] This diagram shows the overall configuration of the control device according to the first embodiment. [Figure 2] This figure shows an example of the data structure of a pronunciation dictionary according to the first embodiment. [Figure 3] This figure shows the flow of registering a word in a pronunciation dictionary according to the first embodiment. [Figure 4] This figure shows specific examples of each process in the registration process according to the first embodiment. [Figure 5] This figure shows specific examples of each process in the registration process according to the first embodiment. [Figure 6] This figure shows the configuration of a word registration device according to the second embodiment. [Modes for carrying out the invention]
[0012] <First Embodiment> The method for registering words in the pronunciation dictionary according to the first embodiment will be explained in detail below with reference to Figures 1 to 5.
[0013] (Overall configuration of the control system) Figure 1 shows the overall configuration of the control device according to the first embodiment. The steering control device 9 shown in Figure 1 is used for steering a ship, which is an example of a moving object. This steering control device automatically controls the steering of the ship (without the intervention of the pilot) in accordance with steering commands issued by the issuing officer.
[0014] The steering control device 9 includes a voice recognition engine 1 for recognizing spoken steering commands. The voice recognition engine 1 receives the spoken voice of the steering command and converts it into text information. The steering control device 9 interprets the meaning of the spoken content (steering command) based on the text information converted by the voice recognition engine 1 and performs steering control. The voice recognition engine 1 in this embodiment may utilize existing technology.
[0015] The speech recognition engine 1 includes a pronunciation dictionary 10. The pronunciation dictionary 10 is an information table in which words (text information) to be recognized and their pronunciations (acoustic models) are registered in association with each other. By referring to this pronunciation dictionary 10, the speech recognition engine 1 identifies a word whose input speech matches the acoustic model and converts it into text information.
[0016] (Example of pronunciation dictionary) FIG. 2 is a diagram showing an example of the data configuration of the pronunciation dictionary according to the first embodiment. As shown in FIG. 2, the pronunciation dictionary 10 associates a word (text) with a pronunciation (acoustic model). The "word" in FIG. 2 is text information, and in addition to arithmetic numbers such as "2" and "3", words such as "menjo" and "ippai" are registered. The "pronunciation" in FIG. 2 is the pronunciation corresponding to each "word". When there are multiple pronunciations for one word, all of them are associated. For example, for the word "2" as a number, multiple pronunciations such as "ni" (NI), "nii" (NII), and "huta" (HUTA) are associated.
[0017] (Flow of pronunciation dictionary registration work) FIG. 3 is a diagram showing the flow of the work of registering a word in the pronunciation dictionary according to the first embodiment. FIGS. 4 and 5 are diagrams showing specific examples of each process in the registration work according to the first embodiment. Hereinafter, the flow of the word registration method according to the present embodiment will be described in detail while referring to FIGS. 3 to 5.
[0018] The processing flow shown in FIG. 3 shows the flow of the work of registering a word to be recognized by the speech recognition engine 1 in the pronunciation dictionary 10 before the operation start of the operation control device 9.
[0019] As shown in Figure 3, first, the worker surveys all the phrases used as steering commands for ships (for example, about 100 different phrases) and extracts words to be registered (words to be registered as candidates) (Step S01). Here, for example, frequently used phrases such as "hard to port" and "hard to starboard" are added as registration candidates as single words. Also, since "full" is used in other phrases as well, it is added as a registration candidate as a separate word from "hard to port" and "hard to starboard". Furthermore, as mentioned above, for words that have multiple pronunciations, such as "2" (ni, ni, futa), all of the pronunciations are associated and added as registration candidates. In the following, the list of all words and their pronunciation combinations that were listed as potential registration candidates in Step S01 will also be referred to as the "potential registration word group."
[0020] Next, the operator determines whether there are multiple words in the list of candidate words for registration that have matching vowel arrangement patterns (step S02). If there are multiple words with matching vowel arrangement patterns (step S02; YES), the operator replaces one of those words with another word formed by combining it with the words before and after it in the dialogue (step S03).
[0021] Here, "vowel arrangement pattern" refers to the types of vowels ("A", "I", "U", "E", and "O"), their number, and their order in the pronunciation of each word. "Vowel arrangement pattern matches" means that the types, number, and order of vowels in the pronunciation of each word are exactly the same.
[0022] The processes described in steps S02 to S03 above will be explained in more detail with reference to Figure 4.
[0023] As shown in Figure 4, suppose the list of candidate words for registration includes the word "8" (hachi: HACHI) and the word "helm" (helm: KAJI). In this case, these two words, "8" and "helm," are words whose vowel arrangement pattern matches, "A" → "I".
[0024] In this case, the worker selects one of the target words (either "8" or "rudder") and replaces it with a single word formed by combining it with the words before and after it in the dialogue. For example, if the worker selects "rudder," they examine the words before and after "rudder" in all the dialogue used for steering commands.
[0025] As a result, suppose the word "helm" is used in two ways in all the steering commands: "steer" and "steer." In this case, the operator removes the word "helm" (kaji: KAJI) from the list of candidate words to register, and instead adds the two words "steer" (kajikire: KAJIKIRE) and "steer" (kajitore: KAJITORE) to the list of candidate words to register. Here, the word "steer" is formed by combining "helm" with "cut" which is in a related context in the dialogue, and the word "steer" is formed by combining "helm" with "take" which is in a related context in the dialogue. As described above, the word "kaji" (helm) is replaced by the two words "kajikire" (helm change) and "kajitore" (helm take). Note that the two newly added words, "kajikire" and "kajitore," do not have matching vowel arrangement patterns.
[0026] Returning to Figure 3, the operator then examines the list of candidate words obtained through steps S01 to S03 and determines whether there are any pairs of words used together in dialogue where the last vowel of the first word matches the first vowel of the second word (step S04). If there are pairs of words where the last vowel of the first word matches the first vowel of the second word (step S04; YES), the operator adds the combined word to the list of candidate words (step S05).
[0027] The processes described in steps S04 to S05 above will be explained in more detail with reference to Figure 5.
[0028] As shown in Figure 5, suppose the list of candidate words for registration includes the words "speed" (hayasa) and "3" (san). These two words, "speed" and "3," may be used together in a line of dialogue. For example, the line "speed 30" is composed of the words "speed" (hayasa) and "3" (san) connected in that order. Furthermore, the final vowel "A" of the first word "speed" (hayasa) matches the first vowel "A" of the second word "3" (san).
[0029] In this case, the worker combines the words "speed" and "3" into a single word, "speed3" (HAYASASAN), and adds it to the list of candidate words to register.
[0030] Similarly, suppose the list of candidate words for registration includes the word "5" (go) and the word "degree" (do). These two words, "5" and "degree," may be used together in a line of dialogue. For example, the line "15 degrees" is composed of the words "5" (go) and "degree" (do) connected in that order. Furthermore, the final vowel "O" of the first word "5" (go) matches the initial vowel "O" of the second word "degree" (do).
[0031] In this case, the worker combines the words "5" and "degree" into a single word, "5 degrees" (godo), and adds this to the list of candidate words to register.
[0032] Returning to Figure 3, the operator then registers the words included in the list of candidate words obtained through steps S02-S03 and S04-S05 into the pronunciation dictionary. Then, the actual recognition rate for each word is checked (step S06). For each word, determine whether the recognition rate is below 100%. For words where the recognition rate is below 100% (Step S07; NO), customize the pronunciation dictionary individually for each word (Step S08). Examples of individual customization for each word are as follows: (Example) In the case of misrecognition where a number not found in the dictionary's dialogue (steering commands) appears (for example, the dictionary dialogue "Hard to starboard" is misrecognized as "Hard to starboard 5"), a logic process is applied to restrict the use of numbers so that they are only used with units such as "degrees".
[0033] The registration process is completed when a 100% recognition rate is achieved for all words (Step S07; YES).
[0034] (Effect, Action) As described above, the method for registering a pronunciation dictionary according to this embodiment is characterized in that, when there are multiple words with the same vowel arrangement pattern in the candidate word group for registration, one of the words with the same vowel arrangement pattern is combined with the words before and after it in the dialogue to form a single word (first combining process) (steps S02 to S03 in Figure 3).
[0035] Here, we investigated the tendency of misrecognition by speech recognition engine 1 (an existing speech recognition engine) and found that words with the same vowel arrangement pattern tend to be misrecognized by each other. For example, the word "8" (hachi: HACHI) and the word "helm" (kaji: KAJI) were found to be easily misrecognized by the speech recognition engine. Therefore, by doing as described above, it is possible to avoid registering words that are easily misrecognized by each other in the pronunciation dictionary. For example, the words "8" (hachi: HACHI) and "helm" (kaji: KAJI) are easily misrecognized, but the relationship between "8" (hachi: HACHI) and "helm break" (kajikire: KAJIKIRE) or "steer" (kajitore: KAJITORE) is not one that is easily misrecognized by speech recognition engine 1 because their vowel arrangement patterns are different.
[0036] Furthermore, the method for registering a pronunciation dictionary according to this embodiment is characterized in that, when two words are used together in a line of dialogue, if the last vowel of the first word matches the first vowel of the second word, a process is performed to combine the two words belonging to that pair into a single word (a second combining process) (steps S04 to S05 in Figure 3).
[0037] In this study, we investigated the tendency for misrecognition by speech recognition engine 1 and found that when two words are used together, misrecognition is more likely to occur when the last vowel of the first word matches the first vowel of the second word. For example, when the phrase "speed 3" is included in the dialogue, speech recognition engine 1 tends to have difficulty recognizing it as a combination of the words "speed" and "3". Similarly, when the phrase "5 degrees" is included in the dialogue, speech recognition engine 1 tends to have difficulty recognizing it as a combination of the words "5" and "degree". Therefore, by doing as described above, the phrase "speed 3" that appears in the dialogue is recognized as a single word, "speed 3" (haysasan). As a result, the speech recognition engine 1 will recognize the phrase "speed 3" as a single word, rather than trying to recognize it as a combination of two words, "speed" and "3," thus suppressing misrecognition.
[0038] Furthermore, in the method for registering a pronunciation dictionary according to this embodiment, steps S02 to S03 (first merging process) in Figure 3 are executed first, followed by steps S04 to S05 (second merging process). By executing the processes in steps S02 to S03 first, the number of words targeted for processing in steps S04 to S05 can be reduced. Therefore, the overall processing load of the processing flow in Figure 3 can be reduced.
[0039] (Modified version of the first embodiment) In the first embodiment described above, both steps S02-S03 (first merging process) and steps S04-S05 (second merging process) in Figure 3 were included, but other embodiments are not limited to this configuration. For example, in the word registration method according to other embodiments, only one of steps S02-S03 (first merging process) or steps S04-S05 (second merging process) may be performed.
[0040] Furthermore, although the first embodiment described above involved replacing words in step S03 (first merging process) in Figure 3, other embodiments are not limited to this configuration. For example, in other embodiments, step S03 in Figure 3 may involve adding words that have been combined into a single word to the list of candidate words for registration.
[0041] <Second Embodiment> In the first embodiment, the process flow shown in Figure 3 was described as being executed by an operator and the registration work being performed manually. However, in other embodiments, the system is not limited to this configuration. That is, the process flow shown in Figure 3 may be executed automatically by the device. A word registration device according to the second embodiment will be described with reference to Figure 6.
[0042] (Configuration of the word registration device) Figure 6 shows the configuration of a word registration device according to the second embodiment. As shown in Figure 6, the word registration device 2 is a device that registers words in the pronunciation dictionary 10 of the speech recognition engine 1 used for steering (operating) a ship (mobile object). The word registration device 2 has a typical computer configuration, consisting of a CPU 20, memory 21, input / output interface 22, and recording medium 23.
[0043] The CPU 20 is the processor that controls the overall operation of the word registration device 2, and performs various functions by operating according to a predetermined program. The functions of the CPU 20 will be described later. Memory 21 is the so-called main memory, and it is where programs and other data necessary to operate the CPU 20 are loaded. The input / output interface 22 is a communication interface for exchanging information with external devices. In this embodiment, the input / output interface 22 accepts input information indicating the content of all the lines used as steering commands. The recording medium 23 is a high-capacity storage device such as an HDD or SSD.
[0044] The CPU 100 operates according to the program, performing functions as an extraction unit 201, a first merging unit 202, a second merging unit 203, and a registration unit 204.
[0045] The extraction unit 201 extracts words contained in the lines spoken during steering (operation) (step S01 in Figure 3).
[0046] The first combining unit 202, if there are multiple words among the extracted words that have the same vowel arrangement pattern, combines one of those words with the preceding and succeeding words in the dialogue into a single word (steps S02-S03 in Figure 3).
[0047] The second combining unit 203 combines two words belonging to a pair of words used to connect in a line of dialogue into a single word if the last vowel of the first word matches the first vowel of the second word (steps S04-S05 in Figure 3).
[0048] The registration unit 204 registers the extracted words and the words that have been combined into a single word in the pronunciation dictionary 10 (step S06 in Figure 3).
[0049] In the above-described embodiment, the processes of various operations performed by the word registration device 2 are stored in program form on a computer-readable recording medium, and the various operations are performed by a computer reading and executing this program. The computer-readable recording medium refers to magnetic disks, magneto-optical disks, CD-ROMs, DVD-ROMs, semiconductor memory, etc. Alternatively, this computer program may be distributed to a computer via a communication line, and the computer that receives the distribution may execute the program.
[0050] The above program may be intended to implement some of the functions described above. Furthermore, it may be a so-called differential file (differential program) that can implement the above functions in combination with a program already recorded in the computer system.
[0051] As described above, several embodiments relating to this disclosure have been explained, but all of these embodiments are presented as examples and are not intended to limit the scope of the invention. These embodiments can be carried out in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their variations are included in the scope and spirit of the invention, as well as in the claims and their equivalents.
[0052] <Note> The word registration method and word registration device described in each embodiment can be understood, for example, as follows:
[0053] (1) In the first embodiment, the word registration method is a method for registering words in a pronunciation dictionary 10 of a speech recognition engine 1 used for steering (operating) a ship (mobile body), and comprises the steps of: extracting words contained in lines spoken during steering; if there are multiple words among the extracted words that have the same vowel arrangement pattern, performing a first combining process to combine one of the words with the preceding and succeeding words in the line into a single word; and registering the extracted words and the combined single word in the pronunciation dictionary.
[0054] (2) In a second embodiment, the word registration method further includes a second merging process in which, when the last vowel of the first word and the first vowel of the second word of a pair of words used together in a line of dialogue match, the two words belonging to that pair are combined into a single word.
[0055] (3) In a third embodiment, the word registration method performs a second merging process after performing a first merging process.
[0056] (4) In a fourth embodiment, the word registration device 2 is a device for registering words in the pronunciation dictionary of a speech recognition engine used for controlling a mobile body, and comprises: an extraction unit 201 that extracts words contained in lines spoken during operation; a first combining unit 202 that, if there are multiple words among the extracted words that have the same vowel arrangement pattern, combines one of the words with the preceding and succeeding words in the line to form a single word; and a registration unit 204 that registers the extracted words and the combined word in the pronunciation dictionary.
[0057] (5) In a fifth embodiment, the word registration method is a method for registering words in a pronunciation dictionary of a speech recognition engine used for controlling a mobile body, comprising the steps of: extracting words contained in lines spoken during the control; combining two words belonging to a pair of words used together in the lines into one word when the last vowel of the first word matches the first vowel of the second word; and registering the extracted words and the combined word into the pronunciation dictionary. [Explanation of symbols]
[0058] 9. Control System 1. Speech Recognition Engine 10 Pronunciation Dictionary 2. Word Registration Device
Claims
1. A control device for maneuvering a mobile object, An extraction unit that extracts words contained in the lines spoken during the aforementioned operation, If there are multiple words among the extracted words that have the same vowel arrangement pattern, a first combining processing unit combines one of the words with the preceding and succeeding words in the dialogue to form a single word. A registration unit that registers the extracted words and words combined into a single word into a pronunciation dictionary, A speech recognition engine that takes the spoken dialogue uttered during piloting as input, identifies words registered in the pronunciation dictionary from the spoken dialogue, and converts them into text information. A means for interpreting the meaning of spoken content based on text information converted by the aforementioned speech recognition engine and performing steering control, A control system equipped with a flight control device.
2. The system further includes a second combining processing unit that, when the last vowel of the first word and the first vowel of the second word in a pair of words used together in the aforementioned dialogue match, combines the two words belonging to that pair into a single word. The control device according to claim 1.
3. After the first combining unit performs the process of combining into a single word, the second combining unit performs the process of combining into a single word. The control device according to claim 2.
Citation Information
Patent Citations
Vocaburary registration aid for voice recognition equipment
JP1988292197A
Object sound processor and transport equipment system using same, and object sound processing method
JP2005338286A
Voice recognition apparatus, voice recognition apparatus for picking, and voice recognition method
JP2012037820A
Crane
JP2019214466A
Information provision system
WO2016042600A1