Differential decimation apparatus, method, and program
Through the text acquisition, pronunciation string conversion and mark string conversion of the differential extraction device, the problem of unknown words being registered incorrectly is solved, and the accuracy and efficiency of dictionary registration are improved.
Patent Information
- Application Number
- CN202111008156.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-11-04
- Filing Date
- 2021-08-31
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2041-08-31
AI Technical Summary
In the prior art, unknown words are mistakenly registered as dictionary candidates, resulting in unnecessary dictionary registration.
The difference extraction device utilizes a combination of text acquisition, pronunciation string conversion, tag string conversion and comparison unit to extract the difference between the input tag string and the output tag string, thereby preventing the registration of unknown words with correct tags among unknown words.
The registration of unknown words that are correctly marked even if not registered is effectively prevented, thereby improving the accuracy and efficiency of dictionary registration and reducing unnecessary dictionary registration.
Smart Images

Figure CN114519998B_ABST
Abstract
Description
[0001] Reference of Related Applications Based on Priority Application
[0002] This application is based on Japanese Patent Application No. 2020-184610 filed on November 4, 2020, the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0003] Embodiments of the present application relate to a differential extraction device, method, and program. BACKGROUND
[0004] Generally, a technology for assisting a user's dictionary registration work by searching for an unknown word that is not registered in a dictionary as a candidate for dictionary registration is being developed. As such a technology, for example, a method is known in which a compound word is extracted from a result obtained by morphological analysis of a text, and if the compound word is not registered in a constructed dictionary, the compound word is regarded as an unknown word.
[0005] This method is generally not particularly problematic, but according to the present inventor's research, sometimes an unknown word that becomes a correct tag even if not registered is also extracted as a candidate for dictionary registration. In this case, a word that does not need to be registered is registered. SUMMARY
[0006] The present application is to solve the problem of providing a differential extraction device, method, and program that can prevent registration of an unknown word that becomes a correct tag even if not registered among unknown words.
[0007] The differential extraction device of the embodiment has a text acquisition unit, a pronunciation string conversion unit, a tag string conversion unit, and a comparison unit. The text acquisition unit acquires a text in which an input tag string is described. The pronunciation string conversion unit converts the input tag string into a pronunciation string. The tag string conversion unit converts the pronunciation string into an output tag string. The comparison unit compares the input tag string and the output tag string to extract a difference.
[0008] According to the differential extraction device of the above structure, it is possible to prevent registration of an unknown word that becomes a correct tag even if not registered among unknown words. BRIEF DESCRIPTION OF DRAWINGS
[0009] Figure 1 is a block diagram illustrating the structure of the differential extraction device of the first embodiment.
[0010] Figure 2 is a schematic diagram illustrating conversion from a pronunciation string to a tag string in the first embodiment.
[0011] Figure 3 is a schematic diagram for explaining the comparison unit in the first embodiment.
[0012] Figure 4 is a schematic view illustrating a display screen in the first embodiment.
[0013] Figure 5 is a flowchart for explaining the operation in the first embodiment.
[0014] Figure 6 is a schematic view for explaining the operation in the first embodiment.
[0015] Figure 7 is a schematic view illustrating syllables of Japanese in the first embodiment.
[0016] Figure 8 is a schematic view illustrating a pronunciation state sound score vector in the first embodiment.
[0017] Figure 9 is a block diagram illustrating a feature quantity conversion section of a modification example of the first embodiment.
[0018] Figure 10 is a flowchart for explaining the operation in the modification example of the first embodiment.
[0019] Figure 11 is a block diagram illustrating a structure of a difference extraction device of the second embodiment.
[0020] Figure 12 is a flowchart for explaining the operation in the second embodiment.
[0021] Figure 13 is a schematic view for explaining a word presumption section in the second embodiment.
[0022] Figure 14 is a schematic view illustrating a display screen in the second embodiment.
[0023] Figure 15 is a schematic view showing an instruction example in the second embodiment.
[0024] Figure 16 is a block diagram illustrating a structure of a difference extraction device of the third embodiment.
[0025] Figure 17 is a flowchart for explaining the operation in the third embodiment.
[0026] Figure 18 is a schematic view illustrating a display screen in the third embodiment.
[0027] Figure 19 is a schematic view showing a registration example of a word registration section in the third embodiment.
[0028] Figure 20is a diagram showing an example of a display at the time of registration reflection of the third embodiment.
[0029] Figure 21 is a diagram showing a display screen and a registration screen of the third embodiment.
[0030] Figure 22 is a block diagram illustrating a hardware structure of the differential extraction device of the fourth embodiment.
[0031] Symbol explanation
[0032] 1: differential extraction device; 10: text acquisition section; 20: pronunciation string conversion section; 21: morpheme analysis section; 22: reading sound addition processing section; 30: mark string conversion section; 31: feature quantity conversion section; 31a: sound synthesis section; 31b: sound feature quantity calculation section; 31c: sound score calculation section; 32: conversion section; 33: storage section; 34: language model; 35: word dictionary; 40: comparison section; 50: word estimation section; 60: word category determination section; 61: unknown word determination section; 62: mark fluctuation determination section; 70: display control section; 71: display; 80: instruction section; 81: mouse; 90: word registration section; 101: text read-in button; 101a: open button; 102: input mark string display screen; 103: output mark string display screen; 104, 600 to 602: display attribute; 400: cursor; 401: word candidate screen; 402: word candidate; 403: range correction button; 700: word registration button; 701, 800: word registration screen; 702: mark input frame; 703: pronunciation registration frame; 704: part-of-speech registration frame; 705, 802: registration button; 801: effective display. DETAILED DESCRIPTION
[0033] Hereinafter, the device, method, and storage medium of the embodiments will be described with reference to the drawings. In the following description, the case where the differential extraction device is mounted on a voice recognition system and used for extracting a word registered in a word dictionary for voice recognition will be described as an example. Further, in order to easily understand the use, the differential extraction device can also be called by any name such as a word extraction device, a word extraction support device, a dictionary registration device, or a dictionary registration support device.
[0034] <First Embodiment>
[0035] Figure 1 is a block diagram illustrating the structure of the differential extraction device of the first embodiment. The differential extraction device 1 is provided with a text acquisition section 10, a pronunciation string conversion section 20, a mark string conversion section 30, a comparison section 40, and a display control section 70.
[0036] Here, the text acquisition unit 10 acquires text containing the input token string. The acquired text is sent to the pronunciation string conversion unit 20 and the comparison unit 40. For example, the text acquisition unit 10 may select and open a document file in a memory (not shown) in response to an operator's operation, thereby acquiring the text containing the input token string from the document file. Alternatively, for example, the text acquisition unit 10 may acquire text inputted via a keystroke or pasted from another document file in response to a user's keyboard or mouse operation. Furthermore, the text acquisition unit 10 may also function as a token string acquisition unit that acquires the input token string.
[0037] The pronunciation string conversion unit 20 converts the input tag string acquired by the text acquisition unit 10 into a pronunciation string. For example, the pronunciation string conversion unit 20 parses the acquired input tag string and converts the input tag string into a pronunciation string based on the obtained parsing result. The converted pronunciation string is sent to the tag string conversion unit 30. Such a pronunciation string conversion unit 20 may also include a morpheme analysis unit 21 and a pronunciation addition processing unit 22. The pronunciation string is a text representing the pronunciation of the input tag string. For example, when the input tag string is "学習" (pronunciation; gakushuu: meaning; study), the pronunciation string becomes "ガクシュウ" (gakushuu: study).
[0038] The morphological analysis unit 21 analyzes the input token string acquired by the text acquisition unit 10. For example, the morphological analysis unit 21 segments the input token string into words and performs morphological analysis to infer the part of speech of each word. Furthermore, the term "word" in morphological analysis can also be referred to as "morpheme." That is, morphological analysis involves segmenting the input token string into morphemes and inferring the part of speech of each morpheme. The morphological analysis unit 21 can also be referred to as the "input token string analysis unit" or "analysis unit."
[0039] The pronunciation adding processing unit 22 adds pronunciation to each word based on the result of morpheme analysis and converts it into a pronunciation string. The pronunciation adding processing unit 22 can also use, for example, a morpheme dictionary (not shown) to add pronunciation to each word. The morpheme dictionary is used for morpheme analysis and is a dictionary that records the entry (word), pronunciation, part of speech, and variation form for each morpheme. In addition, without limitation thereto, the pronunciation adding processing unit 22 can also use the word dictionary 35 described later to add pronunciation to each word. The word dictionary 35 is a dictionary that stores word tags, pronunciation strings (pronunciations), and parts of speech in association with each other.
[0040] The mark string conversion section 30 converts the phonetic string converted by the phonetic string conversion section 20 into an output mark string. For example, the mark string conversion section 30 analyzes the phonetic string converted by the phonetic string conversion section 20, and converts the phonetic string into an output mark string according to the analysis result obtained. The converted output mark string is sent to the comparison section 40. Such a mark string conversion section 30 can also have a feature quantity conversion section 31, a conversion section 32, and a storage section 33, for example. The storage section 33 can also have a language model 34 and a word dictionary 35.
[0041] Here, the feature quantity conversion section 31 converts the phonetic string into a sound score vector. Here, the feature quantity conversion section 31 can perform either of (1) a process of directly converting the phonetic string into a sound score vector and (2) a process of converting the phonetic string into a sound signal, converting the sound signal into a sound score vector. In the first embodiment, a case where the process of (1) above is used will be described as an example. In a modification of the first embodiment, the process of (2) above will be described.
[0042] Here, the sound score vector is also referred to as a phonetic string feature quantity vector, and is a feature vector in which the phonetic sequence becomes a correct answer in the conversion section 32. In addition, the feature quantity conversion section 31 can also be referred to as a sound score conversion section.
[0043] The conversion section 32 converts the sound score vector into an output mark string using the language model 34 and the word dictionary 35. In detail, the conversion section 32 generates a phonetic string from the sound score vector, and converts the generated phonetic string into an output mark string using the language model 34 and the word dictionary 35 as exemplified above. In addition, the conversion section 32 can also be referred to as an output mark string conversion section. Figure 2
[0044] The storage section 33 stores the language model 34 and the word dictionary 35 used for sound recognition.
[0045] As the language model 34, a model made from the same statistical information as a sound recognition engine whose sound recognition result is to be confirmed is used. As one example, as the language model, an n-gram language model (n is a natural number of 1 or more) determined in accordance with the appearance probability of one word of language model learning data can be used. As the language model, in addition to the 1-gram language model, a 2-gram language model, a 3-gram language model, a 4-gram language model, a 5-gram language model, and other language models can be used. In the case where the 1-gram language model is used, the appearance probability of each word of the language model learning data is used as the appearance probability of each word of the language model. In the case where the 2-gram language model is used, the appearance probability of each word pair of the language model learning data is used as the appearance probability of each word pair of the language model. In the case where the 3-gram language model is used, the appearance probability of each word sequence of three words of the language model learning data is used as the appearance probability of each word sequence of three words of the language model. In the case where the 4-gram language model is used, the appearance probability of each word sequence of four words of the language model learning data is used as the appearance probability of each word sequence of four words of the language model. In the case where the 5-gram language model is used, the appearance probability of each word sequence of five words of the language model learning data is used as the appearance probability of each word sequence of five words of the language model. Figure 2 In the example, a 2-gram language model (n=2) is used as language model 34, in which the probability of occurrence of a certain word (the second word) is determined by the number of n-1 known words (the first word) immediately preceding it. Alternatively, a language model modeled using recurrent neural networks (RNNs) can be used. Furthermore, weighted finite-state transducer (WFST) speech recognition technology can be used.
[0046] In the word dictionary 35, the labels of words to be recognized as sound, the pronunciation strings corresponding to the labels, and the word's part of speech information are registered in association. For example, for a word like "責価" (hyooka: evaluation), the label for the word is "責価" (hyooka: evaluation). For example, for a word like "責価" (hyooka: evaluation), the pronunciation string corresponding to the label is "ヒョオカ" (hyooka: evaluation).
[0047] The comparison unit 40 is as follows Figure 3 As shown, the input tag string acquired by the text acquisition unit 10 and the output tag string obtained from the input tag string via the pronunciation string conversion unit 20 and the tag string conversion unit 30 are compared to extract the difference. In addition, the term "extracting the difference" can also be replaced by "detecting the difference" or "determining the difference". The comparison unit 40 sends the difference extraction result to the display control unit 70. The difference extraction result includes, for example, an input tag string that can distinguishably contain the difference and an output tag string that can distinguishably contain the difference. In addition, without limitation to this, the difference extraction result may also include difference determination information (for example, the difference mark "○○" and the position (the xxth character from the beginning of the text)) that determines the difference between the input tag string and the output tag string and the input tag string and the output tag string.
[0048] The display control unit 70 displays the input tag string, output tag string, and comparison result sent from the comparison unit 40 on the display. Figure 4 As illustrated, the display control unit 70 displays a text read button 101, an input tag string display screen 102, and an output tag string display screen 103 on the display 71. The text read button 101 is a button for causing the text acquisition unit 10 to acquire text. The input tag string display screen 102 is a screen configured with an input tag string including a difference. The output tag string display screen 103 is a screen configured with an output tag string including a difference. The input tag string display screen 102 and the output tag string display screen 103 may also be referred to as an input tag string display area and an output tag string display area, respectively. The difference is displayed distinguishably from the tag strings other than the difference according to the display attribute 104. InFigure 4 In the present embodiment, the display attribute 104 indicates an underlined solid line. However, the display attribute 104 is not limited to this, and an attribute capable of discriminating the presence or absence of a difference can be appropriately used. For example, as the display attribute, various display attributes such as the color of a character, the type of a font, the size, the color of a background, and the like can be used.
[0049] Next, the operation of the difference extraction device configured as described above will be described using the flowchart of FIG. 10 and the schematic diagram of FIG. 11. Figure 5 Figures 6 to 8 The operation of the difference extraction device corresponds to the difference extraction method.
[0050] In step ST10, the text acquisition section 10 acquires a text in which an input token string is described. Specifically, for example, the display control section 70 causes the text read-in button 101, the input token string display screen 102, and the output token string display screen 103 to be displayed on the display 71 as shown in FIG. 1. Further, in the case before the text acquisition, unlike the case shown in FIG. 1, the input token string display screen 102 is a blank. The text read-in button 101 is a button for causing the text acquisition section 10 to perform a process of causing a file selection screen 101b including a list of files that can be selected and an open button 101a to be displayed on the display 71 by the user's operation. The open button 101a is a button for causing the text acquisition section 10 to perform a process of opening a file selected in the list and acquiring a text in the opened file by the user's operation. Figure 6 Figure 6 In this case, the text acquisition section 10 causes the file selection screen 101b to be displayed on the display 71 by the operation of the text read-in button 101. In addition, the text acquisition section 10 selects a file in the list by the user's operation, and acquires a text in which an input token string is described from the selected file by the operation of the open button 101a. The acquired text is displayed on the input token string display screen 102, and is sent to the pronunciation string conversion section 20.
[0051] In this case, the text acquisition section 10 causes the file selection screen 101b to be displayed on the display 71 by the operation of the text read-in button 101. In addition, the text acquisition section 10 selects a file in the list by the user's operation, and acquires a text in which an input token string is described from the selected file by the operation of the open button 101a. The acquired text is displayed on the input token string display screen 102, and is sent to the pronunciation string conversion section 20.
[0052] In step ST20, the pronunciation string conversion section 20 analyzes the input token string acquired in step ST10, and converts it into a pronunciation string. Specifically, the pronunciation string conversion section 20 analyzes the input token string, and converts the input token string into a pronunciation string according to the analysis result. Such step ST20 is performed by the morpheme analysis section 21 and the reading addition processing section 22 as steps ST21 to ST22.
[0053] In step ST21, the morphological analysis unit 21 performs morphological analysis on the acquired input token string, segmenting the morphemes and inferring the part of speech. Specifically, the morphological analysis unit 21 segments the input token string into words and performs morphological analysis to infer the part of speech of each word. For example, if the acquired token string is "評価実験" (hyooka jikken: evaluation experiment), it is divided into "評価" (hyooka: evaluation) and "実験" (jikken: experiment), and each is inferred to be a noun.
[0054] In step ST22, the pronunciation addition processing unit 22 receives the results from the morpheme analysis unit 21, adds pronunciation to each morpheme, and supplies the resulting pronunciation string to the token string conversion unit 30. Specifically, based on the results of the morpheme analysis, the pronunciation addition processing unit 22 adds pronunciation to each word and converts it into a pronunciation string. For example, if the morpheme analysis unit 21 outputs "讟価" (hyooka: evaluation) and "実験" (jikken: experiment) as results, the pronunciation addition processing unit 22 adds pronunciation to "ヒョオカ" (hyooka: evaluation) and "ジッケン" (jikken: experiment), respectively. Then, "ヒョオカジッケン" (hyooka: evaluation experiment) is supplied to the token string conversion unit 30. This concludes step ST20, which includes steps ST21 and ST22.
[0055] In step ST30, the labeled string conversion unit 30 receives the utterance string converted in step ST20 as input, analyzes the utterance string, converts it into an output labeled string, and sends the output labeled string to the comparison unit 40. Step ST20 is executed by the feature quantity conversion unit 31 and the conversion unit 32 as steps ST31 to ST33.
[0056] In step ST31, the feature conversion unit 31 converts the pronunciation string obtained in step ST20 into a sound score vector. Specifically, the feature conversion unit 31 generates a pronunciation string feature vector from the acquired pronunciation string. The pronunciation string feature vector is a feature vector such that the pronunciation sequence becomes a correct answer in the subsequent conversion unit 32. For example, in a DNN-HMM sound recognition engine that uses DNN and HMM, the sound interval is cut into 1 frame at regular intervals. In addition, for the cut frames, DNN is used to calculate the pronunciation state output probability vector (pronunciation state sound score vector) of the pronunciation sequence. In addition, DNN is the abbreviation of deep neural network. HMM is the abbreviation of hidden Markov model.
[0057] Here, the pronunciation state sound score vector is explained. Here, the unit of pronunciation is described as a syllable. In the case of Japanese, in addition to the so-called 50 sounds, there are also voiced sounds ("Ga" (ga), "Za" (za), etc.), half-voiced sounds ("Pa" (pa), etc.), and diphthongs ("Kya" (kya), "Ja" (ja), etc.). Furthermore, the consonant "N" (n, m) and the geminate consonant "T" (tt, kk, pp) are also treated as one syllable, and the long vowel " " is treated by replacing the preceding vowel. Here, the syllables of Japanese are set as Figure 7 The 102 exemplified are explained. Each pronunciation is usually expressed by an HMM of about 3 states, but in order to simplify the explanation, each pronunciation is explained as one state. The pronunciation state sound score vector in this case is a 102-dimensional vector that indicates the likelihood of the syllable to which the value of each element of the vector corresponds. That is, as the pronunciation state sound score vector, for each syllable, one 102-dimensional vector is output from the feature quantity conversion section 31 for the pronunciation string input to the feature quantity conversion section 31. However, this is not limiting, and each pronunciation can be expressed by 3 states, and a 306-dimensional pronunciation state sound score vector can be used for each syllable.
[0058] As the conversion method to the pronunciation state sound score vector, the element (output probability) corresponding to the syllable of the conversion target can be set to 1, and the other elements can be set to 0. For example, in the case where "Hyooka" (hyooka: evaluation) is input as a pronunciation string, as shown in Figure 8 As exemplified, for the input "Hyo" (hyo), a vector in which only the element corresponding to "Hyo" (hyo) is set to 1 and all the others are set to 0 can be output. Similarly, for "O" (o), a vector in which only the element corresponding to "O" (o) is set to 1 and all the others are set to 0 can be output. The same applies to "Ka" (ka). Furthermore, this pronunciation state sound score vector column has the highest likelihood with respect to a pronunciation string such as "Hyo" (hyo), "O" (o), and "Ka" (ka). Therefore, when this pronunciation state sound score vector string is supplied to the conversion section 32 (for example, a DNN-HMM decoder), the conversion section 32 converts the sound score vector into a pronunciation string and converts the pronunciation string into a token string. In detail, if a pronunciation string identical to the input pronunciation string is present in the word dictionary 35, the conversion section 32 outputs the pronunciation string identical to the input with respect to the pronunciation string, and outputs a token string determined in dependence on the language model 34 with respect to the token string.
[0059] Further, the method of creating the sound score vector of the pronunciation state is not limited to this, and an arbitrary ratio can be used for outputting, such as 10.0 for the element corresponding to the state and 5.0 for the other elements. Alternatively, noise can be added to the sound score vector of the pronunciation state, and it can be determined whether the desired result is output under stricter conditions. Alternatively, in HMM sound recognition using a Gaussian Mixture Model (GMM), a vector in which the average of the GMM representing each pronunciation state string is used as an element can be used as the sound score vector of the pronunciation state. However, in this case, a language model and a sound model for a GMM-HMM sound recognition engine are used when the string transformation is performed.
[0060] In step ST32, the transformation unit 32 transforms the sound score vector of the pronunciation state obtained in step ST31 into a pronunciation string. Specifically, in the following, a case in which the value of each element of the sound score vector of the pronunciation state is a 102-dimensional vector indicating the likelihood of the corresponding syllable will be described as an example. Further, the unit of pronunciation will be described as a syllable.
[0061] The transformation unit 32 estimates the corresponding syllable from the sound score vector of the pronunciation state. The syllable of Japanese is expressed not only by the so-called 50 sounds but also by voiced sounds ("ga", "za", etc.), half-voiced sounds ("pa", etc.), and diphthongs ("kya", "ja", etc.). Further, the consonant "n" and the consonant "tt", "kk", and "pp" are treated as one syllable, and the long vowel " " is treated by replacing the preceding vowel. Here, the syllable of Japanese will be described as Figure 7 The 102 syllables exemplified will be described. Further, in the present specification, the syllable is expressed by a katakana, but is not limited to this, and can be expressed by a hiragana. In the sound score vector of the pronunciation state, the value of each element indicates the likelihood of each syllable, so the syllable whose value indicating the likelihood of the corresponding syllable is large is estimated. For example, when the value of only the element corresponding to the syllable "hyo" is 1 and the values of all the other elements are 0 in the element of the 102-dimensional sound score vector of the pronunciation state, the sound score vector of the pronunciation state is transformed into the syllable "hyo". Further, the value of the element of the sound score vector of the pronunciation state is described as being only 0 and 1, but is not limited to this. As the sound score vector of the pronunciation state, an arbitrary vector in which the value of each element indicates the likelihood can be used. That is, such a sound score vector of the pronunciation state is transformed into the syllable whose likelihood is high in the same order, and the pronunciation string is generated. In this way, the transformation unit 32 transforms the sound score vector of the pronunciation state into the syllable, and generates the pronunciation string.
[0062] In step ST33, the conversion section 32 converts the pronunciation string obtained in step ST32 into an output token string while referring to the language model 34 and the word dictionary 35 in the storage section 33. That is, the conversion section 32 refers to the word dictionary 35 to guess candidates of tokens, that is, candidates of words, corresponding to the pronunciation string. In addition, the conversion section 32 uses the language model 34 to select a word that is appropriate as an article from the candidates of words guessed in the word dictionary 35 while considering the association of the words before and after, and generates a token string.
[0063] Here, the 2-gram language model expressed using the appearance probability of two words is used, and in the learning data of the language model, the combination of "hyooka jikken" appears more often than the combinations of "hyooka jikken" and "hyooka jikken". Figure 2 One example of the operation of step ST33 will be described in detail. Here, in the word dictionary 35, the token "hyooka" is registered for the pronunciation string "ヒョオカ" (hyooka), and the tokens "jikken" and "jikken" are registered for the pronunciation string "ジッケン" (jikken). In addition, the following case will be described as an example: the 2-gram language model expressed using the appearance probability of two words is used, and in the learning data of the language model, the combination of "hyooka jikken" appears more often than the combinations of "hyooka jikken" and "hyooka jikken".
[0064] The conversion section 32 first refers to the word dictionary 35 for the pronunciation string "hyooka jikken". As a result, one candidate of "hyooka" is obtained for the pronunciation string "hyooka", and two candidates of "jikken" are obtained for "jikken".
[0065] Next, the conversion section 32 uses the language model 34 to determine the appropriate combination of "hyooka" and "jikken". In the case of the language model exemplified above, the appearance probability of "hyooka jikken" is higher than that of "hyooka jikken", and therefore, for the pronunciation string "hyooka jikken", the token string "hyooka jikken" is determined. In the case of this example, the determined token string "hyooka jikken" is as follows: Figure 2 Figure 4 The same as the input tag string "hyooka jikken" (evaluation experiment), the indicated tag string is correct even if not registered to the word dictionary 35. Further, not limited to this, sometimes as Figure 3 As shown, the decided tag string "shinsoo gakushuu" (deep learning) is different from the input tag string "shinsoo gakushuu" (deep learning). In this case, the difference "shinsoo" (deep) is extracted at the comparison section 40 at the later stage.
[0066] In such a conversion from a pronunciation string to a tag string, it is possible to use a Viterbi algorithm using an occurrence probability of n-gram (n is a natural number of 1 or more) with respect to the input pronunciation string. Further, the search algorithm is not limited to the Viterbi algorithm, and other algorithms such as a tree trellis search algorithm can be used. In addition, the conversion section 32 supplies the converted output tag string to the comparison section 40. Thereby, the step ST30 including the steps ST31 to ST33 ends.
[0067] In the step ST40, the comparison section 40 compares the input tag string acquired in the step ST10 and the output tag string supplied in the step ST30, and extracts the difference. For example as Figure 3 As shown, the input tag string acquired by the text acquisition section 10 is converted to a pronunciation string by the pronunciation string conversion section 20, and converted to an output tag string by the tag string conversion section 30. Then, when compared by the comparison section 40, "shinsoo" (deep) is extracted as the difference.
[0068] In the step ST70, the display control section 70 displays the input tag string display screen 102 including the input tag string and the output tag string display screen 103 including the output tag string on the display 71, for example as Figure 4 As shown, the display control section 70 displays the input tag string display screen 102 including the input tag string and the output tag string display screen 103 including the output tag string on the display 71. In addition, the display control section 70 displays the difference in a state that can be distinguished from other tags by the display attribute 104 on the display 71 in both the input tag string display screen 102 and the output tag string display screen 103. In this state, the tag including the difference can be registered to the word dictionary 35 appropriately according to the operation of the keyboard or the mouse or the like by the user. In addition, it is also possible to register to the word dictionary 35 after the processing of the word presumption or the like described later.
[0069] As described above, according to the first embodiment, the text acquisition section acquires a text in which an input tag string is described. The pronunciation string conversion section converts the input tag string to a pronunciation string. The tag string conversion section converts the pronunciation string to an output tag string. The comparison section compares the input tag string and the output tag string, and extracts the difference.
[0070] With such a configuration, the unknown words among the unknown words included in the input token string that do not become correct tokens in the output token string are extracted as the difference. In other words, the unknown words among the unknown words included in the input token string that become correct tokens in the output token string are not extracted as the difference. Thus, registration of the unknown words that become correct tokens even if not registered can be prevented. In addition, the difference from the output token string can be extracted even from a small number of input token strings. In addition, the input token string and the output token string that is converted from the input token string via the pronunciation string are compared, and the portion that does not become a correct token is extracted as the difference, so a difference that is useful for sound recognition can be extracted. In addition, the user performs the dictionary registration work on the extracted difference, so registration of unnecessary words such as token fluctuation can be prevented. In addition, the words that should be registered to the word dictionary can be easily prompted to the user at the time of creating the word dictionary. In addition, the user performs the dictionary registration work, so the word dictionary can be improved according to the field of the token string for each user.
[0071] In addition, according to the first embodiment, the token string conversion section can also have a feature quantity conversion section, a storage section, and a conversion section. The feature quantity conversion section can also convert the pronunciation string into a sound score vector. The storage section can also store a language model for sound recognition and a word dictionary. The conversion section can also generate a pronunciation string from the sound score vector, and convert the generated pronunciation string into an output token string using the language model and the word dictionary. In this case, in addition to the aforementioned effects, since the output token string is obtained using the language model for sound recognition and the word dictionary, a more appropriate difference can also be extracted as a token including an unknown word that is not in the word dictionary.
[0072] In addition, according to the first embodiment, the pronunciation string conversion section can also have a morpheme analysis section and a reading addition processing section. The morpheme analysis section can also divide the input token string into words, and perform morpheme analysis that estimates the part of speech of each word. The reading addition processing section can also add a reading to each word according to the result of the morpheme analysis, and convert into a pronunciation string. In this case, in addition to the aforementioned effects, for example, compared to a case where a pronunciation string is converted using information other than a reading such as stress and pause, it is possible to easily convert into a pronunciation string.
[0073] <Variant of the First Embodiment>
[0074] The variant of the first embodiment is a manner in which the feature quantity conversion section 31 does not perform processing that directly converts the pronunciation string into a sound score vector, but performs processing that converts into a sound score vector after converting the pronunciation string into a sound signal.
[0075] In conjunction therewith, the feature quantity conversion section 31, as shown in FIG. 4, converts the pronunciation string into a sound signal, and then converts the sound signal into a sound score vector. Figure 9The illustrated sound synthesis section 31a, the sound feature amount calculation section 31b, and the sound score calculation section 31c.
[0076] Here, the sound synthesis section 31a synthesizes a sound signal from the phonetic string transformed by the phonetic string transformation section 20. The synthesized sound signal is sent to the sound feature amount calculation section 31b. Further, the "sound signal" is also referred to as a "sound waveform signal". For example, the sound synthesis section 31a generates a sound waveform signal in accordance with the input phonetic string.
[0077] The sound feature amount calculation section 31b calculates a sound feature vector from the sound signal synthesized by the sound synthesis section 31a. For example, the sound feature amount calculation section 31b calculates a sound feature vector representing a spectrum in a predetermined frame unit from the sound signal. The calculated sound feature vector is sent to the sound score calculation section 31c.
[0078] The sound score calculation section 31c calculates a sound score vector from the sound feature vector calculated by the sound feature amount calculation section 31b. For example, the sound score calculation section 31c estimates the likelihood of each phoneme from the sound feature vector, and calculates a phonetic state sound score vector. The calculated sound score vector is sent to the aforementioned transformation section 32.
[0079] The other structures are the same as those of the first embodiment.
[0080] Next, the operation of the modification example configured as described above will be described using the flowchart of Figure 10 In the following description, the operation of the step ST31 of transforming a phonetic string into a sound score vector will be described. That is, as with the foregoing, the processing of the steps ST10 to ST20 is executed, and in the step ST30, the processing of the step ST31 is started. The step ST31 includes the steps ST31-1 to ST31-3.
[0081] In the step ST31-1, the sound synthesis section 31a synthesizes a sound signal from the phonetic string transformed by the phonetic string transformation section 20. Here, the sound synthesis section 31a can use various known methods capable of generating a sound waveform signal from an arbitrary phonetic string. For example, a method of storing waveform data in a phoneme unit, and selecting and connecting the waveform data in accordance with the input phonetic string can be used. As for the pitch information representing the pitch of the sound, the waveform data can be directly connected without change, or the pitch of the waveform data can be corrected by predicting a natural pitch change by a known technique. Alternatively, instead of storing the waveform data, a spectrum parameter sequence in a phoneme unit can be stored, and a sound signal can be synthesized using a source filter model. Or, a DNN that predicts a spectrum parameter sequence from a sequence of phonemes can be used. In any case, the sound synthesis section 31a synthesizes a sound signal from the phonetic string, and sends the sound signal to the sound feature amount calculation section 31b.
[0082] In step ST31-2, the sound feature quantity calculation section 31b calculates a sound feature vector from the sound signal after the synthesis in step ST31-1. For example, in the sound feature quantity calculation section 31b, a sound feature vector sequence is calculated from a sound waveform signal by the same processing as that used in the sound recognition processing. First, the sound feature quantity calculation section 31b performs a fast Fourier transform on the input sound data, for example, with a frame length of 10 ms and a frame shift of 5 ms, and converts it into a spectrum. Next, the sound feature quantity calculation section 31b calculates the sum of the power spectrum of each frequency domain according to the specifications of a predetermined bandwidth, converts it into a filter bank feature vector, and outputs it to the sound score calculation section 31c as a sound feature vector. As the sound feature vector, in addition to this, various sound feature vectors such as mel frequency cepstral coefficients (MFCC) can be used.
[0083] In step ST31-3, the sound score calculation section 31c calculates a sound score vector from the sound feature vector calculated in step ST31-2. For example, the sound score calculation section 31c uses a DNN to estimate a sound score vector of the pronunciation state with the sound feature vector as input, and outputs it. As for the processing of the sound score calculation section 31c, various publicly known methods used in sound recognition can also be used. Instead of a DNN based on full connection, a convolutional neural network (CNN), a long short-term memory (LSTM), or the like can also be used. In any case, the sound score calculation section 31c calculates a sound score vector from the sound feature vector, and outputs it to the conversion section 32. Based on the above, the step ST31 including steps ST31-1 to ST31-3 ends.
[0084] Hereinafter, the processing after step ST32 is executed as described above.
[0085] As described above, according to the modification of the first embodiment, the feature quantity conversion section is provided with a sound synthesis section, a sound feature quantity calculation section, and a sound score calculation section. The sound synthesis section synthesizes a sound signal from a pronunciation string. The sound feature quantity calculation section calculates a sound feature vector from the sound signal. The sound score calculation section calculates a sound score vector from the sound feature vector.
[0086] Thus, with the structure of converting a pronunciation string into a sound signal and then into a sound score vector, in addition to the effects of the first embodiment, a more appropriate sound score vector can be supplied to the conversion section using a language model and a word dictionary for sound recognition.
[0087] To supplement, according to the modification example, the outputted pronunciation state sound score vector is similar to the vector generated in the actual sound recognition processing, and thus it is possible to generate a label string closer to the sound recognition result. The pronunciation state sound score vector of the modification example not only has a tendency for the value of the element corresponding to the input syllable to increase, but also has a tendency for the value of the element corresponding to a similar syllable to increase, unlike the aforementioned pronunciation state sound score vector (pronunciation state output probability vector) in which only 0 and 1 are used as elements. That is, in the pronunciation state sound score vector described in the first embodiment, only the value of the element corresponding to the input syllable is set to 1. In contrast, with respect to the sound score vector described in the modification example, the value of the element corresponding to the input syllable and the value of the element corresponding to a syllable similar to the input syllable each increase, and thus it is possible to be similar to the vector generated in the actual sound recognition processing.
[0088] <Second Embodiment>
[0089] Next, the second embodiment will be described using Figures 11 to 15 . The second embodiment is different from the first embodiment or the modification example thereof in that the difference extracted by the comparison section 40 is additionally processed. For example, the second embodiment is different from the first embodiment in which only the difference is extracted and displayed in that the extracted difference is converted into a word-unit difference and displayed. In addition, in the second embodiment, the range of the displayed word candidates is corrected, and thus an improvement in the quality of word extraction can be expected.
[0090] Figure 11 is a block diagram illustrating the structure of the difference extraction apparatus 1 of the second embodiment, and with respect to the same constituent elements as those described above, the same symbols are attached, and detailed description thereof will be omitted. Here, only the different parts will be described. The following embodiments also omit the repeated description.
[0091] The difference extraction apparatus 1 of the second embodiment is different from the structure shown in Figure 1 in that it further includes a word estimation section 50 and an instruction section 80.
[0092] Here, the word estimation section 50 estimates a label of a word candidate including the difference extracted by the comparison section 40 in the input label string, based on the analysis result of the input label string by the morpheme analysis section 21. Here, the analysis result of the input label string is, for example, the result of the morpheme analysis by the morpheme analysis section 21.
[0093] In conjunction therewith, the display control section 70 causes the input label string including the word candidate estimated by the word estimation section 50 to be displayed on the display 71.
[0094] The indicator unit 80 indicates the range of the tokens that include at least a portion of the word candidates in the input token string displayed on the display 71. For example, the indicator unit 80 may indicate the range of the tokens based on a keyboard or mouse (not shown) operation performed by the user. Furthermore, without limitation thereto, the indicator unit 80 may indicate the range of the tokens based on an operation of another input device such as a touch panel.
[0095] Next, use Figure 12 Flowchart and Figures 13 to 15 The schematic diagram of FIG. 1 illustrates the operation of the differential extraction device constructed as described above.
[0096] Now, similarly to the above, steps ST10 to ST40 are executed to extract the difference between the input tag string and the output tag string.
[0097] In step ST50, the word inference unit 50 infers the token of the word candidate containing the difference in the input token string based on the analysis result of the input token string. Specifically, the word inference unit 50 extracts a character string that can be inferred to be a word formed by connecting adjacent morphemes of the word of the difference extracted by the comparison unit 40, and outputs it as a word candidate. Specifically, the word inference unit 50 extracts the character string that can be inferred to be a word formed by connecting adjacent morphemes of the word of the difference extracted by the comparison unit 40, and outputs it as a word candidate. Figure 13 As shown in the example, let's assume that the word to be differentiated is "deep" (shinsoo: deep layer), and check whether it forms a word together with the words before and after it. In this case, before the differentiation, there is the word "wa" (wa), and after the differentiation, there is the word "learning" (gakushuu: learning), so there is a possibility that it forms a word with the following word. Therefore, in the word estimation unit 50, "deep learning" (shinsoo gakushuu: deep learning) is estimated as a word candidate. As a judgment of a character string that constitutes a word in this way, for example, a rule such as "the connecting part of 'noun-general' is estimated as a word" is used. In addition, in the word estimation unit 50, it is not limited to using a rule of a single morpheme analysis result, and other rules can also be used to link multiple morphemes that appear frequently adjacent to each other based on the results of a large number of morpheme analyses to estimate them as word candidates.
[0098] In step ST71, the display control unit 70 Figure 14As illustrated, the mark string acquired by the text acquisition section 10 is displayed on the input mark string display screen 102, and the mark string output by the mark string conversion section 30 is displayed on the output mark string display screen 103. In the display control section 70, the mark including the difference is displayed on the display 71 using the display attribute 104, based on the difference extracted by the comparison section 40 and the word candidate guessed by the word guessing section 50. In this state, the mark including the difference can be registered in the word dictionary 35 as appropriate according to the operation of the keyboard or mouse or the like by the user. Further, it is also possible to register in the word dictionary 35 after the next step ST80.
[0099] In the step ST80, the instruction section 80 uses the cursor 400, the word candidate screen 401, the word candidate 402, and the range correction button 403. In the instruction section 80, the range of the word can be changed according to the operation of the user by changing the range of the display attribute 104 of the display control section 70.
[0100] Specifically, the instruction section 80, as illustrated, Figure 15 When the cursor 400 is aligned over the display attribute 104 of the input mark string display screen 102, the word candidate screen 401 is opened, and the word candidate 402 is displayed (step ST80-1).
[0101] The instruction section 80 selects the candidate by aligning the cursor 400 over the word candidate 402, and changes the range of the display attribute 104 of the input mark string display screen 102 and the output mark string display screen 103 (step ST80-2). For example, the instruction section 80 aligns the cursor 400 over the mark "ペンローズ" (penroozu: Penrose) of the word candidate, and selects "ムーア·ペンローズ" (muua penroozu: Moore-Penrose) from among the word candidates 402 in the word candidate screen 401. Thus, the instruction section 80 changes the range of the display attribute 104 from the mark "ペンローズ" (penroozu: Penrose) of the word candidate to the range "ムーア·ペンローズ" (muua penroozu: Moore-Penrose) including the entire mark. Further, not limited thereto, the instruction section 80 can align the cursor 400 over the mark "ペンローズ" (penroozu: Penrose) of the word candidate, and select "ペン" (pen: pen) or "ローズ" (roozu: rose) from among the word candidates 402 in the word candidate screen 401. Thus, the instruction section 80 changes the range of the display attribute 104 from the mark "ペンローズ" (penroozu: Penrose) of the word candidate to the range "ペン" (pen: pen) or "ローズ" (roozu: rose) including a part of the mark.
[0102] Alternatively, the instruction section 80 selects the range correction button 403, and changes the range of the display attribute 104 using the cursor 400. For example, the instruction section 80 selects the range correction button 403, moves the cursor 400 according to the operation of the mouse 81 by the user, and expands the range of "penroozu" to select the range of "muuapenroozu". Without being limited to this, the instruction section 80 can also select the range correction button 403, move the cursor 400 according to the operation of the mouse 81 by the user, and reduce the range of the display attribute 104 from the mark "penroozu" of the word candidate to the range "pen" or "roozu" including a part of the mark.
[0103] As a result of the step ST80-2 or ST80-2a, the display attribute 104 of the input mark string display screen 102 becomes "muua penroozu" (step ST80-3), and the display attribute 104 of the output mark string display screen 103 becomes "muua Penrose". In addition, in the case where the range of the display attribute 104 of the input mark string display screen 102 is "pen" or "roozu", the range of the display attribute 104 of the output mark string display screen 103 is "Pen" or "rose".
[0104] As described above, according to the second embodiment, the analysis section analyzes the input mark string. The word estimation section estimates the mark of the word candidate including the difference in the input mark string, according to the analysis result of the input mark string. Thus, with the structure capable of estimating the mark of the word candidate including the difference, in addition to the effect of the first embodiment, even in the case where the compound word of the difference and the noun connected to the difference is an unknown word, the unknown word can be estimated as the word candidate.
[0105] In addition, according to the second embodiment, the display control section displays the input mark string including the word candidate on the display. The instruction section instructs the range of the mark including at least a part of the word candidate in the displayed input mark string. Thus, with the structure capable of correcting the range of the estimated word candidate, the qualitative improvement of the word extraction can be expected.
[0106] <Third Embodiment>
[0107] Next, the use of the word candidate will be described. Figures 16 to 21, which illustrates the third embodiment. In the third embodiment, the word candidate is displayed using the display attribute corresponding to the word class, which is determined for the word candidate presumed in the second embodiment. In addition, the third embodiment can register the displayed word candidate in the word dictionary 35, and reflect the result of the registration to the display.
[0108] Figure 16 is a block diagram showing the process of the differential extraction device 1 of the third embodiment. The differential extraction device 1 has the word class determination section 60 and the word registration section 90 in addition to the structure shown in Figure 11
[0109] Here, the word class determination section 60 determines the word class of the word candidate presumed by the word presumption section 50. For example, the word class determination section 60 can determine the word class of the word candidate presumed by the word presumption section 50 as an unknown word using the unknown word determination section 61. Alternatively, for example, the word class determination section 60 can determine the word class of the word candidate presumed by the word presumption section 50 as a mark fluctuation using the mark fluctuation determination section 62. In addition, it is not limited thereto, and various classes indicating word marks can be used as the word class determination section 60. For example, the word class determination section 60 can presume various classes such as proper nouns, verbs, and the like.
[0110] If the mark of the word candidate presumed by the word presumption section 50 is not registered in the word dictionary 35, the unknown word determination section 61 determines the mark of the word candidate as an unknown word.
[0111] If the mark of the word candidate presumed by the word presumption section 50 and the mark within the output mark string corresponding to the mark of the word candidate are different marks of the same word, the mark fluctuation determination section 62 determines the two marks as a mark fluctuation. The determination of the mark fluctuation can be performed, for example, according to whether the two marks are in a different mark dictionary. The different mark dictionary is a dictionary in which different marks of the same word are described. The "different mark dictionary" can also be referred to as "different mark information" or "mark fluctuation determination information".
[0112] In addition, the display control section 70 displays the mark of the word candidate in the display 71 using the display attribute corresponding to the word class determined by the word class determination section 60.
[0113] The word registration section 90 registers the mark of the range indicated by the indication section 80 in the word dictionary 35.
[0114] Next, the flowchart of Figures 17 to 21 and the block diagram of Figure 18 will be described.The schematic diagram of FIG. 1 illustrates the operation of the differential extraction device constructed as described above.
[0115] Now, steps ST10 to ST50 are executed in the same manner as described above, and labels of word candidates including the difference are estimated.
[0116] In step ST60, the word type determination unit 60 executes the unknown word determination unit 61 and the token fluctuation determination unit 62 in parallel. The unknown word determination unit 61 determines that the word candidate estimated in step ST50 is an unknown word if the token is not registered in the word dictionary 35. For example, if the token "ペンローズ" (Penrose) is estimated as a word candidate based on the difference between the input token string and the output token string, the token "ペンローズ" (Penrose) is not included in the word dictionary 35 and is therefore determined to be an unknown word.
[0117] In the token fluctuation determination unit 62, if the token of the candidate word estimated in step ST50 and the token in the output token string corresponding to the candidate word token are different tokens of the same word, the token fluctuation is determined. For example, in the case of the input token string "所" (tokoro: place, part) and the corresponding output token string "ところ" (tokoro: part, place), the candidate word is estimated to be "所" (tokoro: place, part) based on the difference between the two. If the estimated candidate word token "所" (tokoro: place, part) and the token in the corresponding output token string "ところ" (tokoro: part, place) are in different token dictionaries, they are different tokens of the same word and are therefore determined to be token fluctuation.
[0118] In step ST72, the display control unit 70 displays Figure 19As exemplified, according to the word candidate presumed to be the extracted difference, and further according to the word category, the mark of the difference is displayed on the display 71 using the display attribute 600 to 602 corresponding to the word category. In this example, "penroozu" is determined to be an unknown word, and "tokoro" is determined to be a marked fluctuation, so "penroozu" is displayed using the display attribute 600 of double lines, "tokoro" is displayed using the display attribute 602 of dotted lines, and the other words are displayed using the display attribute 601 of solid lines. In this example, the display attribute 601 is set to the display attribute of double lines, the display attribute 602 of dotted lines, and the other display attribute 601, but is not limited thereto, and any character modification can be used according to the word category. As a modification example of the display attribute, various categories such as the concentration of highlighting, the character size, the font, the color, the boldface, the italic, the arrangement of predetermined marks (for example, black triangles) before and after the character, and the like can be appropriately used.
[0119] After the step ST72, the step ST80 is appropriately executed. Further, if there is no operation by the user, the step ST80 is omitted.
[0120] In the step ST90, the word registration unit 90, for example, displays the word candidate screen 401 on the display 71. The word candidate screen 401 is a screen including the word candidate 402, the range correction button 403, and the word registration button 700. The word registration unit 90 displays the word candidate screen 401 when the word registration button 700 is pressed. Figure 20 The word registration processing is executed using the word candidate screen 401 and the word registration screen 701 as exemplified. The word candidate screen 401 is a screen including the word candidate 402, the range correction button 403, and the word registration button 700. The word registration screen 701 is a screen displayed by the operation of the word registration button 700, and includes the mark input frame 702, the pronunciation registration frame 703, the word category registration frame 704, and the registration button 705. For example, when the cursor 400 is aligned over the display attribute 601 corresponding to the operation of the mouse 81 by the user, the word candidate screen 401 is opened, and when the word registration button 700 is pressed, the word registration screen 701 is opened. In the word registration screen 701, the mark, the pronunciation, and the word category of the word with respect to the range of the display attribute 601 are respectively input to the mark input frame 702, the pronunciation registration frame 703, and the word category registration frame 704, and when the registration button 705 is pressed, the word is registered in the word dictionary 35. In this example, the mark input frame 702, the pronunciation registration frame 703, and the word category registration frame 704 are input, but the mark, the pronunciation, and the word category can be automatically input after the word registration button 700 is pressed.
[0121] In addition, the word registration unit 90 can also be configured to, for example, automatically register the word in the word dictionary 35 when the word candidate screen 401 is opened. Figure 20As illustrated, the registered word is reflected to the display screen, and the mark string conversion section 30, the comparison section 40, the word estimation section 50, the word category determination section 60, and the display control section 70 are executed again manually or automatically. In Figure 21 The lower layer illustrates the updated input mark string display screen 102 and the output mark string display screen 103. After "shinsoo gakushuu" (deep learning) is registered in the word dictionary 35 by the word registration section 90, such word registration reflection processing is executed, and "shinsoo gakushuu" (deep learning) is displayed in the output mark string display screen 103. Therefore, the comparison section 40 does not perform the difference extraction, and the display attribute 601 is not displayed.
[0122] In addition, the word registration section 90 can also cause the word registration screen 800 for collectively registering a plurality of words to be displayed in the display 71 as illustrated. Figure 21 The word registration screen 800 displays the word candidates estimated by the word estimation section 50 and the words of the input mark string of the input mark string display screen 102 corresponding to the difference of the word candidates of the range indicated by the indication section 80 as a plurality of words to be registered in the word dictionary 35. At the active display 801 within the word registration screen 800, the word to be subjected to the word registration can be specified. By pressing the registration button 802 within the word registration screen 800, the word registration section 90 can collectively register the words made active at the active display 801 in the word dictionary 35.
[0123] Further, in Figure 21 , the active display 801 is a check box, but is not limited thereto, and various display modes can be used. For example, instead of the check box, various display modes such as a circle mark, a cross mark, a hatching, or the like can be used. In addition to this, in the example illustrated in Figure 22 , the mark input box 702, the pronunciation registration box 703, and the category registration box 704 are automatically input, but can be manually input by the user. In any case, by the registration to the word dictionary 35, the step ST90 is ended.
[0124] As described above, according to the third embodiment, the word category determination section 60 determines the word category of the word candidate. Therefore, in addition to the effect of the second embodiment, it is possible to distinguish whether it is the word category requiring the registration of the word candidate before the word candidate is registered.
[0125] In addition, according to the third embodiment, the display control section 70 can also cause the mark of the word candidate to be displayed in the display using the display attribute corresponding to the word category. In this case, it is possible to support the judgment of whether the registration of the word candidate is required before the word candidate is registered by the user.
[0126] In addition, according to the third embodiment, the word registration unit 90 can also register the indicated range of the tag in the word dictionary. In this case, it is possible to register the tag confirmed by the user in the word dictionary.
[0127] In addition, according to the third embodiment, it can also be that if the tag of the word candidate is not registered in the word dictionary, the unknown word determination unit 61 in the word type determination unit 60 determines the tag of the word candidate as an unknown word. In this case, it is possible to accurately detect an unknown word that is not registered in the word dictionary among the word candidates.
[0128] In addition, according to the third embodiment, it can also be that if the tag of the word candidate and the tag within the output tag string corresponding to the tag of the word candidate are different tags of the same word, the tag fluctuation determination unit 62 in the word type determination unit 60 determines the two tags as tag fluctuations. In this case, it is possible to detect a word in which tag fluctuations that do not need to be newly registered in the word dictionary among the word candidates.
[0129] <4th Embodiment>
[0130] Figure 1 is a block diagram illustrating the hardware structure of the differential extraction device of the fourth embodiment. The fourth embodiment is a specific example of the first to third embodiments, and is a way of implementing the differential extraction device 1 by a computer.
[0131] The differential extraction device 1 has a CPU (Central Processing Unit) 2, a RAM (Random Access Memory) 3, a program memory 4, an auxiliary storage device 5, and an input / output interface 6 as hardware. The CPU 2 communicates with the RAM 3, the program memory 4, the auxiliary storage device 5, and the input / output interface 6 via a bus. That is, the differential extraction device 1 of the present embodiment is implemented by a computer with such a hardware structure.
[0132] CPU 2 is an example of a general-purpose processor. RAM 3 is used in CPU 2 as a work memory. RAM 3 includes a volatile memory such as SDRAM (Synchronous Dynamic Random Access Memory), and the like. Program memory 4 stores a program for implementing each part corresponding to each embodiment. The program can also be configured, for example, as a program for causing a computer to implement each function as follows. [1] A function of acquiring a text in which an input token string is described. [2] A function of converting the input token string into a pronunciation string. [3] A function of converting the pronunciation string into an output token string. [4] A function of extracting a difference by comparing the input token string and the output token string. In addition, as the program memory 4, for example, a ROM (Read-Only Memory), a part of auxiliary storage device 5, or a combination thereof is used. Auxiliary storage device 5 stores data non-temporally. Auxiliary storage device 5 includes a non-volatile memory such as an HDD (hard disc drive) or an SSD (solid state drive).
[0133] Input-output interface 6 is an interface for connecting with other devices. Input-output interface 6 is used for connection with a keyboard, mouse 81, and display 71, for example.
[0134] The program stored in program memory 4 includes computer executable commands. The program (computer executable commands) causes a computer to perform a predetermined process when executed by CPU 2 as a processing circuit. For example, the program causes CPU 2 to perform a series of processes described with respect to each part of Figure 9 , Figure 11 , Figure 16 and Figure 5 when executed by CPU 2. For example, the computer executable commands included in the program cause a computer to perform a difference extraction method when executed by CPU 2. The difference extraction method can also include each step corresponding to each function of [1] to [4] described above. In addition, the difference extraction method can also appropriately include each step shown in Figure 10 , Figure 12 , Figure 17 and Figure 22 .
[0135] The program can be provided to the differential extraction device 1 as a computer in a state stored in a storage medium readable by the computer. In this case, for example, the differential extraction device 1 further has a drive (not shown) that reads out data from the storage medium, and acquires the program from the storage medium. As the storage medium, for example, a magnetic disk, an optical disk (CD-ROM, CD-R, DVD-ROM, DVD-R, or the like), a magneto-optical disk (MO or the like), a semiconductor memory, or the like can be appropriately used. The storage medium can also be referred to as a non-transitory computer readable storage medium. In addition, the program can be stored in a server on a communication network, and the differential extraction device 1 downloads the program from the server using the input / output interface 6.
[0136] The processing circuit that executes the program is not limited to a general-purpose hardware processor such as the CPU 2, and can use a dedicated hardware processor such as an ASIC (Application Specific Integrated Circuit). The term "processing circuit (processing unit)" includes at least one general-purpose hardware processor, at least one dedicated hardware processor, or a combination of at least one general-purpose hardware processor and at least one dedicated hardware processor. In the example shown, the CPU 2, the RAM 3, and the program storage 4 correspond to the processing circuit.
[0137] According to the device, method, and storage medium of at least one embodiment described above, the text in which the input token string is recorded is acquired, the input token string is converted into a pronunciation string, the pronunciation string is converted into an output token string, the input token string and the output token string are compared, and the difference is extracted, so that registration of an unknown word that becomes a correct token even if not registered in the unknown word can be prevented.
[0138] Furthermore, several embodiments of the present application have been described above, but these embodiments are presented by way of example only, and are not intended to limit the scope of the application. These embodiments can be implemented in various other ways, and various omissions, substitutions, and changes can be made without departing from the spirit of the application. These embodiments and modifications thereof, as well as the application encompassed by the scope and spirit of the application, are included within the scope of the application as recited in the patent claims and the scope equivalent thereto.
[0139] Furthermore, the above-described embodiments can be summarized as the following technical solutions.
[0140] (Technical Solution 1)
[0141] A differential extraction device includes:
[0142] a text acquisition unit that acquires a text in which an input token string is recorded;
[0143] a pronunciation string conversion section that converts the input token string into a pronunciation string;
[0144] a token string conversion section that converts the pronunciation string into an output token string; and
[0145] a comparison section that extracts a difference by comparing the input token string and the output token string.
[0146] (TECHNICAL SOLUTION 2)
[0147] The difference extraction device according to Technical Solution 1, further comprising:
[0148] a parsing section that parses the input token string; and
[0149] a word speculation section that speculates a token of a word candidate including the difference in the input token string, based on a result of the parsing of the input token string.
[0150] (TECHNICAL SOLUTION 3)
[0151] The difference extraction device according to Technical Solution 2, further comprising:
[0152] a display control section that causes the input token string including the word candidate to be displayed on a display; and
[0153] an indication section that indicates a range of a token including at least a part of the word candidate in the displayed input token string.
[0154] (TECHNICAL SOLUTION 4)
[0155] The difference extraction device according to Technical Solution 2, wherein
[0156] the difference extraction device further comprises a word category determination section that determines a word category of the word candidate.
[0157] (TECHNICAL SOLUTION 5)
[0158] The difference extraction device according to Technical Solution 4, wherein
[0159] the difference extraction device further comprises a display control section that causes the input token string including the word candidate to be displayed on a display,
[0160] the display control section causes the token of the word candidate to be displayed on the display using a display attribute corresponding to the word category.
[0161] (TECHNICAL SOLUTION 6)
[0162] The difference extraction device according to Technical Solution 3, wherein
[0163] The differential extraction device further includes a word registration unit that registers the token of the indicated range in a word dictionary.
[0164] (7)
[0165] The differential extraction device according to (4), wherein
[0166] The word type determination unit includes an unknown word determination unit that determines the token of the word candidate as an unknown word if the token of the word candidate is not registered in the word dictionary.
[0167] (8)
[0168] The differential extraction device according to (4), wherein
[0169] The word type determination unit includes a token fluctuation determination unit that determines two tokens as token fluctuations if the token of the word candidate and the token within the output token string corresponding to the token of the word candidate are different tokens of the same word.
[0170] (9)
[0171] The differential extraction device according to (1), wherein
[0172] The token string conversion unit includes:
[0173] a feature quantity conversion unit that converts the pronunciation string into a sound feature vector;
[0174] a storage unit that stores a language model for sound recognition and a word dictionary; and
[0175] a conversion unit that generates a pronunciation string from the sound feature vector and converts the generated pronunciation string into the output token string using the language model and the word dictionary.
[0176] (10)
[0177] The differential extraction device according to (9), wherein
[0178] The feature quantity conversion unit includes:
[0179] a sound synthesis unit that synthesizes a sound signal from the pronunciation string;
[0180] a sound feature quantity calculation unit that calculates a sound feature vector from the sound signal; and
[0181] The sound score calculation unit calculates a sound score vector based on the sound feature vector.
[0182] (Technical Solution 11)
[0183] According to the differential extraction device described in Technical Solution 1, wherein
[0184] The pronunciation string conversion unit includes:
[0185] The morpheme analysis unit divides the input token string into words and performs morpheme analysis to infer the part of speech of each word; and
[0186] The reading addition processing unit converts the each word into the pronunciation string by adding a reading according to the result of the morpheme analysis.
[0187] (Technical Solution 12)
[0188] A differential extraction method includes:
[0189] Obtaining a text in which an input token string is recorded;
[0190] Converting the input token string into a pronunciation string;
[0191] Converting the pronunciation string into an output token string; and
[0192] Comparing the input token string and the output token string to extract a difference.
[0193] (Technical Solution 13)
[0194] A storage medium stores a program for causing a computer to implement the following functions:
[0195] A function of obtaining a text in which an input token string is recorded;
[0196] A function of converting the input token string into a pronunciation string;
[0197] A function of converting the pronunciation string into an output token string; and
[0198] A function of comparing the input token string and the output token string to extract a difference.
Claims
1. A differential extraction device comprising: A text acquisition unit that acquires a text containing an input markup string; a pronunciation string conversion unit, converting the input mark string into a pronunciation string; a mark string conversion unit, converting the pronunciation string into an output mark string; a comparing unit that compares the input tag string and the output tag string and, based on a result of the comparison, determines a portion of the output tag string that is different from a portion of the input tag string as a difference to extract the difference, wherein the portion is located at corresponding positions in the input tag string and the output tag string, respectively; A parsing unit, configured to parse the input tag string; a word estimation unit that estimates a token of a word candidate containing the difference in the input token string based on a result of parsing the input token string; as well as A word type determination unit determines the word type of the word candidate, wherein The word estimation unit extracts a character string estimated to constitute a word by linking the difference and at least one character adjacent to the difference from the input token string, and outputs the extracted character string as a token of the estimated word candidate. The word type determination unit includes a mark fluctuation determination unit that uses a different mark dictionary. If two marks, including the mark of the word candidate and the mark in the output mark string corresponding to the mark of the word candidate, are in the different mark dictionary, the two marks are determined to be mark fluctuations, wherein the different mark dictionary is a dictionary that records different marks of the same word with the same meaning. The tag string conversion unit includes: a feature quantity conversion unit that converts the utterance string into a sound score vector; a storage unit storing a language model and a word dictionary for voice recognition; and a conversion unit that generates a second pronunciation string from the acoustic score vector and converts the generated second pronunciation string into the output label string using the language model and the word dictionary; The acoustic score vector is a 102-dimensional vector representing the likelihood of each syllable contained in the Japanese pronunciation string. The pronunciation of each syllable is expressed in one state. Japanese includes 102 syllables. The difference extraction device further includes a display control unit that performs the following display control processing: displays the mark of the word candidate on a display according to the display attribute of the word candidate, and displays the input mark string including the word candidate on a display.
2. The differential extraction device according to claim 1, wherein: Also features: The indicating unit indicates a range of the symbols including at least a part of the word candidates in the displayed input symbol string.
3. The differential extraction device according to claim 1, wherein: The display control unit displays the mark of the word candidate on a display using the display attribute corresponding to the word type.
4. The differential extraction device according to claim 2, wherein: The difference extraction device further includes a word registration unit that registers the designated range of symbols in a word dictionary.
5. The differential extraction device according to claim 1, wherein: The word type determination unit includes an unknown word determination unit configured to determine that the word candidate's symbol is an unknown word if the symbol of the word candidate is not registered in a word dictionary.
6. The differential extraction device according to claim 1, wherein: The feature quantity conversion unit includes: a sound synthesis unit for synthesizing a sound signal from the pronunciation string; an acoustic feature calculation unit for calculating an acoustic feature vector based on the sound signal; and The sound score calculation unit calculates a sound score vector based on the sound feature vector.
7. The differential extraction device according to claim 1, wherein: The pronunciation string conversion unit includes: a morpheme analysis unit that divides the input token string into words and performs morpheme analysis to estimate the part of speech of each word; and The pronunciation adding processing unit adds pronunciation to each of the words based on the result of the morphological analysis to convert the words into the pronunciation string.
8. A differential extraction method comprising: Get the text containing the input markup string; Converting the input token string into a pronunciation string; Converting the pronunciation string into an output token string; comparing the input tag string and the output tag string and, based on a result of the comparison, determining a portion of the output tag string that is different from a portion of the input tag string as a difference to extract the difference, wherein the portion is located at corresponding positions in the input tag string and the output tag string, respectively; Parsing the input tag string; Inferring a tag of a word candidate containing the difference in the input tag string based on the parsing result of the input tag string; as well as Determine the word type of the word candidate, When estimating the labels of the word candidates, the differential extraction method further includes: extracting a character string estimated to constitute a word by linking the difference and at least one character adjacent to the difference from the input tag string, and outputting the extracted character string as a tag of the estimated word candidate, When determining the word type of the word candidate, the differential extraction method further includes: using a different tag dictionary, if two tags including the tag of the word candidate and the tag in the output tag string corresponding to the tag of the word candidate are in the different tag dictionary, then the two tags are determined to be tag fluctuations, wherein the different tag dictionary is a dictionary that records different tags of the same word with the same meaning, When the pronunciation string is transformed into an output mark string, the difference extraction method includes: Converting the utterance string into a sound score vector; storing a language model and a word dictionary for voice recognition in a storage unit; and generating a second pronunciation string from the acoustic score vector, and transforming the generated second pronunciation string into the output token string using the language model and the word dictionary; The acoustic score vector is a 102-dimensional vector representing the likelihood of each syllable contained in the Japanese pronunciation string. The pronunciation of each syllable is expressed in one state. Japanese includes 102 syllables. The difference extraction method further includes executing the following display control processing: displaying the mark of the word candidate on a display according to the display attribute of the word candidate, and displaying the input mark string including the word candidate on a display.
9. A storage medium storing a program for causing a computer to implement the following functions: Function to obtain the text containing the input markup string; A function of converting the input markup string into a pronunciation string; A function of converting the pronunciation string into an output markup string; a function of comparing the input tag string and the output tag string and, based on a result of the comparison, determining a portion of the output tag string that is different from a portion of the input tag string as a difference to extract the difference, wherein the portion is located at corresponding positions in the input tag string and the output tag string, respectively; A function for parsing the input markup string; inferring the function of the token containing the difference word candidate in the input token string based on the parsing result of the input token string; as well as A function for determining the word type of the word candidate, Among them, the function of inferring the mark of the word candidate also includes: A function of extracting a character string estimated to constitute a word by linking the difference and at least one character adjacent to the difference from the input character string, and outputting the extracted character string as a tag of the estimated word candidate, Among them, the function of determining the word type of the word candidate also includes: using a different mark dictionary, if two marks including the mark of the word candidate and the mark in the output mark string corresponding to the mark of the word candidate are in the different mark dictionary, then the two marks are determined to be a mark fluctuation function, wherein the different mark dictionary is a dictionary that records different marks of the same word with the same meaning, The function of converting the pronunciation string into an output markup string includes: A function for converting the utterance string into a sound score vector; A function of storing a language model and a word dictionary for voice recognition in a storage unit; and A function of generating a second pronunciation string from the acoustic score vector and converting the generated second pronunciation string into the output token string using the language model and the word dictionary; The acoustic score vector is a 102-dimensional vector representing the likelihood of each syllable contained in the Japanese pronunciation string. The pronunciation of each syllable is expressed in one state. Japanese includes 102 syllables. The program is further configured to cause a computer to implement a display control function of displaying the mark of the word candidate on a display according to a display attribute of the word candidate and displaying the input mark string including the word candidate on a display.
Citation Information
Patent Citations
Manufacturing method of light-emitting module and light-emitting module
JP2020184610A
Testing and tuning of automatic speech recognition systems using synthetic inputs generated from its acoustic models
US20060085187A1
Searching device, searching method, and program
US20130006629A1