Voice recognition device, voice recognition method, and recording medium

The speech recognition device and method enhance accuracy by integrating context information to predict semantic content, addressing the limitations of existing technologies in accurately recognizing speech.

WO2026004103A1PCT designated stage Publication Date: 2026-01-02NEC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/023550
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-28
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing speech recognition technologies struggle to improve accuracy by leveraging context information related to the semantic content of the speech being recognized.

Method used

A speech recognition device and method that incorporates context information, including related strings and context feature information, to enhance recognition accuracy by predicting the content of the speech.

Benefits of technology

Improves speech recognition accuracy by utilizing context information to predict the semantic content, reducing the need for manual context preparation and enhancing recognition performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024023550_02012026_PF_FP_ABST
    Figure JP2024023550_02012026_PF_FP_ABST
Patent Text Reader

Abstract

This voice recognition device comprises: a context reception means that receives input of context information including at least one of a related character string related to a query and context feature information extracted from the related character string; and a voice recognition means that recognizes a voice using the context information and outputs a recognition result for the voice.
Need to check novelty before this filing date? Find Prior Art

Description

Voice recognition device, voice recognition method, and recording medium

[0001] The present disclosure relates to the technical fields of a voice recognition device, a voice recognition method, and a recording medium.

[0002] A speech recognition technique has been proposed that, for example, selects from input text a plurality of candidate character strings to be recognized as words and phrases, and for each selected candidate character string, generates a plurality of candidate pronunciations for the candidate character strings by combining predetermined pronunciations with each character contained in the candidate character string. Data associating each of the generated candidate pronunciations with each candidate character string is combined with language model data that records a numerical value indicating the frequency with which each word or phrase appears in the text to generate frequency data indicating the frequency of appearance for each pair of character string representing a word and pronunciation. The input speech is recognized based on the generated frequency data, and recognition data is generated in which, for each of a plurality of words and phrases contained in the input speech, character strings representing the words and phrases are associated with pronunciations. A combination of candidate character strings and pronunciations is selected and output from the combinations included in the recognition data (see Patent Document 1).

[0003] Japanese Patent Application Laid-Open No. 2008-216756

[0004] An object of this disclosure is to provide a speech recognition device, a speech recognition method, and a recording medium that aim to improve upon the techniques related to the prior art documents mentioned above.

[0005] One aspect of the speech recognition device includes a context receiving means for receiving input of context information including at least one of a related string related to a query and context feature information extracted from the related string, and a speech recognition means for recognizing speech using the context information and outputting a recognition result of the speech.

[0006] One aspect of the speech recognition method is a computer-executed speech recognition method that includes accepting input of context information including at least one of a related string related to a query and context feature information extracted from the related string, recognizing speech using the context information, and outputting a recognition result of the speech.

[0007] One aspect of the recording medium has recorded thereon a computer program for causing a computer to execute a speech recognition method that includes accepting input of context information including at least one of related strings related to a query and context feature information extracted from the related strings, recognizing speech using the context information, and outputting a recognition result of the speech.

[0008] 1 is a block diagram showing an example of the configuration of a speech recognition device according to an embodiment. FIG. 2 is a flowchart showing an example of the operation of the speech recognition device according to an embodiment. FIG. 3 is a block diagram showing the configuration of a speech recognition device according to an embodiment. FIG. 4 is a block diagram showing an example of the operation of the speech recognition device according to an embodiment. FIG. 5 is a block diagram showing an example of the configuration of a speech recognition device according to an embodiment. FIG. 6 is a block diagram showing an example of the operation of the speech recognition device according to an embodiment. FIG. 7 is a block diagram showing an example of the configuration of a speech recognition device according to an embodiment. FIG. 8 is a block diagram showing an example of the operation of the speech recognition device according to an embodiment.

[0009] Hereinafter, embodiments of a speech recognition device, a speech recognition method, and a recording medium will be described with reference to the drawings. [1: First Embodiment]

[0010] A first embodiment of a voice recognition device, a voice recognition method, and a recording medium will be described with reference to Figures 1 and 2. In the following, the first embodiment of a voice recognition device, a voice recognition method, and a recording medium will be described using a voice recognition device 10.

[0011] 1, the speech recognition device 10 includes a context receiving unit 11 and a speech recognition unit 12. The operation of the speech recognition device 10 will be described with reference to the flowchart of FIG.

[0012] 2, the context receiving unit 11 receives input of context information (step S11). The context information includes at least one of a character string related to the query (referred to as a "related character string") and context feature information extracted from the related character string.

[0013] A query may be a request for information related to the speech to be recognized, a request for information related to the content of the speech to be recognized, or a request for information different from the content of the speech to be recognized.

[0014] Context refers to the continuity of words in a character string, and the continuity of semantic content within the flow of the character string. In this embodiment, context information may be information related to the semantic content indicated by the speech to be recognized. Related character strings may be character strings related to the semantic content indicated by the speech to be recognized. Context feature information extracted from related character strings may be information related to the semantic content indicated by the speech to be recognized.

[0015] The speech recognition unit 12 recognizes the speech based on the context information (step S12). The speech recognition unit 12 outputs the speech recognition result (step S13). The speech recognition result may be a character string.

[0016] In this way, the speech recognition device 10 performs a speech recognition method that includes accepting input of context information including at least one of related strings related to the query and context feature information extracted from the related strings, recognizing speech using the context information, and outputting the speech recognition result.

[0017] The speech recognition device 10 described above may be realized by a computer reading a computer program recorded on a recording medium. In this case, the computer program may cause the computer to execute a speech recognition method including, when a query is input, accepting input of context information including at least one of related strings related to the query and context feature information extracted from the related strings, recognizing speech using the context information, and outputting a speech recognition result. [Technical Effect]

[0018] The speech recognition device 10 according to the present disclosure accepts input of context information, and can recognize speech using the context information. This improves speech recognition accuracy. [2: Second Embodiment]

[0019] A second embodiment of a voice recognition device, a voice recognition method, and a recording medium will be described with reference to FIGS. 3 to 7. In the following, the second embodiment of a voice recognition device, a voice recognition method, and a recording medium will be described using a voice recognition device 20. Note that, for the second embodiment, descriptions that overlap with the description of the first embodiment will be omitted as appropriate. [2-1: Configuration of the voice recognition device 20]

[0020] The configuration of the voice recognition device 20 will be described with reference to Fig. 3. Fig. 3 is a block diagram showing the configuration of the voice recognition device 20.

[0021] 3, the speech recognition device 20 includes a calculation device 21, a storage device 22, and a communication device 23. The speech recognition device 20 may further include an input device 24 and an output device 25. However, the speech recognition device 20 does not necessarily include at least one of the input device 24 and the output device 25. The calculation device 21, the storage device 22, the communication device 23, the input device 24, and the output device 25 may be connected via a data bus 26.

[0022] The arithmetic device 21 includes at least one processor (i.e., one processor or multiple processors) as hardware. The processor may include, for example, a processor conforming to a von Neumann computer architecture. The processor conforming to the von Neumann computer architecture may include at least one of a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit). The processor may include, for example, a processor conforming to a non-von Neumann computer architecture. The processor conforming to the non-von Neumann computer architecture may include at least one of an FPGA (Field Programmable Gate Array) and an ASIC (Application Specific Circuit).

[0023] The arithmetic device 21 reads a computer program 221 including at least one of computer program code and computer program instructions. For example, the arithmetic device 21 may read the computer program 221 stored in the storage device 22. For example, the arithmetic device 21 may read the computer program 221 stored in a computer-readable, non-transitory recording medium using a recording medium reading device (not shown) included in the voice recognition device 20. The computer program 221 read from the recording medium may be stored in the storage device 22. The arithmetic device 21 may acquire (i.e., download or read) the computer program 221 from a device (not shown) located outside the voice recognition device 20 via the communication device 23 (or another communication device). The downloaded computer program 221 may be stored in the storage device 22.

[0024] The arithmetic device 21 executes the loaded computer program 221. As a result, logical functional blocks for executing information processing to be performed by the speech recognition device 20 are realized within the arithmetic device 21. In other words, the arithmetic device 21, together with the storage device 22 or the like in which the computer program 221 is recorded (in other words, together with the storage device 22 and the computer program 221 recorded in the storage device 22 or the like), can function as a controller or a computer for realizing logical functional blocks for executing processing to be performed by the speech recognition device 20. In other words, the at least one processor included in the arithmetic device 21, the memory (recording medium) included in the storage device 22 or the like, and the computer program 221 are configured so that the speech recognition device 20 performs the information processing to be performed by the speech recognition device 20.

[0025] A computational model that can be constructed by machine learning may be implemented in the computational device 21 by the computational device executing the computer program 221. An example of a computational model that can be constructed by machine learning is a computational model including a neural network (so-called artificial intelligence (AI)). In this case, learning of the computational model may include learning of parameters of the neural network (e.g., at least one of a weight and a bias). The computational device 21 may execute a context output process and a speech recognition process using the computational model. In other words, the operation of executing the context output process and the speech recognition process may include the operation of executing the context output process and the speech recognition process using the computational model. Note that a computational model that has been constructed by offline machine learning using training data may be implemented in the computational device 21. Furthermore, the computational model implemented in the computational device 21 may be updated by online machine learning on the computational device 21. Alternatively, the calculation device 21 may perform the context output processing and the speech recognition processing using a calculation model implemented in a device external to the calculation device 21 (i.e., a device provided outside the speech recognition device 20) in addition to or instead of the calculation model implemented in the calculation device 21.

[0026] The recording medium for recording the computer program 221 executed by the arithmetic device 21 may be at least one of a CD-ROM, CD-R, CD-RW, flexible disk, MO, DVD-ROM, DVD-RAM, DVD-R, DVD+R, DVD-RW, DVD+RW, Blu-ray (registered trademark), or other optical disk, a magnetic medium such as a magnetic tape, a magneto-optical disk, a semiconductor memory such as a USB memory, or any other medium capable of storing a program. The recording medium may include a device capable of recording a computer program (for example, a general-purpose device or a dedicated device in which the computer program 221 is implemented in a state in which it can be executed in at least one of the forms of software and firmware). Furthermore, each process or function included in the computer program 221 may be realized by a logical processing block realized within the arithmetic device 21 when the arithmetic device 21 (i.e., processor) executes the computer program 221, or may be realized by hardware such as a predetermined gate array (FPGA (Field Programmable Gate Array), ASIC (Application Specific Integrated Circuit)) provided in the arithmetic device 21, or may be realized in a form that mixes logical processing blocks and partial hardware modules that realize some elements of the hardware.

[0027] The storage device 22 includes at least one memory capable of storing desired data. In other words, the storage device 22 includes at least one memory containing desired data. For example, the storage device 22 may store a computer program 221 executed by the arithmetic device 21. In this case, the storage device 22 (memory) may be used as the above-mentioned recording medium for recording the computer program 221 executed by the arithmetic device 21. The storage device 22 may temporarily store data used by the arithmetic device 21 when the arithmetic device 21 is executing the computer program 221. The storage device 22 may store data to be stored long-term by the voice recognition device 20. The storage device 22 may include at least one of a RAM (Random Access Memory), a ROM (Read Only Memory), a hard disk device, a magneto-optical disk device, an SSD (Solid State Drive), and a disk array device. In other words, the storage device 22 may include a non-transitory recording medium.

[0028] The communication device 23 may be capable of communicating with devices external to the voice recognition device 20. The communication device 23 may perform wired communication or wireless communication.

[0029] The input device 24 is a device capable of accepting information input to the voice recognition device 20 from outside. The input device 24 may include an operation device (e.g., a keyboard, a mouse, a touch panel, etc.) that can be operated by a user of the voice recognition device 20. The input device 24 may include a recording medium reading device capable of reading information recorded on a recording medium that is detachable from the voice recognition device 20, such as a USB (Universal Serial Bus) memory. Note that when information is input to the voice recognition device 20 via the communication device 23 (in other words, when the voice recognition device 20 obtains information via the communication device 23), the communication device 23 may function as an input device.

[0030] The output device 25 is a device capable of outputting information to the outside of the voice recognition device 20. The output device 25 may output visual information such as characters or images, auditory information such as voice, or tactile information such as vibration, as the information. The output device 25 may include, for example, at least one of a display, a speaker, a printer, and a vibration motor. The output device 25 may be capable of outputting information to a recording medium detachable from the voice recognition device 20, such as a USB memory. Note that when the voice recognition device 20 outputs information via the communication device 23, the communication device 23 may function as the output device.

[0031] FIG. 3 shows an example of logical functional blocks realized in the arithmetic device 21 to perform speech recognition. As shown in FIG. 3, a context output unit 211 and a recognition unit 212 may be realized in the arithmetic device 21. However, the context output unit 211 may be realized outside the arithmetic device 21. The context output unit 211 may have a character string output unit 2111 and a vector extraction unit 2112. The recognition unit 212 may have a speech acquisition unit 2121, a context acceptance unit 2122, and a speech recognition unit 2123. The "context acceptance unit 2122" is a component corresponding to the "context acceptance and output unit 11" in the first embodiment described above, and the "speech recognition unit 2123" is a component corresponding to the "speech recognition unit 12" in the first embodiment described above. [2-2: Speech recognition method executed by speech recognition device 20]

[0032] The speech recognition method executed by the speech recognition device 20 will be described with reference to Fig. 4. Fig. 4 is a block diagram showing an example of the flow of information in the speech recognition method executed by the speech recognition device 20.

[0033] As shown in FIG. 4 , when a query is input, the string output unit 2111 outputs a related string related to the query. The query may be a request for information related to the speech to be recognized. The query may also be a request for information related to the content of the speech to be recognized. The content of the speech may be the content of the utterance. The content of the speech may be referred to as a "topic." For example, the query may be content such as "What was the topic of the speech?"

[0034] When the query is a request for information related to the content of the speech to be recognized, the related string may be a string related to the content of the speech to be recognized. For example, when attempting to recognize the content of speech in a meeting, the related string may be a string summarizing the content of the agenda to be spoken. The related string may also be text contained in some or all of the materials used in the meeting. For example, when speech related to a meeting is to be recognized, the string output unit 2111 may output a string contained in the agenda of the relevant meeting as the related string.

[0035] The query may be a request for information different from the content of the speech to be recognized. When the query is a request for information different from the content of the speech to be recognized, the related string may have information different from the content of the speech to be recognized. For example, the related string may be a string related to the type of speech language to be recognized (e.g., Japanese, English, etc.). The related string may also be a string related to the caller of the speech to be recognized. In this case, the related string may be a string related to the gender, age, etc. of the speaker. The related string may also be a string related to the environment in which the speech to be authenticated is recorded. In this case, the related string may be a string related to the place where it is spoken. Specifically, the related string may be, for example, "conference room," "inside a car," etc.

[0036] The related character string may be a character string containing information that can be used when recognizing the speech of the recognition target. For example, the related character string may be a character string containing information that can improve the accuracy of recognizing the speech of the recognition target.

[0037] A query may be a request for multiple different pieces of information. If the query is a request for multiple types of information, the associated string may be a string that includes multiple types of information.

[0038] The string output unit 2111 may output related strings related to the query using a Retrieval Augmented Generation (RAG) technique. The string output unit 2111 may obtain information related to the query from an arbitrary information source and output related strings related to the query. The string output unit 2111 may be connected to an arbitrary information source online.

[0039] The vector extraction unit 2112 extracts context feature information from the related strings. The context feature information may be a sentence vector (sentence embedding) extracted from the related strings. The sentence vector extracted from the related strings is called a "context vector." Because the context vector is extracted from the related strings, it contains information that can be used to recognize the recognition target, just like the related strings. The vector extraction unit 2112 may extract information that can be used to improve recognition accuracy from relatively long related strings and output it as a context vector.

[0040] The context vector extracted by the vector extraction unit 2112 is output to the context reception unit 2122. The context reception unit 2122 receives input of the context vector as context information.

[0041] The speech acquisition unit 2121 acquires speech. The speech acquisition unit 2121 may acquire an entire group of speech all at once. The group of speech may be a group of speech to be recognized. The group of speech may be, for example, speech included in a single recorded audio file. The group of speech may be, for example, speech from the beginning to the end of an event to be recognized (a meeting, a speech, etc.).

[0042] Alternatively, the voice acquisition unit 2121 may acquire a portion of a group of voices. The voice acquisition unit 2121 may acquire a portion of a group of voices sequentially.

[0043] The speech recognition unit 2123 recognizes speech using a context vector as context information and outputs the speech recognition result. The speech recognition unit 2123 performs recognition processing on the speech to be recognized using the context indicated by the context vector. The speech recognition unit 2123 outputs a character string that is the recognition result of the speech to be recognized.

[0044] The speech recognition unit 2123 receives context information and speech to be recognized as input, and outputs a character string that is the recognition result of the speech to be recognized. The speech recognition unit 2123 may receive a context vector and speech to be recognized as input, and output a character string that is the recognition result of the speech to be recognized.

[0045] The speech recognition unit 2123 may perform recognition processing on all of a group of speeches at once, and in this case, the speech recognition unit 2123 may output all of the recognition results for the group of speeches at once.

[0046] The speech recognition unit 2123 may sequentially perform recognition processing on the entire block of speech. When the speech acquisition unit 2121 sequentially acquires portions of a block of speech, the speech recognition unit 2123 may sequentially perform recognition processing on the portions of speech acquired by the speech acquisition unit 2121. In this case, the speech recognition unit 2123 may sequentially output recognition results for the portions of speech.

[0047] Alternatively, the speech recognition unit 2123 may recognize speech to be recognized for each predetermined recognition unit. The predetermined recognition unit may be, for example, one character or one word. In this case, the speech recognition unit 2123 may output a recognition result of the speech for each recognition unit. [2-3: Modification 1]

[0048] 5 shows a first variation of logical functional blocks implemented in the calculation device 21 of the speech recognition device 20′ for performing speech recognition. The context output unit 211′ may include a character string output unit 2111. The recognition unit 212′ may include a speech acquisition unit 2121, a context acceptance unit 2122, a speech recognition unit 2123, and a vector extraction unit 2112′.

[0049] 6 , the related character strings output by the character string output unit 2111 are input to the context receiving unit 2122. The context receiving unit 2122 receives the related character strings as context information. The vector extraction unit 2112′ extracts a context vector from the related character strings received by the context receiving unit 2122.

[0050] The speech recognition unit 2123 may receive a related character string and a speech to be recognized, and output a character string that is a recognition result of the speech to be recognized. In this case, the speech recognition device 20 does not need to include the vector extraction unit 2112. [2-4: Modification 2]

[0051] 7 shows a second variation of the speech recognition method. The context output unit 211" may not include the string output unit 2111 and the vector extraction unit 2112. In other words, the context output unit 211" may output a context vector extracted from a related string related to a query, without outputting a related string related to the query and extracting a context vector from the related string. The context output unit 211" may, for example, obtain a context vector extracted in a device (external device) provided outside the speech recognition device 20 from the external device and output the context vector.

[0052] The speech recognition device 20 may terminate speech recognition when it has completed recognition of all acquired speeches to be recognized. Alternatively, the speech recognition device 20 may terminate speech recognition when it detects that the speaker has finished speaking. Alternatively, the speech recognition device 20 may terminate speech recognition when it detects various operations by the user (for example, operation of a recognition termination button). [2-5: Technical Effects]

[0053] The speech recognition device 20 according to this disclosure can improve the accuracy of speech recognition processing by using related character strings related to the content of the speech to be recognized. By using context information, it becomes possible to predict the content of the speech, and highly accurate speech recognition can be performed.

[0054] The context information does not have to be information that specifies the character string itself of the speech recognition result. The speech recognition device 20 can improve the accuracy of speech recognition without receiving the character string itself of the speech recognition result.

[0055] Furthermore, since the context information input to the speech recognition device 20 is information output by the context output unit 211, there is no need to manually create the context information. Therefore, compared to using a user dictionary in which words are registered, for example, preparing the context information is easier. [3: Third Embodiment]

[0056] A third embodiment of a voice recognition device, a voice recognition method, and a recording medium will be described with reference to Figures 8 and 9. Hereinafter, the third embodiment of a voice recognition device, a voice recognition method, and a recording medium will be described using a voice recognition device 30. Note that, for the third embodiment, descriptions that overlap with the descriptions of the first and second embodiments will be omitted as appropriate. Note that, in the drawings, parts common to the first and second embodiments are designated by the same reference numerals.

[0057] 8, the arithmetic unit 21 included in the speech recognition device 30 includes, as logical functional blocks, a context output unit 311 and a recognition unit 312. The recognition unit 312 may include a speech acquisition unit 2121, a context acceptance unit 2122, and a speech recognition unit 3123. [3-1: Speech recognition method executed by the speech recognition device 30]

[0058] The speech recognition method executed by the speech recognition device 30 will be described with reference to Fig. 9. Fig. 9 is a block diagram showing an example of the flow of information in the speech recognition method executed by the speech recognition device 30.

[0059] 9, when a query is input, the context output unit 311 outputs context information. The context receiving unit 2122 receives input of context information.

[0060] The voice acquisition unit 2121 acquires voice. The voice acquisition unit 2121 may acquire the entire group of voice all at once. Alternatively, the voice acquisition unit 2121 may acquire a portion of the group of voice. The voice acquisition unit 2121 may sequentially acquire the portions of the group of voice.

[0061] The speech recognition unit 3123 may sequentially recognize a group of speech parts and output a recognition result for the group of speech parts. The group of speech parts may be a predetermined recognition unit of speech. The predetermined recognition unit may be, for example, one character or one word.

[0062] The speech recognition unit 3123 in the third embodiment outputs the speech recognition result to the outside of the recognition unit 312, provides it as confirmable information, and uses it to sequentially recognize the next part of the speech. The speech recognition unit 3123 recognizes the speech using the recognition result of the part of the speech that is a group, in addition to the context information accepted by the context acceptance unit 2122. [3-2: Technical Effects]

[0063] The speech recognition device 30 according to this disclosure uses at least the recognition result of the immediately preceding speech, i.e., information indicating the continuity of the semantic content in the flow of character strings up to that point, and therefore can recognize speech more appropriately. [4: Fourth Embodiment]

[0064] A fourth embodiment of a voice recognition device, a voice recognition method, and a recording medium will be described with reference to Figures 10 and 11. Hereinafter, the fourth embodiment of a voice recognition device, a voice recognition method, and a recording medium will be described using a voice recognition device 40. Note that, for the fourth embodiment, descriptions that overlap with the descriptions of the first to third embodiments will be omitted as appropriate. Note that, in the drawings, parts common to the first to third embodiments are designated by the same reference numerals.

[0065] 10, the arithmetic unit 21 included in the speech recognition device 40 includes, as logical functional blocks, a context output unit 311 and a recognition unit 412. The recognition unit 412 may include a first conversion unit 4124 and a second conversion unit 4125 in addition to a speech acquisition unit 2121, a context acceptance unit 2122, and a speech recognition unit 4123. [4-1: Speech recognition method executed by the speech recognition device 40]

[0066] The speech recognition method executed by the speech recognition device 40 will be described with reference to Fig. 11. Fig. 11 is a block diagram showing an example of the flow of information in the speech recognition method executed by the speech recognition device 40.

[0067] 11 , when a query is input, the context output unit 311 outputs context information. The context receiving unit 2122 receives input of context information. The voice acquiring unit 2121 acquires voice.

[0068] The first conversion unit 4124 converts the speech into speech feature information. The first conversion unit 4124 may convert the speech into speech feature information suitable for recognition by the speech recognition unit 4123.

[0069] The second conversion unit 4125 converts the context information into text feature information. The second conversion unit 4125 may convert the context information into text feature information suitable for recognition by the speech recognition unit 4123. The second conversion unit 4125 may convert the context information into a sentence vector of a predetermined length. The second conversion unit 4125 may convert the context information into a sentence vector of a length suitable for recognition by the speech recognition unit 4123. In other words, the sentence vector of a predetermined length may be a sentence vector of a length suitable for recognition by the speech recognition unit 4123. The sentence vector of a predetermined length may be an embedding sequence corresponding to the number of fixed-length tokens.

[0070] The second conversion unit 4125 may convert the related character strings as context information into sentence vectors of a predetermined length. Alternatively, the second conversion unit 4125 may convert the context vectors as context information into sentence vectors of a predetermined length. In this case, the length of the context vector may be different from the predetermined length. Furthermore, the context vectors as context information may be sentence vectors suitable for conversion into sentence vectors of a predetermined length.

[0071] Furthermore, the second conversion unit 4125 converts the recognition result of the group of speech parts into text feature information. The second conversion unit 4125 may convert the recognition result of the group of speech parts into text feature information suitable for recognition by the speech recognition unit 4123. The speech recognition unit 4123 may receive as input text feature information obtained by converting speech feature information and context information, and text feature information obtained by converting the recognition result of the group of speech parts, and output a speech recognition result.

[0072] The second conversion unit 4125 may output combined feature information obtained by combining text feature information obtained by converting context information and text feature information obtained by converting a recognition result of a group of speech segments. In this case, the speech recognition unit 4123 may receive the speech feature information and the combined feature information and output a speech recognition result. [4-2: Technical Effects]

[0073] The speech recognition device 40 according to the present disclosure converts the context information relating to the semantic content of the speech to be recognized into a desired length, and therefore can reduce the processing load even when the context information is relatively long. [5: Supplementary Note]

[0074] Some or all of the above embodiments may be described as, but are not limited to, the following supplementary notes. [Supplementary Note 1] A speech recognition device comprising: a context receiving means for receiving input of context information including at least one of related strings related to a query and context feature information extracted from the related strings; and a speech recognition means for recognizing speech using the context information and outputting a recognition result of the speech. [Supplementary Note 2] The speech recognition device according to Supplementary Note 1, comprising: a context output means for outputting the context information when the query is input. [Supplementary Note 3] The speech recognition device according to Supplementary Note 2, wherein the context output means has: a string output means for outputting the related strings when the query is input; and an extraction means for extracting the context feature information from the related strings, and outputs the context feature information as the context information, and the context receiving means receives input of the context feature information. [Supplementary Note 4] The speech recognition device according to Supplementary Note 2, wherein the context output means has string output means for outputting the related string when a query is input, and outputs the related string as the context information, the context receiving means receives input of the related string, and the speech recognition means has extraction means for extracting the context feature information from the related string. [Supplementary Note 5] The speech recognition device according to Supplementary Note 1, wherein the speech recognition means recognizes the speech for each predetermined recognition unit and outputs a recognition result of the speech for each predetermined recognition unit. [Supplementary Note 6] The speech recognition device according to Supplementary Note 1, wherein the speech recognition means sequentially recognizes parts of a group of speech and outputs a recognition result for the part of the group of speech. [Supplementary Note 7] The speech recognition device according to Supplementary Note 6, wherein the speech recognition means recognizes the speech using the recognition result for the part of the group of speech in addition to the context information. [Supplementary Note 8] The speech recognition device according to Supplementary Note 7, wherein the speech recognition means has second conversion means for converting the context information and each of the recognition results for the group of speech portions into text feature information, and recognizes the speech using the text feature information. [Supplementary Note 9] The speech recognition device according to Supplementary Note 8, wherein the second conversion means converts the context information into text feature information of a predetermined length.[Supplementary Note 10] The speech recognition device according to Supplementary Note 1, wherein the speech recognition means has first conversion means for converting the speech into speech feature information, wherein the speech feature information and the context information are input, and output a recognition result of the speech. [Supplementary Note 11] The speech recognition device according to Supplementary Note 2, wherein the context output means, when the query is input, outputs context feature information extracted from the related string. [Supplementary Note 12] A speech recognition method executed by a computer, comprising: accepting input of context information including at least one of related strings related to a query and context feature information extracted from the related strings; recognizing speech using the context information, and outputting a recognition result of the speech. [Supplementary Note 13] A recording medium having recorded thereon a computer program for causing a computer to execute a speech recognition method, comprising: accepting input of context information including at least one of related strings related to a query and context feature information extracted from the related strings; recognizing speech using the context information, and outputting a recognition result of the speech.

[0075] Furthermore, some or all of the configurations described in Supplementary Notes 2 to 11, which are dependent on Supplementary Note 1, and Supplementary Notes 12 and 13, may be dependent in the same manner as Supplementary Notes 2 to 10. Furthermore, not limited to Supplementary Notes 1, 12, and 13, some or all of the configurations described as Supplements may be dependent on various hardware, software, various recording means for recording software, or systems, within the scope of each of the above-mentioned embodiments.

[0076] This disclosure may be modified as appropriate within the scope of the claims and the entire specification without departing from the gist or concept of the invention, and the speech recognition device, speech recognition method, and program that involve such modifications are also included in the technical concept of this disclosure.

[0077] 10, 20, 30, 40 Speech recognition device 11, 211, 311 Context output unit 12, 212, 312, 412 Recognition unit 2111 Character string output unit 2112 Vector extraction unit 2121 Speech acquisition unit 2122 Context reception unit 2123, 3123, 4123 Speech recognition unit 4124 First conversion unit 4125 Second conversion unit

Claims

1. A speech recognition device comprising: a context receiving means for receiving input of context information including at least one of related strings related to a query and context feature information extracted from the related strings; and a speech recognition means for recognizing speech using the context information and outputting a recognition result of the speech.

2. The speech recognition device according to claim 1, further comprising a context output means for outputting the context information when the query is input.

3. The speech recognition device according to claim 2, wherein the context output means includes: a string output means for outputting the related string when the query is input; and an extraction means for extracting the context feature information from the related string, and outputs the context feature information as the context information; and the context receiving means receives input of the context feature information.

4. The speech recognition device according to claim 2, wherein the context output means has a string output means for outputting the related string when the query is input, and outputs the related string as the context information; the context acceptance means accepts input of the related string; and the speech recognition means has an extraction means for extracting the context feature information from the related string.

5. The speech recognition device according to claim 2, wherein the context output means outputs context feature information extracted from the related character string when the query is input.

6. The speech recognition device according to claim 1, wherein said speech recognition means recognizes said speech for each predetermined recognition unit and outputs a recognition result of said speech for each predetermined recognition unit.

7. The speech recognition device according to claim 1, wherein said speech recognition means sequentially recognizes a group of speech parts and outputs a recognition result for said group of speech parts.

8. The speech recognition device according to claim 7, wherein the speech recognition means recognizes the speech using the recognition result of the part of the speech block in addition to the context information.

9. A speech recognition device according to claim 8, further comprising second conversion means for converting the context information and the recognition results of the chunks of speech into text feature information, and wherein the speech recognition means recognizes the speech using the text feature information.

10. The speech recognition device according to claim 9, wherein the second conversion means converts the context information into text feature information of a predetermined length.

11. A speech recognition device according to claim 1, further comprising a first conversion means for converting the speech into speech feature information, wherein the speech recognition means receives the speech feature information and the context information and outputs a recognition result of the speech.

12. A computer-implemented speech recognition method comprising: accepting input of context information including at least one of a related string related to a query and context feature information extracted from the related string; recognizing speech using the context information; and outputting a recognition result of the speech.

13. A recording medium having recorded thereon a computer program for causing a computer to execute a speech recognition method, the method including: accepting input of context information including at least one of related strings related to a query and context feature information extracted from the related strings; recognizing speech using the context information; and outputting the recognition result of the speech.

Citation Information

Patent Citations

  • Voice recognition system

    JP2008015439A

  • Information processing device, information processing method, and program

    WO2021171820A1