Information processing method, program, information processing device, and information processing system
The method improves speech recognition accuracy by estimating the speaker's speech style from video data and using a corresponding language model, addressing the issue of varying speech manners in different situations.
Patent Information
- Application Number
- JP2021193077
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-11-29
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2041-11-29
AI Technical Summary
Conventional speech recognition technologies fail to accurately reflect the varying speech manner of a person based on the situation, leading to decreased accuracy in speech recognition.
An information processing method that estimates the speaker's speech style from video data and selects or generates a language model corresponding to that style for improved speech recognition accuracy.
Enhances speech recognition accuracy by using a language model tailored to the speaker's speech style, accounting for different situations such as formal, informal, or spoken language.
Smart Images

Figure 0007790111000001 
Figure 0007790111000002 
Figure 0007790111000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing method, a program, an information processing device, and an information processing system. [Background technology]
[0002] Conventionally, in speech recognition for recognizing speech recorded together with video, a technology has been known in which a language model is selected based on the result of image recognition of the video, and speech recognition is performed using the selected language model. Specifically, for example, it is known to select a language model using attributes such as the gender and age of a person speaking in the video as the image recognition result, and perform speech recognition using the selected language model. Summary of the Invention [Problem to be solved by the invention]
[0003] However, even for the same person, the manner of speech varies depending on the situation when the speech is made, such as the person with whom the speech is made, the scene of the conversation, etc. Therefore, when a language model is selected based on the attributes of a person, as in the conventional technology described above, the manner of speech according to the situation may not be reflected in the language model, which may result in a decrease in the accuracy of speech recognition.
[0004] The disclosed technology has been developed in consideration of the above circumstances, and aims to improve the accuracy of speech recognition. [Means for solving the problem]
[0005] The disclosed technology is an information processing method by a computer, the computer receiving input of video data including a video of a speaker, When the video data is input, the device outputs estimation information indicating the probability that each speech style matches the speaker's speech style for each of a plurality of language models stored in a storage unit, and when a setting is made to select a language model according to the speaker's speech style, the device selects a language model corresponding to the speech style with the highest probability from the plurality of language models stored in the storage unit as the language model according to the speaker's speech style, and when a setting is made to generate a new language model according to the speaker's speech style, the device generates a new language model according to the speaker's speech style based on the estimation information, and provides the language model according to the selected speaker's speech style or the new language model to a speech recognition unit that performs speech recognition, and generates the language model according to the selected speaker's speech style or the new language model. and displaying text data resulting from speech recognition of audio data included in the video data using the above-mentioned method on a display device. [Effects of the Invention]
[0006] The accuracy of voice recognition can be improved. [Brief explanation of the drawings]
[0007] [Figure 1] FIG. 1 illustrates an example of an information processing system according to a first embodiment. [Figure 2] FIG. 2 illustrates an example of a hardware configuration of an information processing device. [Figure 3] FIG. 2 illustrates an example of a hardware configuration of a terminal device. [Figure 4] FIG. 2 is a diagram illustrating functions of the information processing apparatus according to the first embodiment. [Figure 5] FIG. 2 is a diagram illustrating a functional configuration of a terminal device according to the first embodiment. [Figure 6] 10 is a flowchart illustrating processing of the information processing apparatus of the first embodiment. [Figure 7] FIG. 2 is a first diagram illustrating processing by the information processing apparatus according to the first embodiment. [Figure 8] FIG. 10 is a second diagram illustrating the processing of the information processing apparatus of the first embodiment. [Figure 9] 10 is a flowchart illustrating processing by an information processing apparatus according to a second embodiment. [Figure 10] 10 is a flowchart illustrating a process of an information processing apparatus according to a third embodiment. [Figure 11] FIG. 10 illustrates an example of a system configuration according to a fourth embodiment. [Figure 12] FIG. 13 illustrates an example of a system configuration according to a fifth embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0008] (First embodiment) The first embodiment will be described below with reference to the drawings. Fig. 1 is a diagram showing an example of an information processing system according to the first embodiment.
[0009] The information processing system 100 of this embodiment includes an information processing device 200 and a terminal device 300, and the information processing device 200 and the terminal device 300 are connected via a network or the like.
[0010] In the information processing system 100 of this embodiment, the information processing device 200 has a voice recognition processing unit 220. That is, the information processing system 100 of this embodiment is an example of a voice recognition system.
[0011] When the information processing device 200 receives input of video data including voice data from the terminal device 300, the voice recognition processing unit 220 estimates, from the video data, the speaker's speaking style corresponding to the situation in which the utterance was made.
[0012] The information processing device 200 then performs speech recognition on the audio data included in the video data using a language model corresponding to the estimated speaking style, and outputs text data that is the result of the speech recognition. The text data may be output to the terminal device 300 and displayed on the terminal device 300, for example.
[0013] In other words, the speaker's manner of speaking is the manner of speaking that includes the words and manner of speaking used by the speaker. In the following description, the manner of speaking of the speaker may be referred to as the speaker's speaking style.
[0014] In this embodiment, the situations in which speech occurs specifically include, for example, a situation in which a speaker is reading a text printed on paper or displayed on a display, a situation in which multiple speakers who are close friends are enjoying a conversation, a situation in which multiple speakers are in a superior-subordinate relationship and the subordinate is reporting to the superior, and a situation in which the speakers are meeting each other for the first time.
[0015] The speaking style of a speaker may include, for example, written speech, spoken speech, and informal speech.
[0016] Written language is the language used in writing. Spoken language is the language used in conversation, and is more polite than informal speech. Informal speech is familiar language or everyday conversational language.
[0017] For example, if the situation is one in which a sentence is being read aloud, the speaker's speech style is likely to be the written language used when writing sentences. Also, if the situation is one in which the speakers are close to each other, the speaker's speech style is likely to be informal colloquial. Also, if the situation is one in which the speakers are meeting for the first time, the speaker's speech style is likely to be colloquial.
[0018] For example, when the information processing device 200 of this embodiment estimates that the speaker's speaking style is likely to be written language based on the speech situation shown in the video data, it performs speech recognition using a language model corresponding to written language.
[0019] The terminal device 300 of this embodiment transmits, for example, video data including audio data to the information processing device 200, and may be a smartphone, a tablet terminal, etc. Furthermore, the terminal device 300 may be, for example, an imaging device that acquires video data, or may be the imaging device itself.
[0020] Furthermore, the terminal device 300 of the present embodiment may be a spherical imaging device and may be installed, for example, at the center of a conversation. In this way, by using the spherical imaging device as the terminal device 300, it is possible to capture video data of all speakers participating in the conversation.
[0021] In this way, the information processing device 200 of this embodiment estimates the speaker's speaking style from the speech situation indicated by the video data, and performs speech recognition using a language model corresponding to the speaking style. Also, the information processing device 200 of this embodiment performs speech recognition using video data, which is time-series information that takes into account the situation before and after the utterance to be recognized.
[0022] Therefore, for example, in this embodiment, even if the same person speaks in different situations, speech recognition can be performed using a language model appropriate to the situation in which the utterance was made, thereby improving the accuracy of speech recognition.
[0023] 1, the information processing system 100 includes one terminal device 300, but the present invention is not limited to this. The information processing system 100 may include a plurality of terminal devices 300, and the terminal device 300 that transmits voice data to the information processing device 200 and the terminal device 300 that receives text data output from the information processing device 200 may be separate terminal devices.
[0024] In addition, in this embodiment, voice data may be received from the terminal device 300, and text data may be output to a display device included in the information processing device 200. In addition, in this embodiment, voice data may be directly input to the information processing device 200, and text data may be output to the terminal device 300. In this embodiment, voice data may be directly input to the information processing device 200, and text data may be output to a display device included in the information processing device 200.
[0025] 1, the information processing device 200 has the voice recognition processing unit 220, but the present invention is not limited to this. The voice recognition processing unit 220 may be realized by a plurality of information processing devices 200.
[0026] Next, the hardware configurations of the information processing device 200 and the terminal device 300 will be described with reference to Figures 2 and 3. Figure 2 is a diagram showing an example of the hardware configuration of the information processing device.
[0027] The information processing device 200 is constructed by a computer, and as shown in FIG. 2, includes a CPU 201, a ROM 202, a RAM 203, an HD 204, an HDD (Hard Disk Drive) controller 205, a display 206, an external device connection I / F (Interface) 208, a network I / F 209, a bus line B1, a keyboard 211, a pointing device 212, a DVD-RW (Digital Versatile Disk Rewritable) drive 214, and a media I / F 216.
[0028] Of these, the CPU 201 controls the overall operation of the information processing device 200. The ROM 202 stores programs used to drive the CPU 201, such as an IPL. The RAM 203 is used as a work area for the CPU 201. The HD 204 stores various data such as programs. The HDD controller 205 controls the reading and writing of various data from and to the HD 204 under the control of the CPU 201.
[0029] The display (display device) 206 displays various types of information such as a cursor, menus, windows, characters, or images. The external device connection I / F 208 is an interface for connecting various types of external devices. In this case, the external devices are, for example, USB (Universal Serial Bus) memories, printers, etc. The network I / F 209 is an interface for data communication using a communication network. The bus line B1 is an address bus, a data bus, etc. for electrically connecting each component such as the CPU 201 shown in FIG. 2.
[0030] The keyboard 211 is a type of input means having multiple keys for inputting characters, numbers, various instructions, etc. The pointing device 212 is a type of input means for selecting and executing various instructions, selecting a processing target, moving a cursor, etc. The DVD-RW drive 214 controls reading and writing of various data from and to a DVD-RW 213, which is an example of a removable recording medium. Note that this is not limited to a DVD-RW, and may be a DVD-R, etc. The media I / F 216 controls reading and writing (storing) of data from and to a recording medium 215, such as a flash memory.
[0031] 3 is a diagram showing an example of the hardware configuration of a terminal device 300 according to this embodiment. The terminal device 300 includes a CPU 301, a ROM 302, a RAM 303, an EEPROM 304, a CMOS sensor 305, an image sensor I / F 306, an acceleration / direction sensor 307, a media I / F 309, and a GPS receiver 311.
[0032] Of these, the CPU 301 is an arithmetic processing unit that controls the overall operation of the terminal device 300. The ROM 302 stores the CPU 301 and programs used to drive the CPU 301, such as the IPL. The RAM 303 is used as a work area for the CPU 301. The EEPROM 304 reads or writes various data, such as smartphone programs, under the control of the CPU 301. The ROM 302, RAM 303, and EEPROM 304 are examples of storage devices of the terminal device 300.
[0033] The CMOS (Complementary Metal Oxide Semiconductor) sensor 305 is a type of built-in imaging means that captures an image of a subject (mainly a self-portrait) to obtain video data under the control of the CPU 301. Note that instead of a CMOS sensor, an imaging means such as a CCD (Charge Coupled Device) sensor may also be used.
[0034] The imaging element I / F 306 is a circuit that controls the driving of the CMOS sensor 305. The acceleration / azimuth sensor 307 is a variety of sensors, such as an electronic magnetic compass that detects geomagnetism, a gyrocompass, and an acceleration sensor. The media I / F 309 controls the reading and writing (storage) of data from and to a recording medium 308, such as a flash memory. The GPS receiver 311 receives GPS signals from GPS satellites.
[0035] The terminal device 300 also includes a long-distance communication circuit 312, an antenna 312a of the long-distance communication circuit 312, a CMOS sensor 313, an image sensor I / F 314, a microphone (sound collection device) 315, a speaker 316, a sound input / output I / F 317, a display (display device) 318, an external device connection I / F (Interface) 319, a short-distance communication circuit 320, an antenna 320a of the short-distance communication circuit 320, and a touch panel 321.
[0036] Of these, the long-distance communication circuit 312 is a circuit that communicates with other devices via a communication network. The CMOS sensor 313 is a type of built-in imaging means that captures an image of a subject and obtains video data under the control of the CPU 301. The imaging element I / F 314 is a circuit that controls the driving of the CMOS sensor 313. The microphone 315 is a built-in circuit that converts sound into an electrical signal. The speaker 316 is a built-in circuit that converts an electrical signal into physical vibrations to generate sound such as music or voice. The audio input / output I / F 317 is a circuit that processes the input and output of audio signals between the microphone 315 and the speaker 316 under the control of the CPU 301. It should be noted that either the CMOS sensor 305 or the CMOS sensor 313 may be disposed near the display 318 and the other may be disposed on the back surface of the terminal device 300 .
[0037] The display 318 is a type of display means such as a liquid crystal display or organic electroluminescence (EL) display that displays an image of a subject, various icons, etc. The external device connection I / F 319 is an interface for connecting various external devices. The short-range communication circuit 320 is a communication circuit such as NFC (Near Field Communication) or Bluetooth (registered trademark). The touch panel 321 is a type of input means that allows a user to operate the terminal device 300 by pressing the display 318. The display 318 is an example of a display unit that the terminal device 300 has.
[0038] In this embodiment, the terminal device 300 is a smartphone or a tablet terminal, but is not limited to this. The terminal device 300 may be a general computer having a hardware configuration similar to that of the information processing device 200 shown in FIG.
[0039] Next, functions of the information processing device 200 of this embodiment will be described with reference to Fig. 4. Fig. 4 is a diagram for explaining functions of the information processing device of the first embodiment.
[0040] The information processing device 200 of this embodiment includes a speech recognition processing unit 220 , an input receiving unit 230 , a speech period detection unit 231 , a speech acquisition unit 232 , and a speech style estimation unit 233 .
[0041] The speech recognition processing unit 220 includes a speech recognition unit 221 , a language model providing unit 222 , a language model storage unit 223 , and an output text generating unit 224 .
[0042] A plurality of language models are stored in the language model storage unit 223. Specifically, a first language model 241, a second language model 242, and a third language model 243 are stored in the language model storage unit 223.
[0043] The first language model 241 is a language model corresponding to written language, the second language model 242 is a language model corresponding to spoken language, and the third language model 243 is a language model corresponding to informal spoken language.
[0044] Each of the first language model 241, the second language model 242, and the third language model 243 of this embodiment receives as input the speech recognition result output by the speech recognition unit 221, and outputs to the language model providing unit 222 the appearance probability of the next phoneme, character, or word for each processing unit of the speech recognition unit 221.
[0045] The input receiving unit 230 receives input of video data transmitted from an external device. Note that the video data in this embodiment includes audio data and moving image data.
[0046] The speech section detection unit 231 detects sections in the video data where speech is being made based on the audio data included in the input video data. The speech section detection unit 231 also outputs the audio data for each detected speech section to the audio acquisition unit 232, and outputs the video data to the speech style estimation unit 233.
[0047] The speech acquisition unit 232 acquires the speech data output from the speech period detection unit 231, extracts speech features from the speech data, and outputs the extracted features to the speech recognition unit 221 of the speech recognition processing unit 220. In other words, the speech acquisition unit 232 outputs the speech data acquired during a predetermined period that is set as a speech period to the speech recognition unit 221.
[0048] MFCC is known as a speech feature, but LPC (Linear Predictive Coding), FBANK (Log Mel-Filterbank Coefficients), etc. may also be used.
[0049] The speech style estimation unit 233 determines the situation in which the speaker made an utterance from the video data output from the speech section detection unit 231, and estimates the speaker's speech style based on the situation of the utterance. Then, the speech style estimation unit 233 generates speech style estimation information indicating the estimation result, and outputs the speech style estimation information to the language model provision unit 222 of the speech recognition processing unit 220. In other words, the speech style estimation unit 233 generates speech style estimation information for each predetermined period that is defined as a speech section, and outputs it to the language model provision unit 222. The speech style estimation information will be described in detail later.
[0050] The speech recognition unit 221 performs speech recognition based on the speech features input from the speech acquisition unit 232 and the language model selected by the language model provision unit 222, and outputs the recognition result. Specifically, the speech recognition unit 221 sequentially outputs text data, which is the result of the speech recognition for each utterance section, to the output text generation unit 224 and each language model stored in the language model storage unit 223.
[0051] The processing unit of the speech recognition in the speech recognition unit 221 of this embodiment may be a phoneme, a character, a word, etc. The processing unit of the speech recognition in the speech recognition unit 221 is the same as the processing units of the first language model 241, the second language model 242, and the third language model 243 stored in the language model storage unit 223. The language model providing unit 222 selects a language model corresponding to the speaker's speaking style from a plurality of language models stored in the language model storage unit 223, based on the speaking style estimation information input from the speaking style estimation unit 233. Then, the language model providing unit 222 provides the selected language model to the speech recognition unit 221.
[0052] In this embodiment, the language model to be provided can be specified in advance by inputting a language model selection signal to the language model providing unit 222. Specifically, for example, if the speech style of a speaker appearing in video data is specified in advance, the language model to be provided to the speech recognition unit 221 may be specified in advance by the language model selection signal. The language model selection signal may be input in advance by, for example, a user of the information processing system 100.
[0053] In this way, speech recognition can be performed using the specified language model without being affected by the estimation result of the speech style estimation unit 233.
[0054] The output text generation unit 224 generates output text data by chronologically concatenating text data that is the result of speech recognition for processing units, such as phonemes, characters, and words, output from the speech recognition unit 221. Then, the output text generation unit 224 outputs the generated output text data to the terminal device 300.
[0055] Here, we will explain the speech style estimation unit 233. The speech style estimation unit 233 of this embodiment is machine-learned in advance using a pair of learning data in which video data is input data and the speech style of a speaker in the video (wording used in the speaker's speech) is used as correct answer data.
[0056] In other words, the speech style estimation unit 233 of this embodiment is a trained model that takes video data as input data and outputs speech style estimation information for multiple speech styles, including the probability that each speech style matches the speech style of the speaker in the video.
[0057] In this embodiment, for example, the input data included in the learning data is video data captured of a presenter giving a presentation to a large audience using a screen on which presentation materials are projected. In this case, the correct answer data for the video data (input data) during the period in which the presenter is speaking to the audience is a spoken speaking style. Furthermore, the correct answer data for the video data during the period in which the presenter is reading from the screen or materials in front of him / her is a written speaking style. Furthermore, the correct answer data for the video data during the period in which one of the audience members stands up and speaks is a spoken speaking style.
[0058] In this embodiment, for example, the input data included in the training data is video data captured of a meeting with a small number of participants in a conference room or the like. In this case, correct answer data for video data captured during a time period in which a speaker is reading a document is a written speaking style, and correct answer data for video data captured during a time period in which the atmosphere is formal is a spoken speaking style. Also, correct answer data for video data captured during a time period in which a speaker is having a casual conversation with a colleague is a casual spoken speaking style.
[0059] Furthermore, when a superior is included among the participants in a meeting, the correct answer data corresponding to the video data of the meeting taken during the time period when the superior is speaking will be an informal speaking style, while the correct answer data corresponding to the video data of the time period when participants other than the superior are speaking to the superior will be an informal speaking style.
[0060] Furthermore, in this embodiment, when the speakers appearing in the video data that serves as input data are limited (for example, only members of a certain department of a certain company), learning may be performed to recognize each member and to learn their speaking habits and personal relationships. By training the speech style estimation unit 233 in this way, the accuracy of estimating the speaker's speech style can be improved.
[0061] Next, the speech style estimation information of this embodiment will be described.
[0062] The speech style estimation information of this embodiment includes information indicating the probability of each of a plurality of speech styles. Specifically, for example, the speech style estimation information includes the probability that the input video data will be a speech style using written language (first speech style), a speech style using spoken language (second speech style), and a speech style using casual spoken language (third speech style).
[0063] The probability for each speech style is the probability that a plurality of speech styles matches the speech style of a speaker in a video. The plurality of speech styles may be determined in advance.
[0064] An example of speech style estimation information is, for example, "the probability of the first speech style is 5%, the probability of the second speech style is 70%, and the probability of the third speech style is 25%." In this case, it can be seen that the speech style of the speaker in the video is most likely to be a speech style using spoken language (the second speech style).
[0065] 4, the information processing device 200 is provided with the input receiving unit 230, the speech segment detection unit 231, the speech acquisition unit 232, and the speech recognition processing unit 220, but is not limited to this. Some or all of these units may be provided in a device other than the information processing device 200 that is capable of communicating with the information processing device 200. In other words, the information processing device 200 may be realized by a plurality of information processing devices.
[0066] Next, the functional configuration of the terminal device 300 will be described with reference to Fig. 5. Fig. 5 is a diagram illustrating the functional configuration of the terminal device of the first embodiment.
[0067] The terminal device 300 of this embodiment includes a video data acquisition unit 330 , an output unit 340 , and a communication unit 350 .
[0068] The video data acquisition unit 330 acquires video data. Specifically, the video data acquisition unit 330 acquires video data captured by an imaging device or the like included in the terminal device 300.
[0069] The output unit 340 outputs various types of information from the terminal device 300. Specifically, the output unit 340 causes the display 318 to display text data received from the information processing device 200.
[0070] The communication unit 350 controls communication between the terminal device 300 and the information processing device 200. Specifically, the communication unit 350 transmits video data from the terminal device 300 to the information processing device 200, and receives text data from the information processing device 200.
[0071] Next, the processing of the information processing device 200 of this embodiment will be described with reference to Fig. 6. Fig. 6 is a flowchart illustrating the processing of the information processing device of the first embodiment.
[0072] In the information processing device 200 of this embodiment, the input receiving unit 230 receives input of video data from the terminal device 300 (step S601). Subsequently, the information processing device 200 detects a speech period from the video data using the speech period detection unit 231 (step S602).
[0073] Specifically, for example, the speech section detection unit 231 detects a first period from a certain timing to another timing as the speech section of a first speaker, and detects a second period following the first period as the speech section of a second speaker.
[0074] Next, the information processing device 200 extracts audio data and video data from the video data using the speech section detection unit 231, outputs the audio data to the speech recognition unit 221, and outputs the video data to the speech style estimation unit 233 (step S603).
[0075] Following step S603, the information processing device 200 causes the speech recognition unit 221 to extract speech features from the speech data (step S604), and proceeds to step S607, which will be described later.
[0076] Furthermore, following step S603, the information processing device 200 inputs the video data to the speech style estimation unit 233 and acquires speech style estimation information (step S605).
[0077] Next, the information processing device 200 inputs the speech style estimation information to the language model providing unit 222, selects a language model to be provided to the speech recognition unit 221 from the language models stored in the language model storage unit 223, and provides the selected language model to the speech recognition unit 221 (step S606).
[0078] Specifically, the language model providing unit 222 refers to the probability of each language model included in the speech style estimation information, selects the language model with the highest probability from among the language models stored in the language model storage unit 223, and provides it to the speech recognition unit 221.
[0079] Next, the information processing device 200 performs speech recognition for each processing unit in the speech recognition unit 221 based on the notified language model and speech feature amount, and outputs text data for each processing unit to the output text generation unit 224 (step S607).
[0080] Next, the information processing device 200 generates output text data by collecting the text data of the processing units using the output text generating unit 224, and outputs the generated text data to the terminal device 300 (step S608).
[0081] In this embodiment, the speech style according to the speaker's speech situation is estimated based on video data of the speaker, and a language model corresponding to the estimated speech style is selected and used for speech recognition.
[0082] Therefore, according to this embodiment, the speaker's speaking style can be estimated based on time-series information (video data) that takes into account the circumstances before and after the utterance to be recognized. Therefore, the speaking style according to the circumstances in which the speaker made the utterance can be used for speech recognition, thereby improving the accuracy of speech recognition.
[0083] The processing of the information processing device 200 of this embodiment will be specifically described below with reference to Figures 7 and 8. Figure 7 is a first diagram illustrating the processing of the information processing device of the first embodiment.
[0084] In FIG. 7, the video data G1 received by the input receiving unit 230 is a captured image of two people P1 and P2 who are relatively close, enjoying a conversation while eating and drinking.
[0085] The speech section detection unit 231 detects the speech sections of person P1 and person P2 from this video data G1, outputs the audio data of the detected sections to the audio acquisition unit 232, and outputs the video data to the speech style estimation unit 233.
[0086] When video data of an utterance section in which the speaker is person P1 is input, the utterance style estimation unit 233 outputs utterance style estimation information for the utterance section.
[0087] 7, person P1 and person P2 have a relatively close relationship. Therefore, the speech style estimation unit 233 generates speech style estimation information that is most likely to be the third language model 243, next most likely to be the second language model 242, and least likely to be the first language model 241, and outputs the generated speech style estimation information to the language model providing unit 222.
[0088] When this speech style estimation information is input, the language model providing unit 222 provides the third language model 243 with the highest probability to the speech recognition unit 221 .
[0089] The voice recognition unit 221 performs voice recognition based on the voice feature amount of the voice data of the person P1 and the third language model 243, and outputs text data as a result of the voice recognition.
[0090] Next, when video data of an utterance section in which the speaker is person P2 is input, the utterance style estimation unit 233 outputs utterance style estimation information for the utterance section. At this time, just like in the case of person P1, the utterance style estimation unit 233 generates utterance style estimation information in which the third language model 243 is most likely to be the third language model, the second language model 242 is next most likely to be the second language model, and the first language model 241 is least likely to be the first language model, and outputs the generated utterance style estimation information to the language model providing unit 222.
[0091] When this speech style estimation information is input, the language model providing unit 222 provides the third language model 243 with the highest probability to the speech recognition unit 221 .
[0092] The voice recognition unit 221 performs voice recognition based on the voice feature amount of the voice data of the person P2 and the third language model 243, and outputs text data as a result of the voice recognition.
[0093] Next, the information processing device 200 generates output text data that combines the two pieces of text data, and causes the terminal device 300 to display the output text data.
[0094] A screen 71 shown in FIG. 7 is an example of a screen on which output text data is displayed on the terminal device 300.
[0095] The screen 71 includes display areas 72 and 73. The display area 72 displays text data T1, which is the result of speech recognition performed on the speech data of person P1 using the third language model 243. The display area 73 displays text data T2, which is the result of speech recognition performed on the speech data of person P2 using the third language model 243.
[0096] That is, in this embodiment, the text data resulting from speech recognition is displayed on the screen 71 for each speaker.
[0097] On the screen 71, the text data T1 is "This sweet is delicious," which is written in casual speech. On the screen 71, the text data T2 is "It is delicious after all," which is written in casual speech.
[0098] 8 is a second diagram illustrating the processing of the information processing device of the first embodiment. In FIG. 8, video data G2 received by input receiving unit 230 is an image of person P3 giving a presentation while displaying materials on display 85.
[0099] The speech section detection section 231 detects the speech section of person P3 from this video data G2, outputs the audio data of the detected section to the audio acquisition section 232, and outputs the video data to the speech style estimation section 233.
[0100] When video data of an utterance section in which the speaker is person P3 is input, the utterance style estimation unit 233 outputs utterance style estimation information for the utterance section.
[0101] 8, person P3 is reading a sentence displayed on the display 85. Therefore, the speech style estimation unit 233 generates speech style estimation information that is most likely to be the first language model 241, next most likely to be the second language model 242, and least likely to be the third language model 243, and outputs the generated speech style estimation information to the language model providing unit 222.
[0102] When this speech style estimation information is input, the language model providing unit 222 provides the first language model 241 with the highest probability to the speech recognition unit 221 .
[0103] The speech recognition unit 221 performs speech recognition based on the speech feature amount of the speech data of the person P3 and the first language model 241, and outputs text data as a result of the speech recognition.
[0104] A screen 81 shown in FIG. 8 is an example of a screen on which output text data is displayed on the terminal device 300.
[0105] The screen 81 displays text data T3, which is the result of performing speech recognition on the speech data of person P3 using the first language model.
[0106] On the screen 81, the text data T3 is "In this case, the result will be like this," which is written language.
[0107] In this way, in this embodiment, the speech style, which indicates the language and speaking manner of the speaker, is estimated based on video data (video data) capturing the situation in which the speaker is speaking, and speech recognition is performed using a language model corresponding to the speech style.
[0108] Therefore, for example, even if person P1 in Figure 7 and person P3 in Figure 8 are the same person, speech recognition can be performed using a language model that corresponds to the speaking style that corresponds to the situation in which the speech was made, thereby improving the accuracy of speech recognition.
[0109] In this embodiment, for example, when text data that is the result of speech recognition is displayed, the terminal device 300 may display the speech style that corresponds to the language model used for the speech recognition.
[0110] Specifically, for example, on the screen 71 shown in FIG. 7, information such as "A language model corresponding to casual speech was used" may be displayed in association with the text data T1 and T2.
[0111] By displaying such information, it is possible to present the estimation result of the speaking style to the user viewing the screen 71.
[0112] In addition, in this embodiment, the speech style estimation unit 233 outputs information including the probability of each of a plurality of speech styles as speech style estimation information, but is not limited to this. For example, the speech style estimation unit 233 may output only the speech style with the highest probability as speech style estimation information.
[0113] (Second embodiment) The second embodiment will be described below with reference to the drawings. The second embodiment differs from the first embodiment in that a language model to be provided to the speech recognition unit 221 is generated using speech style estimation information and a plurality of language models stored in the language model storage unit 223. In the following description of the second embodiment, differences from the first embodiment will be described, and components having similar functional configurations to those of the first embodiment will be assigned the same reference numerals as those used in the description of the first embodiment, and their description will be omitted.
[0114] Fig. 9 is a flowchart illustrating the processing of the information processing apparatus of the second embodiment. The processing from step S901 to step S905 in Fig. 9 is the same as the processing from step S601 to step S605 in Fig. 6, and therefore the description will be omitted.
[0115] Following step S904 in FIG. 9, the information processing device 200 causes the language model providing unit 222 to generate a language model to be provided to the speech recognition unit 221 based on the speech style estimation information output from the speech style estimation unit 233 (step S906).
[0116] The processing of the language model providing unit 222 of this embodiment will be described below.
[0117] The language model providing unit 222 of this embodiment accepts input of speech style estimation information, generates a new language model using the probability of each speech style included in the speech style estimation information and the language model corresponding to each speech style, and provides the new language model to the speech recognition unit 221.
[0118] Specifically, for example, suppose that the language model providing unit 222 receives input of speech style estimation information indicating that the speech style of the speech section is 5% to be a speech style using written language (first speech style), 70% to be a speech style using spoken language (second speech style), and 25% to be a speech style using casual spoken language (third speech style).
[0119] In this case, for example, when text data resulting from speech recognition of words in a processing unit by the speech recognition unit 221 is input to the first language model 241, the probability that word A will appear in the next processing unit is 5%. Similarly, when this text data is input to the second language model 242, the probability that word A will appear in the next processing unit is 10%, and when this text data is input to the third language model 243, the probability that word A will appear in the next processing unit is 20%.
[0120] In this case, the language model providing unit 222 of this embodiment adds the results of multiplying the probability of each speech style included in the speech style estimation information by the appearance probability of word A in the language model corresponding to each speech style. Then, the language model providing unit 222 generates a language model in which the appearance probability of word A in the next processing unit is equal to the calculated value, and provides the language model to the speech recognition unit 221.
[0121] In this case, the language model providing unit 222 determines that the occurrence probability of word A in the next processing unit is 5% probability of being the first speaking style × 5% probability of occurrence in the first language model 241 + 70% probability of being the second speaking style × 10% probability of occurrence in the second language model 242 + 25% probability of being the third speaking style × 20% probability of occurrence in the third language model 243 =12.25% A language model is generated so that the following is provided to the speech recognition unit 221.
[0122] After step S906, information processing apparatus 200 of this embodiment proceeds to step S907. Steps S907 and S908 are similar to steps S607 and S608 in Fig. 6, and therefore a description thereof will be omitted.
[0123] In this embodiment, a language model generated based on the probability of each speech style included in the speech style estimation information and the occurrence probability of words in each processing unit in each language model is used for speech recognition. Therefore, in this embodiment, even if there is an error in the estimation result by the speech style estimation unit 233, the influence of the error can be suppressed.
[0124] Furthermore, in this embodiment, speech recognition is performed using a language model that reflects the occurrence probability of words for each language model corresponding to each speaking style. Therefore, in this embodiment, even when the probabilities of multiple speaking styles included in the speaking style estimation information are close to each other and it is difficult to estimate the speaking style, the accuracy of speech recognition can be improved.
[0125] (Third embodiment) A third embodiment will be described below with reference to the drawings. The third embodiment differs from the first embodiment in that a language model is selected or generated according to settings for the information processing device. In the following description of the third embodiment, differences from the first embodiment will be described, and components having similar functional configurations to those of the first embodiment will be assigned the same reference numerals as those used in the description of the first embodiment, and descriptions thereof will be omitted.
[0126] The information processing device 200 of this embodiment can set the language model providing unit 222 to select a language model stored in the language model storage unit 223 according to the estimation result of the speech style estimation information, or to generate a new language model.
[0127] The language model providing unit 222 of this embodiment selects or generates a language model according to the settings when speech style estimation information is input from the speech style estimation unit 233. Furthermore, if neither language model selection nor language model generation is set, the language model providing unit 222 of this embodiment provides the speech recognition unit 221 with the language model specified by the language model selection signal.
[0128] Fig. 10 is a flowchart illustrating the processing of the information processing apparatus of the third embodiment. The processing from step S1001 to step S1003 in Fig. 10 is the same as the processing from step S601 to step S603 in Fig. 6, and therefore the description thereof will be omitted.
[0129] In the information processing device 200 of this embodiment, when speech style estimation information is input, the language model providing unit 222 determines whether or not a setting for selecting a language model has been made (step S1004).
[0130] In step S1004, if the setting for selecting a language model is made, the information processing device 200 proceeds to steps S604 and S605 in FIG.
[0131] In step S1004, if the setting for selecting a language model has not been made and the setting for generating a language model has been made, the information processing device 200 proceeds to steps S904 and S905 in FIG.
[0132] In step S1005, if the setting to generate a language model is not made, the speech recognition processing unit 220 selects a language model from the multiple language models stored in the language model storage unit 223 in accordance with the language model selection signal (step S1006).
[0133] Next, the information processing device 200 causes the speech acquisition unit 232 to extract speech features from the speech data output from the speech period detection unit 231 (step S1007), and proceeds to step S1008.
[0134] The processing in steps S1008 and S1009 in FIG. 10 is similar to the processing in steps S607 and S608 in FIG. 6, and therefore a description thereof will be omitted.
[0135] In this manner, in this embodiment, it is possible to set whether a language model is selected according to the probability of each speech style included in the speech style estimation information, or whether a language model is newly generated using the speech style estimation information.
[0136] In addition, in this embodiment, for example, a predetermined threshold may be set for the probability of each speech style included in the speech style estimation information, and whether or not a setting for selecting a language model has been made may be determined according to the relationship between the probability of each speech style included in the speech style estimation information and the predetermined threshold.
[0137] Specifically, in this embodiment, if the highest probability value of each speech style included in the speech style estimation information is equal to or greater than a predetermined threshold, it may be determined that a setting to select a language model has been made.
[0138] In other words, in this embodiment, if the highest probability of each speech style included in the speech style estimation information is less than a predetermined threshold, which may be, for example, 50%, a new language model is generated.
[0139] In this embodiment, by doing so, when the probabilities of each speech style included in the speech style estimation information are close to each other, a language model that reflects the appearance probability of words in the language model corresponding to each speech style can be used for speech recognition. Therefore, according to this embodiment, the accuracy of speech recognition can be improved.
[0140] (Fourth embodiment) The fourth embodiment will be described below with reference to the drawings. In the fourth embodiment, an example of a specific usage scene of the information processing system according to the first to third embodiments is shown.
[0141] Fig. 11 is a diagram showing an example of a system configuration of the fourth embodiment, in which any one of the first to third embodiments is used in a remote conference system.
[0142] The teleconference system 100A of this embodiment includes an information processing device 200, a hemispherical imaging device 400, and an electronic whiteboard 500, all of which are connected via a network.
[0143] In this embodiment, the hemispherical imaging device 400 and the electronic whiteboard 500 may be installed in geographically separate locations. Specifically, for example, the hemispherical imaging device 400 may be installed in a conference room of an office located in A city, A prefecture, and the electronic whiteboard 500 may be installed in a conference room of an office located in B city, B prefecture.
[0144] The hemispherical imaging device 400 captures hemispherical image data in a conference room. The hemispherical imaging device 400 also includes a sound collection device to acquire audio data of speeches made in the conference room. The hemispherical imaging device 400 may also include a communication device that transmits video data including the hemispherical image data and audio data to the information processing device 200.
[0145] The electronic whiteboard 500 is an example of a display device that has, for example, a large display with a touch panel, detects coordinates on the board indicated by a user, connects the coordinates, and displays strokes. The electronic whiteboard 500 is also sometimes called an electronic information board or an electronic whiteboard.
[0146] The information processing device 200 of this embodiment includes a voice recognition processing unit 220, and converts speech of each speaker participating in a conference into text data based on video data acquired in a conference room where the hemispherical imaging device 400 is installed, for example, and displays the text data on the electronic whiteboard 500. Note that the electronic whiteboard 500 may display the video data acquired by the hemispherical imaging device 400 together with the text data.
[0147] In this embodiment, by using the information processing device 200 in this manner, for example, in a conference where multiple people with different positions are speaking, speech recognition can be performed using a language model corresponding to the speaking style according to the position of the person speaking.
[0148] Therefore, according to this embodiment, for example, even when only text data that is the result of speech recognition is displayed on the interactive whiteboard 500, the speaking style of each speaker is reflected in the text data. Therefore, in this embodiment, for example, viewers of the interactive whiteboard 500 can understand the speaker's position and human relationships.
[0149] 11, the video data is captured by the hemispherical imaging device 400, but the present invention is not limited to this. The video data may be acquired by a general imaging device, a spherical imaging device, or the like.
[0150] (Fifth embodiment) The fifth embodiment will be described below with reference to the drawings. The fifth embodiment shows an example of a specific usage scene of the information processing systems of the first to third embodiments.
[0151] Fig. 12 is a diagram showing an example of a system configuration of the fifth embodiment, in which any one of the first to third embodiments is used in a counseling support system.
[0152] The counseling support system 100B of this embodiment may be introduced, for example, in a place where counseling is provided by a counselor such as a clinical psychologist. The counseling support system 100B of this embodiment includes an information processing device 200, an imaging device 700, and a terminal device 800, which are connected to each other via a network.
[0153] The imaging device 700 may be set in, for example, a counseling room where counseling is conducted, and acquires video data including a person receiving counseling. In other words, the imaging device 700 acquires video data including moving images and audio of the person receiving counseling. The imaging device 700 then transmits the acquired video data to the information processing device 200. In the following description, the person receiving counseling may be referred to as a client.
[0154] The video data acquired by the imaging device 700 may include an image of the counselor and audio data indicating the counselor's speech.
[0155] The terminal device 800 of this embodiment may be, for example, a terminal device carried by a counselor. The terminal device 800 may be, for example, a tablet-type terminal device having a display.
[0156] When the information processing device 200 acquires video data from the imaging device 700, it estimates the speaking style of the client from the image of the client during counseling and converts the client's voice data into text data. Then, when the information processing device 200 receives a request to play the video data from the terminal device 800, it causes the terminal device 800 to display the text data superimposed on the video data.
[0157] The speech style of a client during counseling is likely to reflect the client's psychological state. For example, if the text data converted from the client's voice data is written, it can be seen that the client is in a tense state. Also, for example, if the text data converted from the client's voice data is spoken, it can be seen that the client is in a relatively relaxed state.
[0158] Furthermore, for example, if the text data converted from the client's voice data changes from spoken to written language, it can be determined that the topic of counseling has changed to one that makes the client nervous.
[0159] In this embodiment, the voice data of the client during counseling is converted into text data that corresponds to the client's speaking style and is presented to the counselor, thereby assisting the counselor in understanding the client's psychological state. Also, in this embodiment, by having the counselor understand the client's psychological state during counseling, the counselor can learn whether appropriate topics were selected, etc.
[0160] In this embodiment, the counseling support system 100B is applied to counseling by a clinical psychologist or the like, but is not limited to this. The counseling support system 100B may also be used, for example, for job-seeking consultations for students or job seekers, or for interviews within an organization such as a company.
[0161] Each function of the above-described embodiments can be realized by one or more processing circuits. Here, the term "processing circuit" in this specification includes a processor programmed to perform each function by software, such as a processor implemented by an electronic circuit, as well as devices such as an ASIC (Application Specific Integrated Circuit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), or a conventional circuit module designed to perform each function described above.
[0162] Additionally, the devices described in the embodiments are merely representative of one of several computing environments for implementing the embodiments disclosed herein.
[0163] In one embodiment, information processing apparatus 200 includes multiple computing devices, such as a server cluster, configured to communicate with each other over any type of communication link, including a network, shared memory, etc., to perform the processes disclosed herein. Similarly, information processing apparatus 200 may include multiple computing devices configured to communicate with each other.
[0164] Furthermore, the information processing device 200 can be configured to share the disclosed processing steps in various combinations. For example, a process executed by the information processing device 200 can be executed by another information processing device. Similarly, a function of the information processing device 200 can be executed by another information processing device. Furthermore, the elements of the information processing device and the other information processing device may be integrated into a single information processing device or may be separated into multiple devices.
[0165] Although the present invention has been described above based on the embodiments, the present invention is not limited to the requirements shown in the above embodiments. These requirements can be changed without departing from the spirit of the present invention, and can be appropriately determined depending on the application form. [Explanation of symbols]
[0166] 100 Voice Recognition System 200 Information processing device 220 Speech recognition processing unit 221 Voice Recognition Unit 222 Language Model Provider 223 Language Model Memory 224 Output Text Generation Unit 230 Input Reception Unit 231 Speech Activity Detection Unit 232 Voice Acquisition Unit 233 Speaking Style Estimation Unit 300 Terminal Device [Prior art documents] [Patent documents]
[0167] [Patent Document 1] Japanese Patent Application Laid-Open No. 2004-333738
Claims
1. An information processing method by a computer, comprising: Accepts input of video data including video of the speaker, When the video data is input, for each of a plurality of speech styles corresponding to a plurality of language models stored in a storage unit, output estimation information indicating the probability that each speech style matches the speech style of the speaker; When a setting is made to select a language model according to the speech style of the speaker, a language model corresponding to the speech style with the highest probability is selected as the language model according to the speech style of the speaker; generating a new language model according to the speech style of the speaker based on the estimated information when a setting is made to generate a new language model according to the speech style of the speaker; providing a language model corresponding to the speech style of the selected speaker or the new language model to a speech recognition unit that performs speech recognition; An information processing method, comprising: displaying on a display device text data resulting from speech recognition of audio data contained in the video data using a language model corresponding to the speech style of the selected speaker or the new language model.
2. An information processing method by a computer, comprising: Accepts input of video data including video of the speaker, When the video data is input, for each of a plurality of speech styles corresponding to a plurality of language models stored in a storage unit, output estimation information indicating the probability that each speech style matches the speech style of the speaker; selecting a language model according to the manner of speech of the speaker based on the estimated information when a setting is made to select a language model according to the manner of speech of the speaker; generating a new language model using the plurality of language models and a probability for each speech mode included in the estimated information when a setting is made to generate a new language model according to the speech mode of the speaker; providing a language model corresponding to the speech style of the selected speaker or the new language model to a speech recognition unit that performs speech recognition; An information processing method, comprising: displaying on a display device text data resulting from speech recognition of audio data contained in the video data using a language model corresponding to the speech style of the selected speaker or the new language model.
3. The computer When the highest probability value among the probabilities that each of the plurality of speech modes matches the speech mode of the speaker is less than a predetermined threshold value, The information processing method according to claim 1 or 2, further comprising determining that a setting is set to generate a new language model according to the speaker's speech style.
4. The computer The information processing method according to claim 1 , wherein the text data is displayed on the display device for each utterance of the speaker.
5. The computer 5. The information processing method according to claim 4, further comprising the step of displaying information indicating a speech style corresponding to the language model provided to the speech recognition unit on the display device together with the text data.
6. Accepting input of video data including a video of a speaker, When the video data is input, for each of a plurality of speech styles corresponding to a plurality of language models stored in a storage unit, output estimation information indicating the probability that each speech style matches the speech style of the speaker; When a setting is made to select a language model according to the speech style of the speaker, a language model corresponding to the speech style with the highest probability is selected from among a plurality of language models stored in a storage unit as the language model according to the speech style of the speaker; generating a new language model according to the speech style of the speaker based on the estimated information when a setting is made to generate a new language model according to the speech style of the speaker; providing a language model corresponding to the speech style of the selected speaker or the new language model to a speech recognition unit that performs speech recognition; A program that causes a computer to execute a process of displaying on a display device text data resulting from speech recognition of the audio data contained in the video data using a language model that corresponds to the speech style of the selected speaker, or the new language model.
7. Accepting input of video data including a video of a speaker, When the video data is input, for each of a plurality of speech styles corresponding to a plurality of language models stored in a storage unit, output estimation information indicating the probability that each speech style matches the speech style of the speaker; selecting a language model according to the manner of speech of the speaker based on the estimated information when a setting is made to select a language model according to the manner of speech of the speaker; When a setting is made to generate a new language model according to the manner of speech of the speaker, generating a new language model using the probability for each mode of speech included in the estimation information and the plurality of language models; providing a language model corresponding to the speech style of the selected speaker or the new language model to a speech recognition unit that performs speech recognition; A program that causes a computer to execute a process of displaying on a display device text data resulting from speech recognition of the audio data contained in the video data using a language model that corresponds to the speech style of the selected speaker, or the new language model.
8. an input receiving unit that receives input of video data including a video of a speaker; an estimation unit that, when the video data is input, outputs estimation information indicating the probability that each speech style corresponds to a speech style of the speaker for each of a plurality of language models stored in a storage unit; a model providing unit that, when a setting is made to select a language model according to the speaker's speech style, selects a language model corresponding to the speech style with the highest probability as a language model according to the speaker's speech style, and, when a setting is made to generate a new language model according to the speaker's speech style, generates a new language model according to the speaker's speech style based on the estimated information; an output text generation unit that provides a language model corresponding to the speech style of the selected speaker or the new language model to a speech recognition unit that performs speech recognition, and displays text data resulting from speech recognition of the audio data included in the video data using the language model corresponding to the speech style of the selected speaker or the new language model on a display device.
9. an input receiving unit that receives input of video data including a video of a speaker; an estimation unit that, when the video data is input, outputs estimation information indicating the probability that each speech style corresponds to a speech style of the speaker for each of a plurality of language models stored in a storage unit; a model providing unit that selects a language model according to the speaker's speech style based on the estimated information when a setting is made to select a language model according to the speaker's speech style, and that generates a new language model according to the speaker's speech style using the probability for each speech style included in the estimated information and the plurality of language models when a setting is made to generate a new language model according to the speaker's speech style; an output text generation unit that provides a language model corresponding to the speech style of the selected speaker or the new language model to a speech recognition unit that performs speech recognition, and displays text data resulting from speech recognition of the audio data included in the video data using the language model corresponding to the speech style of the selected speaker or the new language model on a display device.
10. 10. An information processing system comprising: the information processing device according to claim 8; and an imaging device that captures the video data.
11. The information processing system according to claim 10 , wherein the imaging device is a spherical imaging device or a semi-spherical imaging device.
Citation Information
Patent Citations
Device and method for voice recognition using video information
JP2004333738A
Voice synthesizer, voice quality generating device, and program
JP2005257747A
Style detecting device for speech, its method and its program
JP2007219286A
Speech recognizer, speech recognition method, and program for speech recognition
JP2008026721A
Utterance mode detection device, and utterance mode detection method
JP2015087557A