Voice information processing device and voice information processing method
By detecting speech interruptions in a speech information processing device and automatically outputting the spoken speech, the problem of users having difficulty remembering the content after the interruption is solved, and the effect of simplifying the operation is achieved.
Patent Information
- Application Number
- CN202010999526.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-22
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2040-09-22
Smart Images

Figure CN114255757B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a voice information processing device and a voice information processing method, and is particularly suitable for converting a user's spoken voice into text. Background Art
[0002] Conventional voice information processing devices have been developed that take a user's spoken voice, convert it into text, and then send the text as a message on a chat application or as an email. Using these devices, users can send a text message of their desired content to another party simply by speaking, without having to use their hands.
[0003] Furthermore, Patent Document 1 describes a technology that, when an interruption occurs while a phone number is being input on a telephone, temporarily saves the data processed so far to nonvolatile memory and restores the data after the interruption is completed. Furthermore, Patent Document 2 describes a technology that, in a digital broadcast receiving system, when a signal loss occurs while recording a received signal to a recording / reproduction device, generates a loss information signal corresponding to the loss time and records it to the recording / reproduction device. The generated and recorded loss information signal then outputs video or audio for the lost portion.
[0004] Prior art literature
[0005] Patent Literature
[0006] [Patent Document 1] Japanese Patent Application Laid-Open No. 2007-151188
[0007] [Patent Document 2] Japanese Patent Application Laid-Open No. 2003-319280 Summary of the Invention
[0008] Problems to be solved by the invention
[0009] When a user is using the aforementioned conventional speech information processing device to convert their speech into text, they may sometimes interrupt their speech mid-sentence for some reason. After the interruption, when the reason for the interruption is resolved and they resume speaking for text conversion, the user often does not accurately remember the content of the sentence spoken so far and cannot accurately determine where to start speaking and what content to say. In such cases, the user is forced to temporarily cancel the sentence they have spoken and converted into text and restart speaking from the beginning, which is a cumbersome operation for the user.
[0010] The present invention has been made to solve such a problem, and an object of the present invention is to enable the user to complete text conversion of a desired sentence without performing complicated operations when an utterance for text conversion is interrupted during the utterance.
[0011] Means for solving problems
[0012] In order to solve the above-mentioned problems, in the present invention, the user's speech is sequentially textualized during the period when the user's speech to be textualized is accepted, that is, the speech acceptance period, and when it can be considered that the user's speech is interrupted, the speech content that the user has already said during the speech acceptance period is automatically output through voice.
[0013] Effects of the Invention
[0014] According to the present invention constructed as described above, if a situation occurs in which a user's speech can be considered interrupted, the content of the user's speech so far is automatically output as voice. Therefore, by listening to the output voice, the user can understand the content of the sentence they have spoken so far and recognize where they are and where they should resume speaking. This allows the user to resume speaking from the middle of a sentence without canceling the already texted sentence. Therefore, according to the present invention, if a speech is interrupted midway during texting, the user can complete the texting of the desired sentence without performing complicated operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 This is a block diagram showing a configuration example of a speech information processing device according to the first embodiment of the present invention.
[0016] Figure 2 This is a flowchart showing an example of the operation of the speech information processing device according to the first embodiment of the present invention.
[0017] Figure 3 This is a block diagram showing an example of the functional configuration of a speech information processing device according to the second embodiment of the present invention.
[0018] Figure 4 This is a flowchart showing an example of the operation of the speech information processing device according to the second embodiment of the present invention.
[0019] Description of reference numerals:
[0020] 1.1A Voice Information Processing Device
[0021] 10 Voice output unit
[0022] 11 Voice input unit
[0023] 12.12A Voice Information Processing Department
[0024] 14 cameras. DETAILED DESCRIPTION
[0025] <First embodiment>
[0026] Hereinafter, a first embodiment of the present invention will be described with reference to the drawings. Figure 1 It is a block diagram showing an example of the functional structure of the voice information processing device 1. The voice information processing device 1 involved in this embodiment is a device installed in a vehicle. The voice information processing device 1 has the function of providing users with a text chat environment for multiple people to send and receive text messages. In particular, the voice information processing device 1 involved in this embodiment has the following functions: during text chat, the voice spoken by the occupant using the device (hereinafter referred to as the "user") during the voice acceptance period (described later) is input, the sentence represented by the input voice is textualized, and the textualized sentence is sent as a message. By utilizing this function, the user can create and send a message to the other party in the text chat without manual input. Hereinafter, the vehicle equipped with the voice information processing device 1 will be referred to as the "this vehicle."
[0027] like Figure 1 As shown, a microphone 2 and a speaker 3 are connected to the voice information processing device 1. The microphone 2 is located in a position where it can receive the speech of a user riding in the vehicle. The microphone 2 receives the speech and outputs a speech signal representing the received speech. The speaker 3 is located inside the vehicle, receives the speech signal, and plays speech based on the received speech signal.
[0028] like Figure 1 As shown, the voice information processing device 1 includes a voice output unit 10, a voice input unit 11, and a voice information processing unit 12 as functional components. Each of the above functional modules 10 to 12 can be composed of hardware, DSP (Digital Signal Processor), or software. For example, when composed of software, each of the above functional modules 10 to 12 is actually composed of a computer's CPU, RAM, ROM, etc., and is implemented through program operations stored in a recording medium such as RAM, ROM, hard disk, or semiconductor memory. In the functional structure, the voice output unit 10 inputs a voice signal, drives the speaker 3 based on the input voice signal, and causes the speaker 3 to play voice based on the voice signal.
[0029] The following describes the operation of the voice information processing device 1 when the user's spoken words are converted into text and sent as a message, while in chat mode. Chat mode allows the user to engage in text chats with a desired party (or multiple parties) using the voice information processing device 1. The user switches to chat mode by operating the operating mechanism (which may be a touch panel) of the voice information processing device 1 or by voice command. At this time, the user also appropriately configures settings required for sending messages, such as specifying the party with whom the text chat is to be conducted.
[0030] To convert a desired sentence into text using the voice information processing device 1 and send it as a message, the user speaks a predetermined message start word consisting of fixed words, then speaks the sentence they wish to convert into text, and then speaks a predetermined message end word consisting of fixed words. An example of a message start word is a word like "message start," and an example of a message end word is a word like "message end." In other words, in this embodiment, the period from the utterance of the message start word until the utterance of the message end word is the period during which the speech of the sentence to be converted into text is accepted. This period corresponds to the "speech acceptance period."
[0031] Furthermore, in this embodiment, after the user says the message end word, if they wish to send the sentence spoken during the voice acceptance period as a message, they say the message send word. An example of the message send word is the phrase "message send." In response to the user saying the message send word, the sentence spoken by the user is sent to the other party.
[0032] When operating in chat mode, the voice input unit 11 inputs the voice signal output by the microphone 2, performs analog / digital conversion processing including sampling, quantization, and encoding on the voice signal, and performs other signal processing to generate voice data (hereinafter referred to as "input voice data"), which is then buffered in the buffer 13. The buffer 13 is a storage area formed in a working area such as RAM. The input voice data is data of a voice waveform obtained by sampling at a predetermined sampling period (for example, 16 kHz).
[0033] The voice information processing unit 12 analyzes the input voice data buffered in the buffer 13 at any time and monitors whether the voice waveform of the message start word appears in the input voice data. In this embodiment, the voice pattern of the message start word (= the pattern of the voice waveform when the message start word is spoken) is pre-registered. Multiple voice patterns can also be registered. The voice information processing unit 12 compares the voice waveform of the input voice data with the voice pattern related to the message start word at any time, and calculates the similarity using a prescribed method. When the similarity becomes greater than a certain level, it is determined that the waveform of the message start word appears in the input voice data. In the following, the voice information processing unit 12 detecting the waveform of the message start word in the input voice data is appropriately expressed as "the voice information processing unit 12 detects the message start word."
[0034] If the message start word is detected, the voice information processing unit 12 performs voice recognition at any time using the input voice data buffered in the buffer 13 as an object, textualizes the sentences recorded in the input voice data, and describes them as text in the sentence data stored in the storage unit not shown. This processing is referred to as "textualization processing" below. In addition, the textualization of the input voice data is appropriately performed by implementing morpheme analysis, syntactic structure analysis, semantic structure analysis, etc. based on existing technologies related to natural language processing. Artificial intelligence technology can also be used in some technologies. In addition, it can also be configured so that the voice information processing unit 12 cooperates with an external device to perform textualization processing. For example, it can also be configured so that textualization processing is performed in collaboration with a cloud server that provides a service for textualization of voice data.
[0035] In parallel with the text conversion process, the voice information processing unit 12 continuously analyzes the input voice data buffered in the buffer 13 and monitors whether the voice waveform of the message end word appears in the input voice data. This monitoring is performed based on the pre-registered voice pattern of the message end word using the same method as the monitoring of the voice waveform of the message start word described above.
[0036] If the voice waveform of the message end word is detected in the input voice data, the voice information processing unit 12 terminates the text conversion process. Thereafter, the voice information processing unit 12 continuously analyzes the input voice data buffered in the buffer 13 to monitor whether the voice waveform of the message start word appears in the input voice data. This monitoring is performed based on the pre-registered voice pattern of the message start word using the same method as the monitoring of the voice waveform of the message start word described above.
[0037] When the voice waveform of the message sending word is detected in the input voice data, the voice information processing unit 12 sends a message to a predetermined server via the network N in accordance with a protocol for the text described in the sentence data.
[0038] Furthermore, in parallel with the text conversion process, the voice information processing unit 12 continuously analyzes the input voice data buffered in the buffer 13, monitoring whether the voice waveform of the cancel word appears in the input voice data. The cancel word is, for example, the phrase "message canceled." This monitoring is performed based on the pre-registered voice pattern of the cancel word, using the same method as monitoring the voice waveform of the message start word mentioned above. If the voice waveform of the cancel word is detected in the input voice data, the voice information processing unit 12 cancels the text conversion process and deletes the text recorded in the sentence data to that point. Thereafter, the voice information processing unit 12 resumes monitoring the voice waveform of the message start word in the input voice data.
[0039] Furthermore, the voice information processing unit 12 performs the following processing during text conversion, that is, from the time the message start word is detected until the text conversion process is completed or canceled. Specifically, it determines whether the user has not spoken for a specified period of time or longer. If the user has not spoken for a specified period of time or longer, this means the following. For example, suppose the user has spoken "Hello." In this case, it means that after saying "Hello," the user has not spoken for a specified period of time or longer.
[0040] The voice information processing unit 12 analyzes the input voice data and determines that the user has not spoken for a period of time exceeding the predetermined time if the sound pressure value of the voice waveform exceeds a first threshold value (a threshold value for determining speech), then falls below a second threshold value (a threshold value for determining non-speech, which may be the same value as the first threshold value), and the state below the second threshold value continues for a predetermined time period or longer. However, any method of determination may be used.
[0041] When it is detected that the user has not spoken for a period exceeding a predetermined time, the voice information processing unit 12 performs the following processing. Specifically, the voice information processing unit 12 causes the voice output unit 10 to output, via voice, the sentence represented by the text recorded so far in the sentence data (= the text already generated in the texting process). Hereinafter, the voice output by the voice output unit 10 in this manner will be referred to as "texted voice," and the voice information processing unit 12 causing the voice output unit 10 to output texted voice will be simply expressed as "the voice information processing unit 12 outputs texted voice." Texted voice corresponds to the "voice corresponding to the content of the speech already spoken by the user during the voice acceptance period" in the technical proposal.
[0042] The processing of the voice information processing unit 12 will be described in detail. The voice information processing unit 12 generates voice data for outputting the sentence represented by the text recorded in the sentence data as voice. The voice data is generated appropriately using existing technologies such as speech synthesis technology. The voice information processing unit 12 then outputs a voice signal based on the voice data to the voice output unit 10, causing the voice based on the voice data to be played from the speaker 3.
[0043] The speech information processing unit 12 then continues texting. If a user utters a sentence for texting, it converts the utterance into text. If the user utters a sentence and then remains silent for a predetermined period of time or longer, it outputs the texted speech again. The speech information processing unit 12 also concurrently detects message end and cancel words.
[0044] According to the above configuration, the voice information processing device 1 operates, for example, in the following manner. For example, it is assumed that the user says the following sentence: "I am driving there now. I just passed point A. The scheduled arrival time is 13:00. The roads are congested, so I may be late. Contact me when I am almost there." It is also assumed that after the user says the message start words, at the moment of saying the sentence "I am driving there now. I just passed point A.", the user interrupts speaking for some reason. An example of the reason is that the vehicle approaches an intersection or starts to stop, so it is necessary to concentrate on driving, or it is necessary to pay a fee at a toll booth on the road. In addition, through the text processing of the voice information processing unit 12, the part spoken by the user is textualized, and the text is recorded in the sentence data.
[0045] In this case, if a predetermined time or longer passes without further utterance after the sentence "...just passed point A." is spoken, the speech information processing unit 12 of the speech information processing device 1 according to this embodiment automatically causes the speech output unit 10 to output speech corresponding to the generated text. In this example, the sentence "Currently driving to point A. Just passed point A." is output via speech.
[0046] As a result of the above processing, the following effects are achieved. That is, after interrupting the speech for texting, when resuming the speech for texting, the user needs to start speaking again from the sentence after the part just spoken. However, the user may not accurately remember the content of the sentence spoken so far, and may not know exactly where to start speaking and what kind of sentence to say. In this example, that is to say, although the user should start speaking from the place where "the scheduled time is...", he or she may not know exactly where to say and where to start speaking. In such a case, although it is also possible to cancel the voice input so far by saying a cancellation word, the texting of the desired sentence and the sending as a message can be performed from the beginning, but such operations are cumbersome for the user.
[0047] On the other hand, this embodiment has the following advantages. Specifically, if a user continues not speaking for a considerable period of time, it can be considered that the user's speech has been interrupted. This is because, generally, when a user speaks a series of sentences that they wish to send as a message for the purpose of texting, they will not stop speaking for an unnecessarily long period of time in the middle of their speech.
[0048] Furthermore, according to the speech information processing device 1 involved in this embodiment, when a situation occurs in which a user's speech can be considered to have been interrupted, the sentence that the user has already texted is automatically output as speech. Therefore, by listening to the outputted speech, the user can understand the content of the sentence that they have spoken and texted so far. As a result, the user can restart speaking from the sentence in progress without canceling the already texted sentence. Therefore, according to this embodiment, if a speech intended for texting is interrupted in the middle of speech, the user can complete the texting of the desired sentence without performing complicated operations.
[0049] Furthermore, the user may completely forget that they have spoken for text after interrupting their speech for texting. In such a case, according to this embodiment, the textualized sentence based on the user's speech is automatically output as voice, thereby prompting the user to realize that they are in the middle of speaking (of course, they can also recognize the content of the sentence they have already spoken).
[0050] Next, a method for processing speech information performed by the speech information processing device 1 will be described using a flowchart. Figure 2 The flowchart shows the operation of the voice information processing unit 12 when the chat mode is turned on. Figure 2As shown, the voice information processing unit 12 analyzes the input voice data buffered in the buffer 13 at any time and monitors whether the voice waveform of the message start word appears in the input voice data (step SA1). If it appears (step SA1: Yes), the voice information processing unit 12 starts text conversion (step SA2).
[0051] Next, the voice information processing unit 12 monitors whether the voice waveform of the message end word appears in the input voice data (step SA3), whether the voice waveform of the cancel word appears in the input voice data (step SA4), and whether the period of no speech has exceeded a predetermined time (step SA5). If the voice waveform of the message end word appears in the input voice data (step SA3: Yes), the voice information processing unit 12 ends the text conversion process (step SA6) and monitors whether the voice waveform of the message send word appears in the input voice data (step SA7). If the voice waveform of the message send word appears in the input voice data (step SA7: Yes), the voice information processing unit 12 sends the message to the text described in the sentence data (step SA8).
[0052] When the voice waveform of the cancellation word appears in the input voice data (step SA4 : Yes), the voice information processing unit 12 cancels the text conversion process (step SA9 ) and returns the processing procedure to step SA1 .
[0053] If the period of no speech exceeds the predetermined time (step SA5: YES), the speech information processing unit 12 causes the speech output unit 10 to output the sentence represented by the text recorded so far in the sentence data (= the text generated in the text conversion process) by speech (step SA10). The speech information processing unit 12 then returns the processing procedure to step SA3.
[0054] <Modification of the first embodiment>
[0055] In the first embodiment described above, if a period of no speech during the speech acceptance period exceeds a predetermined time, the speech information processing unit 12 causes the speech output unit 10 to output the sentence represented by the text generated so far as speech (text-converted speech). In this regard, the speech information processing unit 12 may also cause the speech output unit 10 to output speech based on speech data stored in the buffer 13 (= the recorded speech of the user's speech) instead of the text-converted speech. In this configuration, the speech output instead of the text-converted speech (= the recorded speech of the user's speech) corresponds to the "speech corresponding to the content of the speech already uttered by the user during the speech acceptance period" in the technical claims.
[0056] In this case, for example, the voice information processing unit 12 cuts out voice data corresponding to the portion spoken by the user during the voice acceptance period from the input voice data stored in the buffer 13, and outputs a voice signal based on the cut-out voice data to the voice output unit 10. This modification is also applicable to the second embodiment (including modifications of the second embodiment) described later.
[0057] <Second embodiment>
[0058] Next, a second embodiment will be described. Figure 3 This is a block diagram illustrating an example of the functional configuration of the voice information processing device 1A according to this embodiment. In the following description of the second embodiment, identical elements to those in the first embodiment are denoted by identical reference numerals, and detailed descriptions thereof are omitted. Furthermore, in this embodiment, for ease of explanation, the user using the voice information processing device 1A is assumed to be the driver. However, this is for ease of explanation, and passengers other than the driver may also be users of the voice information processing device 1A.
[0059] By comparison Figure 1 and Figure 3 As can be seen, the voice information processing device 1A according to this embodiment includes a voice information processing unit 12A in place of the voice information processing unit 12 according to the first embodiment. Furthermore, a camera 14 is connected to the voice information processing device 1A according to this embodiment. Camera 14 is positioned so that it can capture an image of the user's upper body, including their face, when the user is seated in the driver's seat. Camera 14 captures images at a predetermined interval and outputs image data based on the captured images to voice information processing unit 12A.
[0060] The voice information processing unit 12 according to the first embodiment causes the voice output unit 10 to output voice related to a texted sentence (texted voice) when the user does not speak for a predetermined time or longer during the voice reception period. On the other hand, the voice information processing unit 12A according to this embodiment outputs texted voice when the user does not speak for a predetermined time or longer after the user moves their face so as to view the outside through the side window.
[0061] To elaborate, upon detecting the message start word, the voice information processing unit 12A uses existing recognition technology to identify images of the upper body of a person (= the upper body of the user) in the photographic image data input from the camera 14 at a predetermined period. The unit then continuously analyzes the upper body image to monitor whether the user is performing an action while viewing the outside through the side window (the action of turning the face toward the side window and looking out). This monitoring is performed based on existing action recognition technology. Obviously, this monitoring can also be performed using a model learned through deep learning or other machine learning methods.
[0062] Then, when the voice information processing unit 12A detects that the user has performed an action of observing the outside through the side window, and if the user remains silent for a predetermined time or longer after the detection, the voice information processing unit 12A automatically outputs text-converted voice.
[0063] This embodiment has the following effects. Specifically, when a situation occurs in which the user's speech can be considered interrupted, the textualized sentence that the user has previously used is automatically output as speech, thereby achieving the same effects as the first embodiment. Furthermore, if the user observes the outside through the side window and does not speak for textualization for a predetermined period of time, it can be more reliably estimated that the driver interrupted the speech for textualization by looking at the scenery outside, compared to a situation in which the period of inaction lasts for a predetermined period of time. Therefore, according to this embodiment, compared to the first embodiment, textualized speech can be output in situations where speech interruption is more reliably estimated.
[0064] Next, regarding the operation of the speech information processing device 1A according to this embodiment, Figure 4 The flowchart of Figure 4 In the flowchart of Figure 2 The same processing as in the flowchart is assigned the same step number and its description is omitted. Figure 4 As shown, the speech information processing device 1A according to this embodiment performs the same Figure 2The processing of step SA5 is different from that of step SB1. That is, in step SB1, the voice information processing unit 12A monitors whether the period of time during which the driver's face moves so as to observe the outside through the side window and then does not speak becomes longer than a predetermined time. Then, if the period of time during which the driver's face moves so as to observe the outside through the side window and then does not speak becomes longer than a predetermined time in step SB1 (step SB1: yes), the processing sequence proceeds to step SA1. In addition, in the second embodiment, the voice information processing unit 12A can also be configured to output a recorded voice of the user's speech instead of the text-converted voice, as described in the modification of the first embodiment.
[0065] <First Modification of Second Embodiment>
[0066] Next, the first variant of the second embodiment is described. In the second embodiment, the voice information processing unit 12A outputs text-based voice when the user (driver) moves his / her face in a manner of observing the outside through the side window and the user does not speak for a predetermined time. In this regard, the voice information processing unit 12A involved in this variant performs the following processing. In addition, this variant is based on the premise that a car navigation device is provided in this vehicle. That is, the voice information processing unit 12A involved in this variant outputs text-based voice based on the input from the camera 14 when the user moves his / her face in a manner of observing the display screen of the car navigation device and the user does not speak for a predetermined time.
[0067] Citation Figure 4 The flowchart of FIG. 1 illustrates the operation of the voice information processing device 1A according to this variation. In step SB1, the voice information processing unit 12A monitors whether the user does not speak for a predetermined time after the user's face moves so as to view the display screen of the car navigation device.
[0068] If the user does not speak for texting for a predetermined period of time after viewing the display screen of the car navigation device, it can be reliably estimated that the driver interrupted the speech for texting due to viewing the display screen.
[0069] In addition, in the second embodiment and this variation, an example of a configuration is described in which the voice information processing unit 12A monitors whether the user's face moves in a specified manner based on the photographic results of the camera 14, and if the user does not speak for a specified period of time after moving in the specified manner, the voice output unit 10 outputs the voice associated with the generated text. However, this configuration example is not limited to the illustrated case. As an example, the configuration may also monitor whether the user's face moves in a manner that allows the user to observe other passengers other than the user and perform corresponding processing. Alternatively, the configuration may monitor whether the user's face moves in a manner that allows the user to observe components provided on the vehicle, such as rearview mirrors or side mirrors, and perform corresponding processing. In addition, in the first variation of the second embodiment, the voice information processing unit 12A may also be configured to output a recorded voice of the user's speech instead of the texted voice, as described in the variation of the first embodiment.
[0070] <Second Modification of Second Embodiment>
[0071] Next, a second variant of the second embodiment will be described. The voice information processing unit 12A involved in this variant monitors whether the user's face shows an expression of concentration on driving based on the photographic results of the camera 14, and outputs textualized voice if the user does not speak for a predetermined period of time after showing this expression. This monitoring is performed based on existing facial expression recognition technology. Obviously, this monitoring can also be performed using a model learned through deep learning or other machine learning methods. In addition, it is assumed that the driver is concentrating on driving and this is reflected in his expression just before the vehicle enters an intersection, while entering an intersection, while parked in a parking lot, or while driving on a congested road.
[0072] Citation Figure 4 The flowchart of FIG. 1 illustrates the operation of the voice information processing device 1A according to this modification. In step SB1 , the voice information processing unit 12A monitors whether the user does not speak for a predetermined time after the user's face becomes focused on driving.
[0073] If the user does not speak for texting for a predetermined period of time after adopting an expression indicating that the user is concentrating on driving, it can be reliably estimated that the driver interrupted the speech for texting due to being concentrating on driving. Therefore, the configuration of this variation can achieve the same effects as the second embodiment.
[0074] In addition, in this modified example, an example of the following structure is described: the voice information processing unit 12A monitors whether the user's face has a specified expression based on the photographic result of the camera 14, and after the specified expression is formed, if the user does not speak for a specified time, the voice output unit 10 outputs the voice related to the generated text. However, this example of the structure is not limited to the illustrated case. As an example, it can also be configured to monitor whether the user's face has an expression of surprise and perform corresponding processing. In addition, in the second modified example of the second embodiment, the voice information processing unit 12A can also be configured to output the recorded voice of the user's speech instead of the textualized voice, as already described in the modified example of the first embodiment.
[0075] <Third Modification of Second Embodiment>
[0076] Next, we will describe a third variation of the second embodiment. In this variation, the voice information processing unit 12A monitors whether the user has started yawning based on the image capture results from the camera 14. Upon detecting the start of a yawn, the voice information processing unit 12A monitors whether the yawn has ended and, upon detecting that the yawn has ended, outputs text-converted speech. Furthermore, yawn onset / end detection is performed using existing image recognition technology. This monitoring can also be performed using a model learned through deep learning or other machine learning methods.
[0077] References Figure 4 The operation of the speech information processing device 1A according to the present modification will be described with reference to the flowchart. In step SB1, the speech information processing unit 12A monitors whether the user has finished yawning after starting to yawn.
[0078] While the user is yawning, it can be considered that the user has interrupted the speech for text conversion due to this situation. Therefore, according to the configuration of this variation, the same effect as the second embodiment can be achieved. Furthermore, in the third variation of the second embodiment, the voice information processing unit 12A can also be configured to output the recorded voice of the user's speech instead of the text converted speech, as described in the variation of the first embodiment.
[0079] <Fourth Modification of Second Embodiment>
[0080] Next, the fourth variant of the second embodiment will be described. The voice information processing unit 12A involved in this variant monitors whether the user has started a phone call based on the photographic results of the camera 14. When the voice information processing unit 12A detects that the user has started a phone call, it monitors whether the call has ended, and outputs text-converted voice when it detects that the call has ended. Furthermore, the voice information processing unit 12A stops texting the voice data based on the voice input by the voice input unit 11 while the user is on the phone (= the period from when the start of the call is detected until when the end of the call is detected). It is assumed that the user is talking on the phone using his or her own mobile phone in the car.
[0081] In addition, the detection of the start / end of a user's call is performed based on existing image recognition technology. This monitoring can obviously also be performed using a model learned through deep learning or other machine learning methods. In addition, in this modified example, the voice information processing unit 12A detects the start / end of a call based on the photography results of the camera 14, but the method for performing this detection is not limited to the exemplified method. As an example, the voice information processing device 1A can be connected to a mobile phone in a communicative manner, and a predetermined signal is sent from the mobile phone to the voice information processing device 1A when a call starts and ends, and the voice information processing unit 12A detects the start / end of a call based on the predetermined signal.
[0082] Citation Figure 4 The operation of the voice information processing device 1A according to this modification will be described with reference to the flowchart. In step SB1, the voice information processing unit 12A monitors whether the user has started a telephone call and the call has ended.
[0083] If the user is currently on a phone call, it can be considered that the user has interrupted the speech intended for text conversion due to this situation. Therefore, according to this variation, the same effect as the second embodiment can be achieved. Furthermore, the speech input by the speech input unit 11 during the call is not speech intended for text conversion, but speech used for the phone call and should not be subject to text conversion. In this way, according to this variation, speech that should not be subject to text conversion can be prevented from being texted.
[0084] Furthermore, in the fourth variant, the voice information processing unit 12A may be configured to perform the following processing when causing the voice output unit 10 to output text-converted voice. Specifically, the voice information processing unit 12A may be configured to cause the voice output unit 10 to output, as voice, a sentence indicating that text conversion of the voice input by the voice input unit 11 has been stopped while the user is currently on a call, in response to the text-converted voice (the voice corresponding to the utterance). For example, the voice information processing unit 12A may perform the following processing. First, the voice information processing unit 12A outputs the text-converted voice. Next, the voice information processing unit 12A outputs, as voice, a sentence containing the following: "Voice during a telephone call is not text-converted. Subsequent content can be input." The voice data serving as the source of this voice is prepared in advance. The illustrated processing is merely an example; for example, the voice information processing unit 12A may be configured to output the text-converted voice after first outputting the sentence indicating that text conversion of the voice input by the voice input unit 11 has been stopped while the user is currently on a call.
[0085] In the fourth modification, the voice information processing unit 12A may be configured to output a recorded voice of the user's speech instead of the text-converted voice, as described in the modification of the first embodiment.
[0086] Furthermore, in the third and fourth modified examples, an example configuration is described in which the voice information processing unit 12A detects a predetermined state in which the user is unable to speak for text conversion, and then outputs text-converted speech when the predetermined state is resolved. However, this configuration example is not limited to the example. For example, the voice information processing unit 12A may also be configured to detect when the user starts and ends eating, and output text-converted speech when the user ends eating.
[0087] While the embodiments of the present invention (including variations) have been described above, each of the embodiments is merely an example of a specific embodiment of the present invention and is not to be construed as limiting the technical scope of the present invention. In other words, the present invention can be implemented in various forms without departing from its main purpose or essential features.
[0088] For example, in the first embodiment described above, the voice information processing unit 12 transmits text as a message in a text chat. However, text transmission is not limited to the methods exemplified in the various embodiments. For example, text transmission can also be performed by email. Furthermore, text transmission does not only mean sending to a specific party but also broadly encompasses the concept of transmitting text to external devices, such as sending text to a server or a specific host device. For example, sending text sentences to a message posting website or forum website in accordance with a protocol is also included in text transmission. The same applies to the second embodiment.
[0089] In the first embodiment, the speech information processing device 1 is installed in a vehicle. However, the speech information processing device 1 does not necessarily need to be installed in a vehicle. The same applies to the second embodiment. In other words, the present invention is widely applicable to speech information processing devices that convert user speech into text.
[0090] Furthermore, in each of the above embodiments, the voice acceptance period begins when the user speaks the message start word. Alternatively, the voice acceptance period may begin when the user performs a predetermined operation on a touch screen or other input mechanism, or when the user performs a predetermined gesture in a configuration capable of detecting gestures. This also applies to the message end word, message send word, and cancel word.
Claims
1. A voice information processing device, characterized in that: have: Voice input unit, inputs voice; a voice output unit for outputting voice; and The voice information processing unit converts the voice input by the voice input unit into text during the voice acceptance period, wherein the voice acceptance period is the period from when the user accepts the spoken voice to be converted into text, and is the period from when the user finishes speaking the message start word or when the user performs a prescribed operation on the input device or performs a prescribed gesture, until when the user starts speaking the message end word or when the user performs a prescribed operation on the input device or performs a prescribed gesture. The voice information processing unit sequentially converts the user's speech into text during the voice acceptance period, and when it can be considered that the user's speech is interrupted, the voice output unit automatically outputs the speech content that the user has spoken during the voice acceptance period by voice. The voice information processing unit causes the voice output unit to output voice corresponding to the content of the speech uttered by the user during the voice acceptance period when a period during which the user has not spoken a word during the voice acceptance period exceeds a predetermined time.
2. The speech information processing device according to claim 1, wherein The voice information processing device is connected to a camera for photographing the user's face. The voice information processing unit monitors whether the user's face moves in a specified manner based on the camera's photographic results, and after moving in the specified manner, if the user does not speak for more than the specified time, causes the voice output unit to output voice corresponding to the speech content.
3. The speech information processing device according to claim 2, wherein The voice information processing device is provided in a vehicle. The prescribed manner is a manner in which the user's face views the outside through the side window.
4. The speech information processing device according to claim 2, wherein The voice information processing device is installed in a vehicle equipped with a car navigation system. The predetermined manner is the manner in which the user's face appears on the display screen of the car navigation device.
5. The speech information processing device according to claim 1, wherein The voice information processing device is connected to a camera for photographing the user's face. The voice information processing unit monitors whether the user's face has a specified expression based on the camera's photographic results. After the user has the specified expression, if the user does not speak for more than the specified time, the voice output unit outputs voice corresponding to the speech content.
6. The speech information processing device according to claim 5, wherein The voice information processing device is provided in a vehicle. The prescribed expression is an expression in which the user's face is focused on driving.
7. The speech information processing device according to claim 1, wherein The voice information processing unit detects a predetermined state in which the user cannot speak for text conversion, and causes the voice output unit to output voice corresponding to the speech content when the predetermined state is released after the predetermined state is reached.
8. The speech information processing device according to claim 7, wherein: The voice information processing device is connected to a camera for photographing the user's face. The voice information processing unit detects that the user starts yawning based on a result of photography by the camera, and causes the voice output unit to output voice corresponding to the utterance when the yawning ends after the user starts yawning.
9. The speech information processing device according to claim 7, wherein: The voice information processing unit detects when a user starts a phone call, and when the call ends, causes the voice output unit to output a voice corresponding to the spoken content, while the user is still on the phone, stopping texting the voice input by the voice input unit.
10. The speech information processing device according to claim 9, wherein When the voice information processing unit causes the voice output unit to output the voice corresponding to the speech content, the voice information processing unit causes the voice output unit to output as voice a sentence indicating that textualization of the voice input by the voice input unit has been stopped while the user is on a call, corresponding to the voice corresponding to the speech content.
11. The speech information processing device according to any one of claims 1 to 10, wherein: When causing the voice output unit to output the utterance content in voice, the voice information processing unit causes the voice output unit to output the sentence represented by the generated text as voice.
12. The speech information processing device according to any one of claims 1 to 10, wherein: The voice information processing unit causes the voice output unit to output the speech content in voice, and outputs a recorded voice of the user's speech.
13. A method for processing speech information, characterized in that: The steps include: The speech information processing unit of the speech information processing device converts speech input by the speech input unit of the speech information processing device into text during a speech acceptance period, wherein the speech acceptance period is a period from when the user accepts the spoken speech to be converted into text, and from when the user finishes speaking a message start word or performs a prescribed operation on an input device or performs a prescribed gesture, until when the user starts speaking a message end word or performs a prescribed operation on the input device or performs a prescribed gesture; and The voice information processing unit of the voice information processing device sequentially converts the user's speech into text during the voice acceptance period, and when it can be considered that the user's speech is interrupted, the voice output unit of the voice information processing device automatically outputs the speech content that the user has spoken during the voice acceptance period by voice during the voice acceptance period. The voice information processing unit causes the voice output unit to output voice corresponding to the content of the speech uttered by the user during the voice acceptance period when a period during which the user has not spoken a word during the voice acceptance period exceeds a predetermined time.
Citation Information
Patent Citations
Digital broadcasting reception system
JP2003319280A
Mobile phone
JP2007151188A
Information processing device and information processing method
CN110140167A
Voice recognition system
US20090187406A1