Speech processing apparatus and speech processing method
By setting up a voice command list storage unit in the voice processing device, specific statements can be automatically identified and removed, solving the problem of users having to repeatedly cancel and re-speak, thus improving the user experience.
Patent Information
- Application Number
- CN201910783144.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-08-23
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2039-08-23
AI Technical Summary
Existing voice processing devices cannot automatically identify and remove specific statements that users do not want to send during voice reception, causing users to have to repeatedly cancel and resend the voice.
By setting up a voice command list storage unit in the voice processing device, specific sentences can be identified and automatically removed, and only the text of the article excluding the specific sentences can be sent.
It enables the automatic removal of specific sentences during voice reception, reducing repetitive operations for users and improving user experience.
Smart Images

Figure CN112489640B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a speech processing apparatus and a speech processing method, and more particularly to a technology suitable for use in a speech processing apparatus and a speech processing method that causes a user's voice to be converted into text and transmitted. Background Technology
[0002] Previously, devices for performing speech recognition on user-generated speech were known. For example, Patent Document 1 describes a speech processing device that has both speech recognition and speech synthesis functions, and can interrupt speech data being read aloud via speech synthesis to execute speech recognition, and can interrupt one function to execute the other.
[0003] In addition, conventional devices capable of speech recognition include those that input a user's voice, convert the input voice into text, and send the text as a message or email in a chat application. Using such devices, users can send text containing desired content to another party simply by speaking, without using their hands.
[0004] Existing technical documents
[0005] Patent documents
[0006] Patent Document 1: Japanese Patent Application Publication No. 10-161846 Summary of the Invention
[0007] In the aforementioned devices that convert and transmit user-generated speech into text, the following problems have existed: Conventional devices convert all speech emitted by the user during the speech reception period into text. Therefore, if, during this period, the user needs to emit a specific statement that they do not want to be transmitted as text, and the user emits such a statement, the final text output will contain text containing that specific statement. For example, some conventional devices have the function of performing specific processing corresponding to a specific statement when the user emits it. Furthermore, using such conventional devices requires the device to perform specific processing. In such cases, the user emits the specific statement, resulting in the final text output containing the text of that specific statement. When the final text output contains text containing a statement that the user does not want to be transmitted as text, the user needs to temporarily cancel the text transmission and repeat the process, which is inconvenient for the user.
[0008] This invention was made to solve this problem, and its purpose is to eliminate the need for repetitive work when, during the process of receiving the voice of an article intended to be texted, there are statements that the user does not want to be texted and sent aloud.
[0009] To address the aforementioned issues, this invention transmits the text of the article represented by the voice input during voice reception, excluding specific sentences. According to this invention, instead of transmitting the entire article represented by the voice input during voice reception, specific sentences are automatically removed, and the text of the article with those sentences removed is transmitted. Therefore, even if a user utters a specific sentence during voice reception that they do not want to be transmitted as text, that specific sentence is automatically removed from the final transmitted text, eliminating the need for the user to temporarily cancel text transmission and re-enter the speech. Attached Figure Description
[0010] Figure 1 This is a block diagram illustrating a functional configuration example of the voice processing apparatus according to the first embodiment of the present invention.
[0011] Figure 2 This is an example image showing a chat room screen.
[0012] Figure 3 It is a diagram used to illustrate the relationship between the various periods that include the speech reception period.
[0013] Figure 4 This is a diagram representing an example of the content of full-voice text data.
[0014] Figure 5 This is a flowchart illustrating an example of the operation of the voice processing device according to the first embodiment of the present invention.
[0015] Figure 6 This is a flowchart illustrating an example of the operation of the voice processing device according to the first embodiment of the present invention.
[0016] Figure 7 This is a flowchart illustrating an example of the operation of the voice processing device according to the first embodiment of the present invention.
[0017] Figure 8 This is a flowchart illustrating an example of the operation of the voice processing device according to the first embodiment of the present invention.
[0018] Figure 9 This is a block diagram illustrating a functional configuration example of the voice processing apparatus according to the second embodiment of the present invention.
[0019] Figure 10This is a diagram illustrating the situation where a message is appended to the message bar.
[0020] Figure 11 This is a flowchart illustrating an example of the operation of the voice processing device according to the second embodiment of the present invention.
[0021] Figure 12 This is a block diagram illustrating a functional configuration example of the voice processing apparatus according to the third embodiment of the present invention.
[0022] Explanation of reference numerals in the attached figures:
[0023] 1. 1A, 1B Voice Processing Devices
[0024] 11. Voice Input Section
[0025] Article Processing Department (13, 13A, 13B)
[0026] 14. Specific Processing Execution Control Department
[0027] 15. Article Submission Department
[0028] 16. Voice Command Overview (Storage Department) Detailed Implementation
[0029] <First Embodiment>
[0030] The first embodiment of the present invention will be described below based on the accompanying drawings. Figure 1 This is a block diagram illustrating an example of the functional configuration of the voice processing device 1. The voice processing device 1 according to this embodiment is a vehicle-mounted device that provides a user interface for text chat, allowing multiple users to converse via text messages. Specifically, the voice processing device 1 according to this embodiment has the following function: during text chat, inputting voice messages uttered by a occupant using the device (hereinafter referred to as "user") within a specified period transforms the text represented by the input voice into text, and then sending the text message as a message (hereinafter referred to as "message voice input function"). By utilizing the message voice input function, the user can generate and send messages to other parties in text chat without using hand input.
[0031] Furthermore, the voice processing device 1 according to this embodiment has the following function: when a user issues one of a plurality of pre-prepared voice commands, it recognizes the issued voice command, executes specific processing corresponding to the voice command, or causes other devices to execute specific processing (hereinafter referred to as "voice command receiving function"). The voice commands are prepared in advance, and the user is aware of each voice command and the specific processing executed when each voice command is issued. In this embodiment, the voice command is set to at least include the statement "Make the wipers work." This voice command is specifically referred to as the "wiper drive instruction command." The wiper drive instruction command is a voice command that instructs the user to start driving the wipers; if the user issues the wiper drive instruction command, then "drive the wipers" is executed as the corresponding specific processing.
[0032] The voice commands are not limited to those illustrated in this embodiment. For example, a voice command could be a statement such as "Find a nearby convenience store" to search for a specific type of facility based on specific criteria, or a statement such as "Go home" to search for a route home. Regarding these voice commands, the voice processing device 1 causes a navigation device (not shown) to perform corresponding processing. The vehicle on which the voice processing device 1 is mounted will be referred to as "this vehicle" below.
[0033] In the following description, during periods other than the voice reception period described later, it is assumed that voice commands are appropriately received, and the processing performed by the voice processing device 1 in accordance with the voice commands during periods other than the voice reception period is omitted from the description.
[0034] like Figure 1 As shown, the voice processing device 1 is connected to the microphone 2 and the touch screen 3. The microphone 2 is positioned to pick up the voice of the user installed in the vehicle. The microphone 2 picks up the voice and outputs the voice signal of the picked-up voice.
[0035] The touchscreen 3 has a display panel such as an LCD panel or an OLED panel, and a touch sensor arranged overlapping the display panel. It displays images in the display area and detects touch operations on the touch detection area. Various screens related to text chat are displayed on the touchscreen 3. The touchscreen 3 is located in the center of the dashboard or similar position, allowing the user to visually identify the display area and perform touch operations on the touch detection area.
[0036] like Figure 1As shown, the voice processing device 1 includes a general control unit 10, a voice input unit 11, a voice data analysis unit 12, a text processing unit 13, a specific processing execution control unit 14, and a text sending unit 15. Each of the above functional modules 10-15 can be constructed using hardware, a DSP (Digital Signal Processor), or software. For example, in the case of software construction, each of the above functional modules 10-15 is actually constructed using a computer's CPU, RAM, ROM, etc., and operates by means of a program stored in a recording medium such as RAM, ROM, hard disk, or semiconductor memory.
[0037] like Figure 1 As shown, the voice processing device 1 includes a voice command list storage unit 16 (equivalent to "storage unit" in the claims) as a storage medium. The voice command list storage unit 16 stores voice command list data 17. The voice command list data 17 describes the text of the voice command statement (hereinafter referred to as "voice command statement") for each voice command. In this embodiment, the voice command list data 17 describes at least the text of a voice command statement such as "Turn on the wipers." for a wiper drive instruction command.
[0038] The following describes the actions of the voice processing device 1 when it switches from a non-chat mode to a chat mode, and when the user's spoken words are converted into text and sent as a message. Chat mode refers to a mode in which the voice processing device 1 provides a text chat user interface, allowing the user to engage in text chat with a desired recipient.
[0039] The overall control unit 10 controls all parts of the voice processing device 1 in a comprehensive manner through the firmware of the voice processing device 1, the applications running on the firmware, and the functions of other programs. The overall control unit 10 can control the touch screen 3 to display various images, and can also detect the position coordinates of the touch screen 3 when a touch operation is performed.
[0040] When the overall control unit 10 is in a mode other than chat mode, it displays a button (icon) on the touch screen 3 to indicate the transition to chat mode. If the user touches the button, the overall control unit 10 launches a text chat-related application and displays various text chat-related screens on the touch screen 3 using the functions of the application and its accompanying programs.
[0041] When a user wishes to engage in text chat with a specific person (or multiple people), they can perform specific touch operations on a designated screen to enable or disable the voice input function, and further instruct the user to open a message exchange chat room as the venue for exchanging messages with that person. The overall control unit 10 enables or disables the voice input function based on the user's selection, and also displays the message exchange chat room screen 20 indicated by the user on the touch screen 3 based on the user's instruction. Hereinafter, the message exchange chat room screen 20 will be referred to as "chat room screen 20".
[0042] Figure 2 (A) is a simplified illustration of the chat room screen 20 according to this embodiment. Figure 2 As shown in (A), in chat room screen 20, message bars 21 recording messages from the target are displayed in chronological order on the left side of the screen, and message bars 21 recording messages from the user are displayed in chronological order on the right side of the screen. All the messages recorded in a message bar 21 become the unit of message sent in a conversation. In addition, the voice input function for messages is clearly indicated to be enabled or disabled in chat room screen 20.
[0043] When the overall control unit 10 opens a message exchange chat room while the message voice input function is enabled, it outputs a function activation instruction to the voice input unit 11 and the voice data analysis unit 12, instructing them to start the message voice input function. Conversely, when the message voice input function is disabled, or the message exchange chat room is closed via user instructions, the overall control unit 10 outputs a function deactivation instruction to the voice input unit 11 and the voice data analysis unit 12, instructing them to deactivate the message voice input function.
[0044] The voice input unit 11 inputs the voice spoken by the user. The processing of the voice input unit 11 will be described in detail below.
[0045] During the period from when the function enable instruction is input from the general control unit 10 until the function disable instruction is input (hereinafter referred to as the "message voice input period"), the voice input unit 11 performs the following processing: It inputs a voice signal output from the microphone 2, performs analog-to-digital conversion processing on the voice signal including sampling, quantization, and encoding, performs other signal processing to generate voice data, and buffers it in the buffer 18. The buffer 18 is storage formed in a working area such as RAM. The voice data is voice waveform data sampled at a predetermined sampling period (for example, 16kHz).
[0046] The voice data analysis unit 12 detects the start and end of the voice reception period (described later). Additionally, the voice data analysis unit 12 detects the presence of message transmission words in the voice data buffered in the buffer 18. The processing of the voice data analysis unit 12 is explained in detail below.
[0047] Figure 3 This is a graph clearly showing the relationships between the periods of voice input, the period when the start word of the message is spoken, the period when the end word of the message is spoken, the period when the send word of the message is spoken, and the period of voice reception, on a time axis. On the axis, time progresses from left to right. First, using... Figure 3 This describes the message start word, message end word, message sending word, and the duration of voice reception.
[0048] In this embodiment, when the user wishes to convert an article into text and send it as a message using the voice input function, the user speaks a message start word consisting of a fixed phrase, then speaks the article that the user wishes to convert into text, and finally speaks a message end word consisting of a fixed phrase. An example of a message start word is a phrase like "message start," and an example of a message end word is a phrase like "message end." In other words, in this embodiment, the period from the end of the message start word's pronunciation to the beginning of the message end word's pronunciation is the period for receiving the voice of the article that the user wishes to convert into text (i.e., the article the user wishes to convert into text). Furthermore, this period is equivalent to the "voice reception period."
[0049] Furthermore, in this embodiment, as described later, after the user pronounces the end-of-message, the text of the pre-determined message to be sent is displayed in the message bar 21 of the chat room screen 20. The user confirms the content of the message displayed in the message bar 21, and if the content is satisfactory and the user wishes to send the message, pronounces the message sending phrase. An example of a message sending phrase is the statement "Message sent." The message is sent to the recipient in response to the user's pronouncement of the message sending phrase.
[0050] exist Figure 3 The process begins during the timed TS message voice input period and ends during the timed TE message voice input period. The speech begins with the start word of the timed T1 message following the timed TS message and ends with the start word of the timed T2 message following the timed T1 message. Similarly, the speech begins with the end word of the timed T3 message following the timed T2 message and ends with the end word of the timed T4 message following the timed T3 message. Finally, the speech begins with the send word of the timed T5 message following the timed T4 message and ends with the send word of the timed T6 message following the timed T5 message. Figure 3In the case of the example, the period from time T2 to time T3 is equivalent to the voice reception period, during which the text spoken by the user becomes a textual object.
[0051] Furthermore, if a message voice input period begins (i.e., a function activation instruction is input from the general control unit 10), the voice data analysis unit 12 continuously analyzes the voice data buffered in the buffer 18, monitoring whether the voice waveform of the message start word appears in the voice data. In this embodiment, the voice pattern of the message start word (i.e., the pattern of the voice waveform when the message start word is pronounced) is pre-registered. Multiple voice patterns may also be registered. The voice data analysis unit 12 continuously compares the voice waveform of the voice data with the voice pattern of the message start word and calculates the similarity. If the similarity is a certain level or higher, it determines that the waveform of the message start word has appeared in the voice data.
[0052] Upon detecting the presence of a message start word's speech waveform in the speech data, the speech data analysis unit 12 determines the end position of the message start word's speech waveform in the speech data (hereinafter referred to as the "start word end position"). The start word end position is determined by timing relative to the start of the speech reception period. Figure 3 The position corresponding to timing T2). The position of the speech waveform in the speech data, for example, the timing that begins during the message speech input period (T2). Figure 3 The timing (TS) is used as the starting point (cycle 0) and represented as cycle 0. For example, if the speech data is sampled at a sampling period of 16kHz, the start word end position is represented as "cycle 16324". After determining the start word end position, the speech data analysis unit 12 outputs the information indicating the start word end position to the document processing unit 13.
[0053] Then, the voice data analysis unit 12 continuously analyzes the voice data buffered in the buffer 18, monitoring whether a voice waveform of a message end word appears in the voice data. This monitoring is performed based on a pre-registered voice pattern of message end words, using the same method as described above for monitoring whether a voice waveform of a message start word appears. If a voice waveform of a message end word is detected in the voice data, the voice data analysis unit 12 determines the start position of the voice waveform of the message end word in the voice data (hereinafter referred to as the "end word start position"). The end word start position is determined by timing the end of the voice reception period (…). Figure 3 The position corresponding to the timing T3). After determining the start position of the end word, the speech data analysis unit 12 outputs the information indicating the start position of the end word to the document processing unit 13.
[0054] Then, the voice data analysis unit 12 continuously analyzes the voice data buffered in the buffer 18, monitoring whether the voice waveform of a message sending word appears in the voice data. This monitoring is performed based on the pre-registered voice pattern of the message sending word, using the same method as described above for monitoring whether the voice waveform of a message start word appears. If the voice waveform of a message sending word is detected in the voice data, the voice data analysis unit 12 outputs a transmission instruction information notifying the overall control unit 10 that the voice waveform of a message sending word has appeared in the voice data. Then, the voice data analysis unit 12 resumes monitoring whether the voice waveform of a message start word appears in the voice data.
[0055] The above processing is performed by the voice data analysis unit 12. As a result, when the user pronounces the start word of the message, information indicating the end position of the start word is immediately output to the document processing unit 13, and when the user pronounces the end word of the message, information indicating the start position of the end word is immediately output to the document processing unit 13. In addition, when the user pronounces the send word of the message, a transmission instruction information for notifying the appearance of the voice waveform of the send word in the voice data is immediately output to the overall control unit 10.
[0056] The text processing unit 13 outputs the text of the article after removing specific sentences from the article represented by the speech input unit 11 during the speech reception period, which is the period during which the speech of the article to be converted into text is received. In particular, after the speech reception period ends, the text processing unit 13 according to this embodiment extracts the text of specific sentences from the text of all the articles represented by the speech input unit 11 during the speech reception period, deletes them, and outputs the text of the article after deletion. At this time, the text processing unit 13 extracts the text of specific sentences from the article that has been converted into text based on the consistency between the text of specific sentences and the text of specific sentences stored in advance in the speech command list storage unit 16. The processing of the text processing unit 13 will be described in detail below.
[0057] When the speech data analysis unit 12 outputs information indicating the end position of the message start word, the document processing unit 13 inputs this information. Then, when the speech data analysis unit 12 outputs information indicating the start position of the message end word, the document processing unit 13 inputs this information as well. If the document processing unit 13 inputs information indicating the start position of the message end word, it identifies the end position of the message start word and the start position of the message end word in the speech data buffered in the buffer 18 based on this information and the previously input information indicating the end position of the message start word. Furthermore, the document processing unit 13 acquires the speech data buffered in the buffer 18 that falls within the range from the end position of the message start word to the start position of the message end word (hereinafter referred to as "processing target speech data").
[0058] After acquiring the speech data of the processing target, the document processing unit 13 performs speech recognition on the speech data of the processing target, so that the article recorded in the speech data of the processing target becomes text, generating text data 23 (hereinafter referred to as "full speech text data 23") that describes the article that has become text. When generating full speech text data 23, the document processing unit 13 divides the article into sentences and records them in the data. A sentence is the smallest unit of expression that is the complete expression of content in the elements of an article. Sentences in Japanese and Chinese are basically ended by a period ".", and in English they are basically ended by a period ".".
[0059] Figure 4 The figures are examples representing the content of the full-speech text data 23. For example, suppose a user speaks the text "Thank you today. Please continue to support me in the future." during speech reception, and this text becomes the subject of the full-speech text data 23. In this case, such as Figure 4 As shown in (A), the elements "Thank you for today." constitute a sentence, and then the elements "Please continue to take care of me in the future" constitute another sentence. Among them, in Figure 4 (A) (described later) Figure 4 In (B) as well, the separator “ / ” is conveniently used as a symbol to separate statements.
[0060] As described above, the voice processing device 1 according to this embodiment has a voice command receiving function, and the user knows that by uttering a certain voice command from among the pre-prepared voice commands, a specific process corresponding to that voice command can be executed. Therefore, the user sometimes utters a voice command during the voice receiving period in order to execute a specific process.
[0061] For example, suppose a user wants an article like "Thank you today. Please continue to support me in the future." to be sent as a message. After the initial message is spoken, the article "Thank you today." is spoken first. Alternatively, suppose the user notices that the rain outside is getting heavier and considers turning on the windshield wipers. In this case, to make the wipers work as quickly as possible, the user will speak "Turn on the windshield wipers." before speaking "Please continue to support me in the future." In this scenario, if... Figure 4 As shown in (B), in the full-voice text data 23, the state records the three statements: "Thank you for today.", "Make the wipers work.", and "Please continue to take care of me in the future."
[0062] Specifically, by performing lexical analysis, grammatical structure analysis, and semantic structure analysis based on existing technologies related to natural language processing, full-speech text data 23 is appropriately generated based on speech recognition. Artificial intelligence technology may also be used as a part of the technology. Alternatively, the document processing unit 13 may be configured to generate full-speech text data 23 in conjunction with an external device. For example, it may be configured such that the speech processing device 1 can access a network, and the document processing unit 13 sends speech data to a server that has the function of generating full-speech text data 23 based on speech data, and receives full-speech text data 23 in response.
[0063] After generating the full-voice text data 23, the text processing unit 13 refers to the voice command list data 17 stored in the voice command list storage unit 16 and performs the following processing. That is, the text processing unit 13 determines whether each sentence of the text recorded in the full-voice text data 23 is consistent with a voice command statement recorded in the voice command list data 17. As described above, in this embodiment, the voice command list data 17 records the text of the statement "Make the wipers work." as a voice command statement. Therefore, in this embodiment, the text processing unit 13 determines at least whether each sentence recorded in the full-voice text data 23 is consistent with a voice command statement such as "Make the wipers work."
[0064] When determining whether there is consistency between the statements recorded in the full-voice text data 23 and the voice command statements recorded in the voice command list data 17, in this embodiment, the text processing unit 13 considers not only the consistency of the strings but also the consistency of the intent in the intent interpretation processing based on natural language processing to determine whether there is consistency between the statements recorded in the full-voice text data 23 and the voice command statements. Therefore, except for the case where the strings of each statement are completely identical, in the case of intent interpretation processing, the consistency between the statements is determined. Figure 1 Even when textual differences arise within a certain range, the document processing unit 13 sometimes determines that the sentences are consistent.
[0065] The text processing unit 13 maintains the state of the text data 23 for statements that do not match any of the voice command statements. On the other hand, the text processing unit 13 deletes statements from the text data 23 that match a particular voice command statement. For example, if the content of the text data 23 is... Figure 4In the case shown in (B), the second statement, "Make the wipers work," is consistent with the voice command statement. Therefore, in this case, the text processing unit 13 deletes the second statement from the full-voice text data 23. As a result of the above processing, when a statement consistent with a certain voice command statement is recorded in the full-voice text data 23, the text processing unit 13 extracts and deletes that statement from the full-voice text data 23.
[0066] After performing the processing of determining the consistency with the voice command statement and deleting the statements that have consistency, the text processing unit 13 outputs the processed full-voice text data 23 (hereinafter referred to as "processed voice text data") to the overall control unit 10. Furthermore, if the text processing unit 13 determines that a certain statement in the full-voice text data 23 has consistency with a certain voice command statement, it outputs the recognition information used to identify the voice command corresponding to that voice command statement to the specific processing execution control unit 14.
[0067] The specific processing execution control unit 14 performs specific processing based on specific statements removed by the document processing unit 13, or causes a device with the function of performing specific processing to perform specific processing. The processing of the specific processing execution control unit 14 will be described in detail below.
[0068] When the specific processing execution control unit 14 receives recognition information for a voice command from the document processing unit 13, it performs the following processing: The specific processing execution control unit 14 recognizes the voice command (i.e., the voice command issued by the user) based on the input recognition information and performs processing to achieve a state where specific processing corresponding to the recognized voice command has been performed. For each voice command, the processing performed by the specific processing execution control unit 14 is predetermined, and by performing the predetermined processing, the specific processing execution control unit 14 achieves a state where specific processing corresponding to the voice command has been performed.
[0069] For example, when identification information indicating a wiper drive instruction is input, the specific processing execution control unit 14 outputs a control instruction indicating the drive of the wipers to the control unit that controls the drive of the wipers. The control unit inputs the control instruction and begins to drive the wipers.
[0070] When processed voice-text data is input from the text processing unit 13, the overall control unit 10 performs the following processing: It generates a message bar 21 on the chat room screen 20 and displays the text described in the input processed voice-text data as a message in the message bar 21. For example, when displaying... Figure 2 In the chat room screen 20 shown in (A), the input was... Figure 4In the case of processed speech-text data with the same content as (A) full speech-text data 23, the overall control unit 10 shall, as follows: Figure 2 As shown in (B), a message bar 21 is generated, and a message is displayed in the message bar 21. The user confirms the content of the message in the message bar 21. If the content is correct, the user pronounces the message sending word, thereby instructing the user to send the message. If the message sending word is pronounced, the voice waveform of the message sending word appears in the voice data, and the voice data analysis unit 12 outputs the sending instruction information to the overall control unit 10.
[0071] After displaying the message, the overall control unit 10 monitors whether a transmission instruction has been entered from the voice data analysis unit 12. If a transmission instruction has been entered, the overall control unit 10 instructs the document transmission unit 15 to send the message.
[0072] Based on instructions from the overall control unit 10, the article sending unit 15 sends a message to the designated server according to the protocol. As a result, the article after removing voice command statements from the article spoken by the user during voice reception becomes text and is sent as a message to the other party.
[0073] As explained above, the voice processing device 1 according to this embodiment transmits text from the text represented by voice input during the voice reception period, after removing voice command statements (specific statements). According to this structure, instead of transmitting the entire text represented by voice input during the voice reception period as text, the voice command statements are automatically removed, and the text of the text after removing the voice command statements is transmitted. Therefore, even if the user issues a voice command during the voice reception period that they do not want to be transmitted as text, the voice command is automatically removed from the text that is ultimately transmitted as text, eliminating the need for the user to temporarily cancel text transmission and perform repetitive operations such as re-speaking.
[0074] Next, a flowchart will be used to explain the operation of the voice processing device 1. Figure 5 The flowchart FA is a flowchart illustrating an example of the operation of the general control unit 10 related to the display of the chat room screen 20 and the output of function activation and deactivation instructions. In the processing of the general control unit 10 described using flowchart FA, it is assumed that when the user instructs to open the message exchange chat room, they select to enable the message voice input function.
[0075] like Figure 5As shown in the flowchart FA, the user performs a specified touch operation on the touch screen 3, selects to enable the voice input function, and instructs to open a message exchange chat room for text chat with the desired object (step SX1). Corresponding to the instruction in step SX1, the overall control unit 10 enables the voice input function based on the user's selection (step SA1) and displays the corresponding chat room screen 20 on the touch screen 3 (step SA2).
[0076] Next, the overall control unit 10 outputs a function activation instruction to the voice input unit 11 and the voice data analysis unit 12 (step SA3). Then, the overall control unit 10 monitors whether the message voice input function is turned off, or whether the message exchange chat room is closed according to the user's instructions, etc. (step SA4). If the message voice input function is turned off, or the message exchange chat room is closed, a function deactivation instruction is output to the voice input unit 11 and the voice data analysis unit 12 (step SA5).
[0077] Figure 6 Flowchart FB is a flowchart illustrating an example of the operation of the voice input unit 11. The voice input unit 11 repeatedly executes the processing shown in flowchart FB. Figure 6 As shown, the voice input unit 11 monitors whether the function enable instruction information output by the general control unit 10 in step SA3 of flowchart FA has been input (step SB1). If the function enable instruction information has been input (step SB1: Yes), the voice input unit 11 begins to generate voice data based on the voice signal input from the microphone 2 and buffers the voice data into the buffer 18 (step SB2). Next, the voice input unit 11 monitors whether the function disable instruction information has been input from the general control unit 10 (step SB3). If the function disable instruction information has been input (step SB3: Yes), the voice input unit 11 ends the generation and buffering of voice data (step SB4). After the processing in step SB4, flowchart FB ends.
[0078] Figure 7 , Figure 8 This is a flowchart illustrating the operation examples of the voice data analysis unit 12, the document processing unit 13, the specific processing execution control unit 14, the overall control unit 10, and the document sending unit 15. Figure 7 In the flowchart, FC represents the action example of the speech data analysis unit 12, and FD represents the action example of the document processing unit 13. Additionally, in... Figure 8 In the flowchart, FE represents an example of the action of the specific processing execution control unit 14, FF represents an example of the action of the general control unit 10, and FG represents an example of the action of the article sending unit 15.
[0079] like Figure 7As shown in flowchart FC, the voice data analysis unit 12 monitors whether the function enable instruction information (step SC1) summarizing the output of the control unit 10 in step SA3 of flowchart FA has been input. If it has been input (step SC1: Yes), the voice data analysis unit 12 monitors whether the function disable instruction information (step SC2) summarizing the output of the control unit 10 in step SA5 of flowchart FA has been input, and monitors whether the voice waveform of the message start word appears in the voice data (step SC3). If the function disable instruction information has been input (step SC2: Yes), the processing of flowchart FC ends. If the voice waveform of the message start word appears (step SC3: Yes), the voice data analysis unit 12 determines the end position of the start word and outputs the information indicating the end position of the start word to the document processing unit 13 (step SC4).
[0080] Next, the voice data analysis unit 12 monitors whether a voice waveform indicating the end of a message appears in the voice data (step SC5). If a voice waveform appears (step SC5: Yes), the voice data analysis unit 12 determines the start position of the end word and outputs information indicating the start position of the end word to the document processing unit 13 (step SC6).
[0081] Next, the voice data analysis unit 12 monitors whether the voice waveform of the message sending word appears in the voice data (step SC7). If the voice waveform is detected (step SC7: Yes), the voice data analysis unit 12 outputs a transmission instruction to the overall control unit 10 (step SC8). After the processing in step SC8, the flowchart FC ends. The voice data analysis unit 12 repeats the processing of the flowchart FC.
[0082] like Figure 7 As shown in the flowchart FD, the text processing unit 13 monitors whether information indicating the end position of the message start word has been input (step SD1). If it has been input (step SD1: Yes), the text processing unit 13 monitors whether information indicating the start position of the message end word has been input (step SD2). If it has been input (step SD2: Yes), the text processing unit 13 obtains the processing target speech data based on the information input in step SD1 and the information input in step SD2 (step SD3).
[0083] Next, the document processing unit 13 performs speech recognition on the speech data to be processed, turning the document recorded in the speech data into text and generating full-speech text data 23 (step SD4). Next, the document processing unit 13 refers to the speech command list data 17 stored in the speech command list storage unit 16 and performs the following processing: The document processing unit 13 determines whether each sentence in the full-speech text data 23 matches a speech command statement in the speech command list data 17, deletes the matching sentence from the full-speech text data 23, and generates processed speech text data (step SD5). Next, the document processing unit 13 outputs the processed speech text data to the overall control unit 10 (step SD6). Furthermore, if the document processing unit 13 determines that a sentence in the full-speech text data 23 matches a speech command statement, it outputs the recognition information related to that speech command statement to the specific processing execution control unit 14 (step SD7). After the processing in step SD7, the flowchart FD ends. Article Processing Department 13 Repeated Execution Flowchart FD Processing.
[0084] like Figure 8 As shown in flowchart FE, the specific processing execution control unit 14 monitors whether recognition information has been input (step SE1). If it has been input (step SE1: Yes), the specific processing execution control unit 14 recognizes the voice command based on the input recognition information and executes processing to achieve the state after the specific processing corresponding to the recognized voice command has been executed (step SE2). After the processing in step SE2, flowchart FE ends. The specific processing execution control unit 14 repeats the processing of flowchart FE.
[0085] like Figure 8 As shown in flowchart FF, the overall control unit 10 monitors whether processed voice-text data has been input from the text processing unit 13 (step SF1). If input has been received (step SF1: Yes), the overall control unit 10 generates a message bar 21 in the chat room screen 20 and displays the text described in the input processed voice-text data as a message in the message bar 21 (step SF2). Next, the overall control unit 10 monitors whether a sending instruction message has been input from the voice data analysis unit 12 (step SF3). If input has been received (step SF3: Yes), the overall control unit 10 instructs the text sending unit 15 to send the message (step SF4). After the processing in step SF4, flowchart FF ends. The overall control unit 10 repeats the processing of flowchart FF.
[0086] like Figure 8As shown in flowchart FG, the document sending unit 15 monitors for any instructions from the overall control unit 10 (step SG1). If an instruction is received (step SG1: Yes), the document sending unit 15 sends the message displayed in the message bar 21 (step SG2). After processing in step SG2, flowchart FG ends. The overall control unit 10 repeats the processing of flowchart FG.
[0087] <Second Implementation>
[0088] Next, the second embodiment will be described. In the following description of the second embodiment, elements that are the same as those in the first embodiment will be given the same reference numerals, and their detailed descriptions will be omitted. Figure 9 This is a block diagram illustrating a functional configuration example of the voice processing device 1A according to this embodiment. For example... Figure 9 As shown, the voice processing apparatus 1A according to this embodiment replaces the text processing unit 13 according to the first embodiment by having a text processing unit 13A, and replaces the general control unit 10 according to the first embodiment by having a general control unit 10A. In the processing of the voice processing apparatus 1A according to this embodiment, the processing of the text processing unit 13A and the processing of the general control unit 10A after the start of the message voice input period differ from those in the first embodiment. Hereinafter, the processing of the voice processing apparatus 1A after the start of the message voice input period will be described.
[0089] Similar to the first embodiment, the speech data analysis unit 12 continuously analyzes the speech data buffered in the buffer 18, and outputs information indicating the end position of the start word and information indicating the start position of the end word to the document processing unit 13A based on the analysis results. If information indicating the end position of the start word is input, the document processing unit 13A subsequently performs speech recognition and language parsing on the speech data buffered in the buffer 18, and monitors whether sentences appear in the document represented by the speech data.
[0090] For example, let's say it's user-to-user Figure 4 The article shown in (B) was spoken. In this case, when the speech data corresponding to the first sentence, "Thank you today," is buffered to buffer 18, the article processing unit 13A detects the presence of a "sentence" based on the results of speech recognition and language parsing. Similarly, the article processing unit 13A detects the presence of "sentences" when the speech data corresponding to the second and third sentences is buffered. However, since the analysis by the article processing unit 13A takes time, there may be a time lag between the timing of buffering the speech data corresponding to a sentence to buffer 18 and the timing of detecting the presence of a sentence in the article represented by the speech data.
[0091] The document processing unit 13A monitors whether sentences appear in the document represented by the speech data until information indicating the start position of the end word is input from the speech data analysis unit 12. If information indicating the start position of the end word is input from the speech data analysis unit 12, the document processing unit 13A identifies the start position of the end word in the speech data, discards the analysis results for the speech data after the start position of the end word, and does not consider the speech data after the start position of the end word as objects of monitoring whether a "sentence" has appeared.
[0092] When a statement is detected, the document processing unit 13A performs the following processing: It refers to the voice command list data 17 stored in the voice command list storage unit 16 and determines whether the detected statement matches any of the voice command statements recorded in the voice command list data 17. If the detected statement does not match any voice command statement, the document processing unit 13A outputs the text data describing the detected statement to the overall control unit 10A.
[0093] On the other hand, if the appearing statement is consistent with a certain voice command statement, the text processing unit 13A discards the appearing statement. In this case, the text data describing the appearing statement is not output to the overall control unit 10A. Furthermore, if the appearing statement is consistent with a certain voice command statement, the text processing unit 13A outputs the recognition information of the voice command corresponding to the voice command statement determined to be consistent with the appearing statement to the specific processing execution control unit 14. The processing of the specific processing execution control unit 14 when the recognition information is input is as described in the first embodiment.
[0094] Each time text data describing a statement is input from the text processing unit 13A, the overall control unit 10A appends the text of that statement from the text data to the message bar 21 of the chat room screen 20. The message bar 21 is generated appropriately. Furthermore, the overall control unit 10A monitors whether a sending instruction is input from the voice data analysis unit 12. If a sending instruction is input, the overall control unit 10A instructs the text sending unit 15 to send a message. The text sending unit 15, as in the first embodiment, sends a message (containing all the statements described in the message bar 21) in accordance with the instruction.
[0095] After performing the above processing, regarding the various sentences included in the text spoken by the user during the voice reception period, for sentences that are inconsistent with the voice command sentences, the text of the sentence is output to the overall control unit 10A through the text processing unit 13A. On the other hand, for sentences that are consistent with the voice command sentences, the text of the sentence is not output to the overall control unit 10A through the text processing unit 13A, and the text of the sentence is not sent through the text sending unit 15.
[0096] Figure 10 This diagram illustrates the situation where a message is appended to the message bar 21 through the processing of the general control unit 10A. Now, let's assume the content of the chat room screen 20 is as follows: Figure 10 As shown in (A), in this state, the user... Figure 4 The text in (B) is amplified. If the user amplifies a statement such as "Thank you for today." in the first statement, a message bar 21A is generated accordingly, and the statement is displayed in message bar 21A. Figure 10 (B)). Next, if the user speaks the second statement, "Make the wipers work.", this statement is not appended to message bar 21A because it matches the voice command statement. Then, if the user speaks the third statement, "Please take care of me next time too.", this statement is displayed in message bar 21A accordingly. Figure 10 (C)). Then, if the user pronounces the message sending words, the entire text, including all the statements recorded in message bar 21A, is sent as a message through text sending unit 15.
[0097] According to the structure of this embodiment, it has the same effect as the first embodiment. That is, instead of making the entire text represented by the voice input during the voice reception period into text and sending it, the voice command is automatically removed, and the text of the article after removing the voice command is sent. Therefore, even if the user issues a voice command that they do not want to be made into text and sent during the voice reception period, the voice command is automatically removed from the article that is eventually sent as text, and the user does not need to temporarily cancel the text transmission and perform the repetitive operation of re-speaking.
[0098] Next, an example of the operation of the voice processing device 1A according to this embodiment will be explained using a flowchart. Figure 11 Flowchart FH is a flowchart representing an example of the operation of the document processing unit 13A, and flowchart FI is a flowchart representing an example of the operation of the overall control unit 10A.
[0099] like Figure 11As shown in flowchart FH, the document processing unit 13A monitors whether information indicating the end position of the start word has been input (step SH1). If it has been input (step SH1: Yes), the document processing unit 13A monitors whether information indicating the start position of the end word has been input (step SH2), and monitors whether a "statement" appears in the document represented by the voice data (step SH3). If a statement appears (step SH3: Yes), the document processing unit 13A refers to the voice command list data 17 stored in the voice command list storage unit 16 and determines whether the appearing statement is consistent with one of the voice command statements recorded in the voice command list data 17 (step SH4).
[0100] If the appearing statement does not match any voice command statement (step SH4: No), the text processing unit 13A outputs the text data describing the appearing statement to the overall control unit 10A (step SH5). After processing in step SH5, the processing steps return to step SH2. On the other hand, if the appearing statement matches a certain voice command statement (step SH4: Yes), the processing steps return to step SH2. In this case, the text data describing the appearing statement is not output to the overall control unit 10A.
[0101] If information indicating the start position of the ending word is input (step SH2: Yes), the document processing unit 13A identifies the start position of the ending word in the speech data and discards the analysis results for the speech data after the start position of the ending word (step SH6). After processing in step SH6, flowchart FH ends. The document processing unit 13A repeats the processing of flowchart FH.
[0102] like Figure 11 As shown in flowchart FI, the overall control unit 10A monitors whether a transmission instruction message has been input from the voice data analysis unit 12 (step SI1) and whether text data has been input from the document processing unit 13A (step SI2). If text data has been input (step SI2: yes), the overall control unit 10A appends the text of the statement described in the text data to the message bar 21 (step SI3). After processing in step SI3, the processing steps return to step SI1. On the other hand, if a transmission instruction message has been input, the overall control unit 10A instructs the document sending unit 15 to send a message (step SI4). After processing in step SI4, flowchart FI ends. The overall control unit 10A repeats the processing of flowchart FI.
[0103] <Third Implementation>
[0104] Next, the third embodiment will be described. In the following description of the third embodiment, elements that are the same as those in the first embodiment will be given the same reference numerals, and their detailed descriptions will be omitted. Figure 12 This is a block diagram illustrating a functional configuration example of the voice processing device 1B according to this embodiment. For example... Figure 12 As shown, the voice processing device 1B differs from the voice processing device 1 of the first embodiment in that it has an article processing unit 13B instead of the article processing unit 13 of the first embodiment, and does not have a specific processing execution control unit 14.
[0105] In this embodiment, in addition to the voice processing device 1B, the vehicle is also equipped with a specific processing execution device (not shown). The specific processing execution device is a device independent of the voice processing device 1B. The specific processing execution device has the following functions: when a user issues a voice command, it recognizes the issued voice command and executes specific processing corresponding to the recognized voice command, or executes specific processing corresponding to the voice command by causing other devices under its control to perform such specific processing. In an environment where such a specific processing execution device is provided, the user, similar to in the first embodiment, may issue a voice command during voice reception to cause the specific processing execution device (or devices under its control) to perform specific processing.
[0106] The text processing unit 13B of the speech processing apparatus 1B according to this embodiment performs the following processing as a different process from that of the text processing unit 13 according to the first embodiment. That is, when the text processing unit 13 according to the first embodiment determines that a certain sentence in the full-speech text data 23 is consistent with a certain speech command sentence, it outputs the recognition information of the speech command corresponding to the speech command sentence to the specific processing execution control unit 14. On the other hand, the text processing unit 13B according to this embodiment does not perform such recognition information output.
[0107] According to the structure of this embodiment, which is the same as the first embodiment described above, even if the user issues a voice command that they do not want to be sent as text during the reception of voice, the voice command is automatically removed from the text that is eventually sent as text, and the user does not need to temporarily cancel the text transmission and repeat the work of re-speaking.
[0108] The above three embodiments have been described, but each of these embodiments is merely a specific example of implementing the present invention and should not be construed as limiting the scope of the present invention. That is, the present invention can be implemented in various forms without departing from its spirit or main features.
[0109] For example, in the embodiments described above, the article sending unit 15 sends text as a message in a text chat, but the sending of text is not limited to the methods exemplified in each embodiment. For example, text can also be sent via email. Furthermore, text sending does not mean sending to only a specific recipient; it broadly includes the concept of transmitting text to external devices, such as sending text to a server or a specific host device. For example, sending text to a news submission website or forum website according to a protocol is also included in text sending.
[0110] Furthermore, in the first embodiment, the voice processing device 1 is installed in a vehicle to prevent the transmission of text containing voice commands that are spoken inside the vehicle. However, the voice processing device 1 does not necessarily have to be installed inside the vehicle, and the object to be prevented from transmitting text is not limited to voice commands. The object to be prevented from transmitting text may, for example, be a pre-registered statement that is unsuitable to be text and transmitted. The same applies to the second and third embodiments.
[0111] Furthermore, in the embodiments described above, the voice reception period begins when the user sends a message start word. Alternatively, the voice reception period can begin when the user performs a specified touch operation on the touchscreen 3, or when the user performs a specified gesture in a gesture-detecting structure. The same applies to both the message end word and the message send word. In particular, regarding the message end word, the voice reception period can also end when the user remains silent for a certain period.
Claims
1. A voice processing device, characterized in that, have: The voice input unit is used to input the voice commands given by the user. The voice data analysis unit detects the start and end of the voice reception period, which refers to the period from the end of the message start word to the start of the message end word, during which the voice of the article to be turned into text is received. The text processing unit, based on speech recognition, outputs the text of the article after removing specific statements instructing the execution of specific processing from the article represented by the speech input unit during the speech reception period; The article sending unit sends the text output by the article processing unit; as well as The specific processing execution control unit executes the specific processing based on the specific statement removed by the article processing unit, or causes a device having the function of executing the specific processing to execute the specific processing.
2. The voice processing device as described in claim 1, characterized in that, After the speech reception period ends, the text processing unit extracts the text of the specific sentence from the text formed by all the speech inputs of the speech input unit during the speech reception period, deletes it, and outputs the text of the deleted article.
3. The voice processing device as described in claim 2, characterized in that, The article processing unit extracts the text of the specific statement from the article that has become text based on the consistency between the text of the specific statement and the text of the specific statement stored in advance by the storage unit.
4. The voice processing device as described in claim 1, characterized in that, During the voice reception period, the text processing unit continuously performs voice recognition on the voice input unit and monitors whether a sentence appears. Each time a sentence appears, it determines whether the sentence is the specific sentence. If it is not the specific sentence, it outputs the text of the sentence. On the other hand, if it is the specific sentence, it does not output the text of the sentence.
5. The voice processing device as described in claim 4, characterized in that, The text processing unit determines whether the text is the specific statement based on the consistency between the text and the text of the specific statement stored in advance by the storage unit.
6. A speech processing method, characterized in that, Includes the following steps: The voice data analysis unit of the voice processing device detects the start and end of the voice reception period, which refers to the period from the end of the message start word to the start of the message end word, during which the voice of the article to be turned into text is received. The text processing unit of the speech processing device outputs, based on speech recognition, the text of the article after removing specific statements instructing the performance of specific processing from the article represented by the speech input unit of the speech processing device during the speech reception period; The step of the text sending unit of the voice processing device sending the text output by the text processing unit; as well as The specific processing execution control unit of the speech processing device executes the specific processing based on the specific statement removed by the article processing unit, or causes a device with the function of executing the specific processing to execute the steps of the specific processing.
Citation Information
Patent Citations
Speech processor and information processing method used for the same
JP1998161846A
Voice information interaction method and device and equipment thereof
CN107680589A
Device and method for processing information and program storage medium
JP2001117884A