Information processing device, display device, information processing system, and information processing program

JP2026148275APending Publication Date: 2026-09-17SHARP KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025036741
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2026-09-17

AI Technical Summary

Benefits of technology

【0010】 本開示の一態様によれば、ユーザが理解しやすいように音声データを出力することが可能となる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026148275000001_ABST
    Figure 2026148275000001_ABST
Patent Text Reader

Abstract

This technology provides the ability to output audio data in a way that is easy for users to understand. [Solution] The information processing device (100) includes a phrase analysis unit (203) that analyzes phrases of text generated as a response to the user's voice, a voice data generation unit (210) that generates voice data for voice output from the text, and a timing control unit (206) that controls the output timing of the voice data corresponding to the phrases analyzed by the phrase analysis unit (203).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an information processing apparatus, a display apparatus, an information processing system, and an information processing program. [Background Art]

[0002] Conventionally, systems that generate appropriate answers to user questions have been proposed. In this question answering system, when a question document is input, the document is divided into predetermined sections, and it is determined whether each divided section is important. The display device displays example answers for important sections. Accordingly, even when there are a plurality of question contents in one document, accurate answers can be provided to each respective question. [Prior Art Literature] [Patent Literature]

[0003] [Patent Literature 1] Japanese Unexamined Patent Publication No. 2005-092271 [Summary of the Invention] [Problem to be Solved by the Invention]

[0004] However, in the above-described system, interfaces such as response speed and response display method are not user-centric.

[0005] An object of one aspect of the present disclosure is to provide a technique capable of outputting audio data in a manner that is easy for a user to understand. [Means for Solving the Problem]

[0006] In order to solve the above problem, an information processing apparatus according to an aspect of the present disclosure includes: a phrase analysis unit that analyzes phrases of a sentence generated as a response to a user's voice; an audio data generation unit that generates audio data for audio output from the sentence; and a timing control unit that controls output timing of the audio data corresponding to the phrases analyzed by the phrase analysis unit.

[0007] To solve the above problems, an information processing system according to one aspect of the present disclosure is an information processing system including an information processing device and a display device that is communicably connected to the information processing device, wherein the information processing device includes a phrase analysis unit that analyzes phrases of a sentence generated as a response to a user's voice, a voice data generation unit that generates voice data for voice output from the sentence, a text data generation unit that generates text data for display from the sentence, and a timing control unit that controls the output timing of the voice data and the text data based on the phrases analyzed by the phrase analysis unit, and the display device includes a display unit that displays the text data generated by the information processing device.

[0008] To solve the above problems, an information processing program according to one aspect of the present disclosure is an information processing program for causing a computer to function as an information processing device, wherein the computer functions as: a phrase analysis unit that analyzes phrases of text generated as a response to a user's voice; a voice data generation unit that generates voice data for voice output from the text; and a timing control unit that controls the output timing of voice data corresponding to the phrases analyzed by the phrase analysis unit.

[0009] Each aspect of the information processing device described herein may be implemented by a computer. In this case, an information processing program that enables the computer to implement the information processing device by operating the computer as each part (software element) of the information processing device, and a computer-readable recording medium on which the program is recorded, also fall within the scope of this disclosure. [Effects of the Invention]

[0010] According to one aspect of this disclosure, it is possible to output audio data in a way that is easy for the user to understand. [Brief explanation of the drawing]

[0011] [Figure 1] This is a block diagram showing an example configuration of a display device according to one embodiment of the present disclosure. [Figure 2] This figure shows an example of the appearance of a display device. [Figure 3] This is a block diagram showing an example of the hardware configuration of an information processing device according to one embodiment of the present disclosure. [Figure 4] This is a block diagram showing an example of the functions of a control unit of an information processing device according to one embodiment of the present disclosure. [Figure 5] This flowchart shows an example of processing performed by an information processing device according to one embodiment of this disclosure. [Figure 6] This figure shows an example of a screen displayed by the display unit. [Figure 7] This diagram illustrates the changes in the contents of the variables that hold the content of the retrieved string and the variables that hold the position of punctuation marks. [Modes for carrying out the invention]

[0012] Hereinafter, one embodiment of the present disclosure will be described in detail with reference to the drawings. In the drawings, identical or substantially identical components will be denoted by the same reference numerals, and detailed descriptions will not be repeated.

[0013] (Display device 1) Figure 1 is a block diagram showing an example configuration of the display device 1 according to this embodiment. The display device 1 according to this embodiment is a voice-interactive display device and, as shown in Figure 1, comprises a sensing unit 21, an operation signal receiving unit 22, an information processing device 100, a display unit 30, and an audio output unit 40. The display device 1 is typically a stationary display device, but may also be a wall-mounted display device, or a smartphone, tablet terminal, PC (Personal Computer), etc. Stationary or wall-mounted display devices are equipped with fixing parts for fixing to a structure. For example, a stationary display device is equipped with fixing parts for fixing to a stand which is a structure, and a wall-mounted display device is equipped with fixing parts for fixing to a wall which is a structure.

[0014] Furthermore, the components of the display device 1 are not limited to being configured as a single device, but may be configured as a system distributed across multiple devices connected via a communication network. For example, the information processing device 100 may be configured as a set-top box or as a cloud server.

[0015] (Sensing unit 21) For example, the sensing unit 21 includes one or more imaging units (cameras), and uses these imaging units to capture images of the user. In this case, the sensing unit 21 supplies sensing data, including imaging data showing the images captured by the imaging units, to the information processing device 100.

[0016] As another example, the sensing unit 21 is equipped with one or more microphones that collect the user's voice. In this case, the sensing unit 21 supplies sensing data, including voice data indicating the voice collected by the microphones, to the information processing device 100. The voice collected by the microphones is mainly intended for human speech, but is not limited to this, and refers to sound in general that propagates through a medium such as air.

[0017] (Operation signal receiving unit 22) The operation signal receiving unit 22 receives an operation signal from the remote control device 2 used by a user to remotely operate the display device 1. The operation signal is, for example, a signal indicating a result of the user operating a user interface displayed on the display unit 30 using the remote control device 2. The operation signal receiving unit 22 supplies the received operation signal to the information processing device 100. Note that a device capable of operating the display device 1 is not limited to the remote control device 2, and the display device 1 may be operable by a keyboard or the like (not shown). Furthermore, the configuration may be such that by installing a dedicated application on a smartphone or tablet terminal owned by the user, the smartphone or tablet terminal and the display device 1 can communicate with each other. In this case, the user can remotely operate the display device 1 using the smartphone or tablet terminal instead of the remote control device 2. At this time, the operation signal receiving unit 22 supplies the operation signal received from the smartphone or tablet terminal to the information processing device 100. Furthermore, the user can also remotely operate the display device 1 using voice. Although the above description explains that the user's uttered voice acquired by the sensing unit 21 is supplied to the information processing device 100, the present invention is not limited thereto. The user's uttered voice acquired by the sensing unit 21 may be supplied to the operation signal receiving unit 22. According to such a configuration, the operation signal receiving unit 22 acquires the user's uttered voice from the sensing unit 21, converts the acquired uttered voice into predetermined data, and thus can receive the user's uttered voice as an operation signal. This enables remote operation via voice.

[0018] (Information processing device 100) The information processing device 100 generates a display image to be displayed on the display unit 30 and output audio to be output from the audio output unit 40 based on at least one of sensing data supplied from the sensing unit 21 and an operation signal supplied from the operation signal receiving unit 22.

[0019] The display image is, for example, an image including an image of a partner for voice dialogue with the user (hereinafter referred to as an avatar), text corresponding to the utterance content of the avatar, and the like in the display device 1. Further, the display image may include data generated by generative AI (Artificial Intelligence) such as a large language model (LLM) through dialogue with the user.

[0020] The output audio includes, in addition to audio related to the display image, various notification sounds, operation sounds for operating the user interface displayed on the display unit 30, and the like. The output audio may include audio obtained by converting text data generated by generative AI such as a large language model (LLM) through dialogue with the user.

[0021] Details of the information processing device 100 will be described later.

[0022] (Display Unit 30) The display unit 30 displays a display image generated by the information processing device 100. As an example, the display unit 30 is configured to include a display panel and a driver that drives the display panel based on image data of the display image. As an example, the display panel is a liquid crystal panel or an organic EL panel.

[0023] (Audio Output Unit 40) The audio output unit 40 is a speaker that outputs output audio generated by the information processing device 100.

[0024] (Other Components) The display device 1 may further include a content receiving unit for receiving content data and supplying the received content data to the information processing device 100. The content data may include, but is not limited to, encoded video data, encoded audio data, and related information associated with the video data. The encoded video data may, but is not limited to, data encoded by various video encoding technologies such as MPEG2, MPEG4, H.264, and H.265 (e.g., TS (Transport Steam)). The related information may, but is not limited to, at least one of the following: program information related to the content, program guide data including the program information, data broadcasting information provided in relation to the content, and explanatory information related to the content. The related information may, for example, be extracted from the content data without decoding processing by the aforementioned video encoding technology such as MPEG2. The content receiving unit may also be configured to acquire content data from the Internet via wireless or wired communication. Furthermore, the content receiving unit may be configured to acquire content data from broadcast waves and may include a tuner for selecting one of the multiple channels included in the broadcast wave. When the content receiving unit includes a tuner, the display device 1 may also be referred to as a television receiver (or simply a television).

[0025] (Example of the appearance of display device 1) Figure 2 shows an example of the external appearance of the display device 1. The display device 1 shown in Figure 2 is a stationary display device equipped with legs, and is installed, for example, in a room such as a living room in a house. As an example, the display device 1 shown in Figure 2 is equipped with one sensing unit 21 above the display unit 30, and one operation signal receiving unit 22 and two audio output units 40 below the display unit 30.

[0026] (Hardware configuration of the information processing device 100) Next, an example of the hardware configuration of the information processing device 100 will be described with reference to Figure 3. Figure 3 is a block diagram showing an example of the hardware configuration of the information processing device 100.

[0027] As shown in Figure 3, the information processing device 100 comprises a control unit 101, a memory 102, a storage unit 103, and a communication interface 104. These components are interconnected via a bus 105 so that they can communicate with one another.

[0028] Memory 102 can be implemented as, for example, RAM (Random Access Memory), DRAM (Dynamic Random Access Memory), or other volatile memory. Storage unit 103 is composed of ROM (Read-Only Memory), HDD (Hard Disk Drive), SSD (Solid State Drive), etc., and stores sensing data supplied from sensing unit 21 and operation signals supplied from operation signal receiving unit 22. Storage unit 103 also stores a program for processing the sensing data and operation signals.

[0029] Furthermore, the memory unit 103 stores system data 50, such as AI prompts and knowledge / learning data. Details of this system data 50 will be described later.

[0030] Furthermore, the memory unit 103 may store various databases, such as personal management information and system configuration information. Personal management information includes, for example, information such as each user's characteristics, preferences, tastes, and usage history. System configuration information includes information such as software application customization specifications, software versions, hardware specifications, and hardware versions.

[0031] The control unit 101 is composed of, for example, one or more processors. The control unit 101 can be implemented as, for example, a CPU (Central Processing Unit), GPU (Graphics Processing Unit), MPU (Micro Processor Unit), FPGA (Field-Programmable Gate Array), or other device that performs arithmetic processing. The control unit 101 reads a program from the storage unit 103 and executes the program using the memory 102 as a working area. Although Figure 3 shows one control unit 101, it is not limited to this, and multiple control units 101 may be provided.

[0032] The communication interface 104 is implemented as hardware such as a network adapter, various communication software, or a combination thereof, and is configured to enable wireless or wired communication over a communication network.

[0033] (Functions of the control unit 101) Next, an example of the functions of the control unit 101 will be described with reference to Figure 4. Figure 4 is a block diagram showing an example of the functions of the control unit 101. The control unit 101 reads a program from the storage unit 103 and executes the program using the memory 102 as a work area, thereby functioning as the speech recognition unit 201, the text generation unit 202, the phrase analysis unit 203, the output data generation unit 204, the speech synthesis unit 205, and the timing control unit 206. The functions of each of these units may, for example, be distributed across multiple devices connected via a communication network. In this case, the functions of each of these units may be configured as a system consisting of multiple devices. Also, what is described as "~unit" here may be rephrased as "~circuit," "~device," or "~equipment," or as "~step," "~procedure," or "~process." That is, what is described as "~unit" may be implemented by a program stored in the storage unit 103 as described above, or it may be implemented by hardware such as elements, devices, boards, and wiring only, or by a combination of software and hardware. These functions will be described below.

[0034] The speech recognition unit 201 recognizes the user's voice input by the sensing unit 21 and converts it into text data. The speech recognition unit 201 includes, for example, a speech compression circuit and performs stepwise processing of feature transformation (speech compression), phoneme extraction, pattern matching, and language model shaping to convert the user's voice input by the sensing unit 21 into text data.

[0035] The text generation unit 202 is implemented, for example, by Generative Artificial Intelligence (Generative Artificial Intelligence). The text generation unit 202 takes the text data converted by the speech recognition unit 201 as an AI prompt and generates text data (sentences) that respond to the user's voice. This text data may, for example, be an answer to a question from the user.

[0036] The phrase analysis unit 203 analyzes the text generated by the text generation unit 202 and extracts phrases from the text. In this embodiment, a phrase refers to a unit that is pronounced as a single syllable without being divided into phrases. For example, a sentence divided into phrases is called a phrase. Examples of punctuation marks include characters such as ",", ".", "、", "。", "!", and "?".

[0037] The output data generation unit 204 generates output data from the text generated by the text generation unit 202. The output data generation unit 204 includes an audio data generation unit 210 and a text data generation unit 211.

[0038] The voice data generation unit 210 generates voice output data (hereinafter also referred to as voice data) from the string generated by the text generation unit 202. The voice synthesis unit 205 performs voice synthesis using the voice output data generated by the voice data generation unit 210 and outputs it to the voice output unit 40, thereby speaking as a response to the user's voice.

[0039] The text data generation unit 211 generates display data from the string generated by the text generation unit 202. The display unit 30 outputs the text data generated by the text data generation unit 211 to the screen and displays it as a response to the user's voice.

[0040] The timing control unit 206 controls the timing of the audio data output by the audio output unit 40 and the timing of the display data (text data) output by the display unit 30. For example, the timing control unit 206 causes the audio data to be output to the audio output unit 40 for each phrase of the sentence that has been divided by the phrase analysis unit 203. The timing control unit 206 also causes the text data to be displayed on the display unit 30 sequentially in accordance with the timing of the text output (generation timing) of the text generated by the text generation unit 202.

[0041] The memory unit 103 stores information such as text data (AI prompt) generated by the speech recognition unit 201, text data as a response to user speech generated by the text generation unit 202, sentence segments analyzed by the sentence segment analysis unit 203, speech data generated by the speech data generation unit 210, and display data generated by the text data generation unit 211 as system data 50.

[0042] (Processing flow of the information processing device 100) Figure 5 is a flowchart showing an example of processing performed by an information processing device 100 according to one embodiment of the present disclosure. First, the control unit 101 initializes the acquired string content holding variable (S101). The acquired string content holding variable is a variable for holding strings generated by the text generation unit 202 in the order of output timing. The acquired string content holding variable is empty in the initialized state.

[0043] Next, the control unit 101 initializes the punctuation position holding variable (S102). The punctuation position holding variable is a variable used to hold the position of punctuation detected by the phrase analysis unit 203. The punctuation position holding variable is 0 in the initial state. Each time the phrase analysis unit 203 detects a punctuation mark, the value of the punctuation position holding variable is updated.

[0044] Next, the text generation unit 202 takes the text data converted by the speech recognition unit 201 as an AI prompt and starts generating text data (sentences) that respond to the user's voice. After that, the processes in steps S103 to S111 are repeated until the end of the sentence.

[0045] In step S104, the text data generation unit 211 obtains the string generated by the text generation unit 202 and generates display data from this string. The display unit 30 then outputs the display data corresponding to the string generated by the text data generation unit 211 to the screen and displays it as a response to the user's voice.

[0046] Furthermore, the segment analysis unit 203 obtains the string generated by the text generation unit 202 and adds that string to the variable that holds the content of the obtained string (S105). Then, the segment analysis unit 203 determines whether the obtained string contains punctuation marks such as ",", ".", "、", "。", "!", "?", etc. (S106).

[0047] If the acquired string contains punctuation (S106, Yes), the segment analysis unit 203 extracts the string from the acquired string content variable to the punctuation mark obtained, using the punctuation mark position variable (S107). The audio data generation unit 210 generates audio data from the string extracted by the segment analysis unit 203.

[0048] The speech synthesis unit 205 performs speech synthesis using the speech data generated by the speech data generation unit 210 and outputs it to the speech output unit 40, thereby speaking as a response to the user's voice (S108). Then, the phrase analysis unit 203 updates the punctuation position holding variable (S109).

[0049] Furthermore, if the acquired string does not contain punctuation (S106, No), the phrase analysis unit 203 determines whether or not the string has been acquired up to the end of the sentence (S110). If the string has been acquired up to the end of the sentence (S110, Yes), it extracts the string from the punctuation position variable to the last string from the acquired string content variable. The audio data generation unit 210 generates audio data from the string extracted by the phrase analysis unit 203.

[0050] The speech synthesis unit 205 then performs speech synthesis using the speech data generated by the speech data generation unit 210 and outputs it to the speech output unit 40, thereby speaking as a response to the user's voice (S108). The phrase analysis unit 203 then updates the punctuation position holding variable (S109) and terminates the process.

[0051] Furthermore, if the string has not been retrieved to the end of the sentence (S110, No), the control unit 101 repeats the process in steps S103 to S111 until the retrieval of the last string of the sentence is complete.

[0052] Figure 6 shows an example of a screen displayed by the display unit 30. The text data generation unit 211 generates display data from the string generated by the text generation unit 202. The display unit 30 outputs the text data generated by the text data generation unit 211 to the screen and displays it as a response to the user's voice. As shown in Figure 6, the first screen 30-1 displays only the first string "こ" generated by the text generation unit 202.

[0053] On the next screen 30-2, the next string generated by the text generation unit 202, "ん", is displayed, and the whole word "こん" is displayed. This process is repeated, and on the final screen 30-n, the whole word "こんにちは" is displayed. In this way, the timing control unit 206 displays the text data on the display unit 30 in accordance with the timing of the text output by the text generation unit 202. Alternatively, the timing control unit 206 may display the text generated by the text generation unit 202 on the display unit 30 one character at a time.

[0054] Figure 7 is a diagram illustrating the changes in the contents of the retrieved string content storage variable and the punctuation position storage variable. Figure 7 is a diagram illustrating the changes in the contents of the retrieved string content storage variable and the punctuation position storage variable for each loop iteration of steps S103 to S111 of the flowchart shown in Figure 5. First, the retrieved string content storage variable and the punctuation position storage variable are initialized.

[0055] In loop 1, the phrase analysis unit 203 obtains the string "お" from the text generation unit 202 and adds "お" to the variable that holds the content of the obtained string. At this time, since the obtained string does not contain punctuation, the phrase analysis unit 203 maintains the variable that holds the punctuation position as is. Then, the display unit 30 displays "お".

[0056] In loop 2, the phrase analysis unit 203 acquires the character string "hayou" from the sentence generation unit 202, and adds "hayou" to the acquired character string content holding variable. At this time, since the acquired character string does not contain any punctuation marks, the phrase analysis unit 203 maintains the punctuation position holding variable as it is. Then, the display unit 30 displays "ohayou".

[0057] In loop 3, the phrase analysis unit 203 acquires the character string ", kyou no" from the sentence generation unit 202, and adds ", kyou no" to the acquired character string content holding variable. The display unit 30 displays "ohayou, kyou no". Here, since the phrase analysis unit 203 finds the punctuation mark ",", the phrase analysis unit 203 extracts the character string "ohayou," up to the punctuation mark "," as a phrase. The audio output unit 40 outputs "ohayou" as the speech content. Since the acquired character string contains a punctuation mark, the phrase analysis unit 203 updates the punctuation position holding variable to "5".

[0058] In loop 4, the phrase analysis unit 203 acquires the character string "tenki wa," from the sentence generation unit 202, and adds "tenki wa," to the acquired character string content holding variable. The display unit 30 displays "ohayou, kyou no tenki wa,". Here, since the phrase analysis unit 203 finds the punctuation mark ",", the phrase analysis unit 203 extracts the character string "kyou no tenki wa," up to the punctuation mark "," as a phrase. The audio output unit 40 outputs "kyou no tenki wa" as the speech content. Since the acquired character string contains a punctuation mark, the phrase analysis unit 203 updates the punctuation position holding variable to "12".

[0059] (Effects) As described above, according to an embodiment of the present disclosure, the following operational effects can be obtained.

[0060] In the information processing apparatus 100, the timing control unit 206 controls the output timing of audio data corresponding to the phrase analyzed by the phrase analysis unit 203. Therefore, the output timing of audio data can be changed according to the phrases of a sentence.

[0061] Furthermore, the timing control unit 206 outputs text data in accordance with the timing of text output by the text generation unit 202, and outputs audio data for each phrase. Therefore, the output timing of audio data and text data can be made different, and audio data and text data can be output in a way that is easy for the user to understand.

[0062] Furthermore, the timing control unit 206 divides the text into phrases based on punctuation and outputs audio data for each divided phrase. Therefore, the audio data can be output in a way that is easy for the user to understand.

[0063] Furthermore, when the phrase analysis unit 203 detects a punctuation mark, the timing control unit 206 controls the output of audio data for the sentence between that point and the punctuation mark immediately before it was detected by the phrase analysis unit 203, thus enabling the output of audio data that is easy for the user to understand.

[0064] Furthermore, the timing control unit 206 outputs text data in accordance with the timing of text output by the text generation unit 202, so that the text data can be output in a way that is easy for the user to read.

[0065] Furthermore, the timing control unit 206 outputs text data one character at a time, making the text data even easier for the user to read.

[0066] [Examples of implementation using software] The functions of the information processing device 100 (hereinafter referred to as "the device") are programs that cause the device to function as a computer, and these programs can be realized by programs that cause each control block of the device (particularly each part included in the control unit 101) to function as a computer.

[0067] In this case, the device includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., memory) as hardware for executing the program. By executing the program using this control device and storage device, the functions described in each of the embodiments are realized.

[0068] The above program may be recorded on one or more computer-readable recording media, not temporary ones. These recording media may or may not be provided by the above device. In the latter case, the program may be supplied to the above device via any wired or wireless transmission medium.

[0069] Furthermore, some or all of the functions of each of the above control blocks can also be realized by logic circuits. For example, an integrated circuit in which logic circuits functioning as each of the above control blocks are formed is also included in the scope of the present invention. In addition, it is also possible to realize the functions of each of the above control blocks by, for example, a quantum computer.

[0070] Furthermore, each process described in the above embodiments may be performed by AI (Artificial Intelligence). In this case, the AI ​​may operate on the control device described above, or it may operate on other devices (for example, an edge computer or a cloud server).

[0071] 〔summary〕 This disclosure includes at least the following aspects:

[0072] (Aspect 1) An information processing device according to Embodiment 1 of the present disclosure includes: a phrase analysis unit that analyzes phrases of text generated as a response to a user's voice; a voice data generation unit that generates voice data for voice output from the text; and a timing control unit that controls the output timing of the voice data corresponding to the phrases analyzed by the phrase analysis unit.

[0073] (Aspect 2) An information processing device according to Embodiment 2 of the present disclosure, in Embodiment 1 described above, further comprises a text data generation unit that generates text data for display from the text, and a timing control unit controls the output of the audio data and the text data such that the output timings of the audio data and the text data are different.

[0074] (Aspect 3) In the information processing device according to Embodiment 3 of the present disclosure, in Embodiment 2 described above, the timing control unit divides the text into phrases based on the punctuation of the text and outputs audio data for each divided phrase.

[0075] (Aspect 4) The information processing device according to aspect 4 of the present disclosure, in aspect 3 above, further comprises a text generation unit that generates the text as a response to the user's voice, a phrase analysis unit that holds the text output by the text generation unit and detects punctuation in the held text, and a timing control unit that, when a punctuation is detected by the phrase analysis unit, controls the output of audio data of the text between the phrase and the punctuation detected immediately before by the phrase analysis unit.

[0076] (Aspect 5) In the information processing apparatus according to aspect 5 of the present disclosure, in aspect 4 described above, the timing control unit outputs the text data in accordance with the timing of the text output by the text generation unit.

[0077] (Aspect 6) In the information processing device according to embodiment 6 of this disclosure, in embodiment 5 described above, the timing control unit outputs the text data one character at a time.

[0078] (Aspect 7) A stationary display device according to aspect 7 of the present disclosure comprises a display unit that displays the text data generated by the information processing device, in any of aspects 2 to 6 described above, and the display unit has a fixing unit that is fixed to a structure.

[0079] (Pattern 8) An information processing system according to aspect 8 of the present disclosure is an information processing system including an information processing device and a display device that is communicably connected to the information processing device, wherein the information processing device includes a phrase analysis unit that analyzes phrases of a text generated as a response to a user's voice, a voice data generation unit that generates voice data for voice output from the text, a text data generation unit that generates text data for display from the text, and a timing control unit that controls the output timing of the voice data and the text data based on the phrases analyzed by the phrase analysis unit, and the display device includes a display unit that displays the text data generated by the information processing device.

[0080] (Aspect 9) An information processing program according to aspect 9 of the present disclosure is an information processing program for causing a computer to function as an information processing device, wherein the computer functions as: a phrase analysis unit that analyzes phrases of text generated as a response to a user's voice; a voice data generation unit that generates voice data for voice output from the text; and a timing control unit that controls the output timing of voice data corresponding to the phrases analyzed by the phrase analysis unit.

[0081] This disclosure is not limited to the embodiments described above, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of this disclosure. [Explanation of Symbols]

[0082] 1 Display device 2 Remote control device 21 Sensing Unit 22 Operation signal receiving unit 30 Display section 40 Audio output section 50 System Data 100 Information Processing Devices 101 Control Unit 102 memory 103 Storage section 104 Communication I / F 201 Speech Recognition Unit 202 Sentence generation section 203 Sentence Analysis Unit 204 Output Data Generation Unit 205 Speech Synthesis Unit 206 Timing Control Unit 210 Audio data generation unit 211 Text Data Generation Unit

Claims

1. A segment analysis unit that analyzes the segments of text generated as a response to the user's voice, A voice data generation unit that generates voice data for voice output from the aforementioned text, An information processing device comprising: a timing control unit that controls the output timing of audio data corresponding to a phrase analyzed by the phrase analysis unit; and a timing control unit that controls the output timing of audio data corresponding to a phrase analyzed by the phrase analysis unit.

2. The information processing device further includes a text data generation unit that generates text data for display from the document, The information processing apparatus according to claim 1, wherein the timing control unit controls the output of the audio data and the text data such that the output timings of the audio data and the text data are different.

3. The information processing apparatus according to claim 2, wherein the timing control unit divides the text into phrases based on punctuation marks in the text and outputs audio data for each divided phrase.

4. The information processing device further includes a text generation unit that generates the text as a response to the user's voice, The phrase analysis unit holds the text output by the text generation unit and detects punctuation in the held text. The information processing apparatus according to claim 3, wherein the timing control unit controls the output of audio data of the sentence between the punctuation mark detected by the phrase analysis unit and the punctuation mark detected immediately before by the phrase analysis unit when the phrase analysis unit detects a punctuation mark.

5. The information processing apparatus according to claim 4, wherein the timing control unit outputs the text data in accordance with the timing of the text output by the text generation unit.

6. The information processing apparatus according to claim 5, wherein the timing control unit outputs the text data one character at a time.

7. The device comprises a display unit that displays the text data generated by the information processing device according to any one of claims 2 to 4, The display unit has a fixing part for fixing to the structure. Display device.

8. An information processing system including an information processing device and a display device that is communicatively connected to the information processing device, The aforementioned information processing device is A segment analysis unit that analyzes the segments of text generated as a response to the user's voice, A voice data generation unit that generates voice data for voice output from the aforementioned text, A text data generation unit that generates text data for display from the aforementioned text, The system includes a timing control unit that controls the output timing of audio data and text data based on the phrases analyzed by the phrase analysis unit, The display device includes a display unit that displays the text data generated by the information processing device, Information processing system.

9. An information processing program that enables a computer to function as an information processing device, The aforementioned computer, A segment analysis unit that analyzes the segments of text generated as a response to the user's voice, A voice data generation unit that generates voice data for voice output from the aforementioned text, A timing control unit controls the output timing of audio data corresponding to the phrases analyzed by the phrase analysis unit, An information processing program that functions as such.

Citation Information

Patent Citations

  • Question-answering method and question-answering device

    JP2005092271A