Display device, display method, program, and recording medium

The integration of voice and image recognition in display devices ensures accurate voice commands and automatic language adaptation, addressing noise interference and manual setting issues.

JP2025119433APending Publication Date: 2025-08-14SHARP KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024014316
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-01
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

Existing display devices struggle with accurate voice recognition due to surrounding noise and require manual language changes, which is time-consuming and often results in incorrect settings.

Method used

Incorporating a microphone for voice input, a camera for image capture, and a voice recognition unit to identify the speaker and recognize voice commands, allowing for automatic language adaptation based on the user's voice and image analysis.

Benefits of technology

Enables accurate operation by voice input, preventing misrecognition from ambient noise and allowing seamless language switching without user intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025119433000001_ABST
    Figure 2025119433000001_ABST
Patent Text Reader

Abstract

To provide a display device that enables appropriate operations when operated by voice input, for example.SOLUTION: A display device is provided, comprising a voice input unit for receiving voice input, an image capturing unit for capturing an image, an identification unit for identifying a speaker on the basis of the image, and a voice recognition unit configured to recognize the voice when the voice was input through the voice input unit and provided that the voice was uttered by the speaker.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a display device and the like. [Background technology]

[0002] For example, as shown in Patent Document 1, a display device is known that can change the setting information of the display device by recognizing the user's voice. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2014-66791 Summary of the Invention [Problem to be solved by the invention]

[0004] An object of the present disclosure is to provide a display device or the like that is capable of performing appropriate operations when performing operations by voice input, for example. [Means for solving the problem]

[0005] The display device of the present disclosure includes an audio input unit that inputs audio, an imaging unit that captures an image, an identification unit that identifies a speaker based on the image, and a voice recognition unit that recognizes the audio when the audio is input from the audio input unit and the audio is the audio uttered by the speaker.

[0006] The display method of the present disclosure includes a voice input step of inputting voice, a photographing step of photographing an image, an identification step of identifying a speaker based on the image, and a voice recognition step of recognizing the voice when the voice is input from the voice input step and the voice is the voice uttered by the speaker.

[0007] The program disclosed herein provides a computer with a voice input function for inputting voice, a photographing function for photographing an image, an identification function for identifying a speaker based on the image, and a voice recognition function for recognizing voice when the voice is input from the voice input function and the voice is produced by the speaker. [Effects of the Invention]

[0008] According to the present invention, for example, it is possible to perform an appropriate operation when performing an operation by voice input. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram illustrating an overview of a display system according to a first embodiment. [Figure 2] FIG. 2 is a diagram illustrating the hardware configuration of the display device according to the first embodiment. [Figure 3] FIG. 2 is a diagram illustrating the configuration of a broadcast control unit in the first embodiment. [Figure 4] FIG. 2 is a diagram illustrating a software configuration according to the first embodiment. [Figure 5] FIG. 3 is a diagram illustrating an example of a setting information DB in the first embodiment. [Figure 6] FIG. 2 is a diagram illustrating a processing flow in the first embodiment. [Figure 7] FIG. 2 is a diagram illustrating an example of operation in the first embodiment. [Figure 8] FIG. 10 is a diagram illustrating a software configuration according to a second embodiment. [Figure 9] 10A is a diagram illustrating an example of a language DB in the second embodiment, and FIG. 10B is a diagram illustrating an example of a language DB. [Figure 10] FIG. 10 is a diagram illustrating a processing flow in the second embodiment. [Figure 11] FIG. 10 is a diagram illustrating an example of operation in the second embodiment. [Figure 12] FIG. 10 is a diagram illustrating a processing flow in the third embodiment. [Figure 13]FIG. 10 is a diagram illustrating an example of operation in the third embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] In general, a display device has a control unit that executes various settings of the display device. The control unit stores information for executing the various settings in a storage unit as setting information. In order to switch the setting information, for example, a display device having a function of recognizing a voice signal uttered by a user and switching the setting information based on the voice recognition is known.

[0011] However, when rewriting setting information using voice recognition, there was a problem that the voice recognition would react to surrounding noise and speaking voices, resulting in incorrect recognition.

[0012] Furthermore, with a typical display device, if you wanted to change the language of the OSD or other display on the display device depending on the user, you had to change the language from the settings each time, which was time-consuming for the user and made it impossible to change the language unless you knew how to set it up in the first place.

[0013] A display device that can solve one or more of the above-mentioned problems will be described below in the following embodiments with reference to the drawings. Note that the following embodiments are merely examples of the invention described in the claims, and the technical scope of the present invention is not limited to the description of the following embodiments.

[0014] [1. First embodiment] The first embodiment will be described below.

[0015] [1.1 Entire display device] 1 is a diagram showing the entire display system 1. The display system 1 includes a display device 10, and may also include an operation device 16 capable of operating the display device 10, and a terminal device 18.

[0016] The display device 10 is a device capable of displaying content, and in this embodiment, for example, is a television capable of receiving broadcast waves and displaying programs. The display device 10 may also be a display capable of displaying video input from an external device. The display device 10 may also be, for example, a projector that projects a display screen onto a screen.

[0017] The display device 10 may have a device for acquiring information for identifying a user who is viewing the display device 10. For example, the display device 10 may have a device for acquiring the voice of the user who is viewing the display device 10, and may have, for example, a microphone 12 as a voice input device. The microphone 12 may be a single microphone or a microphone array consisting of multiple microphones.

[0018] The display device 10 may also have a camera 14 for acquiring an image of a user viewing the display device 10. The camera 14 can, for example, capture an image of an environment including a user viewing the display device 10. One or more cameras 14 may be provided.

[0019] Here, the content may be anything that can be displayed on the display device 10. For example, the content may be a program that is received from terrestrial / BS / CS broadcasting, demodulated, and displayed, a video selected by the user on a video distribution site, or a video input from an external device via HDMI (registered trademark), D-SUB, or the like.

[0020] Furthermore, the display device 10 is usually operated by an operation device 16. The operation device 16 is a device that operates the display device 10 and is connected to the display device 10 wirelessly, for example, via infrared or Bluetooth (registered trademark). The operation device 16 is generally a device that outputs an operation signal corresponding to each manufacturer of the display device 10, and is called a "remote control."

[0021] The terminal device 18 is a device including an information processing device such as a smartphone or tablet owned by a user. For example, by installing an application for operating the display device 10 in the terminal device 18, it becomes possible to provide the user with operations similar to those of the above-mentioned operating device. The terminal device 18 may be connected to the display device 10 via a wireless LAN or directly via short-range wireless communication or the like.

[0022] It is sufficient to have at least one of the operation device 16 and the terminal device 18. Furthermore, the terminal device 18 may be a smartphone or the like that is normally used by the user.

[0023] [1.2 Hardware Configuration] FIG. 2 is a diagram showing the hardware configuration of the display device 10. As shown in FIG.

[0024] The control unit 100 controls the entire display device 10. The control unit 100 realizes various functions by reading and executing various programs stored in a storage device (for example, storage 110 or ROM 120). The control unit 100 may be realized by one or more control devices / arithmetic units (CPUs (Central Processing Units), SoCs (System on a Chip)). The control unit 100 may also be configured by a control circuit.

[0025] The operation control unit 102 receives operations from the user, gives operation instructions to each functional unit, and notifies the control unit 100 of an operation signal corresponding to the received operation. For example, the operation control unit 102 may be a unit receiving an operation signal from a remote control, or may be an operation switch provided on the main body of the display device 10.

[0026] The storage 110 is a non-volatile storage device capable of storing programs and data. For example, the storage 110 may be configured as a storage device such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive). The storage 110 may also be configured as an externally connectable USB memory. The storage 110 may also be, for example, a storage area on the cloud.

[0027] The ROM 120 is a non-volatile memory that can retain programs and data even when the power is turned off.

[0028] The RAM 130 is a main memory that is mainly used when the control unit 100 executes processing. The RAM 130 is a rewritable memory that temporarily stores programs read from the storage 110 or the ROM 120, and data including execution results.

[0029] The broadcast control unit 140 receives broadcast waves transmitted by the broadcast station selected by the user, decodes video and text from the broadcast waves, and outputs them to the display unit 150, and decodes audio from the broadcast waves and outputs them to the audio output unit 165. The configuration of the broadcast control unit 140 will be described in more detail below.

[0030] The display unit 150 is a display device capable of displaying video of a received program and various information. The display unit 150 may be, for example, a device capable of displaying video, such as a liquid crystal display (LCD) or an organic electroluminescence (EL) display. The display unit 150 also includes an interface to which a display device can be connected. For example, the display unit 150 may be configured as an external display device connected via an HDMI (registered trademark) (High-Definition Multimedia Interface), a DVI (Digital Visual Interface), or a Display Port. The display unit 150 may also be, for example, a projection device such as a projector.

[0031] The image capturing unit 155 is an image capturing device, such as a camera, that captures images of the surroundings where the display device 10 is installed. The image capturing unit 155 may be composed of one or more image capturing devices. The image capturing unit 155 outputs the captured image as an image signal. The image capturing unit 155 may also output one or more images as a continuous video.

[0032] The audio input unit 160 is an input device that can input sounds around the area where the display device 10 is installed, such as a microphone. The audio input unit 160 may also be composed of multiple input devices (for example, microphones). The audio input unit 160 is mainly used to input the voice of the user who is viewing the display device, but it can also input general sounds such as environmental sounds.

[0033] The audio output unit 165 outputs audio included in the content. The audio output unit 165 may be a device such as a speaker or headphones. The audio output unit 165 only needs to output sound, and can output general sounds such as music, environmental sounds, etc.

[0034] The communication unit 170 is a communication interface for communicating with other devices. For example, the communication unit 170 may be a network interface connectable to a wireless LAN or a network interface connectable to Ethernet (registered trademark) via a wired connection. The communication unit 170 may also be a communication device connectable to a mobile communication network such as LTE / 4G / 5G / 6G.

[0035] 2 may be configured by an external device connected to the display device 10. For example, the audio input unit 160 may be a microphone connected via USB. Furthermore, the one or more components shown in FIG. 2 may be realized by a terminal device connected via wireless communication such as Bluetooth (registered trademark) or a terminal device connected via a network. For example, a display device, audio input / output device, and image capture device included in a smartphone, tablet, smart speaker, or wearable terminal device connected to the display device 10 may be used.

[0036] The configuration of FIG. 2 may include any necessary components in the embodiment.

[0037] (Broadcast Control Unit) A simple configuration of the broadcast control unit 140 will be described with reference to Fig. 3. For example, broadcast data of digital broadcasts such as terrestrial, BS, or CS is acquired by the tuner unit 200. Then, the OFDM demodulation unit 210 performs OFDM demodulation on the broadcast data, performs error correction, etc., and then outputs TS packets.

[0038] The demultiplexer 220 demultiplexes the TS packets and outputs the video packets to the video decoding unit 230 and the audio packets to the audio decoding unit 260. The demultiplexer 220 also outputs data packets containing information about subtitles (for example, subtitle TS packets) to the subtitle decoding unit 240.

[0039] The video decoding unit 230 decodes the video data from the input video packets and outputs the video data to the image processing unit 250. The subtitle decoding unit 240 decodes the subtitle data from the input data packets and outputs the subtitle data to the image processing unit 250. Here, the subtitle data includes information related to the subtitles. For example, the subtitle data in this embodiment may include subtitle management data included in the subtitle PES data and data on the subtitle text to be displayed.

[0040] The image processing unit 250 outputs to the display unit 150 an image in which subtitle data is superimposed on the video data as needed.

[0041] Furthermore, the audio decoding unit 260 decodes audio data from the input audio packets and outputs the audio to the audio output unit 165 .

[0042] 3 is a schematic illustration of a general broadcast control unit 140, but other configurations are also possible. For example, the video decoding unit 230 and the audio decoding unit 260 may be the same decoding unit.

[0043] [1.3 Software Configuration] The software configuration will be described with reference to FIG.

[0044] The control unit 100 executes a program stored in, for example, a memory unit (storage 110, ROM 120, RAM 130), thereby realizing each function.

[0045] The control unit 100 has a setting unit that performs settings for various operations in the display device 10. In Fig. 4, as an example, a language setting unit 1010 that performs settings related to language is realized.

[0046] The language setting unit 1010 performs settings related to the language in the display device 10. Here, the language may indicate a type of language, such as "Japanese" or "English." Alternatively, the language may indicate a type of characters to be displayed (for example, "traditional characters," "simplified characters," etc.).

[0047] Furthermore, the language settings include settings regarding the language displayed by the display device 10 (output language) and the language input to the display device 10 (input language). For example, the control unit 100 displays the contents such as the user interface on the screen to be displayed and the menu display in the set language based on the language setting. Furthermore, the control unit 100 performs voice recognition from the user based on the set language.

[0048] The language setting unit 1010 may execute a video recognition unit 1012 , a voice recognition unit 1014 , and a language setting reflection unit 1016 .

[0049] The video recognition unit 1012 performs recognition based on the image (or video) captured by the image capture unit 155. For example, the video recognition unit 1012 determines one or more users included in the image based on the image. A user refers to a person who can be recognized from within the image.

[0050] In this specification, a "user" refers to a person who can be recognized from an image captured by the video recognition unit 1012. A "viewer" refers to a user who is viewing the content of the display device 10 (for example, a person whose line of sight is directed toward the display device 10, a person whose face is directed toward the display device 10, etc.). A "speaker" refers to a user who is making a sound.

[0051] In this case, the video recognition unit 1012 may recognize a user, for example, when it recognizes a face or a human shape. Alternatively, the video recognition unit 1012 may recognize a person using, for example, pattern recognition or machine learning. Alternatively, the video recognition unit 1012 may recognize a user, the user's orientation, or the user's movement using any other known method. This enables the video recognition unit 1012 to recognize, for example, the presence or absence of a user, the user's position, the orientation of the user's face, the user's line of sight, the user's mouth movement, the user's body orientation, etc.

[0052] The speech recognition unit 1014 recognizes speech uttered by the user. As a method for recognizing speech, the speech recognition unit 1014 can use methods such as KWS (Keyword Spotting) speech recognition, DNN-HMM speech recognition, end-to-end speech recognition, etc. The speech recognition unit 1014 may also use any other known method. In this way, the speech recognition unit 1014 can recognize speech uttered by the user, and can recognize, for example, language or commands.

[0053] The video recognition unit 1012 and the voice recognition unit 1014 may use an external recognition service. For example, the control unit 100 may transmit video to an external service and use the external service to recognize the user. The control unit 100 may also transmit audio to an external service and use the external service to recognize language or commands based on the audio uttered by the user.

[0054] The display mode control unit 1020 controls switching of a mode for displaying an input signal on the display device 10. For example, the mode for displaying an input signal may be a High-Definition Multimedia Interface (HDMI) mode, a Digital Visual Interface (DVI) mode, an RCA terminal mode, a tuner mode, or the like.

[0055] The OSD control unit 1030 controls the OSD (on-screen display) function of the display device. For example, the OSD may be a display screen that displays setting information such as the subtitle language, the channel currently being received by the tuner, and the volume.

[0056] The command execution unit 1040 executes commands input by the user. Here, a command is an operation or instruction given by the user to cause the display device 10 to perform an operation or process. The command may be input by voice by the user, or may be input from an operation device (e.g., a remote control device) or another connected terminal device (e.g., an information processing device such as a smartphone or tablet). For example, when the user inputs a command such as "Turn up the volume," "Change the channel," or "Turn off the power," the command execution unit 1040 interprets the command and executes the corresponding process. For example, the command execution unit 1040 can increase the volume by executing a command such as "Vol" or "UP."

[0057] The storage 110 reserves an area for a setting information DB storage area 1102 that stores a setting information DB. The setting information DB is a database (DB) that stores various settings related to the display device 10. For example, the control unit 100 executes settings for the display device 10 based on the setting information stored in the setting information DB.

[0058] An example of the setting information stored in the setting information DB will now be described with reference to Fig. 5(a). The setting information stores a setting item (e.g., "language") uniquely assigned to each setting information of the display device, and an arbitrarily input or selected setting value (e.g., "Japanese").

[0059] In this embodiment, the setting items and setting values are stored in Japanese, but they may be stored in other languages. Furthermore, depending on the "language" of the setting items, the contents of the setting values and the corresponding setting values may also be stored in that language.

[0060] The above-described configuration may be provided as needed. For example, the language setting reflecting unit 1016 may not be provided when the language setting is not changed. Furthermore, the video recognition unit 1012 and the voice recognition unit 1014 may be independent from the language setting unit 1010.

[0061] [1.4 Processing flow] The processing flow in this embodiment will be described below with reference to Fig. 6. Note that the following processing will be described as being executed by the control unit 100, but each of the components described in Figs. 2 to 4 may execute the processing of each step.

[0062] First, the control unit 100 determines whether or not a voice is being input from the voice input unit 160 (S100). The control unit 100 temporarily stores the input voice as voice data in a storage unit such as the storage 110 or the RAM 130.

[0063] Next, when a voice is input, the control unit 100 identifies the speaker from the image (S100; Yes→S102).

[0064] The control unit 100 may identify the speaker by, for example, analyzing the movement of the user's lips or the movement of the user's mouth or facial muscles.

[0065] Next, the control unit 100 determines whether the speaker is a viewer (S104). Here, the control unit 100 identifies the viewer by, for example, recognizing that the direction of the user's gaze is directed toward the display device 10, or recognizing that the direction of the user's face or the direction of the user's body is directed toward the display device 10. Furthermore, the control unit 100 may exclude users with their eyes closed from the viewer. In other words, when the user's face is directed toward the display device 10, the control unit 100 recognizes that the user is viewing content.

[0066] Then, when the speaker is a viewer, that is, when a person watching the display device 10 is uttering a voice, the control unit 100 performs voice recognition (S108). The control unit 100 performs voice recognition from the voice data temporarily stored in S100, and recognizes the language uttered by the user.

[0067] Furthermore, the control unit 100 may execute a process according to the result of the voice recognition (S110). For example, the control unit 100 (command execution unit 1040) interprets the command resulting from the voice recognition and executes the corresponding process.

[0068] Note that the above-described operational flow is an example, and the steps may be executed in a different order. For example, the control unit 100 may first recognize television viewers and then identify the speaker from among the viewers. For example, the control unit 100 may first execute S104 to recognize the viewer from among users. Then, the control unit 100 may then execute S102 and execute speech recognition only if the viewer is the speaker.

[0069] Furthermore, although the control unit 100 detects a voice input in S100 and identifies the speaker from the image in S102, the timing of the speech may be determined based on the image. For example, the control unit 100 determines the timing of the user's speech from the image. Then, when the control unit 100 determines that the user has spoken, the control unit 100 may input the speech and perform speech recognition. In this way, the control unit 100 can determine the timing of the user's speech of the device and perform speech recognition on the speech acquired at the timing of the speech.

[0070] [1.5 Example of operation] FIG. 7 is a diagram schematically illustrating the operation of the display device 10 in this embodiment.

[0071] Referring to Fig. 7(a), a camera 14 is arranged on the display device 10. Also, the display device 10 in Fig. 7(a) displays the number of the currently displayed channel as "6ch" in area M12.

[0072] Here, we will explain the case where there are two users, Usr1 as the first user and Usr2 as the second user, in front of the display device 10. Usr1 is viewing content and is facing the direction of the display device 10. Usr2 is not viewing content and is not facing the direction of the display device 10. Note that although Usr1 appears to have his back to the display device 10 in FIG. 7(a), he is considered to be facing the direction of the display device 10 for the sake of simplicity, and this will also apply to the following figures.

[0073] Here, Usr1 utters "Change to channel 3," and Usr2 utters "Change to channel 1" at a different timing than Usr1. Note that Usr2's utterance timing can be either before or after Usr1.

[0074] At this time, the control unit 100 determines that there has been voice input, and the voices uttered by Usr1 and Usr2 are temporarily stored as voice data (S100 in FIG. 6). Here, the control unit 100 can identify the speakers as Usr1 and Usr2 from the image captured by the camera 14 (S102). However, the control unit 100 can further identify the viewer as Usr1 from the image.

[0075] As a result, the control unit 100 performs speech recognition of the speech uttered by the speaker and viewer Usr1 to recognize the language.

[0076] 7(b) is a diagram showing a schematic diagram of the operation after the speech of Usr1 is recognized. The control unit 100 identifies Usr1 as the speaker and performs speech recognition on the speech uttered by Usr1.

[0077] Then, as a result of the voice recognition, the control unit 100 changes the channel of the currently received broadcast wave to "Channel 3." In Fig. 7(b), "3ch" indicating the changed channel is displayed in area M14.

[0078] [1.6 Effects, etc.] In this way, according to this embodiment, the control unit 100 can recognize only the voice of the user watching the display device 10.

[0079] The display device 10 can prevent misrecognition due to, for example, ambient noise or the voice of a user who is not viewing, and can switch setting information by recognizing only the voice from the user who is viewing.

[0080] [2. Second Embodiment] Next, a second embodiment will be described. The second embodiment is an embodiment in which the language setting is reflected in the display device 10 based on the language of the speaker.

[0081] [2.1 Configuration] In the second embodiment, explanations of the same hardware and software configurations as those in the first embodiment will be omitted, and the explanation will focus on the differences from the first embodiment.

[0082] Fig. 8 of the second embodiment replaces Fig. 4 of the first embodiment. Fig. 8 includes a language setting reflecting unit 1016, and secures a storage area for a language DB storage area 1104. Note that the same components as those in Fig. 4 are denoted by the same reference numerals, and their explanations will be omitted.

[0083] The language setting reflecting unit 1016 reflects the language in the setting information DB 1102. For example, when the language is switched by a user operation, the language setting reflecting unit 1016 reflects the language selected by the user in the language of the setting information at that timing.

[0084] In this embodiment, the storage 110 further reserves an area for a language DB 1104 that stores a language DB. The language DB is a database that stores the languages of various countries that correspond to commands and characters to be displayed.

[0085] An example of a language DB will now be described with reference to FIG. 9. The language DB is a database that stores, for example, words corresponding to the languages of various countries. As shown in FIG. 9(a), the language DB stores commands and their corresponding expressions in each language. For example, in FIG. 9(a), words uttered by the user are stored in correspondence with the command "vol" for controlling the volume. Here, keywords such as "volume" in Japanese and "Volume" in English are stored in correspondence with the command "Vol."

[0086] 9(b) is a diagram showing the language DB in a state where only language correspondences are stored. For example, the control unit 100 refers to the language DB when displaying the item names of the OSD display in the set language or when displaying the user interface in the set language.

[0087] [2.2 Processing flow] Fig. 10 is a diagram illustrating the flow of processing in this embodiment. Fig. 10 replaces Fig. 6 of the first embodiment. Instead of S110 in Fig. 6, S200 and S202 are executed.

[0088] In S108, the control unit 100 executes speech recognition and also recognizes the language of the speaker. Note that the speaker here refers to a speaker among the viewers. That is, the target of speech recognition by the control unit 100 is speech uttered by the speaker and the viewers.

[0089] There are various possible methods for the control unit 100 to recognize the language of the speaker from the results of speech recognition. For example, the control unit 100 may recognize the language of the speech uttered by the user from the results of converting speech data into text.

[0090] When the language of the speaker does not match the language set as a setting item, the control unit 100 reflects the setting item in the language of the speaker (S202), which enables the control unit 100 to switch the language displayed on the OSD, for example.

[0091] Furthermore, by switching the language, the control unit 100 can switch the command to be input by voice. For example, the command execution unit 1040 reads the set language stored in the setting information DB, and determines the command corresponding to the set language by referring to the language DB (for example, FIG. 9(a)).

[0092] Furthermore, by changing the set language, the control unit 100 may switch the screen of the user interface including, for example, the OSD display etc. For example, the control unit 100 can refer to a language DB (for example, FIG. 9(b)) and display the OSD display and the user interface display using words and phrases in the set language.

[0093] Furthermore, the control unit 100 may switch the language of the content to be displayed by changing the set language. For example, the control unit 100 may output subtitles in "Japanese" when the set language is "Japanese," and may output subtitles in "English" when the set language is "English."

[0094] 6 is omitted in this embodiment, the control unit 100 may execute S110 after S202. For example, the control unit 100 may execute a command input by voice after switching to the set language.

[0095] In addition, in this embodiment, the speaker's language is recognized after performing voice recognition, but the language may also be recognized based on voice data. For example, the control unit 100 may sample the voice data, identify language-specific features, and recognize the speaker's language by using a language identification model. The control unit 100 may also identify the speaker's language by analyzing the frequency characteristics and acoustic features of the voice. In this case, the control unit 100 may execute S200 and then execute voice recognition after S202.

[0096] [2.3 Example of operation] 11A and 11B are diagrams illustrating an example of the operation of this embodiment. In Fig. 11A, the language setting of the display device 10 is initially set to "Japanese." Therefore, "volume" and language are displayed in area M22 of the display device 10. At this time, the display device 10 accepts voice input from the user in Japanese.

[0097] Here, Usr1 utters the voice in English, "turn up the volume," and Usr2 utters the voice, "Turn down the volume." At this time, when the control unit 100 refers to the image captured by the camera 14, it recognizes that Usr1 is the speaker. Similarly, the control unit 100 also recognizes that Usr1 is a viewer, and therefore recognizes the voice uttered by Usr1.

[0098] As a result, the control unit 100 recognizes the language of the voice uttered by Usr1 as "English" and switches the set language to "English." Furthermore, as shown in FIG. 11(b), the control unit 100 can switch the display from "Volume" to "Volume" in English. Furthermore, the control unit 100 performs processing to increase the volume as a command corresponding to the voice uttered by Usr1.

[0099] [2.4 Effects] As described above, according to this embodiment, the display device 10 can switch the language setting based on the voice uttered by the user.

[0100] 3. Third Embodiment Next, a third embodiment will be described. In the third embodiment, when a speaker holds the operation device 16 or the terminal device 18, the speech of the speaker is recognized.

[0101] In the third embodiment, the description of the same hardware and software configurations as those in the first embodiment will be omitted, and the description will focus on the differences from the first embodiment.

[0102] Fig. 12 is a diagram for explaining the flow of processing in this embodiment, which replaces Fig. 6 in the first embodiment.

[0103] After receiving a voice input, the control unit 100 identifies the speaker from the image (S100; Yes→S102). Here, when the display device 10 identifies the speaker, the control unit 100 determines whether the display device 10 is being operated from the terminal device 18 (S300).

[0104] Here, when the display device 10 is being operated from the terminal device 18 (S300; Yes), the control unit 100 determines whether the speaker is holding the terminal device 18 in his / her hand (S302). For example, it may be determined that the speaker is holding the terminal device 18 by recognizing an image captured by the image capturing unit 155. Here, the state in which the control unit 100 recognizes that the speaker (user) is holding the terminal device 18 may be, for example, when the smartphone is held in the hand, supported by the body, or near the user.

[0105] Also, when the display device 10 is not operated from the terminal device 18 (S300; No), in this case the display device 10 is operated by the operation device 16. Therefore, the control unit 100 determines whether the operation device 16 is in hand (S304).

[0106] Here, if the speaker is holding the terminal device 18 (S302; Yes) or the operating device 16 (S304; Yes), the control unit 100 performs voice recognition and executes processing based on the voice recognition (S108, S110).

[0107] Furthermore, when the speaker does not have the terminal device 18 (S302; No) or the operation device 16 (S304; No), the control unit 100 does not recognize the speaker's voice.

[0108] 13 is a diagram illustrating an example of the operation of this embodiment. For example, Usr1 says, "Change to channel 3," and Usr2 says, "Change to channel 1."

[0109] Here, the display device 10 recognizes the speaker holding the terminal device 18 based on the image captured by the camera 14. For example, in FIG. 13(a), Usr1 is holding the terminal device 18. Therefore, the display device 10 recognizes Usr1 as the speaker and performs voice recognition. Therefore, while "6ch" is selected and displayed as shown in area R32 in FIG. 13(a), in FIG. 13(b) after voice recognition, it can be seen that "3ch" is selected and displayed as shown in area R34.

[0110] In this way, according to this embodiment, it is possible to recognize a speaker based on a camera image. In this case, in this embodiment, for example, when the display device 10 is operated from the terminal device 18 using a smartphone application or the like, the criterion for recognizing voice can be the speaker who holds the terminal device 18.

[0111] [4. Fourth Embodiment] The fourth embodiment will be described below. In the above-described embodiment, the case where the language is switched is described, but in this embodiment, a case where something other than the type of language is switched is described.

[0112] The fourth embodiment has the same hardware and software configurations as the first and second embodiments, and the following description will focus on the differences from the first and second embodiments.

[0113] For example, in the second embodiment, the language DB storage area 1104 is described taking the language of a country as an example of words to be replaced. In this embodiment, for example, words according to age groups may be stored.

[0114] For example, the control unit 100 recognizes the language of the speaker in S108 and S200 of Fig. 10, but may also recognize the age (generation) of the speaker. Then, the control unit 100 sets the setting item "age (generation)" to the recognized age of the speaker.

[0115] For example, when the control unit 100 recognizes that the speaker is 10 years old (child), it displays, for example, "loudness" for the volume. On the other hand, when the control unit 100 recognizes that the speaker is 30 years old (adult), it displays, for example, "volume" for the volume.

[0116] Furthermore, the control unit 100 may change the display mode of the subtitle data to be displayed depending on the age of the speaker, for example. For example, when the control unit 100 recognizes that the speaker's age is 70, the control unit 100 may increase the amount of subtitle data.

[0117] In this way, the display device 10 of this embodiment can appropriately select and display a display according to not only the user's language but also the user's attributes such as age and generation.

[0118] [5. Modifications] The present disclosure is not limited to the above-described embodiments, and various modifications are possible. In other words, embodiments obtained by combining technical means that are appropriately modified within the scope of the present disclosure are also included in the technical scope.

[0119] Although the above-mentioned embodiments are described separately for convenience of explanation, they can be combined to the extent possible. Furthermore, the present invention intends to obtain rights to any of the technologies described in the specification through amendments or divisional applications, etc.

[0120] Although the above-mentioned database (DB) is stored in a storage device of the display device, it may be an external DB. For example, the database may be stored in a cloud. The database may also be provided by an external service.

[0121] In addition, the programs that run on each device in each embodiment are programs that control the CPU, etc. (programs that make a computer function) so as to realize the functions of the above-described embodiments. Information handled by these devices is temporarily stored in a temporary storage device (e.g., RAM) during processing, and then stored in various ROMs and HDDs, and is read, modified, and written by the CPU as needed.

[0122] Here, the recording medium for storing the program may be any of semiconductor media (e.g., ROM, non-volatile memory card, etc.), optical recording media / magneto-optical recording media (e.g., DVD (Digital Versatile Disc), CD (Compact Disc), BD (Blu-ray (registered trademark) Disc), etc.), magnetic recording media (e.g., magnetic tape, flexible disk, etc.), etc.

[0123] Furthermore, when distributing the program on the market, the program can be stored in a portable recording medium and distributed, or transferred to a server computer connected via a network such as the Internet. In this case, the storage device of the server device is also included in the present disclosure.

[0124] Furthermore, the above-mentioned data may not be stored within the device, but may be stored in an external device and called up as needed. For example, the data may be stored in a network attached storage (NAS) or on the cloud.

[0125] The scope of the present disclosure is not limited to the configurations explicitly described in the specification, but also includes combinations of the technologies disclosed in the specification. The configurations of the present disclosure for which a patent is sought are set forth in the appended claims, but it is not intended to exclude them from the technical scope on the grounds that they are not set forth in the claims.

[0126] Furthermore, in the above-mentioned specification, the statements "in the case of" and "when" are given as examples and are not intended to limit the configuration to the described contents. The disclosure also includes configurations that are not in these cases or situations, even if they would be obvious to a person skilled in the art, and the applicant intends to obtain rights to them.

[0127] Furthermore, the processes and data flows described in the specification are not limited to the order in which they are described. For example, the patent also discloses configurations in which some processes are deleted or the order is changed, and the patent holder intends to obtain the rights to such configurations.

[0128] Furthermore, although the functions described in the embodiments are executed by each device, they may be realized by one device or may further utilize an external server.

[0129] Furthermore, each functional block or feature of the device used in the above-described embodiments may be implemented or performed by an electrical circuit, for example, an integrated circuit or multiple integrated circuits. The electrical circuit designed to perform the functions described herein may include a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or a combination thereof. The general-purpose processor may be a microprocessor, or a conventional processor, controller, microcontroller, or state machine. The electrical circuit may be composed of digital circuits or analog circuits. Furthermore, as advances in semiconductor technology emerge that replace current integrated circuits, one or more aspects of the present disclosure may also utilize new integrated circuits based on that technology. [Explanation of symbols]

[0130] 10 Display device 12. Microphone 14 Camera 100 control section 102 Operation control section 110 Storage 120 ROM 130 RAM 140 Broadcast Control Section 150 Display section 155 Photography Department 160 Audio input section 165 Audio output section 170 Communications Department

Claims

1. a voice input unit for inputting voice; an imaging unit that captures an image; an identification unit that identifies a speaker based on the image; a speech recognition unit that recognizes the speech when the speech is input from the speech input unit and the speech is a speech uttered by the speaker; A display device comprising:

2. Further, a setting unit for setting a language for display is provided, the speech recognition unit recognizes the language of the speaker from the speech uttered by the speaker, The display device according to claim 1 , wherein the setting unit sets the language related to the display to the language of the speaker.

3. a determination unit that determines one or more users based on the captured image; The display device according to claim 1 , wherein the identification unit recognizes the line of sight and / or the body direction of the user based on the image, and identifies the speaker from among the users.

4. The display device according to claim 3 , wherein the identification unit identifies, as the speaker, a user whose face is facing the display device.

5. a determination unit that determines one or more users based on the captured image; The display device according to claim 1 , wherein the identification unit recognizes a user who is holding a terminal device or an operation device from among the one or more users based on the image, and identifies the user as the speaker.

6. a voice input step of inputting voice; a capturing step of capturing an image; an identifying step of identifying a speaker based on the image; a speech recognition step of recognizing the speech when the speech is input in the speech input step and the speech is a speech uttered by the speaker; Display methods including.

7. On the computer, A voice input function for inputting voice, A photographing function for photographing an image; an identification function for identifying a speaker based on the image; a voice recognition function that recognizes the voice when the voice is input from the voice input function and the voice is a voice uttered by the speaker; A program to achieve this.

8. A recording medium on which the program of claim 7 is recorded.

Citation Information

Patent Citations

  • Display device

    JP2014066791A