Display device, display method, program, recording medium, and subtitle generating method
The display device adjusts subtitle display modes based on user attributes, particularly age, through facial or voice recognition, ensuring tailored visibility for individual users and groups.
Patent Information
- Application Number
- JP2023214211
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-19
- Publication Date
- 2025-07-01
AI Technical Summary
Existing display devices fail to adjust subtitle display modes based on the attributes of individual users, particularly neglecting the age of viewers and the relationship between content and viewing users.
A display device that acquires video and subtitle data, recognizes users through facial or voice recognition, determines appropriate subtitle display modes based on user attributes, and superimposes subtitles accordingly.
Enables subtitles to be displayed in modes tailored to user attributes, ensuring appropriate visibility and engagement for diverse audiences, especially when multiple users are present.
Smart Images

Figure 2025097797000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a display device and the like.
Background Art
[0002] For example, as shown in Patent Document 1, a display control device including means for adjusting the character size for each user and generating character information is known.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] One object of the present disclosure is to provide, for example, a display device or the like capable of displaying subtitles in an appropriate display mode according to the attributes of the user who views the content.
Means for Solving the Problems
[0005] The display device of the present disclosure includes an acquisition unit that acquires video data and subtitle data from content, a recognition unit that recognizes a user who views the content, a determination unit that determines the attributes of the recognized user, a determination unit that determines a display mode of subtitles to be displayed based on the subtitle data based on the determined attributes of the user, and a display unit that superimposes and displays subtitles based on the subtitle data on the video data based on the determined display mode.
[0006] The display method of the present disclosure includes an acquisition step of acquiring video data and subtitle data from content, a recognition step of recognizing a user who views the content, a determination step of determining the attributes of the recognized user, a determination step of determining a display mode of subtitles to be displayed based on the subtitle data according to the determined attributes of the user, and a display step of superimposing and displaying subtitles based on the subtitle data on the video data based on the determined display mode.
[0007] The program of the present disclosure causes a computer to realize an acquisition function of acquiring video data and subtitle data from content, a recognition function of recognizing a user who views the content, a determination function of determining the attributes of the recognized user, a determination function of determining a display mode of subtitles to be displayed based on the subtitle data according to the determined attributes of the user, and a display function of superimposing and displaying subtitles based on the subtitle data on the video data based on the determined display mode.
[0008] The subtitle generation method of the present disclosure is a subtitle generation method for generating subtitles to be displayed on the video of content, and includes an acquisition step of acquiring video data and subtitle data from the content, a recognition step of recognizing a user who views the content, a determination step of determining the attributes of the recognized user, a determination step of determining a display mode of subtitles to be displayed based on the subtitle data according to the determined attributes of the user, and a generation step of generating subtitles in the determined display mode.
Advantages of the Invention
[0009] According to the present disclosure, for example, it is possible to display subtitles in an appropriate display mode according to the attributes of the viewing user.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Embodiments for Carrying Out the Invention
[0011] Generally, a display device capable of displaying contents such as a program received as a broadcast wave, a video distributed on a video distribution site, or an image output by a computer is known. When displaying the contents, the display device also displays subtitles.
[0012] At this time, even if it is possible to identify the position of the user to be displayed and change the display mode (or style) of the characters according to the position of the user, for example, it may not be possible to change the display mode of the characters according to the attributes of the user.
[0013] In addition, when there are a plurality of viewing users, consideration of which viewing user to determine the character size in accordance with is often not taken into account. In particular, the display mode of the characters has not been determined based on the relationship between the content currently being played and the viewing users.
[0014] Therefore, the display device and the like of the present disclosure can determine and display an appropriate display mode of characters (subtitles) according to the attributes of the viewing users in the display device and the like.
[0015] Hereinafter, the display device of the present disclosure will be described in the following embodiments with reference to the drawings. Note that the following embodiments are illustrative examples of the invention described in the claims, and the technical scope of the present invention is not limited to the description of the following embodiments.
[0016] [1. First Embodiment] Hereinafter, the first embodiment will be described. In the first embodiment, the following will be described as an example.
[0017] The characters and character strings displayed by the display device are obtained from the content, and subtitles (subtitle data) that are superimposed and displayed on the content are taken as an example for display. Note that the displayed characters may be, for example, characters of character broadcasting, characters displayed by a display device such as a program guide, etc. Further, the characters include a character string consisting of one character or a plurality of characters. Further, the characters include Japanese, English, and other characters and symbols that can be displayed by the display device. Further, the characters also include characters based on out-of-character data included in the subtitle data described later.
[0018] The attributes of the user are taken as an example of the age of the user. Note that the attributes of the user may be, for example, gender, race, nationality, etc.
[0019] The display mode of text will be described by taking the text size as an example. Note that the display mode of text may be, for example, the font, color, brightness, etc. of the text.
[0020] [1.1 Entire Display Device] FIG. 1 is a diagram showing the entirety of the display device 10. The display device 10 is a device capable of displaying content. For example, in this embodiment, it is a television capable of receiving broadcast waves and displaying programs. Also, the display device 10 may be a display capable of displaying video input from the outside. Further, the display device 10 may be, for example, a projector that projects a display screen onto a screen.
[0021] The display device 10 may have a device for acquiring information for identifying the user who is viewing. For example, the display device 10 may have, for example, a device for acquiring the voice of the user who is viewing, and may have a microphone 12 as a voice input device, for example. The microphone 12 may be a single microphone or a microphone array composed of a plurality of microphones.
[0022] Also, the display device 10 may have a camera 14 for acquiring an image of the user who is viewing. The camera 14 can, for example, photograph the environment including the user who is viewing the display device 10. One or a plurality of cameras 14 may be arranged.
[0023] Here, the content may be anything that can be displayed on the display device 10. For example, the content may be a program received by terrestrial / BS broadcast / CS broadcast, demodulated, and displayed, a video selected by a user on a video distribution site, or a video input from the outside via HDMI (registered trademark), D-SUB, etc.
[0024] [1.2 Hardware Configuration] The hardware configuration will be described with reference to FIG. 2.
[0025] The control unit 100 controls the entire display device 10. The control unit 100 realizes various functions by reading and executing various programs stored in a storage device (for example, the storage 110 or the ROM 120). The control unit 100 may be realized by one or more control devices / arithmetic devices (CPU (Central Processing Unit), SoC (System on a Chip)). Also, the control unit 100 may be composed of a control circuit.
[0026] The operation control unit 102 receives operations from the user, gives operation instructions to each functional unit, or notifies the control unit 100 of an operation signal corresponding to the received operation. For example, the operation control unit 102 may be an operation signal receiving unit from a remote control, or may be an operation switch provided on the main body of the display device 10.
[0027] The storage 110 is a non-volatile storage device capable of storing programs and data. For example, it may be composed of a storage device such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive). Also, the storage 110 may be configured as a USB memory that can be externally connected. Also, the storage 110 may be, for example, a storage area on the cloud.
[0028] The ROM 120 is a non-volatile memory capable of retaining programs and data even when the power is turned off.
[0029] The RAM 130 is mainly the main memory used by the control unit 100 when executing processing. The RAM 130 is a rewritable memory that temporarily holds programs read from the storage 110 or the ROM 120 and data including execution results.
[0030] The broadcast control unit 140 receives a broadcast wave transmitted by a broadcast station selected by the user, decodes video and characters from the broadcast wave and outputs them to the display unit 150, and decodes audio from the broadcast wave and outputs it to the audio output unit 165. The configuration of the broadcast control unit 140 will be further described later.
[0031] The display unit 150 is a display device capable of displaying the video of the received program and various information. The display unit 150 may be, for example, a device capable of displaying video such as a liquid crystal display (LCD; Liquid Crystal Display) or an organic EL (Organic Electro Luminescence) display. Further, the display unit 150 includes an interface to which a display device can be connected. For example, it may be composed of an external display device connected via HDMI (registered trademark) (High-Definition Multimedia Interface), DVI (Digital Visual Interface), or Display Port. Further, the display unit 150 may be a projection device such as a projector, for example.
[0032] The imaging unit 155 is an imaging device that images the periphery where the display device 10 is installed, and is, for example, a camera or the like. The imaging unit 155 may be composed of one or a plurality of imaging devices. The imaging unit 155 outputs the captured image as an image signal. Further, the imaging unit 155 may output one or a plurality of images as continuous video.
[0033] The voice input unit 160 is an input device capable of inputting the sound around where the display device 10 is installed, and is, for example, a microphone or the like. Further, the voice input unit 160 may be composed of a plurality of input devices (for example, microphones). Further, the voice input unit 160 is mainly used when inputting the voice of the user who is mainly watching, but can also input general sounds such as environmental sounds, for example.
[0034] The voice output unit 165 outputs the voice included in the content. The voice output unit 165 may be, for example, a device such as a speaker or headphones. Further, the voice output unit 165 only needs to output sound, and can output general sounds such as music and environmental sounds, for example.
[0035] The communication unit 170 is a communication interface for communicating with other devices. For example, the communication unit 170 may be a network interface connectable to a wireless LAN, or may be a network interface connectable to Ethernet (registered trademark) by wire. Also, the communication unit 170 may be, for example, a communication device connectable to a mobile communication network such as LTE / 4G / 5G / 6G.
[0036] Here, one or more of the components shown in FIG. 2 may be constituted by an external device connected to the display device 10. For example, the voice input unit 160 may be a microphone connected by USB. Also, one or more of the components shown in FIG. 2 may be realized by a terminal device connected by wireless communication such as Bluetooth (registered trademark), or a terminal device connected via a network. For example, a smartphone, tablet, smart speaker, or wearable terminal device connected to the display device 10 may utilize a display device, voice input / output device, or imaging device that it has.
[0037] Also, the configuration of FIG. 2 only needs to include the configurations necessary in the embodiment. For example, in the first embodiment, it may not have the voice input unit 160.
[0038] (Broadcast control unit) A simple configuration of the broadcast control unit 140 will be described with reference to FIG. 3. For example, broadcast data of digital broadcasts such as terrestrial / BS / CS is acquired by the tuner unit 200. Then, the OFDM demodulation unit 210 demodulates the broadcast data by OFDM, and after error correction and the like are performed, a TS packet is output.
[0039] The separation unit 220 separates the TS packet and outputs the video packet to the video decoding unit 230 and the audio packet to the audio decoding unit 260. Also, the separation unit 220 outputs a data packet (for example, a subtitle TS packet) including information regarding subtitles to the subtitle decoding unit 240.
[0040] The video decoding unit 230 decodes video data from the input video packet and outputs it to the image processing unit 250. Also, the subtitle decoding unit 240 decodes subtitle data from the input data packet and outputs it to the image processing unit 250. Here, the subtitle data is assumed to include information related to subtitles. For example, the subtitle data in this embodiment may include management data of subtitles included in subtitle PES data and data of subtitle texts to be displayed.
[0041] The image processing unit 250 outputs, to the display unit 150, a video obtained by superimposing subtitle data on the video data as necessary.
[0042] Also, the audio decoding unit 260 decodes audio data from the input audio packet and outputs the audio to the audio output unit 165.
[0043] Note that FIG. 3 schematically illustrates a general broadcast control unit 140, and other configurations may be possible. For example, the video decoding unit 230 and the audio decoding unit 260 may be the same decoding unit.
[0044] [1.3 Software Configuration] The software configuration will be described with reference to FIG. 3. Each function is realized when the control unit 100 executes a program stored in, for example, the storage unit (storage 110, ROM 120, RAM 130).
[0045] The user processing unit 1010 executes processing on user information corresponding to the user. For example, it stores user information including identification information for identifying the user. The user processing unit 1010 stores the user information in the user information storage area 1102.
[0046] Here, an example of the user information stored in the user information storage area 1102 will be described with reference to FIG. 5(a). The user information stores a uniquely assigned user ID (for example, "Usr01"), an arbitrarily input or selected user name (for example, "father"), user identification information for identifying the user (for example, "001.jpg"), an age which is an example of the user's attribute (for example, "40"), and subtitle setting information (for example, "size 20 / fast / white").
[0047] Here, in the present embodiment, the user identification information stores the user's face image. For example, when identifying the user by other means, it is sufficient to store the information necessary for identifying the user.
[0048] Also, the subtitle setting information has the display mode of the subtitle set. For example, the size of the subtitle ("size 20"), the speed of the timing at which the subtitle switches (for example, "fast"), and the color of the subtitle (for example, "white") are stored. Note that the subtitle setting information may include other information. For example, the subtitle setting information may further include the position of the subtitle, the font of the subtitle, the specification level of Chinese characters in the subtitle, and the like.
[0049] The target age acquisition unit 1020 acquires the age and age group (hereinafter simply referred to as age) of the viewing user targeted by the content (program). Here, the age of the viewer targeted by the content is referred to as the target age or the targeted age. The process by which the target age acquisition unit 1020 acquires the age (target age) of the viewing user targeted by the program will be described later. The target age acquisition unit 1020 stores it in the program information storage area 1104.
[0050] The program information storage area 1104 stores program information of the received program or the stored program. The program information may be obtained, for example, from program arrangement information (SAI: Service Information) included in the broadcast wave by the display device 10, or may be obtained from the Internet. The program information stores various information related to the program. For example, it includes a program ID, a broadcast time zone, a channel, a program name, program content, etc.
[0051] Since the program information itself is already known, a detailed description is omitted. However, the program information of this embodiment stores at least the content shown in FIG. 5(b). The program information includes a program ID (for example, "CNT001"), a program name (for example, "With Mom"), and a target age "3".
[0052] The subtitle determination unit 1030 determines the display mode and the like of the subtitle to be displayed. Then, the display mode of the subtitle determined by the subtitle determination unit 1030 is output to the control unit 100. The subtitle determination unit 1030 determines the subtitle to be superimposed and displayed on the video based on the display timing (reference time, start time), position (coordinates), color, character size, font, decoration, language selection, display direction (vertical writing, horizontal writing), etc. of the subtitle included in the subtitle data. Also, when the display mode of the subtitle is set for each user, the subtitle determination unit 1030 determines the subtitle to be displayed in the set display mode.
[0053] The face age DB storage area 1106 is a database related to face age. For example, the control unit 100 can estimate the age from the face of the user included in the captured image by referring to the face age DB.
[0054] (User processing unit) Here, the configuration of the user processing unit 1010 will be described with reference to FIG. 6. The user processing unit 1010 acquires the image captured by the imaging unit 155.
[0055] The user processing unit 1010 causes the face detection unit 1012 to detect a face included in an image. Then, the user identification unit 1014 refers to the user information storage area 1102 based on the face detected by the face detection unit 1012 and identifies the user.
[0056] In addition, when the user identification unit 1014 cannot identify the user corresponding to the face detected by the face detection unit 1012 from the user information and stores it as a new user, the user information storage unit 1018 stores the user information in the user information storage area 1102. At this time, the age determination unit 1016 may refer to the face age DB to determine the age. The user information storage unit 1018 may store the age determined by the age determination unit 1016 in the user information. Further, the user information storage unit 1018 may store the age set by the user in the user information.
[0057] [1.4 Process flow] Hereinafter, the process flow in the present embodiment will be described. Although the following processes will be described as being executed by the control unit 100, each configuration described in FIGS. 4 and 6 may execute the processes of each step.
[0058] [1.4.1 User registration process] First, the control unit 100 activates the user management screen based on an operation by the user (S102). By selecting the user management screen, the control unit 100 can newly store, update, or delete user information.
[0059] The control unit 100 starts detecting a face from the image captured by the imaging unit 155 (S104). Here, the control unit 100 may perform face detection according to an operation by the user, or may perform face detection once when the user management screen is activated. Further, the control unit 100 may constantly detect the user's face while this process is being executed.
[0060] Here, when a face is detected in the image, that is, when face information which is an image of the user's face is detected, the control unit 100 temporarily stores the face information (S106; Yes → S108). Here, when two or more faces are detected, the control unit 100 may temporarily store each face. Also, when two or more faces are detected, the control unit 100 may store the face selected by the user from among the multiple detected faces.
[0061] Subsequently, the control unit 100 executes user identification processing based on the stored face image. Here, the control unit 100 determines whether the face image included in the photographing unit 155 matches the face image of the user included in the user information. The method by which the control unit 100 determines whether the face images match is, for example, to detect feature points in each image. Then, when the feature points match by a predetermined ratio or more in the two face images, it is determined that the two images match. Also, the control unit 100 may utilize an external service (for example, an image similarity determination service by AI) to determine the similarity of the two images.
[0062] Here, when the user of the detected face image is not stored in the user information storage area 1102, the control unit 100 stores it as a new user (S112; No → S114). Here, the control unit 100 may prompt the user to input a user name. Also, the control unit 100 may assign a user name according to a predetermined rule.
[0063] Subsequently, the control unit 100 executes age determination processing (S116). The control unit 100 may determine the age from the face image using a face age DB. Also, the control unit 100 may determine the age from the face image by utilizing an external service, for example. Also, the control unit 100 may prompt the user to input an age.
[0064] Subsequently, the control unit 100 determines the display mode of the subtitles (S118). For example, the user determines the display mode of the subtitles.
[0065] For example, as the control unit 100, the following determinations can be made regarding the display mode of subtitles. The following contents may be independent of each other or a combination of multiple display modes may be used.
[0066] (1) Subtitle size. For example, the control unit 100 can set the font size and ratio size as the subtitle size. Also, the control unit 100 may set relative sizes such as small, medium, large, and extra-large as the subtitle size. The control unit 100 may determine the specific font size of the subtitle based on the set relative size.
[0067] (2) Subtitle switching speed. For example, the control unit 100 may be able to set the speed at which subtitles are switched and the display time.
[0068] (3) Subtitle color and font. For example, the control unit 100 may be able to set the color displayed as subtitles, the shape of the font, and the font name.
[0069] Also, the control unit 100 may be able to set one or more display modes for subtitles. Additionally, the control unit 100 may also set other subtitle display modes, such as the position of the subtitles, the language of the subtitles, etc.
[0070] Then, the control unit 100 updates the user information based on the set subtitle display mode (S120). For example, when newly stored as a new user in S114, the control unit 100 sets the set subtitle display mode for the new user. Also, when user information already exists, the control unit 100 updates the user information with the subtitle display mode newly set in S118.
[0071] [1.4.2 Target Age Acquisition Process] Figure 8 shows the process of acquiring the target age of the program being played. The target age refers to the age for which the program is intended for viewing.
[0072] When the user selects a program, the control unit 100 acquires the program's metadata at the start of display or while the program is being displayed (S132). For example, the control unit 100 may acquire the program's metadata from the information contained in the broadcast data, or may acquire the metadata via the Internet.
[0073] The control unit 100 acquires the target age from the metadata (S134). Here, the target age may be a specific age such as 20 years old or 30 years old, or may be an age group. Also, when the target age is 20 years old or older, the control unit 100 may simply set it as 20 years old, or may set the average value among family members as the target age. The control unit 100 stores the acquired target age of the program in the program information.
[0074] [1.4.3 Subtitle display processing] FIG. 9 shows the process of displaying subtitles. Note that although FIG. 9 is described as one process for convenience of explanation, the control unit 100 can execute the processes separately as appropriate.
[0075] When the user attempts to view and selects a program, the control unit 100 acquires the program information (S152). The control unit 100 refers to the program information and acquires the target age of the program (S154).
[0076] The control unit 100 starts the process of detecting the user's face from the image captured by the imaging unit 155 (S156). Here, when there is face information in the captured image (that is, when there is a region containing an image corresponding to a human face), the user corresponding to the face information is recognized (S158; Yes → 160).
[0077] And when the recognized user is stored in the user information (S162; Yes), the control unit 100 acquires the age of the recognized user from the user information (S164). Then, the user repeats the process from S160 until the recognition of all users included in the image is completed.
[0078] When all the ages of all the recognized users have been obtained, the control unit 100 determines the display mode of the subtitles (S168). Here, the method by which the control unit 100 determines the display mode of the subtitles is determined by the following method.
[0079] (1) When no stored user is detected, the control unit 100 determines the display mode of the subtitles, for example, based on the content specified in the subtitle data. Further, the control unit 100 determines based on, for example, the display mode of the subtitles (default display mode) set in the system settings of the display device.
[0080] (2) When there is one recognized user, the control unit 100 determines based on the display mode of the subtitles set for the recognized user.
[0081] (3) When there are a plurality of recognized users, the control unit 100 determines based on the display mode of the subtitles of the user who is farthest from the target age.
[0082] Then, the control unit 100 displays the subtitles in the display mode determined in S168 (S170). Note that the control unit 100 may apply the display mode determined in S168 after confirming with the user whether to enable it.
[0083] [1.5 Operation Example] FIG. 10 is a diagram showing an example of a display screen W10 that stores user information. In the display screen W10, a face image of the user is displayed at G10. Also, a name is displayed at P10, and the age of the user is displayed at P12. Here, when the user selects the button B10, the user name can be changed to an arbitrary name. Also, when the user selects the button B12, the age of the user can be arbitrarily changed. Also, when the user selects the button B14, the control unit 100 can recognize the age from the face image of the user and set the age.
[0084] In addition, P14 can set the display mode of subtitles. The display screen W10 can set the size, switching speed, color, etc. as the display mode of subtitles.
[0085] FIG. 11 is a diagram schematically showing an example of the display of subtitles for a user. For example, in FIG. 11, it will be described as assuming that a program with a current target age of 3 years is being played.
[0086] In FIG. 11(a), Usr03 is watching the program alone. Here, referring to FIG. 5(a), the age of Usr03 is "3 years old". Therefore, in FIG. 11(a), subtitles with a small character size (size10) are displayed in the region R10.
[0087] In FIG. 11(b), Usr03 and Usr04 are watching the program together. Here, referring to FIG. 5(a), the age of Usr04 is "70 years old". Then, the difference between the target age and the age of the viewer is as follows.
[0088] Usr03: User age (3 years) - Target age (3 years) = 0 Usr04: User age (70 years) - Target age (3 years) = 67
[0089] In this case, Usr04 becomes the viewer who is farthest from the target age. Therefore, the control unit 100 displays the subtitles in the display mode set for Usr04. For example, in the region R12 of FIG. 11(b), subtitles with an extra-large character size (size40) are displayed.
[0090] Although not shown in the figure, for example, when Usr01 is also watching, it is as follows.
[0091] Usr01: User age (40 years) - Target age (3 years) = 37 Usr03: User age (3 years) - Target age (3 years) = 0 Usr04: User age (70 years) - Target age (3 years) = 67
[0092] Also in this case, Usr04 becomes the viewer who is farthest from the target age. Note that the viewer who is farthest in age may be defined as the viewer for whom the absolute value of the difference between the target age and the viewer's age is the largest.
[0093] For example, when programs with a target age of 50 years old are viewed by Usr01, Usr03, and Usr04, the differences are as follows.
[0094] Usr01: User age (40 years old) - Target age (50 years old) = -10 Usr03: User age (3 years old) - Target age (50 years old) = -47 Usr04: User age (70 years old) - Target age (50 years old) = 20
[0095] In this case, the viewer who is farthest in age is Usr03. Note that the control unit 100 may display a confirmation screen before reflecting the display mode of the subtitles according to the user who views them.
[0096] FIG. 12(a) is an example of displaying a confirmation screen in the region R14. The confirmation screen displays a screen for confirming "Usr03 and Usr04 have been recognized. Subtitles will be displayed in the settings of Usr04. Is that okay?" Here, when "OK" is selected by the user, the control unit 100 displays the subtitles in the display mode based on Usr04. Also, when "NG" is selected by the user, the control unit 100 displays the subtitles in the initial setting display mode. For example, the control unit 100 displays the subtitles in a display mode based on the subtitle data (a display mode based on information such as the character size included in the subtitle data).
[0097] FIG. 12(b) is an example of displaying a confirmation screen R16 that accepts a cancel operation after the display mode of the subtitles is reflected. For example, the control unit 100 displays the confirmation screen R16 after changing the display mode of the subtitles or based on an arbitrary operation from the user.
[0098] The confirmation screen R16 displays "Subtitles will be shown according to the Usr04 settings. Do you want to revert?" If the user selects "Revert" here, the control unit 100 reverts the display mode of the subtitles to the display mode before the Usr04 settings are applied and displays the subtitles. Thereby, the control unit 100 displays the subtitles in the display mode of the initial settings.
[0099] As described above, according to this embodiment, subtitles in an appropriate display mode can be displayed based on the recognized user. In particular, when multiple users are watching a program, among them, subtitles will be displayed in the display mode of the user whose target age of the program being watched is most different from the age of the user watching, and it is possible to display subtitles in an appropriate manner.
[0100] Note that the process of determining the display mode of the subtitles in this embodiment and displaying the subtitles may operate in real time. For example, in Fig. 11(b), when Usr04 stops watching, the control unit 100 may switch the display of the subtitles to the display mode of the subtitles corresponding to Usr03 shown in Fig. 11(a).
[0101] Also, in the above-described embodiment, an example of storing a face image as user identification information has been described. However, for example, data based on extraction points that extract features specialized for age determination, such as the ratio of eyes to face and wrinkle information, may also be used.
[0102] [2. Second Embodiment] The second embodiment will be described. The second embodiment is an embodiment in which the display mode of the subtitles is determined based on the user's voice instead of the face image of the user used in the first embodiment.
[0103] For the second embodiment, the description of the parts having the same hardware and software configurations as those of the first embodiment will be omitted, and the description will focus on the points different from the first embodiment.
[0104] FIG. 13 of the second embodiment replaces FIG. 6 of the first embodiment. Note that the same components as those in FIG. 6 are denoted by the same reference numerals, and the description thereof will be omitted.
[0105] In this embodiment, the user's voice is input from the voice input unit 160 and output to the voice detection unit 1012B. Then, the user identification unit 1014 compares the input voice with the user identification information stored in the user information storage area 1102. Here, the user information storage area 1102 of this embodiment stores voice data as user identification information.
[0106] Also, the age determination unit 1016 can determine the user's age by using the voice age DB stored in the voice age DB storage area 1106B.
[0107] FIG. 14 is a diagram that replaces FIG. 7. In this embodiment, the control unit 100 starts the process of detecting the user's voice (S202). Then, when there is the user's voice information in the voice data, the control unit 100 temporarily stores the voice information (S204; Yes→S206).
[0108] Then, the control unit 100 executes the user determination process based on the voice information (S110). Here, when the user is not stored, the user is stored in the user information as a new user (S112; Yes→S114). Note that the control unit 100 stores the temporarily stored voice data in the user information storage area 1102.
[0109] As described above, according to this embodiment, the user can be identified and the age can be determined based on the voice input from the user. Note that the user identification information may store the voice data as it is, or may be, for example, a data sequence obtained by extracting feature points specialized for age determination, such as the pitch and hoarseness of the user's voice.
[0110] Also, in S102, while the user is activating the user management screen and performing registration and update, for example, the user may automatically perform user registration or update using a microphone or the like while watching a program.
[0111] Further, the present embodiment and the first embodiment may be combined to determine the age using both the face image and voice of the user. Also, when combining the present embodiment and the first embodiment, when the face image is available, the control unit 100 may determine the display mode of the subtitle using the method of the first embodiment, and when the face image is not available, the control unit 100 may determine the display mode of the subtitle using the method of the second embodiment.
[0112] [3. Third Embodiment] The third embodiment will be described. The third embodiment determines the target age of a program based on the user's viewing. For example, while the user is watching a program, the control unit 100 polls the viewing state of the user at regular intervals. Then, the control unit 100 determines the viewing user by face determination and accumulates the information on the viewing time for each user to obtain the target age for each program.
[0113] This embodiment is an embodiment in which FIG. 8 of the first embodiment is replaced with FIGS. 15 and 16. Hereinafter, it will be described with reference to FIGS. 15 and 16. And for the third embodiment, the description of the parts having the same configuration of hardware and software as those of the first embodiment will be omitted, and the description will be centered on the differences from the first embodiment.
[0114] FIG. 15 shows the process of measuring the viewing time of the user. First, the control unit 100 starts measuring the viewing time (S302). For example, the control unit 100 starts measuring the viewing time by starting a timer.
[0115] When the viewer is watching a program (S304; Yes), the control unit 100 acquires the name of the program being watched (S306). Then, the control unit 100 acquires an image and determines whether there is face information in the image (S308). That is, the control unit 100 detects a face image from the image input from the imaging unit 155. Using this detected face image as face information, the control unit 100 performs user identification processing (S310).
[0116] When the identified user is a registered user (S312; Yes), the viewing time of the user specified in S312 is incremented (S314). And while the user is watching, the viewing time of the user continues to be incremented.
[0117] When the processing for all face information is completed (S316; Yes), the control unit 100 ends this processing. For example, when there are no viewing users in front of the display device 10, when the viewing of the program on the display device 10 ends, etc.
[0118] By the control unit 100 executing the processing of FIG. 15, the viewing time for each user for each program is measured.
[0119] FIG. 16 is a diagram for explaining the processing of updating program information. The control unit 100 first acquires program information (S332). Subsequently, the control unit 100 searches for users who have watched the program (S334). Then, the control unit 100 acquires the ages of all users (S336) and also acquires the viewing time (S338).
[0120] Then, the control unit 100 increments the viewing time for each age. When the processing for all users is completed, the control unit 100 updates the program information (S342; Yes→S344). For example, the control unit 100 may use the age of the user with the longest viewing time in the program as the target age.
[0121] Explanation will be based on a specific example. For example, the viewing times of Program A, Program B, and Program C are as follows.
[0122] User 03 (3 years old): Program A for 2 hours, Program B for 30 minutes, Program C for 40 minutes User 01 (40 years old): Program C for 1 hour User 02 (40 years old): Program C for 30 minutes At this time, the control unit 100 may use the age of the user who has watched the program for the longest time among all programs as the target age.
[0123] For example, Program A: 3 years old (3-year-old users have watched for a total of 2 hours, which is the longest) Program B: 3 years old (3-year-old users have watched for a total of 30 minutes, which is the longest) Program C: 40 years old (3-year-old users have watched for a total of 1 hour and 30 minutes, which is the longest) It will be like this.
[0124] Thus, according to this embodiment, for example, even when the target age of a program cannot be obtained from metadata or the like, the target age can be appropriately obtained based on the viewing history.
[0125] [4. Fourth Embodiment] The fourth embodiment will be described. In the first to third embodiments, the display mode of subtitles was determined based on subtitle data, and subtitles were displayed. In this embodiment, subtitles to be displayed are generated.
[0126] For the fourth embodiment, descriptions of the same parts in terms of hardware and software configurations as those in the above-described embodiments will be omitted, and differences from the first embodiment will be mainly described as an example.
[0127] In the subtitle determination process of FIG. 9, instead of displaying subtitles in S170, the control unit 100 generates display subtitle data as subtitle data to be displayed. The generated subtitle data may be associated with each content (program) and stored in the storage 110.
[0128] When playing content (program), the control unit 100 refers to the display subtitle data stored in the storage 110. Then, the control unit 100 superimposes and displays the subtitle on the video in a display mode based on the display subtitle data.
[0129] Thus, according to this embodiment, subtitle data can be generated. Thereby, for example, when storing content (program) on an external recording medium or on a network, the display device 10 can also store the display subtitle data accordingly. In this way, it becomes possible to display subtitles together with the video on other display devices based on the subtitle data generated in this embodiment.
[0130] [5. Modification Example] The present disclosure is not limited to the above-described embodiments, and various modifications are possible. That is, embodiments obtained by appropriately combining technical means modified within the scope not departing from the gist of the present disclosure are also included in the technical scope.
[0131] In addition, the above-described embodiments are described separately for convenience of explanation, but they can be combined and executed within the possible range. Also, regarding any technology described in the specification, there is an intention to obtain rights in corrections, divisional applications, etc.
[0132] In addition, the programs operating in each device in each embodiment are programs for controlling a CPU, etc. (programs for making a computer function) so as to realize the functions of the above-described embodiments. And the information handled by these devices is temporarily stored in a temporary storage device (for example, RAM) during its processing, and then stored in storage devices such as various ROMs and HDDs, and read out by the CPU as needed for correction and writing.
[0133] Here, as the recording medium for storing the program, any of a semiconductor medium (e.g., ROM, non-volatile memory card, etc.), an optical recording medium or a magneto-optical recording medium (e.g., DVD (Digital Versatile Disc), CD (Compact Disc), BD (Blu-ray (registered trademark) Disc), etc.), a magnetic recording medium (e.g., magnetic tape, flexible disk, etc.) may be used.
[0134] Also, when distributing to the market, the program can be stored in a portable recording medium for distribution, or transferred to a server computer connected via a network such as the Internet. In this case, the storage device of the server device is of course also included in the present disclosure.
[0135] Further, the above-described data does not have to be stored in the device, but may be stored in an external device and appropriately called. For example, the data may be stored in a NAS (Network Attached Storage) or stored in the cloud.
[0136] Note that the scope of the present disclosure is not limited to the configurations explicitly described in the specification, and the combinations of the technologies disclosed in this specification are also included in the scope. Among the present disclosures, the configuration for which a patent is sought is described in the appended claims, but it is not intended to exclude it from the technical scope for the reason that it is not described in the claims.
[0137] Also, in the above-described specification, the descriptions such as "in the case of ~" and "when ~" are explanations as one example, and are not configured to be limited to the described content. For configurations other than these cases and times, those that are obvious to those skilled in the art are also disclosed and have the intention of obtaining rights.
[0138] In addition, with regard to the processes described in the specification and the descriptions of the data flow that involve an order, they are not limited to the described order. For example, configurations in which a part of the process is deleted or the order is swapped are also disclosed and have the intention of obtaining rights.
[0139] Also, although the functions described in the embodiments are described as being executed by respective devices, they may be realized by one device or may further utilize an external server.
[0140] Also, each functional block or various features of the devices used in the above-described embodiments can be implemented or executed by an electric circuit, for example, an integrated circuit or a plurality of integrated circuits. An electric circuit designed to execute the functions described in this specification may include a general-purpose use processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or a combination thereof. The general-purpose use processor may be a microprocessor or may be a conventional type processor, controller, microcontroller, or state machine. The above-described electric circuit may be composed of a digital circuit or may be composed of an analog circuit. Also, when an integrated circuit technology that replaces the current integrated circuit appears due to the progress of semiconductor technology, one or more aspects of the present disclosure can also use a new integrated circuit by such technology.
Description of Reference Numerals
[0141] 10 Display device 100 Control unit 102 Operation control unit 110 Storage 120 ROM 130 RAM 140 Broadcast control unit 150 Display unit 155 Photographing unit 160 Audio input unit 165 Voice output unit 170 Communication unit
Claims
1. An acquisition unit that acquires video data and subtitle data from content; A recognition unit that recognizes a user who views the content; A determination unit that determines the attributes of the recognized user; A determination unit that determines a display mode of subtitles to be displayed based on the subtitle data based on the attributes of the determined user; A display unit that superimposes and displays subtitles based on the subtitle data on the video data based on the determined display mode; A display device comprising the above.
2. The attribute is age, A target age acquisition unit that acquires a target age for which the content is targeted; Based on the age of the user who views the content and the target age, the display mode of the subtitles is determined The display device according to claim 1.
3. The recognition unit can recognize a plurality of users, When the recognition unit recognizes a plurality of users, the determination unit determines the display mode of the subtitles based on the age of the user among the recognized users whose age difference from the target age and the age of the user who views the content is the largest The display device according to claim 2.
4. The display device according to claim 2 or 3, wherein the target age is acquired from information included in the content.
5. The display device according to claim 2 or 3, wherein the target age is determined based on the viewing time during which the user has viewed the content.
6. Further comprising a photographing unit, The display device according to claim 2 or 3, wherein the determination unit determines the age of the user from an image photographed by the photographing unit of the user.
7. Further comprising an audio input unit, The display device according to claim 2 or 3, wherein the determination unit determines the age from the voice of the user input from the audio input unit.
8. An acquisition step of acquiring video data and subtitle data from content; A recognition step of recognizing a user who views the content; A determination step of determining the attributes of the recognized user; A determination step of determining a display mode of subtitles to be displayed based on the subtitle data based on the attributes of the determined user; A display step of superimposing and displaying subtitles based on the subtitle data on the video data based on the determined display mode; A display method including the above.
9. On a computer, An acquisition function of acquiring video data and subtitle data from content; A recognition function of recognizing a user who views the content; A determination function of determining the attributes of the recognized user; A determination function that determines a display mode of subtitles to be displayed based on the subtitle data according to the determined user attributes; A display function that superimposes and displays subtitles based on the subtitle data on the video data according to the determined display mode; A program for realizing the above.
10. A recording medium on which the program of Claim 9 is recorded.
11. In a subtitle generation method for generating subtitles to be displayed on a video of content, An acquisition step of acquiring video data and subtitle data from the content; A recognition step of recognizing a user who views the content; A determination step of determining attributes of the recognized user; A determination step of determining a display mode of subtitles to be displayed based on the subtitle data according to the determined user attributes; A generation step of generating subtitles in the determined display mode; A subtitle generation method including the above steps.
Citation Information
Patent Citations
Display control device and display control method
JP2010107773A