Display device, display method, program, and recording medium

The display device analyzes viewer attributes through imaging and adjusts content display to enhance user experience and advertising relevance by considering the viewer's state.

JP2025185992APending Publication Date: 2025-12-23SHARP KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024094524
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-11
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

Conventional display devices fail to consider the viewing state of the user, leading to ineffective advertising and content display based on customer demographics, as they lack the ability to select and control content according to the viewer's situation.

Method used

A display device equipped with a camera to capture the viewer's image, an estimation unit to analyze attributes such as demographics, posture, and a control unit to select and adjust content display based on these attributes, including screen splitting, subtitle display, content rotation, and quality adjustment.

Benefits of technology

Enhances the convenience and effectiveness of content display by tailoring it to the viewer's attributes, improving advertising relevance and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025185992000001_ABST
    Figure 2025185992000001_ABST
Patent Text Reader

Abstract

To provide a convenient display device by displaying a content in accordance with a viewing state of a viewer.SOLUTION: A display device includes: a display unit that displays a content on a display screen; an acquisition unit 1010 that acquires, as a viewing state of a target person, a target image indicating the target person who is viewing the content from among captured images obtained by photographing the periphery of the display screen; an estimation unit 1020 that estimates an attribute of the target person on the basis of the acquired target image; and a display control unit 1030 that performs selection of a content to be displayed on the display screen or display control concerning a content being displayed on the basis of an estimation result obtained by the estimation unit 1020.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a display device and the like. [Background technology]

[0002] For example, as disclosed in Patent Document 1, a playback device is known that reduces memory consumption and can easily play back a variety of content. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2011-166423 Summary of the Invention [Problem to be solved by the invention]

[0004] An object of the present disclosure is to provide a highly convenient display device by, for example, displaying content according to the viewing state of a person who is viewing the content. [Means for solving the problem]

[0005] The display device of the present disclosure is characterized by comprising a display unit that displays content on a display screen, an acquisition unit that acquires a target image representing a target person watching the content as the viewing state of the target person from a captured image obtained by photographing the area around the display screen, an estimation unit that estimates attributes of the target person based on the acquired target image, and a control unit that selects the content to be displayed on the display screen or performs display control for the content being displayed based on the estimation result of the estimation unit.

[0006] The display method of the present disclosure is characterized by including a display step of displaying content on a display screen, an acquisition step of acquiring a target image representing a target person watching the content as the viewing state of the target person from a captured image obtained by photographing the area around the display screen, an estimation step of estimating attributes of the target person based on the acquired target image, and a control step of selecting the content to be displayed on the display screen or performing display control for the content being displayed based on the estimation result in the estimation step.

[0007] The program disclosed herein is characterized by having a computer realize a display function for displaying content on a display screen, an acquisition function for acquiring a target image representing a target person watching the content as the viewing state of the target person from a captured image obtained by photographing the area around the display screen, an estimation function for estimating the attributes of the target person based on the acquired target image, and a control function for selecting the content to be displayed on the display screen or performing display control for the content currently being displayed based on the estimation result of the estimation function. [Effects of the Invention]

[0008] According to the present disclosure, for example, by displaying content according to the viewing state of a person who is viewing, it is possible to provide a highly convenient display device. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram illustrating an overview of a display device according to a first embodiment. [Figure 2] FIG. 2 is a diagram illustrating the hardware configuration of the display device according to the first embodiment. [Figure 3] FIG. 2 is a diagram illustrating a software configuration according to the first embodiment. [Figure 4] FIG. 3 is a diagram illustrating a specific example of a response variable in the first embodiment. [Figure 5] FIG. 2 is a diagram illustrating a processing flow in the first embodiment. [Figure 6] FIG. 2 is a diagram illustrating a processing flow in the first embodiment. [Figure 7] FIG. 2 is a diagram illustrating a processing flow in the first embodiment. [Figure 8] FIG. 2 is a diagram illustrating a processing flow in the first embodiment. [Figure 9] FIG. 2 is a diagram illustrating an example of operation in the first embodiment. [Figure 10] FIG. 2 is a diagram illustrating an example of operation in the first embodiment. [Figure 11] FIG. 10 is a diagram illustrating a processing flow in the second embodiment. [Figure 12] FIG. 10 is a diagram illustrating a processing flow in the second embodiment. [Figure 13] FIG. 10 is a diagram illustrating an example of operation in the second embodiment. [Figure 14] FIG. 10 is a diagram illustrating an example of operation in the second embodiment. [Figure 15] FIG. 10 is a diagram illustrating a processing flow in the third embodiment. [Figure 16] FIG. 10 is a diagram illustrating a processing flow in the third embodiment. [Figure 17] FIG. 10 is a diagram illustrating a processing flow in the third embodiment. [Figure 18] FIG. 10 is a diagram illustrating an example of operation in the third embodiment. [Figure 19] FIG. 10 is a diagram illustrating a processing flow in the fourth embodiment. [Figure 20] FIG. 10 is a diagram illustrating a processing flow in the fourth embodiment. [Figure 21] FIG. 10 is a diagram illustrating an example of operation in the fourth embodiment. [Figure 22] FIG. 10 is a diagram illustrating an example of operation in the fourth embodiment. [Figure 23] FIG. 13 is a diagram illustrating a processing flow in the fifth embodiment. [Figure 24] FIG. 13 is a diagram illustrating a processing flow in the fifth embodiment. [Figure 25] FIG. 13 is a diagram illustrating an example of operation in the fifth embodiment. [Figure 26] FIG. 20 is a diagram illustrating a processing flow in the sixth embodiment. [Figure 27] FIG. 20 is a diagram illustrating a processing flow in the sixth embodiment. [Figure 28] FIG. 13 is a diagram illustrating an example of operation in the sixth embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] BACKGROUND ART Display devices are known that are capable of displaying not only broadcast content and internet content, but also content for explaining functions and demo content (advertising content) for sales promotion, for example.

[0011] For example, when advertising through the display of content, it is effective to display content according to the class of potential customers (hereinafter referred to as "customer demographics"). However, in conventional configurations, the display of content depends on the characteristics and settings of the display device, so the same content may be displayed to different customer demographics, making it difficult to say that advertising is effective according to customer demographics.

[0012] Furthermore, not only in the above-mentioned advertising activities, but also in relation to the display of content, no consideration has been given to selecting content according to the viewing situation of the viewer or controlling the display of content currently being displayed.

[0013] The display device of the present disclosure that solves the above-mentioned problems will be described in the following embodiments with reference to the drawings. Note that the following embodiments are merely examples of the invention described in the claims, and the technical scope of the present invention is not limited to the description of the following embodiments.

[0014] [1 First Embodiment] [1.1 Overall structure] 1 is a diagram showing the entire display device 10. The display device 10 is a device capable of displaying content, and may be, for example, a television capable of displaying broadcast content based on received broadcast waves or internet content received via the internet on a display screen W10, or may be a display capable of displaying video input from an external device. The display device 10 may also be, for example, a projector that projects the display screen W10 onto a screen.

[0015] The display device 10 has a camera 12 that generates a captured image by capturing an image of the periphery of the display screen W10 (for example, in front of the display device 10 (display screen W10)). One or more cameras 12 may be provided.

[0016] The display device 10 may also have a speaker 14 capable of outputting audio of the content. Furthermore, the display device 10 may also have one or more means capable of acquiring the surrounding situation of the display device 10, such as a sound collection means such as a microphone, or a detection means such as a human sensor, a depth sensor, or a distance sensor.

[0017] The display device 10 has a display screen W10 that displays content. The display screen W10 may have a screen splitting function that allows multiple contents to be displayed simultaneously. For example, if the screen splitting function supports splitting into two screens, it is possible to configure the display to display broadcast content on one screen (parent screen) and internet content, demo content, or the like, as content different from the content displayed on the parent screen on the other screen (child screen).

[0018] Here, the content according to the present disclosure is not particularly limited as long as it can be displayed on the display device 10. For example, the content may be a program received from terrestrial / BS / CS broadcasting, demodulated, and displayed, a video selected by a user on a video distribution site, a video input from an external device via HDMI (registered trademark), D-SUB, or the like, or a still image stored on a USB memory or the like.

[0019] [1.2 Hardware Configuration] The hardware configuration of the display device 10 will be described with reference to FIG.

[0020] The control unit 100 controls the entire display device 10. The control unit 100 realizes various functions by reading and executing various programs stored in a storage device (for example, storage 110 or ROM 120). The control unit 100 may be realized by one or more control devices / arithmetic units (CPUs (Central Processing Units), SoCs (System on Chips)). The control unit 100 may also be configured by a control circuit.

[0021] The operation control unit 102 receives operations from the user, outputs operation instructions to each functional unit, and outputs an operation signal corresponding to the received operation to the control unit 100. For example, the operation control unit 102 may be a unit receiving an operation signal from a remote control, or may be an operation switch provided on the main body of the display device 10.

[0022] The broadcast control unit 104 receives a transmission wave transmitted by a broadcast station selected by the user, decodes video data from the broadcast wave and outputs it to the display unit 140, and decodes audio data from the broadcast wave and outputs it to the audio output unit 150. The broadcast control unit 104 may have, for example, a digital tuner unit (terrestrial / BS / CS, etc.), an OFDM demodulation unit, a DEMUX unit, an MPEG2 decoding unit, etc. Furthermore, there may be multiple broadcast control units 104.

[0023] The storage 110 is a non-volatile storage device capable of storing programs and data. The storage 110 may be configured with a storage device such as a hard disk drive (HDD) or a solid state drive (SSD). The storage 110 may also be a USB memory or the like connectable to a connection terminal (not shown). The storage 110 may also be a storage area provided on a cloud (not shown), for example.

[0024] The ROM 120 is a non-volatile memory that can retain programs and data even when the power is turned off.

[0025] The RAM 130 is a main memory that is mainly used when the control unit 100 executes processing. The RAM 130 is a rewritable memory that temporarily stores programs read from the storage 110 or the ROM 120, and data including execution results.

[0026] The display unit 140 is a display panel capable of displaying images of received programs and various types of information. The display unit 140 may be, for example, a panel capable of displaying images, such as a liquid crystal display (LCD) or an organic electroluminescence (EL) display. The display unit 140 also includes an interface to which the display device 10 can be connected. For example, the display unit 140 may be configured as an external display panel connected via an HDMI (registered trademark) (High-Definition Multimedia Interface), a DVI (Digital Visual Interface), or a Display Port. The display unit 140 may also be, for example, a projection device such as a projector.

[0027] The audio output unit 150 outputs audio included in the content. The audio output unit 150 may be, for example, a device such as the speaker 14 or headphones. The audio output unit 150 only needs to output sound, and can output general sounds such as music, environmental sounds, etc.

[0028] The audio input unit 160 is an input device that can input sounds around the area where the display device 10 is installed, such as a microphone. The audio input unit 160 may also be composed of multiple input devices (for example, microphones). The audio input unit 160 can also input general sounds such as environmental sounds and sounds made by the user, in addition to voices.

[0029] The photographing unit 170 is a photographing device that photographs the surroundings where the display device 10 is installed (for example, in front of the display device 10 (display screen W10)), and is, for example, a camera that can acquire a distance image or a three-dimensional shape of the photographed object. The photographing unit 170 may be composed of one or more photographing devices. The photographing unit 170 outputs the photographed image as an image signal. The photographing unit 170 may also output one or more images as a continuous video.

[0030] The communication unit 180 is a communication interface for communicating with other devices. For example, the communication unit 180 may be a network interface connectable to a wireless LAN or a network interface connectable to Ethernet (registered trademark) via a wired connection. The communication unit 180 may also be a communication device connectable to a mobile communication network such as LTE / 4G / 5G / 6G.

[0031] 2 may be configured as an external device connected to the display device 10. For example, the image capturing unit 170 may be a camera connected via USB.

[0032] 2 may include the components necessary for the embodiment. For example, the audio input unit 160 may not be included in the first embodiment.

[0033] [1.3 Software Configuration] Next, the software configuration will be described with reference to FIG.

[0034] The acquisition unit 1010 acquires, from captured images acquired by controlling the imaging unit 170, an image (target image) representing a person (target person) viewing content displayed on the display unit 140 (display screen W10) as the viewing state of the target person. The acquisition of the target image can be performed, for example, by using a face detection learning model to optimally combine features such as Haar-like features, joint Haar-like features, and sparse features that focus on the contrast between brightness and darkness in partial areas of the target person's face. Note that, in acquiring the target image, in addition to face detection, the target person may be identified based on the posture and posture of the target person as well as audio information acquired via the audio input unit 160.

[0035] When the display device 10 is in a specific operation mode (for example, in-store mode), the acquisition unit 1010 acquires target images in which related parties (store salespeople) are excluded from the target persons. In this case, images showing the faces, postures, etc. of store salespeople are registered in advance, and when an image showing the store salespeople is included in the acquired captured image, the images can be excluded from the target images.

[0036] The estimation unit 1020 estimates attributes of a target person based on the target image acquired by the acquisition unit 1010. The estimation of attributes by the estimation unit 10120 can be performed using a trained model that has been machine-learned using an image representing a person as an explanatory variable and the person's attributes as a target variable. The estimation unit 1020 inputs the target image acquired by the acquisition unit 1010 into the trained model and outputs the analysis result of the target image as an estimation result. Note that the attributes according to the present disclosure refer to properties, characteristics, etc. that represent the target person, and are incidental to the target person.

[0037] Here, a specific example of the objective variables estimated by the estimation unit 1020 according to the first embodiment will be described with reference to the table in FIG. 4. The table illustrated in FIG. 4 shows an example of a combination of objective variables applicable to customer demographics. Here, ID is an identifier for uniquely identifying the customer demographic illustrated in FIG. 4. Age represents the age of the target person to be estimated, and can be selected from, for example, childhood (0-4 years old), boy (5-14 years old), young adult (15-24 years old), middle-aged (25-44 years old), middle-aged (45-64 years old), and elderly (65 years old and above). Gender represents the gender of the target person to be estimated. Accompanying person represents whether the target person to be estimated is accompanied by someone. If there is an accompanying person, a value of Yes can be selected, and if there is no accompanying person, a value of No can be selected. Body type represents the body type of the target person to be estimated, and can be selected from, for example, standard type, thin type, obese type, no selection (No), and the like. The possessions represent the possessions (accessories, etc.) that the person being estimated possesses, and it is possible to specify a specific possession, or simply to indicate whether or not the person possesses the possession by answering Yes or No.

[0038] For example, the customer demographic identified by ID#01 represents a customer demographic into which target persons who satisfy conditions such as age being "middle-aged," gender being "male," accompanying person being "Yes," body type being "No," belongings being "No," etc. Incidentally, the objective variables applicable to the customer demographic according to the present disclosure are not limited to the example shown in Fig. 4, and may be, for example, a customer demographic consisting of a single objective variable such as age or gender, or may be a combination of variables other than the objective variables exemplified in Fig. 4, and there are no limitations on the type or configuration of the objective variables.

[0039] The display control unit 1030 performs screen control of the display screen W10 on the display unit 140, thereby displaying, editing, deleting content from the display screen W10, etc. The display control unit 1030 selects content according to the attribute estimation result output from the estimation unit 1020, and performs display control such as displaying the selected content and editing the content currently being displayed.

[0040] [1.4 Processing flow] Next, the flow of processing according to the first embodiment will be described with reference to the flowchart in Fig. 5. Upon receiving an instruction to start processing, the control unit 100 starts the viewing status acquisition processing (step S10). In the viewing status acquisition processing, the control unit 100 determines whether or not a target image has been acquired from the captured images (step S20).

[0041] If the control unit 100 determines that a target image has been acquired from the captured image, it performs attribute estimation processing based on the acquired target image (step S20; Yes → step S30). On the other hand, if it determines that a target image has not been acquired from the captured image, the control unit 100 returns the processing to step S10 (step S20; No → step S10).

[0042] Based on the estimation result obtained as a result of the attribute estimation process, the control unit 100 selects content to be displayed on the display unit 140 (step S40). Next, the control unit 100 displays the selected content on the display screen W140 and ends the process (step S50 → End).

[0043] Next, the viewing state acquisition process in step S10 of FIG. 5 will be described with reference to the flowchart of FIG.

[0044] When the process starts, the control unit 100 receives an input of an image captured by the camera 12 (image capturing unit 170) (step S11).

[0045] Next, the control unit 100 acquires an image showing a person standing still (in a stationary state) in front of the display unit from the received captured image (step S12).

[0046] Then, the control unit 100 determines whether the operation mode of the device is the storefront mode (step S13). Here, the storefront mode is one type of operation mode of the display device 10, and is an operation mode that operates based on operation instructions from a store salesperson when selling the display device 10 or when promoting the display device 10. The storefront mode can be enabled / disabled through a setting screen (not shown) or the connection of a specific device (e.g., a CAS card). For example, if the use of the display device 10 is mainly for promotional activities using still images stored in a USB memory or the like, the presence or absence of a USB memory attached to or detached from the display device 10 may be used as a trigger for switching the storefront mode between enabled and disabled.

[0047] If it is determined that the operation mode of the device is the in-store mode, the control unit 100 acquires a target image in which the relevant person (store salesperson) is excluded from the target person (step S13; Yes → step S14). On the other hand, if it is determined that the operation mode of the device is not the in-store mode, the control unit 100 skips the process of step S14 and ends the process (step S13; No → "end").

[0048] Next, the attribute estimation process in step S30 of FIG. 5 will be described with reference to the flowchart of FIG.

[0049] When the process starts, the control unit 100 determines whether the operation mode of the device is the in-store mode (step S31). If the control unit 100 determines that the operation mode of the device is the in-store mode, it determines whether a store salesperson is present (step S31; Yes → step S32).

[0050] If it is determined that a store salesperson is present, the control unit 100 selects an image showing a person having a conversation with the store salesperson as a target image for attribute estimation (step S32; Yes → step S33). Then, the control unit 100 estimates attributes of the target person based on the selected target image (step S34 → "end").

[0051] On the other hand, if the control unit 100 determines that the device's operating mode is not the in-store mode (step S31; No) or that a store salesperson is not present (step S32; No), it selects an image depicting a person gazing at the display unit 140 as the target image (step S31; No, step S32; No → step S35). Then, the control unit 100 estimates attributes using the trained model based on the selected target image and ends the process (step S34 → “End”).

[0052] Next, the content selection process in step S40 of FIG. 5 will be described with reference to the flowchart of FIG.

[0053] When the process starts, the control unit 100 estimates whether the attributes of the target person belong to the specific customer demographic exemplified in FIG. 4 (step S41).

[0054] If the control unit 100 estimates that the target person's attributes belong to a specific customer demographic, it determines whether the device's operating mode is the in-store mode (step S41; Yes → step S42). On the other hand, if it determines that the target person's attributes do not belong to a specific customer demographic, it ends the process (step S41; No → "end").

[0055] When the control unit 100 determines that the operation mode of the device is the in-store mode, it selects content based on the attributes from a USB memory or the like attached to the display device 10 and ends the process (step S42; ON → step S43 → "end"). Note that when the in-store mode is ON, content or advertisements recommended by the store may be selected and displayed from the attached USB memory. Furthermore, based on the attributes (customer demographics) of the target person, it may search for content (file names) saved in the USB memory based on a character string representing the attribute, and select and display the content that is found.

[0056] On the other hand, if the control unit 100 determines that the operating mode of the device is not the in-store mode, it selects content based on the attributes from content that can be received via broadcast or communication, and terminates the processing (step S42; No → step S44 → “End”).

[0057] [1.5 Example of operation] Next, an operation example of the first embodiment will be described. Fig. 9 is a diagram schematically illustrating a state in which content C10 recommended to a person P10, who is a target person, is displayed on a display screen W10 in accordance with the attributes (customer demographics) of the person P10.

[0058] 9 shows a situation in which a store salesperson SP10, a person P10, and a juvenile person P12 accompanying person P10 are located around the display device 10 (in front of the display screen W10). In the first embodiment, the store salesperson SP10 is excluded from the target persons for attribute estimation. Therefore, either person P10 or person P12 becomes the target person for attribute estimation.

[0059] Figure 9 shows an example in which the customer demographic for person P10 surrounded by frame F10 (detection window) is estimated, and when the customer demographic for person P10 is estimated to be the customer demographic for ID "#01" in Figure 4 (age "middle-aged", gender "male", accompanying person "Yes", body type "No", belongings "No", etc.), content recommended for the customer demographic for ID "#01" is displayed on display screen W10.

[0060] 5 are continuously performed as long as a target image of the target person can be acquired (as long as a person is present as a subject around the display device 10). On the other hand, as shown in the example of FIG. 9, when there are multiple people (person P10 and person P12) around the display device 10, control may be performed to automatically switch the target person related to attribute estimation to another person (for example, from person P10 to person P12) after displaying content for a certain period of time.

[0061] 10, for example, if the display device 10 has a screen splitting function, it is also possible to configure the notification screen W12 to display a target image (face image) of the target person and a message that recommended content C10 for the target person is being displayed. The target person can understand that recommended content for the target person is being displayed through the display content on the notification screen W12, which can further enhance the sense of immersion in the content C10 displayed on the display screen W10.

[0062] As described above, according to the first embodiment, content according to the attributes of the viewer can be selected and displayed in accordance with the viewing situation of the viewer, so that, for example, if the attribute to be estimated is a customer demographic, effective advertising activities can be carried out according to the customer demographic. Furthermore, as shown in the example of Fig. 4, the objective variable estimated by the estimation unit can be set appropriately according to the target customer demographic, so that the content and advertisements recommended by the store can be optimized.

[0063] [2 Second embodiment] Next, a second embodiment will be described. The second embodiment is a form capable of displaying subtitles for content being displayed on a display screen in accordance with the viewing state (motion state) of a target person.

[0064] The second embodiment can have substantially the same hardware and software configuration as the first embodiment, but differs from the first embodiment in that the estimation unit 1020 estimates, as an attribute, whether or not the target person requires subtitles to be displayed for content, depending on the viewing state (operating state) of the target person acquired by the acquisition unit 1010.

[0065] FIG. 11 is a flowchart illustrating the flow of processing in the second embodiment.

[0066] When the process starts, the control unit 100 determines whether the operation mode of the device is the in-store mode (step S100). If the control unit 100 determines that the operation mode of the device is the in-store mode, it starts the viewing status acquisition process (step S100; Yes → step S110). If the control unit 100 determines that the operation mode of the device is not the in-store mode, it performs display guidance related to the subtitle display function and proceeds to step S110 (step S100; No → step S140 → step S110).

[0067] When the viewing state acquisition process is completed, the control unit 100 performs an attribute estimation process based on the acquired viewing state (motion state of the target person) (step S120). In this case, the control unit 100 estimates, as an attribute of the target person, whether or not the target person requires subtitles to be displayed for the content being displayed (step S130).

[0068] If it is estimated that the target person requires subtitles to be displayed for the content being displayed, the control unit 100 displays the subtitles superimposed on the content displayed on the display screen W10 and terminates the processing (step S130; Yes → step S140 → “End”).

[0069] On the other hand, if it is estimated that the target person does not need subtitles displayed for the content being displayed, the control unit 100 returns the process to step S110 (step S130; No→step S110).

[0070] Next, the viewing state acquisition process in step S110 of FIG. 11 will be described with reference to the flowchart of FIG.

[0071] When the process starts, the control unit 100 acquires a target image representing the motion state of the target person as the viewing state of the target person (step S111).

[0072] The control unit 100 measures the position of the target person's face from the acquired target image (step S112). Next, the control unit 100 measures the position and shape of the target person's hand (step S113). Finally, the control unit 100 measures the distance between the target person's face and hand, and ends the process (step S114).

[0073] The control unit 100 uses a trained model that uses the position of the target person's face, the position and shape of the target person's hands, and the distance between the target person's face and hands as features to estimate whether the target person needs subtitles to be displayed for the content being displayed.

[0074] FIG. 13 is a diagram showing a schematic view of a subtitle display screen W14 displayed on the display screen W10 in accordance with the motion state of a person P14 as a target person.

[0075] In FIG. 13, a person P14 is shown near the display device 10 (in front of the display screen W10) holding his hand over his ear to indicate that he has difficulty hearing the sound output from the display device 10.

[0076] 13, whether or not the target person requires subtitles to be displayed for the currently displayed content is estimated based on the position of the face of person P14 surrounded by frame (detection window) F10, the position and shape of the hand of person P14 surrounded by frame (detection window) F12, and the distance between the position of the face and the hand of person P14. If it is estimated that the target person requires subtitles to be displayed for the currently displayed content, the control unit 100 displays a subtitle display screen W14 on the display screen W10.

[0077] Incidentally, Fig. 14 is a diagram illustrating an example of guidance for the subtitle display function displayed in step S140 of Fig. 11. Fig. 14 shows an example in which a guidance screen W16, displaying the message "Hint: Make a gesture of putting your hand to your ear to turn on subtitles," is superimposed on the display screen W10.

[0078] As described above, according to the second embodiment, subtitles can be displayed for the content being displayed on the display screen according to the viewing state (motion state) of the target person, thereby improving the convenience for the person watching the content.

[0079] [3 Third embodiment] Next, a third embodiment will be described. In the third embodiment, content to be displayed on the display screen is selected according to the viewing state (facial expression, etc.) of the target person.

[0080] The third embodiment can have substantially the same hardware and software configuration as the first embodiment, but differs from the first embodiment in that the estimation unit 1020 estimates the concentration level of the target person as an attribute based on the facial expression of the target person acquired by the acquisition unit 1010, etc.

[0081] FIG. 15 is a flowchart illustrating the flow of processing in the third embodiment.

[0082] When the process starts, the control unit 100 starts a viewing state acquisition process (step S200). When the viewing state acquisition process ends, the control unit 100 performs an attribute estimation process based on the acquired viewing state (facial expression, etc.) (step S210). In this case, the control unit 100 estimates the concentration level of the target person as an attribute of the target person. Note that the concentration level according to the present disclosure is a numerical expression or the like of the target person's concentration (immersion) level in the content being displayed on the display device 10.

[0083] Next, the control unit 100 determines whether the estimated current concentration level exceeds the average value of the accumulated concentration levels (step S220). If it is determined that the current concentration level exceeds the average value, the control unit 100 selects content related to the currently displayed content (step S220; Yes → step S230). On the other hand, if it is determined that the current concentration level is below the average value, the control unit 100 selects other content unrelated to the currently displayed content (step S220; No → step S240).

[0084] Next, the viewing state acquisition process in step S200 in FIG. 15 will be described with reference to the flowchart in FIG.

[0085] When the process starts, the control unit 100 acquires a target image showing the facial expression of the target person, etc., as the viewing state of the target person (step S201).

[0086] The control unit 100 measures the position of the target person's face from the acquired target image (step S202). Next, the control unit 100 determines the target person's facial expression (step S203) and the direction of their gaze (step S204), and ends the process.

[0087] The facial expression of the target person can be determined by detecting fluctuations (changes) in the positions and sizes of each part of the face (e.g., eyes, nose, mouth) and facial muscles. The gaze direction can be determined by detecting fluctuations (changes) in the direction in which the target person's eyes (eyeballs) are facing. Note that a trained model using facial expression and gaze direction as objective variables may be separately prepared, and the facial expression and gaze direction of the target person may be estimated.

[0088] Next, the attribute estimation process in step S210 of FIG. 15 will be described with reference to the flowchart of FIG.

[0089] The control unit 100 estimates the current concentration level of the target person using a trained model that uses the facial expression of the target person and the direction of the target person's gaze as feature quantities (step S211).

[0090] Next, the control unit 100 calculates an average value of the concentration degree from the estimated current concentration degree and the accumulated past concentration degrees, and stores the average value as necessary (step S212).

[0091] Finally, the control unit 100 compares the concentration level estimated in step S211 with the average concentration level calculated in step S212, determines whether the current concentration level exceeds the average accumulated concentration level, and terminates the process (step S213).

[0092] Figure 18 is a schematic diagram showing how the concentration level of a target person P16 is estimated based on the facial expression of the target person, and when the estimated concentration level exceeds the average value, content related to the content being displayed is selected.

[0093] As shown in Fig. 18, the concentration level of a target person P16 surrounded by a frame (detection window) F10 is estimated based on the facial expression of the person P16. If the estimated concentration level of the person P16 exceeds the average value of the accumulated concentration levels, the control unit 100 selects and displays content related to the currently displayed content. In this case, as shown in the example of Fig. 18, it is also possible to superimpose a notification screen W18 on the display screen W10 to notify the user that related content has been selected.

[0094] As described above, according to the third embodiment, the content to be displayed on the display screen can be selected according to the viewing state (facial expression, etc.) of the target person, so that the most suitable content can be displayed without breaking the concentration of the person viewing the content.

[0095] [4 Fourth embodiment] Next, a fourth embodiment will be described. The fourth embodiment is a form that can rotate and display content (hereinafter, sometimes referred to as rotated display) according to the viewing state of a target person (the horizontal angle between both eyes).

[0096] The fourth embodiment can have substantially the same hardware and software configuration as the first embodiment, but differs from the first embodiment in that the estimation unit 1020 estimates the posture of the target person as an attribute based on the acquired viewing state of the target person (the angle of the horizontal line between both eyes).

[0097] FIG. 19 is a flowchart illustrating the flow of processing in the fourth embodiment.

[0098] When the process starts, the control unit 100 determines whether the operation mode of the device is the in-store mode (step S300). If the control unit 100 determines that the operation mode of the device is the in-store mode, it starts the viewing status acquisition process (step S300; Yes → step S310). If the control unit 100 determines that the operation mode of the device is not the in-store mode, it performs display guidance related to the content rotation display function and proceeds to step S310 (step S300; No → step S350 → step S310).

[0099] When the viewing state acquisition process is completed, the control unit 100 performs an attribute estimation process based on the acquired viewing state (the horizontal angle of the target person's eyes) (step S320). In this case, the control unit 100 estimates the posture of the target person as an attribute of the target person.

[0100] Then, the control unit 100 determines whether or not the displayed content needs to be rotated based on the estimated posture (step S330). If it is determined that the displayed content needs to be rotated, the control unit 100 rotates the content displayed on the display screen W10 according to the horizontal angle of the target person's eyes, and ends the process (step S330; Yes → step S340).

[0101] On the other hand, if it is determined that the rotated display of the currently displayed content is not necessary, the control unit 100 returns the process to step S310 (step S330; No→step S310).

[0102] Next, the viewing state acquisition process in step S310 in FIG. 19 will be described with reference to the flowchart in FIG.

[0103] When the process starts, the control unit 100 acquires a target image for which the angle of the horizontal line between the eyes of the target person can be calculated as the viewing state of the target person (step S311).

[0104] The control unit 100 measures the position of the target person's face from the acquired target image (step S312). Next, the control unit 100 sets an imaginary line connecting the vertices of the target person's eyeballs as the horizon (step S313).

[0105] Finally, the control unit 100 measures the angle between the horizontal line set in step S313 and, for example, a vertical line, and ends the process (step S314).

[0106] The control unit 100 estimates the posture of the target person using a trained model that uses the angle between the horizons of both eyes as a feature. Then, the control unit 100 determines whether or not the currently displayed content needs to be rotated based on the estimated posture. If it is determined that the currently displayed content needs to be rotated, the control unit 100 rotates the currently displayed content according to the angle between the horizons of both eyes and displays it.

[0107] FIG. 21 is a diagram showing a schematic representation of the content displayed on the display screen W10 of the display device 10 being rotated in accordance with the angle of the horizontal line between the eyes of person P18 (dotted line in area R10 in the figure).

[0108] 21, whether or not the currently displayed content needs to be rotated is determined based on the angle of the horizontal line between the eyes of person P18 surrounded by frame (detection window) F10. If it is determined that the currently displayed content needs to be rotated, control unit 100 rotates content C10 currently displayed on display screen W10 at a rotation angle according to the angle of the horizontal line between the eyes of person P18.

[0109] Incidentally, Fig. 22 is a diagram illustrating an example of guidance for the subtitle display function displayed in step S350 of Fig. 19. Fig. 23 is an example in which a guidance screen W20, displaying "Hint: The content will rotate if you watch while lying down," is superimposed on the display screen W10.

[0110] As described above, according to the fourth embodiment, the content can be rotated and displayed according to the viewing state of the target person (the angle of the horizontal line between both eyes), thereby improving convenience for people who wish to view content while lying down (relaxed state), for example.

[0111] [5 Fifth embodiment] Next, a fifth embodiment will be described. The fifth embodiment is an embodiment that can control the image quality and sound quality of content displayed on a display screen according to the viewing state (posture) of a target person.

[0112] The fifth embodiment can have substantially the same hardware and software configuration as the first embodiment, but differs from the first embodiment in that the estimation unit 1020 estimates the transition of the target person to a sleep state as an attribute based on the acquired viewing state (posture) of the target person.

[0113] FIG. 23 is a flowchart illustrating the flow of processing in the fourth embodiment.

[0114] When the process starts, the control unit 100 starts the viewing state acquisition process (step S400). When the viewing state acquisition process ends, the control unit 100 performs the attribute estimation process based on the acquired viewing state (posture, bedding use state) (step S410). In this case, the control unit 100 estimates the transition of the target person to a sleep state as an attribute of the target person.

[0115] Then, based on the estimation result, the control unit 100 determines whether the target person has finished getting ready for bed (step S420). If it is determined that the target person has finished getting ready for bed, the control unit 100 controls (suppresses) the image quality (blueness, brightness) and volume (maximum volume) of the content displayed on the display screen W10 by controlling the display unit 140, the audio output unit 150, etc., and ends the process (step S420; Yes → step S430).

[0116] On the other hand, if it is determined that the target person has not completed preparation for bed, the control unit 100 returns the process to step S400 (step S420; No→step S400).

[0117] Next, the viewing state acquisition process in step S400 in FIG. 23 will be described with reference to the flowchart in FIG.

[0118] When the process starts, the control unit 100 acquires a target image from which the posture of the target person can be calculated (step S401).

[0119] The control unit 100 measures the posture of the target person from the acquired target image (step S402). When the target screen is a distance image, the posture of the target person can be determined by determining the position of the center of gravity of each part, such as the head, hands, and feet, of the human body detected from the distance image. When a convolutional neural network is applicable to posture estimation, the position coordinates of each part may be detected from the RGB color image using, for example, a convolutional pose machine or open pose.

[0120] Next, the control unit 100 measures the state of use of the bedding and ends the process (step S403). Note that if bedding such as a futon or pillow is included in the background of the captured image, the state of use of the bedding can be estimated using a trained model that uses the bedding as a feature.

[0121] The control unit 100 uses a trained model that uses the posture of the target person and the state of use of bedding as feature quantities to estimate the target person's transition to a sleep state.

[0122] FIG. 25 is a diagram showing a schematic diagram of how the image quality and volume of content C10 being displayed on the display screen W10 are changed in accordance with the posture of person P20, etc.

[0123] 25, based on the posture of person P20 surrounded by frame (detection window) F10, the transition of person P20 to a sleep state is estimated, and it is determined whether person P20 has finished getting ready to go to bed. If it is determined that person P20 has finished getting ready to go to bed, the control unit 100 controls the image quality and volume of content C10 being displayed on the display screen W10. In this case, as illustrated in FIG. 25, it is also possible to superimpose a notification screen W22 on the display screen W10 (content C10) to notify that the image quality and volume of content C10 have been changed.

[0124] As described above, according to the fifth embodiment, it is possible to control the image quality and sound quality of the content displayed on the display screen according to the viewing state (posture) of the target person, so that the person watching the content can transition to a sleep state with the image quality and volume of the content reduced.

[0125] [6 Sixth Embodiment] Next, a sixth embodiment will be described. The sixth embodiment is a form capable of displaying content that relieves tension according to the viewing state (type, noise, and excitement level) of a target person. In the sixth embodiment, the subject of the target image acquisition is not limited to a person, and therefore, in the following description, the subject will be referred to as an object instead of a target person.

[0126] The sixth embodiment can have substantially the same hardware and software configuration as the first embodiment, but differs from the first embodiment in that the estimation unit 1020 estimates the level of tension of the object based on the viewing state (type, noise, and excitement level) of the acquired object.

[0127] FIG. 26 is a flowchart illustrating the flow of processing in the sixth embodiment.

[0128] When the process starts, the control unit 100 starts the viewing state acquisition process (step S500). When the viewing state acquisition process ends, the control unit 100 performs the attribute estimation process based on the acquired viewing state (type, noise, and excitement level) (step S510). In this case, the control unit 100 estimates the tension level of the object as an attribute of the object.

[0129] Then, the control unit 100 determines whether the estimated tension level exceeds a predetermined threshold (step S520). If it is determined that the tension level exceeds the predetermined threshold, the control unit 100 selects related (tension-relieving) content, displays it on the display screen W10, and ends the process (step S520; Yes → step S530).

[0130] On the other hand, if it is determined that the degree of tension does not exceed the predetermined threshold, the control unit 100 skips step S530 and ends the process (step S520; No).

[0131] Next, the viewing state acquisition process in step S500 in FIG. 26 will be described with reference to the flowchart in FIG.

[0132] When the process starts, the control unit 100 acquires an image representing an object as a target image (step S501). As described above, in the sixth embodiment, the object represented by the target image is not limited to a person, and may be, for example, a pet animal such as a dog, a cat, or a small bird.

[0133] Next, the control unit 100 determines the object represented by the target image (step S502). The object may be determined using a trained model that uses the features of the object as feature quantities.

[0134] Next, the control unit 100 measures the noisiness / excitement level for the classified object and ends the process (step S503). Note that the noisiness / excitement level for an object may also be measured (estimated) using a trained model that uses the object as a feature.

[0135] The control unit 100 estimates the level of tension of the determined object using a trained model that uses the determined object and the noise and excitement level of the object as feature quantities.

[0136] FIG. 28 is a diagram showing a schematic representation of how related (relaxing) content C12 is displayed on the display screen in accordance with the degree of noise and excitement of an object O10.

[0137] 28, the classification result of the object O10 surrounded by the frame (detection window) F10 and the level of tension of the object O10 are used to estimate the level of tension based on the noise and excitement of the object, and it is determined whether the level of tension exceeds a predetermined threshold. If it is determined that the level of tension of the object O10 exceeds the predetermined threshold, the control unit 100 selects content C12 that alleviates tension and displays it on the display screen W10.

[0138] As described above, according to the sixth embodiment, it is possible to display content that reduces tension depending on the viewing state of the target person (type, noise, and excitement level), so that the tension of a person who is in a tense state can be easily reduced.

[0139] [6. Modifications] The present disclosure is not limited to the above-described embodiments, and various modifications are possible. In other words, embodiments obtained by combining technical means that are appropriately modified within the scope of the present disclosure are also included in the technical scope.

[0140] Although the above-mentioned embodiments are described separately for convenience of explanation, they can be combined to the extent possible. Furthermore, the present invention intends to obtain rights to any of the technologies described in the specification through amendments or divisional applications, etc.

[0141] In addition, the programs that run on each device in each embodiment are programs that control the CPU, etc. (programs that make a computer function) so as to realize the functions of the above-described embodiments. Information handled by these devices is temporarily stored in a temporary storage device (e.g., RAM) during processing, and then stored in various ROMs and HDDs, and is read, modified, and written by the CPU as needed.

[0142] Here, the recording medium for storing the program may be any of semiconductor media (e.g., ROM, non-volatile memory card, etc.), optical recording media / magneto-optical recording media (e.g., DVD (Digital Versatile Disc), CD (Compact Disc), BD (Blu-ray (registered trademark) Disc), etc.), magnetic recording media (e.g., magnetic tape, flexible disk, etc.), etc.

[0143] Furthermore, when distributing the program in the market, the program can be stored in a portable recording medium and distributed, or transferred to a server computer connected via a network such as the Internet. In this case, the storage device of the server device is also included in the present disclosure.

[0144] Furthermore, the above-mentioned data may not be stored within the device, but may be stored in an external device and called up as needed. For example, the data may be stored in a network attached storage (NAS) or on the cloud.

[0145] The scope of the present disclosure is not limited to the configurations explicitly described in the specification, but also includes combinations of the technologies disclosed in the specification. The configurations of the present disclosure for which a patent is sought are set forth in the appended claims, but it is not intended to exclude them from the technical scope on the grounds that they are not set forth in the claims.

[0146] Furthermore, in the above-mentioned specification, the statements "in the case of" and "when" are given as examples and are not intended to limit the configuration to the described contents. The disclosure also includes configurations that are not in these cases or situations, even if they would be obvious to a person skilled in the art, and the applicant intends to obtain rights to them.

[0147] Furthermore, the processes and data flows described in the specification are not limited to the order in which they are described. For example, the patent also discloses configurations in which some processes are deleted or the order is changed, and the patent holder intends to obtain the rights to such configurations.

[0148] Furthermore, although the functions described in the embodiments are executed by each device, they may be realized by one device or may further utilize an external server.

[0149] Furthermore, each functional block or feature of the device used in the above-described embodiments may be implemented or performed by an electrical circuit, for example, an integrated circuit or multiple integrated circuits. The electrical circuit designed to perform the functions described herein may include a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or a combination thereof. The general-purpose processor may be a microprocessor, or a conventional processor, controller, microcontroller, or state machine. The electrical circuit may be composed of digital circuits or analog circuits. Furthermore, as advances in semiconductor technology emerge that replace current integrated circuits, one or more aspects of the present disclosure may also utilize new integrated circuits based on that technology. [Explanation of symbols]

[0150] 10 Processing equipment 100 control section 102 Operation control section 104 Broadcast Control Unit 110 Storage 120 ROM 130 RAM 140 Display section 150 Audio output section 160 Audio input section 170 Photography Department 180 Communications Department

Claims

1. a display unit that displays the content on a display screen; an acquisition unit that acquires, from a captured image obtained by capturing an image of the periphery of the display screen, a target image representing a target person viewing the content as a viewing state of the target person; an estimation unit that estimates attributes of the target person based on the acquired target image; a control unit that selects the content to be displayed on the display screen or controls display of the content currently being displayed based on an estimation result by the estimation unit.

2. The estimation unit The display device according to claim 1 , wherein the attribute is estimated using a trained model that inputs the target image and outputs an analysis result of the target image.

3. The acquisition unit 2. The display device according to claim 1, wherein the target image representing a stationary state of the target person in front of the display unit is acquired as the viewing state.

4. The acquisition unit 2. The display device according to claim 1, wherein the target image is acquired by excluding people who are store salespeople from the target people.

5. The estimation unit If there is a person corresponding to the store salesperson, the person who is having a conversation with the store salesperson is set as the target person; 5. The display device according to claim 4, wherein, when there is no person corresponding to the store salesperson, the person gazing at the display unit is used as the target person to estimate the attributes.

6. The acquisition unit acquiring the target image representing the motion state of the target person as the viewing state; The estimation unit As the attribute, it is estimated whether the target person needs subtitles to be displayed for the content; The control unit 2. The display device according to claim 1, wherein the display device displays the content with subtitles added based on the estimation result.

7. The acquisition unit acquiring the target image representing the facial expression of the target person as the viewing state; The estimation unit As the attribute, a concentration level of the target person on the content is estimated; The control unit The display device according to claim 1 , wherein the content to be displayed on the display screen is selected based on the estimation result.

8. The acquisition unit acquiring the target image from which the angle of the horizontal line between both eyes of the target person can be calculated as the viewing state; The estimation unit As the attribute, a posture of the target person is estimated; The control unit 2. The display device according to claim 1, wherein the content displayed on the display screen is rotated at a rotation angle corresponding to the angle of the horizontal line between the eyes based on the estimation result.

9. The control unit It is possible to control audio output accompanying the display of the content, The acquisition unit acquiring the target image for which the posture of the target person can be calculated as the viewing state; The estimation unit As the attribute, a transition of the target person to a sleep state is estimated; The control unit The display device according to claim 1 , wherein the image quality and volume of the content displayed on the display screen are controlled based on the estimation result.

10. The acquisition unit acquiring, as the viewing state, the target image for which the type of the target person and the noise and excitement level of the target person can be calculated; The estimation unit As the attribute, a level of tension of the target person is estimated; The control unit The display device according to claim 1 , wherein the content that alleviates the tension level is selected based on the estimation result.

11. a display step of displaying the content on a display screen; an acquisition step of acquiring a target image representing a target person viewing the content as a viewing state of the target person from a captured image obtained by capturing an image of the periphery of the display screen; an estimation step of estimating attributes of the target person based on the acquired target image; a control step of selecting the content to be displayed on the display screen or controlling the display of the content currently being displayed based on the estimation result in the estimation step.

12. On the computer, a display function for displaying the content on a display screen; an acquisition function for acquiring, from a captured image obtained by capturing an image of the periphery of the display screen, a target image representing a target person viewing the content as a viewing state of the target person; an estimation function for estimating attributes of the target person based on the acquired target image; and a control function for selecting the content to be displayed on the display screen or for controlling the display of the content currently being displayed based on the estimation result of the estimation function.

13. A recording medium on which the program according to claim 12 is recorded.

Citation Information

Patent Citations

  • Reproducing apparatus, reproducing method, program, and recording medium

    JP2011166423A