Artificial intelligence information processing device and artificial intelligence information processing method

By using artificial intelligence information processing devices and neural network technology, combined with sensor information to estimate user gaze, the problem of existing technologies being unable to reflect users' interest in specific scenarios of video content has been solved, achieving more accurate content recommendation and advertising effects.

CN113853598BActive Publication Date: 2025-12-30SONY GROUP CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202080037794.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-05-27
Filing Date
2020-03-09
Publication Date
2025-12-30
Estimated Expiration
2040-03-09

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately reflect users' interests in specific scenes within video content, resulting in poor content recommendation and advertising effectiveness.

Method used

Using an artificial intelligence information processing device, the system estimates user gaze degree through sensor information and neural networks, acquires and infers scene information of user gaze, and uses neural networks to learn the correlation between video and sensor information to generate metadata representing the scene of user gaze.

Benefits of technology

It enables accurate estimation of user gaze patterns in video content, improving the targeting of content recommendations and advertisements, and enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113853598B_ABST
    Figure CN113853598B_ABST
Patent Text Reader

Abstract

The present application provides an artificial intelligence information processing device for generating scene-related information using artificial intelligence, comprising: a gaze degree estimation unit configured to estimate, by artificial intelligence, a gaze degree of a user viewing content based on sensor information; an acquisition unit configured to acquire, based on an estimation result of the gaze degree estimation unit, a video of a scene at which the user is gazing in the content and information related to the content; and a scene information estimation unit configured to estimate, by artificial intelligence, information related to the scene at which the user is gazing based on the video of the scene at which the user is gazing and the information related to the content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology disclosed in this invention relates to an artificial intelligence information processing device and an artificial intelligence information processing method for processing information about content based on artificial intelligence. Background Technology

[0002] Television broadcasting services have been widely used for a long time. Television receivers are now so common that almost every household has one or more installed. Recently, mobile image distribution services using network broadcasting methods, such as Internet Protocol Television (IPTV) and Over-the-Top (OTT), have also been gaining traction.

[0003] Furthermore, a technique for combining television receivers and sensing technology to measure "viewing and listening quality," which indicates the degree of viewer attention to video content, has recently been researched and developed (e.g., see Patent Document 1). Methods for using viewing and listening quality are diverse. For example, the effectiveness of video content and advertising can be evaluated based on the measurement results of viewing and listening quality, and other content and products can be recommended to viewers.

[0004] Existing technical documents

[0005] Patent documents

[0006] Patent Document 1: WO 2017 / 120469

[0007] Patent Document 2: JP 4840393 B

[0008] Patent Document 3: JP 2007-143010 A

[0009] Patent Document 4: JP 2008-236779 A Summary of the Invention

[0010] Technical issues

[0011] The purpose of the technology disclosed in this invention is to provide an artificial intelligence information processing device and method for processing content-related information based on artificial intelligence.

[0012] Solution to the problem

[0013] The first aspect of the technology disclosed in this invention is an artificial intelligence information processing device, comprising:

[0014] The gaze estimation unit is configured to estimate the gaze level of the user viewing the content based on sensor information and artificial intelligence.

[0015] The acquisition unit is configured to acquire video footage of the user-gaze-inducing scene within the content and information about the content, based on the estimation results from the gaze estimation unit.

[0016] The scene information estimation unit is configured to estimate scene information about the user's gaze based on the video of the scene and information about the content, using artificial intelligence.

[0017] The scene information estimation unit uses a neural network that has learned the correlation between video of the scene and information about the content and information about the scene to estimate information that is relevant to the scene being viewed by the user, as an estimate based on artificial intelligence.

[0018] In addition, the gaze estimation unit uses a neural network that has learned the correlation between sensor information and the user's gaze to estimate the gaze that is correlated with the sensor information about the user who is viewing the content, as an estimate based on artificial intelligence.

[0019] Furthermore, a second aspect of the technology disclosed in this specification is an artificial intelligence information processing device, comprising:

[0020] The input unit is configured to receive sensor information about the user viewing the content, and

[0021] The scene information estimation unit is configured to use a neural network that has learned the correlation between sensor information, content, and information about the content and information about the scene being viewed by the user to estimate information relevant to the scene being viewed by the user.

[0022] Furthermore, a third aspect of the technology disclosed in this specification is an artificial intelligence information processing method, comprising:

[0023] The gaze estimation step uses artificial intelligence to estimate the gaze level of a user viewing content based on sensor information.

[0024] The acquisition step, based on the estimation results in the gaze estimation step, is used to acquire video footage of the user-gaze scene within the content and information about the content, as well as...

[0025] The scene information estimation step is used to estimate scene information about the user's gaze based on the video of the scene and information about the content, using artificial intelligence.

[0026] Advantages of the invention

[0027] Based on the technology disclosed in this specification, an artificial intelligence information processing apparatus and method can be provided for estimating metadata about a scene based on artificial intelligence for partial content.

[0028] Furthermore, the effects described in this specification are merely illustrative, and the effects provided by the technology disclosed in this specification are not limited thereto. In addition, the technology disclosed in this specification can also achieve additional effects beyond those described above.

[0029] Other objects, features, and advantages of the technology disclosed in this specification will become clear from the detailed description based on the embodiments and accompanying drawings described below. Attached Figure Description

[0030] Figure 1 This is a diagram illustrating an example configuration of a system used for viewing video content.

[0031] Figure 2 This is a diagram showing an example configuration of a television receiver 100.

[0032] Figure 3 This is a diagram showing an example configuration of the sensor group 300 installed in the television receiver 100.

[0033] Figure 4 This is a diagram illustrating a configuration example of the scene acquisition and scene information estimation system 400.

[0034] Figure 5 This is a diagram illustrating an example configuration of the neural network 500 used in the gaze estimation unit 406.

[0035] Figure 6 This is a diagram illustrating an example configuration of the neural network 600 used in the scene information estimation unit 408.

[0036] Figure 7 This is a flowchart illustrating the processing procedure performed by the scene acquisition and scene information estimation system 400.

[0037] Figure 8 This is a diagram illustrating a variant of the scene acquisition and scene information estimation system 800.

[0038] Figure 9 This is a diagram illustrating an example configuration of the neural network 900 used in the scene information estimation unit 806.

[0039] Figure 10 It is a diagram showing the state of a user watching television, captured by a camera.

[0040] Figure 11 This is a diagram illustrating the state of facial recognition being performed using images captured from a camera.

[0041] Figure 12 This is a diagram showing the state of a gaze scene captured based on inferences about the user's gaze level.

[0042] Figure 13 This is a diagram showing an example configuration of a list screen displaying the scenes the user has been viewing.

[0043] Figure 14 This is a diagram showing an example of the screen configuration that changes after a scene is selected.

[0044] Figure 15 This is a diagram showing an example configuration of a screen that displays relevant content.

[0045] Figure 16 It is a diagram showing the state of the scene that the user has been looking at and how it is reflected on the smartphone screen.

[0046] Figure 17 This is a diagram showing examples of layouts for multiple scenes.

[0047] Figure 18 This is a diagram showing examples of layouts for multiple scenes.

[0048] Figure 19 This is a diagram illustrating an application example of panel speaker technology.

[0049] Figure 20 This is a diagram illustrating a configuration example of an AI system 2000 using the cloud. Detailed Implementation

[0050] In the following description, embodiments of the technology disclosed in this specification will be described in detail with reference to the accompanying drawings.

[0051] A. System Configuration

[0052] Figure 1 This illustration shows an example configuration of a system used for viewing video content.

[0053] The television receiver 100 is equipped with a large screen for displaying video content and speakers for outputting sound. The television receiver 100 includes a tuner for selecting and receiving broadcast signals, or, for example, connects to a set-top box, and thus can use broadcast services provided by television stations. The broadcast signal can be either ground wave or satellite wave.

[0054] Furthermore, the television receiver 100 can also use broadcast mobile image distribution services, such as those using networks like IPTV and OTT. Therefore, the television receiver 100 is equipped with a network interface card and connects to external networks such as the Internet via a router or access point using existing communication standards such as Ethernet (registered trademark) and Wi-Fi (registered trademark).

[0055] A streaming distribution server provides video streaming services over the Internet and broadcast-type mobile image distribution services to a television receiving device 100.

[0056] Furthermore, countless servers are available on the Internet to provide various services. One example of a server is a streaming server. On one side of the television receiver 100, a browser can be enabled, and, for example, a Hypertext Transfer Protocol (HTTP) request can be made for the streaming server to use network services.

[0057] Furthermore, it is assumed that this embodiment also includes an artificial intelligence server (not shown) that provides artificial intelligence functions to clients on the Internet (or the cloud). Here, the artificial intelligence functions are, for example, functions obtained by artificially implementing general functions of the human brain (such as learning, reasoning, data creation, and planning) through software or hardware. Additionally, for example, the artificial intelligence server may be equipped with a neural network that performs deep learning (DL) based on a model that mimics the neural circuits of the human brain. This neural network includes a mechanism by which artificial neurons (nodes) forming a network based on synaptic combinations acquire problem-solving abilities while changing the strength of synaptic combinations according to learning. The neural network can automatically infer problem-solving rules through iterative learning. Meanwhile, the "artificial intelligence server" mentioned in this specification is not limited to a single server device and can take the form of a cloud, for example, providing cloud computing services.

[0058] Figure 2 An example configuration of a television receiver 100 is shown. The television receiver 100 includes: a main control unit 201, a bus 202, a storage unit 203, a communication interface (IF) unit 204, an expansion interface (IF) unit 205, a tuner / demodulation unit 206, a demultiplexer (DEMUX) 207, a video decoder 208, an audio decoder 209, a character overlay decoder 210, a subtitle decoder 211, a subtitle synthesis unit 212, a data decoder 213, a cache unit 214, an application program (AP) control unit 215, a browser unit 216, a sound source unit 217, a video synthesis unit 218, a display unit 219, an audio synthesis unit 220, an audio output unit 221, and an operation input unit 222.

[0059] The main control unit 201 includes, for example, a controller, a read-only memory (ROM) (assuming the ROM includes a rewritable ROM such as an electrically erasable programmable ROM (EEPROM), and random access memory (RAM), and typically controls the overall operation of the television receiver 100 according to a predetermined operating procedure. The controller is a central processing unit (CPU), microprocessor unit (MPU), graphics processing unit (GPU), general-purpose graphics processing unit (GPGPU), etc., and the ROM is a non-volatile memory storing basic operating procedures such as the operating system (OS) and other operating programs. Operating settings required for the operation of the television receiver 100 can be stored in the ROM. The RAM is the working area when the OS and other operating programs are executed. The bus 202 is a data communication path for transmitting / receiving data between the main control unit 201 and each component in the television receiver 100.

[0060] Storage unit 203 is configured as a non-volatile storage device such as flash ROM, solid-state drive (SSD), or hard disk drive (HDD). Storage unit 203 stores the operating procedures and operating settings of the television receiver 100, personal information of the user of the television receiver 100, etc. Furthermore, storage unit 203 stores operating procedures downloaded via the Internet, various types of data created by the operating procedures, etc. In addition, storage unit 203 can also store content such as moving images, still images, and sound obtained through broadcast signals and the Internet.

[0061] The communication interface unit 204 is connected to the Internet via a router (as described above) to perform data transmission / reception to / from each server device or other communication device on the Internet. Furthermore, it is assumed that the communication interface unit 204 acquires data streams of programs transmitted via communication lines. The router can be either a wired connection such as Ethernet (registered trademark) or a wireless connection such as Wi-Fi (registered trademark).

[0062] The tuner / demodulation unit 206 receives broadcast waves, such as terrestrial or satellite broadcasts, via an antenna (not shown) and tunes to (selects) the channel of the service (broadcasting station, etc.) desired by the user, based on the control of the main control unit 201. Furthermore, the tuner / demodulation unit 206 demodulates the received broadcast signal to obtain a broadcast data stream. Simultaneously, to display multiple screens simultaneously, record programs on different channels, etc., a configuration in which multiple tuner / demodulation units (i.e., multiple tuners) are installed in the television receiver 100 can be adopted.

[0063] The demultiplexer 207, based on control signals input to the data streams of the video decoder 208, audio decoder 209, character overlay decoder 210, and subtitle decoder 211, allocates the video data stream, audio data stream, character overlay data stream, and subtitle data stream, which are elements of real-time presentation, to the video decoder 208, audio decoder 209, character overlay decoder 210, and subtitle decoder 211, respectively. The data streams input to the demultiplexer 207 include broadcast data streams according to broadcast services and distribution data streams according to distribution services such as IPTV or OTT. The former is input to the demultiplexer 207 after being selected, received, and demodulated by the tuner / demodulation unit 206, and the latter is input to the demultiplexer 207 after being received by the communication interface unit 204. Furthermore, the demultiplexer 207 reproduces multimedia application or file data that is a component of it and outputs it to the application control unit 215 or temporarily stores it in the cache unit 214.

[0064] Video decoder 208 decodes the video stream input from demultiplexer 207 and outputs video information. Similarly, audio decoder 209 decodes the audio stream input from demultiplexer 207 and outputs audio information. In digital broadcasting, video and audio streams encoded according to standards such as MPEG2 are multiplexed and transmitted or distributed. Video decoder 208 and audio decoder 209 perform decoding processing on the encoded video stream and the demultiplexed video stream by demultiplexer 207 according to standardized decoding methods. Meanwhile, television receiving device 100 may include multiple video decoders 208 and audio decoders 143 to simultaneously decode multiple types of video and audio data streams.

[0065] The character overlay decoder 210 decodes the character overlay data stream input from the demultiplexer 207 and outputs character overlay information. The subtitle decoder 211 decodes the subtitle data stream input from the demultiplexer 207 and outputs subtitle information. The subtitle synthesis unit 212 performs synthesis processing on the character overlay information output from the character overlay decoder 210 and the subtitle information output from the subtitle decoder 211.

[0066] Data decoder 213 decodes the data stream obtained by multiplexing video and audio data in an MPEG-2 TS stream. For example, data decoder 213 notifies main control unit 201 of the decoding results of general event messages stored in the descriptor area of ​​the Program Map Table (PMT), which is one of the Program Specific Information (PSI) tables.

[0067] The application control unit 215 receives control information included in the broadcast data stream from the demultiplexer 207, or obtains control information from a server device on the Internet via the communication interface unit 204, and decodes the control information.

[0068] According to the instructions of the application control unit 215, the browser unit 216 presents multimedia application files or file data as components thereof obtained from a server device on the Internet via the cache unit 214 of the communication interface unit 204. The multimedia application files mentioned here may be, for example, Hypertext Markup Language (HTML) documents, Broadcast Markup Language (BML) documents, etc. Furthermore, the browser unit 216 also performs audio information reproduction of the application by acting on the sound source unit 217.

[0069] The video compositing unit 218 receives video information output from the video decoder 208, subtitle information output from the subtitle compositing unit 212, and application information output from the browser unit 216, and performs appropriate selection or overlay processing. The video compositing unit 218 includes a video RAM (illustration of which is omitted), and drives the display of the display unit 219 based on the video information input to the video RAM. Furthermore, the video compositing unit 218 also performs, as needed, processing based on the control of the main control unit 201 to overlay the Electronic Program Guide (EPG) screen with screen information such as graphics generated by an application executed by the main control unit 201.

[0070] Display unit 219 is a display device configured as, for example, a liquid crystal display (LCD), an organic electroluminescent (EL) display, etc., and presents video information to the user that has undergone selection or overlay processing by video synthesis unit 218. Alternatively, a liquid crystal display device (e.g., see Patent Document 2) that divides a transmissive liquid crystal panel into multiple display areas (blocks) and uses backlighting to individually incident light on each display area can be used as display unit 219. This type of display device has the advantage of controlling the amount of incident light on each display area to expand the dynamic range of brightness of the displayed image.

[0071] The audio synthesis unit 220 receives audio information output from the audio decoder 209 and audio information of the application reproduced by the sound source unit 217, and performs processing such as appropriately selecting or synthesizing audio information.

[0072] Audio output unit 221 is used for outputting audio of program content or data broadcast content selected and received by tuner / demodulation unit 206, as well as outputting audio information (including synthesized audio such as audio instructions or audio proxies) processed by audio synthesis unit 220. Audio output unit 221 is configured as a sound generating element such as a loudspeaker. For example, audio output unit 221 may be a loudspeaker array (multi-channel loudspeaker or super-multi-channel loudspeaker) configured as a combination of multiple loudspeakers, and some or all of the loudspeakers may be externally connected to television receiver 100.

[0073] In addition to cone-shaped speakers, planar speakers (e.g., see Patent Document 3) can also be used for the audio output unit 221. Of course, a speaker array configured as a combination of different types of speakers can also be used as the audio output unit 221. Furthermore, the speaker array can include a speaker array that performs audio output by vibrating the display unit 219 using one or more exciters (actuators) that generate vibrations. The exciters (actuators) can be in the form of being attached to the back of the display unit 219. Figure 19 An example of applying panel speaker technology to a display is shown. The display 1900 is supported at its rear by a bracket 1902. A speaker unit 1901 is attached to the rear of the display 1900. An exciter 1901-1 is provided on the left edge of the speaker unit 1901, and an exciter 1901-2 is provided on the right edge of the speaker unit 1901, forming a speaker array. The exciters 1901-1 and 1901-2 can cause the display 1901 to vibrate based on left and right audio signals to output sound. The bracket 1902 may include a built-in subwoofer for outputting low-frequency sound. Meanwhile, the display 1900 corresponds to a display unit 219 using organic EL elements.

[0074] Return to reference Figure 2 The configuration of the television receiver 100 will be described below. The operation input unit 222 is a command input unit through which a user inputs operation commands for the television receiver 100. The operation input unit 222 includes, for example, a remote control receiver unit that receives commands sent from a remote control (not shown) and operation keys configured with push-button switches. Furthermore, the operation input unit 222 may include a touch panel superimposed on the screen of the display unit 219. Additionally, the operation input unit 222 may include an external input device, such as a keyboard, connected to the expansion interface unit 205.

[0075] The expansion interface unit 205 is an interface group used to expand the functionality of the television receiver 100, and includes, for example, analog video / audio interfaces, universal serial bus (USB) interfaces, memory interfaces, etc. The expansion interface unit 205 may include digital interfaces consisting of DVI terminals, HDMI terminals, display ports, etc.

[0076] In this embodiment, the expansion interface 205 is also used to receive a group of sensors (refer to later). Figure 3The description includes interfaces for sensor signals from various sensors. It is assumed that the sensors include both sensors mounted within the main body of the television receiver 100 and sensors externally connected to the television receiver 100. Externally connected sensors also include sensors built into other consumer electronics (CE) devices and Internet of Things (IoT) devices that exist in the same space as the television receiver 100. The expansion interface 205 can take in sensor signals after signal processing such as noise removal and digital conversion has been performed, or it can take in sensor signals as unprocessed raw (RAW) data (analog waveform signals).

[0077] B. Sensing Function

[0078] One purpose of equipping the television receiver 100 with various sensors is to measure or estimate the degree of fixation (viewing and listening quality) when a user watches video content displayed on the display unit 219. Generally, if the satisfaction with the video content is high, the degree of fixation on a particular scene tends to increase. Accordingly, the degree of fixation can also be referred to as the "degree of satisfaction" with the video content. That is, it is assumed that "degree of fixation" in the present description is synonymous with the expression "degree of satisfaction". Furthermore, when the term "user" is simply referred to in this specification, unless more specifically specified, it is assumed that the user refers to a viewer watching the video content displayed on the display unit 219 (including cases where the viewer plans to watch).

[0079] Figure 3 An example configuration of a sensor group 300 installed in a television receiver 100 is shown. The sensor group 300 includes a camera unit 310, a status sensor unit 320, an environmental sensor unit 330, a device status sensor unit 340, and a user profile sensor unit 350.

[0080] The camera unit 310 includes a camera 311 for capturing a user watching video content displayed on the display unit 219, a camera 312 for capturing video content displayed on the display unit 219, and a camera 313 for capturing the indoor space (or installation environment) where the television receiver 100 is installed.

[0081] Camera 311 is positioned, for example, near the center of the upper edge of the screen of display unit 219, and is used to capture images of the user watching video content. In this embodiment, it is assumed that camera 311 is necessary.

[0082] Camera 312 is configured, for example, to face the screen of display unit 219 and capture the video content being viewed by the user. Alternatively, the user may wear goggles in which camera 312 is mounted. Additionally, it is assumed that camera 312 also includes the function of recording audio from the video content. However, camera 312 is not necessary when the television receiver 100 includes a buffer (described later) for temporarily storing the video and audio streams to be output.

[0083] Camera 313 may be configured as, for example, a panoramic or wide-angle camera, and photographs the indoor space (or installation environment) where the television receiver 100 is mounted. Alternatively, camera 313 may be a camera mounted on a camera stand (platform) that can rotate, for example, about axes of roll, pitch, and yaw. However, camera 310 is not necessary when sufficient environmental data can be obtained through environmental sensor 330 or when environmental data itself is not required.

[0084] The state sensor unit 320 includes one or more sensors that acquire state information about the user's state. The state sensor unit 320 intends to acquire, for example, the user's working state (whether the user is watching video content), the user's behavioral state (such as stillness, walking or running, eyelid opening / closing, gaze direction, or pupil size), mental state (such as the user's level of engagement with or focus on the video content, level of excitement, level of arousal, mood, emotion, etc.), and physiological state as state information. The state sensor unit 320 may include various sensors, such as perspiration sensors, electromyography (EMG) sensors, electrooculography (EOG) sensors, electroencephalogram (EEG) sensors, exhalation sensors, gas sensors, ion concentration sensors, inertial measurement units (IMUs) for measuring user behavior, and voice sensors (microphones, etc.) for collecting user speech.

[0085] The environmental sensor unit 330 includes various sensors that measure information about the environment of an indoor space, such as the one where the television receiver 100 is installed. For example, the environmental sensor 330 includes: a temperature sensor, a humidity sensor, an optical sensor, an illuminance sensor, an airflow sensor, an odor sensor, an electromagnetic wave sensor, a geomagnetic sensor, a global positioning system (GPS) sensor, and a sound sensor (microphone, etc.) for collecting ambient sounds.

[0086] The device status sensor unit 340 includes one or more sensors for acquiring the internal status of the television receiver 100. Alternatively, circuit components such as the video decoder 208 or the audio decoder 209 may have functions such as outputting input signals to the outside and processing the input signals, and may also serve as sensors for detecting the internal status of the device. Furthermore, the device status sensor unit 340 may detect operations performed by the user on the television receiver 100 or other devices, or store the user's past operation history.

[0087] The user profile sensor unit 350 detects profile information of users watching video content via the television receiver 100. The user profile sensor unit 350 does not necessarily need to be configured as a sensor element. For example, user profiles such as the user's age and gender can be detected based on a facial image captured using the camera 311 or the user's speech collected using a voice sensor. Furthermore, user profiles acquired by a multi-functional information terminal (such as a smartphone) carried by the user can be obtained via a link between the television receiver 100 and a smartphone. However, the user profile sensor unit 350 does not need to detect confidential information regarding the user's privacy or secrets. Moreover, it is not necessary to detect the profile of the same user each time the user watches video content, and previously acquired user profile information can be stored, for example, in the EEPROM of the main control unit 201 (as described above).

[0088] Furthermore, through the link between the television receiver 100 and the smartphone, a user-carried multi-functional information terminal (such as a smartphone) can be used as a status sensor unit 320, an environmental sensor unit 330, or a user profile sensor unit 350. For example, sensor information acquired by sensors built into the smartphone and data managed through applications such as healthcare functions (pedometers, etc.), calendars or timebooks / memos, emails, and social networking services (SNS) can be added to the user status data and environmental data.

[0089] C. Scene information estimation based on gaze level

[0090] The television receiver 100 according to this embodiment can be adapted to... Figure 3 The combination of sensing functions shown is used to measure or estimate a user's gaze relative to video content. Various methods exist for utilizing gaze. For example, an information providing device has been proposed that derives a user's tastes and objects of interest based on a total result of attribute information from the content the user has already viewed, and provides the user with viewing support information and value-added information (see Patent Document 4).

[0091] The attribute information mentioned in this article includes, for example, the content type, performers, and keywords, and this attribute information can be extracted from, for example, information associated with the video content (i.e., so-called metadata). Furthermore, metadata is obtained through various acquisition paths, and there are various situations, such as metadata overlaid on the video content, metadata distributed as broadcast data associated with the main version of the broadcast content, and metadata obtained through paths different from the broadcast content (e.g., obtaining the broadcast content's metadata via the internet). In any case, the content's attribute information is often assigned to the entire content on the content creator's or distributor's side.

[0092] The attribute information about content that users are interested in can be used for various applications, such as supporting user viewing (e.g., automatically recording appointments, recommending other content, etc.), providing feedback and content evaluation for content creators or distributors, and promoting the sales of related products.

[0093] However, attribute information about the entire content is usually information that characterizes the entire content (including attribute information in content metadata), and not necessarily attribute information that characterizes individual scenes within the content. For users, information characterizing individual scenes (attribute information or metadata) rather than information characterizing the entire content can be information that users interested in that scene have a higher level of interest in. For example, when content creators or distributors, advertising distributors, etc., perform viewing support, content evaluation, product recommendations, etc., based on the correlation between user gaze measurements of each scene in the content and information characterizing the entire content, there is a possibility that user interests may not be accurately reflected. Therefore, it is conceivable that content or products based on attribute information that the user is not personally interested in might be recommended. Therefore, to achieve more accurate content recommendations, it is necessary to obtain information characterizing each scene that the user is interested in. However, for content creators, it is difficult to attach different types of attribute information to all scenes when creating content.

[0094] Therefore, in this specification, the following will propose a technique for extracting specific scenes that users are watching in video content based on artificial intelligence, inferring metadata (which can be called tags when neural networks are used in artificial intelligence) as information about those specific scenes, and automatically outputting that metadata.

[0095] Figure 4 A configuration example of a scene acquisition and scene information estimation system 400 is shown. Use as needed. Figure 2 The system 400 shown is configured using components in the television receiver 100 or external devices of the television receiver 100 (such as cloud server devices).

[0096] The receiving unit 401 receives video content and metadata associated with the video content. The video content includes broadcast content transmitted from broadcasting stations (radio towers, broadcast satellites, etc.) and streaming content distributed from streaming distribution servers such as OTT services. Furthermore, the metadata received by the receiving unit 401 is assumed to be metadata for the entire content assigned to the content creator or distributor. The receiving unit 401 then separates (demultiplexes) the received signal into a video stream, an audio stream, and metadata, and outputs them to the signal processing unit 402 and the buffer unit 403 in subsequent stages.

[0097] The receiving unit 401 may consist, for example, of the tuner / demodulation unit 206, the communication interface unit 204, and the demultiplexer 207 in the television receiving device 100.

[0098] The signal processing unit 402 comprises, for example, a video decoder 2080 and an audio decoder 209 from the television receiver 100. It decodes the video and audio data streams input from the receiving unit 401 and outputs the video and audio information to the output unit 404. Furthermore, the signal processing unit 402 can output the decoded video and audio data streams to the buffer unit 403.

[0099] The output unit 404 may consist of, for example, the display unit 219 and the audio output unit 221 in the television receiver 100, displaying video information on the screen and outputting audio information through speakers or the like.

[0100] Buffer unit 403 includes a video buffer and an audio buffer, and temporarily stores the video and audio information decoded by signal processing unit 402 only for a certain period of time. This period of time corresponds to, for example, the processing time required to obtain the user's gaze from the video content. Buffer unit 403 may be, for example, RAM or other buffer memory (not shown) in main control unit 201.

[0101] Sensor unit 405 is basically composed of Figure 3 The sensor group 300 shown is configured as follows. However, only the camera 311 that captures the user watching the video content is necessary, while devices such as other cameras, status sensors 320, and environment sensors 330 are optional. For example, Figure 4 The scene acquisition and scene information estimation system 400 shown includes a buffer unit 403, and can therefore identify the scene of the video content that the user is watching, and then record that scene. Therefore, it is not necessary for the camera 312 of the display unit 219 of the television receiver 100 outside the television receiver 100 to identify the scene of the video content.

[0102] The sensor unit 405 outputs the facial image of the user captured by the camera 311 while the user is watching the video content output from the output unit 404 to the gaze estimation unit 406. In addition, the sensor unit 405 can also output images captured by the camera 313, user status information sensed by the status sensor unit 320, and environmental information of the indoor space sensed by the environment sensor unit 330 to the gaze estimation unit 406.

[0103] The gaze estimation unit 406 estimates the gaze level of the user regarding the video content being viewed based on sensor signals output from the sensor unit 405 and using artificial intelligence. In this embodiment, the gaze estimation unit 406 estimates the user's gaze level based primarily on the recognition results of the user's facial image captured by the camera 311 and using artificial intelligence. For example, the gaze estimation unit 406 estimates and outputs the user's gaze level based on image recognition results of facial expressions such as opening the user's pupils or opening the user's mouth. Of course, the gaze estimation unit 406 can also receive sensor signals other than those captured by the camera 311 and estimate the user's gaze level based on artificial intelligence.

[0104] Since the reasoning capabilities of artificial intelligence are provided to the gaze estimation unit 406, a trained neural network can also be used. Figure 5 An example configuration of a gaze estimation neural network 500 used in gaze estimation unit 406 is shown. The gaze estimation neural network 500 includes: an input layer 510 that receives image signals captured by camera 311 and other sensor signals; an intermediate layer 520; and an output layer 530 that outputs the user's gaze. In the illustrated example, the intermediate layer 520 includes multiple intermediate layers 521, 522, ..., and the neural network 500 can perform deep learning. Furthermore, considering the processing of time-series information such as moving images and sounds, the neural network 500 can have a recurrent neural network (RNN) structure including recursive combinations in the intermediate layer 520.

[0105] The input layer 510 includes a stream of moving images (or still images) captured by the camera 311 in its input vector elements. Basically, it assumes that the image signal captured by the camera 311 is input to the input layer 510 while the image signal is RAW data.

[0106] Furthermore, when sensor signals from other sensors besides those captured by camera 311 are used to measure gaze, input nodes corresponding to each sensor signal are added to and arranged in the input layer 510. Additionally, a convolutional neural network (CNN) can be used as input for image signals and audio signals, etc., to perform feature point aggregation processing.

[0107] The output layer includes an output node 530 that outputs the gaze estimation results from the image signal and sensor signal as a gaze level represented by, for example, continuous (or discrete) values ​​from 0 to 100. When the output node 530 outputs a gaze level consisting of continuous values, the scene corresponding to whether the user is "gazing" or "not gazing" can be determined based on whether the output gaze level exceeds a predetermined value. Alternatively, the output node 530 can output discrete values ​​representing whether the user is "gazing" or "not gazing".

[0108] During the learning process of the gaze estimation neural network 500, a combination of facial images or other sensor signals and the expanded amount of the user's gaze is input into the gaze estimation neural network 500, and the weighting coefficients (inference coefficients) of each node in the intermediate layer 520 are updated, increasing the binding strength between the apparent possible gaze of the facial images or other sensor signals and the output nodes, in order to learn the correlation between the user's facial images (which may include other sensor signals) and the user's gaze. Then, in the processing using the gaze estimation neural network 500 (gaze estimation), when facial images captured by camera 311 and other types of sensor information are input into the trained gaze estimation neural network 500, the user's gaze is output with high accuracy.

[0109] For example, implemented in the main control unit 201 Figure 5 The gaze estimation neural network 500 is shown. Therefore, the main control unit 201 may include a processor dedicated to the neural network. Although the gaze estimation neural network 500 can be provided via the cloud on the Internet, it is desirable for the gaze estimation neural network 500 to be arranged in the television receiver 100 to estimate the gaze of the video content in real time.

[0110] For example, a television receiver 100 has been released that includes a gaze estimation neural network 500 trained using an expert teaching database. The gaze estimation neural network 500 can continuously perform learning using algorithms such as backpropagation. Alternatively, the cloud on the Internet can update the gaze estimation neural network 500 in each television receiver 100 installed in the home with the learning results performed based on data collected from a large number of users, as will be described later.

[0111] Reference Figure 4 The description of the scene acquisition and scene information estimation system 400 continues.

[0112] The scene acquisition unit 407 acquires the video stream and audio stream of the part that the user has been looking at (or the part whose look-attenuation level exceeds a predetermined value) as determined by the look-attenuation estimation unit 406 from the buffer unit 403, as well as metadata about the entire content, and outputs the acquired video stream, audio stream, and metadata to the scene information estimation unit 408.

[0113] The video and audio streams acquired by the scene acquisition unit 407 from the buffer unit 403 based on the user's gaze can be considered scenes that the user has gazed at. On the other hand, metadata about the entire content does not necessarily correspond to a specific scene within the content. This is because users can have a strong interest in features that are not actually meaningful in the entire content but are unique to a specific scene (e.g., objects reflected only in a specific scene).

[0114] Therefore, in the scene acquisition and scene information estimation system 400 according to this embodiment, the scene information estimation unit 408 is configured to receive video and audio streams of the scene that the user has been watching from the scene acquisition unit 407, as well as metadata about the entire content, to estimate (seemingly possible) metadata (also called tags) characterizing the scene that the user has been watching based on artificial intelligence, and to output the metadata as information suitable for a specific scene rather than the entire content.

[0115] The trained neural network can be used in the scene information estimation unit 408 to provide reasoning capabilities for artificial intelligence. Figure 6 An example configuration of the scene information estimation neural network 600 used in the scene information estimation unit 408 is shown. The scene information estimation neural network 600 includes: an input layer 610 that receives video and audio streams of the scene the user has viewed, as well as metadata of the entire content; an intermediate layer 620; and an output layer 630 that outputs metadata, which is information characterizing the scene the user has viewed. The intermediate layer 620 includes multiple intermediate layers 621, 622, ... and the neural network 600 is expected to perform deep learning. Furthermore, an RNN structure incorporating recursive combinations of intermediate layers 620, considering time-series information such as video and audio streams, can be employed.

[0116] Input layer 610 includes the decoded video and audio streams stored in buffer unit 403 in the input vector elements. However, when buffer unit 403 stores the video and audio streams in RAW data state, the RAW data is input to input layer 610 as is. Regarding the audio stream, the input waveform signals of consecutive windows along the time axis are included in the input vector elements. There may be overlapping portions between consecutive windows. Additionally, the frequency signals obtained by performing a Fast Fourier Transform (FFT) on the waveform signals of each window can be used as input vector elements.

[0117] In addition, input layer 610 receives metadata, which is information characterizing the entire content. When the content is broadcast program content, the metadata includes, for example, text data such as program name, performer names, program summary, and keywords. Input nodes corresponding to each piece of text data are arranged in input layer 610. Furthermore, CNNs can be used to perform feature point agglomeration processing on inputs such as image signals and audio signals.

[0118] Output is generated from output layer 630, from which (potentially possible) metadata (also called tags) can be inferred. This metadata represents information about the scene the user has viewed. It is assumed that the metadata for each scene, in addition to the metadata of the entire original content such as program name, performer names, program content summary, and keywords, also includes information not included in the metadata of the original content, such as things reflected in the corresponding scene (e.g., the brand names of the clothes and accessories worn by the performers, the name and location of the shop containing the coffee cup held by the performers), the title of the background music, and the performers' lines). Output nodes corresponding to each piece of textual data of the metadata are arranged in output layer 630. Furthermore, output nodes corresponding to the seemingly possible metadata for the video and audio streams input to input layer 610 are triggered.

[0119] In the learning process of the scene information estimation neural network 600, the weighting coefficients (inference coefficients) of each node in the intermediate layer 620, which includes multiple layers, are updated to increase the binding strength of the output nodes with the metadata that appears to be possible for the video and audio streams of the scene that the user has already viewed, in order to learn the correlation between the scene and the metadata. Then, in the processing using the scene information estimation neural network 600, that is, in the scene information estimation process, the metadata that appears to be possible is output with high precision for the video and audio streams of the scene that the user has already viewed.

[0120] For example, implemented in the main control unit 201 Figure 6 The scene information estimation neural network 600 is shown. Therefore, the main control unit 201 may include a processor dedicated to the neural network. For example, a television receiver 100 containing the scene information estimation neural network 600 trained using an expert teaching database has been released. The scene information estimation neural network 600 can perform learning using algorithms such as backpropagation. Furthermore, the cloud on the Internet can update the scene information estimation neural network 600 in each television receiver 100 installed in every household with the learning results performed based on data collected from a large number of users, as will be described later.

[0121] Meanwhile, when real-time performance is not required, the cloud over the internet can provide scene information estimation neural networks 600.

[0122] The user's gaze scene acquired by the scene acquisition unit 407, and the metadata (hereinafter also referred to as "scene metadata") estimated for each scene from the scene information estimation unit 408 as information representing each scene, become the final output of the scene acquisition and scene information estimation system 400.

[0123] There are various output destinations for the user's gaze scenes and scene metadata based on the scene acquisition and scene information estimation system 400. For example, they can be stored in a television receiver 100 containing content already viewed by the user, or they can be uploaded to a server on the Internet. The server serving as the upload destination could be an artificial intelligence server in which a scene information estimation neural network 600 has been constructed, or another server that aggregates the metadata. Furthermore, the user's gaze scenes and scene metadata can be output to an information terminal carried by the user, such as a smartphone. For example, an application linked to the scene acquisition and scene information estimation system 400 (temporarily referred to as a "companion application") can be launched on the smartphone, allowing the user to view the scenes the user has gazed at and the scene metadata.

[0124] Furthermore, there are various methods for using metadata about the scenes a user has viewed. For example, metadata can be used to evaluate the video content a user watches and to recommend other video content. Additionally, metadata can be used for marketing, such as recommending scene-related products to users.

[0125] Figure 7 The flowchart illustrates the process performed in the scene acquisition and scene information estimation system 400 to acquire a scene from video content and output metadata as information characterizing the scene.

[0126] First, receiving unit 401 receives video content and metadata, which is information about the entire content (or attribute data included in the metadata) (step S701). The video content includes broadcast content transmitted from broadcasting stations (radio towers, broadcast satellites, etc.) and streaming content distributed from streaming distribution servers such as OTT services. Furthermore, the metadata received by receiving unit 401 is assumed to be metadata for the entire content assigned to the content creator or distributor.

[0127] Subsequently, the signal processing unit 402 processes the video stream and audio stream of the content received by the receiving unit 401, the output unit 404 outputs the video stream and audio stream, and performs buffering of content and metadata (step S702).

[0128] Then, the gaze estimation unit 406 estimates the gaze of the user who is watching the content presented by the output unit 404 based on artificial intelligence (step S703).

[0129] The gaze estimation unit 406 estimates the user's gaze based primarily on the recognition results of the user's facial image captured by the camera 311, using measurement or artificial intelligence. The gaze estimation unit 406 can also use sensor signals from other sensors besides the image captured by the camera 311 to measure the gaze. Furthermore, for the gaze estimation unit 406, a gaze estimation neural network 500 (see reference) that has learned the correlation between facial images or other types of sensor information and the user's gaze is used. Figure 5 In order to perform estimations based on artificial intelligence.

[0130] Subsequently, when the gaze estimation unit 406 determines that the user has gazed (step S70 is), the scene acquisition unit 407 acquires the video stream and audio stream of the part that the user has gazed from the buffer unit 403 as the scene that the user has gazed, acquires the metadata of the entire content from the buffer unit 403, and outputs the video stream, audio stream, and metadata to the scene information estimation unit 408 (step S705).

[0131] Subsequently, the scene information estimation unit 408 receives the video and audio streams of the scene that the user has been viewing, as well as metadata about the entire content, from the scene acquisition unit 407, and estimates (seemingly possible) metadata representing the scene that the user has been viewing based on artificial intelligence (step S706). The scene information estimation neural network 600, which has learned the correlation between the scene that the user has been viewing and the metadata that seems likely for that scene (see...),... Figure 6 The scene information estimation unit 408 uses this information to perform estimations based on artificial intelligence.

[0132] Then, the scene that the user has viewed and the (possibly possible) metadata representing that scene are output to the predetermined output destination (step S707), and the process ends.

[0133] Figure 8 A variant of the scene acquisition and scene information estimation system 800 is shown. Use as needed. Figure 2 The system 800 shown is configured using components in the television receiver 100 or external devices of the television receiver 100 (such as cloud server devices).

[0134] The configuration and operation of the receiving unit 801, signal processing unit 802, buffer unit 803, output unit 804, and sensor unit 805 are as follows: Figure 4The scene information estimation system 400 shown in the figure has the same configuration and operation, and therefore its detailed description is omitted here.

[0135] The scene information estimation unit 806 uses a single neural network to jointly achieve the following: Figure 4 The processing performed by the gaze estimation unit 406, scene acquisition unit 407, and scene information estimation unit 408 in the scene information estimation system 400 shown.

[0136] Figure 9 An example configuration of a scene information estimation neural network 900 used in a scene information estimation unit 806 is shown. The scene information estimation neural network 900 includes an input layer 910, an intermediate layer 920, and an output layer 930. In the example shown, the intermediate layer 920 includes multiple intermediate layers 921, 922, ..., and the scene information estimation neural network 900 can perform deep learning (DL). Furthermore, a recurrent neural network (RNN) structure can be employed, which includes recursive combinations of time-series information such as moving images and sounds considered in the intermediate layer 920.

[0137] The input layer 910 includes nodes that receive image signals captured by the camera 311 and other types of sensor information, video and audio streams temporarily stored in the buffer unit 803, and metadata of the content.

[0138] Input layer 910 includes the moving image stream (or still image), video stream, and audio stream captured by camera 311, etc., as input vector elements. Essentially, it is assumed that the image and audio signals are input to input layer 910 in a RAW data state. Regarding the audio stream, the input waveform signal of each consecutive window in the time axis direction is used as the input vector to each node of input layer 910. There may be overlapping portions between consecutive windows. Furthermore, the frequency signal obtained by performing FFT processing on the waveform signal of each window can be used as the input vector (as described above). Additionally, a convolutional neural network (CNN) can be used for the input of the image and audio signals, etc., to perform feature point aggregation processing.

[0139] In addition, when metadata consisting of text data such as program name, performer name, program content summary and broadcast content keywords is input into input layer 910, input nodes corresponding to each piece of text data are arranged in input layer 910.

[0140] Simultaneously, output is generated from output layer 930, through which the scenes the user has gazed at and the (potentially possible) metadata (also called tags) representing those scenes can be estimated. It is assumed that the metadata for each scene includes information not included in the metadata of the entire original content, such as things reflected in the corresponding scene (e.g., brand names of clothing and accessories worn by performers, the name and location of the shop containing the coffee cup held by the performer), the title of the background music, and the performer's route, such as program name, performer name, program summary, and keywords. Therefore, output nodes corresponding to each segment of textual data of the metadata are arranged in output layer 930. Then, for the video and audio streams input to input layer 910, output nodes corresponding to the video and audio streams of scenes with high user gaze and the metadata that appears to be possible for those scenes are triggered.

[0141] In the learning process of the scene information estimation neural network 900, the weighting coefficients (inference coefficients) of each node in the intermediate layer 920, which comprises multiple layers, are updated to increase the strength of the combination of the seemingly probable gaze scene and the metadata that the scene seems to be probable with the output node, in order to learn the correlation between the scene and the metadata used for facial images or other sensor signals, as well as the video and audio streams input at each moment. Then, in the processing using the scene information estimation neural network 900, that is, in the scene information estimation processing, for the facial images or other sensor signals and video and audio streams input at each moment, the seemingly probable gaze scene and the metadata that the scene seems to be probable are output with high accuracy.

[0142] For example, implemented in the main control unit 201 Figure 9 The scene information estimation neural network 900 is shown. Therefore, a dedicated processor for the neural network can be included in the main control unit 201 or in a processing circuit separate from the main control unit 201. For example, a television receiver 100 containing a scene information estimation neural network 900 trained using an expert teaching database has been released. The scene information estimation neural network 900 can perform learning using algorithms such as backpropagation. Furthermore, the cloud on the Internet can update the scene information estimation neural network 900 in each television receiver 100 installed in every household with the learning results performed based on data collected from a large number of users, as will be described later.

[0143] Meanwhile, when real-time performance is not required, the cloud over the internet can provide scene information estimation neural networks 900.

[0144] The scene that the user has been looking at and the metadata that seem to be possible for that scene, identified by the scene information estimation unit 806, become the final output of the scene acquisition and scene information estimation system 800.

[0145] D. Feedback of scene information estimation results based on gaze level

[0146] Here, together with the operation of the scene information estimation system 400, a method for feeding back the scene information estimation results from the scene acquisition and scene information estimation system 400 to the user will be described.

[0147] like Figure 10 As shown, when video (broadcast content, streaming mobile images from OTT services, etc.) is displayed on the screen, for example, a camera 311 located near the center of the top edge of the screen of the television receiver 100 continuously captures images of the user.

[0148] The gaze estimation unit 406 identifies faces from the captured images of the camera 311 and measures the user's gaze. Figure 11 The diagram illustrates the state of recognizing two users' facial images 1101 and 1102 from images captured by camera 311. A gaze estimation unit 406 performs gaze estimation on at least one of the facial images 1101 and 1102. To estimate user gaze at content based on artificial intelligence, [the following can be used]. Figure 5 The gaze estimation neural network 500 shown can estimate the user's gaze based on sensor signals from other sensors rather than the captured images from camera 311.

[0149] When the gaze estimation unit 406 estimates that the facial image 1101 or 1102 identified from the captured image of the camera 311 is being gazed, the video stream and audio stream corresponding to the estimated gaze portion and stored in the buffer unit 403 are captured as the scene that the user is gazing at. Figure 12 This demonstrates how the state of a gaze scene is captured based on measurements of the user's gaze.

[0150] The user's gaze scene and metadata representing the scene, output from the scene acquisition and scene information estimation system 400, are uploaded to a server on the Internet. For example, a content recommendation server searches for similar or related content to recommend to the user based on the user's gaze scene and the scene's metadata, and provides the user with information about the recommended content. This type of content recommendation server can use algorithms such as collaborative filtering (CF) or content-based filtering (CBF) to search for similar or related content based on the user's gaze scene. Alternatively, like the AI ​​server mentioned earlier, the content recommendation server can use a neural network that has learned the correlation between metadata and content to extract similar or related content.

[0151] Even if a content recommendation server searches for similar or related content using any algorithm, it does so based on metadata about the specific scenes the user has viewed within the content, rather than the entire content. Therefore, it can recommend content to the user that is less related to metadata than the overall content, such as show titles and performers, but more relevant to details reflected only in the scenes the user has viewed (e.g., the brands of clothing and accessories worn by the performers, such as sunglasses and watches).

[0152] Figure 13 An example configuration of a screen for providing feedback on scenes previously viewed by a user is shown. In the illustrated screen, a list of representative images of scenes previously viewed by a user watching video content via a television receiver 100 is displayed in matrix form. The term "representative image" here can be, for example, a leading image of the scene obtained according to a predetermined algorithm or randomly from the content, or an image captured within the scene. The display format on the screen is not limited to a matrix form and can be loops of virtually infinite length in horizontal or vertical lines. Furthermore, the display format can be configured such that a large number of scenes arranged at arbitrary locations in three-dimensional space can be displayed on a two-dimensional plane of visual display when viewing three-dimensional space from a predetermined viewpoint and the viewpoint can be changed.

[0153] Users can, for example, use the D-pad on a remote control to select representative images of scenes of interest from a list. Then, when the user presses the "OK" button on the remote control while selecting any representative image, the selection of the corresponding scene is confirmed. Simultaneously, the aforementioned remote control operation for selecting a specific scene from the list of representative images can also be indicated via audio, for example, using a microphone through the audio proxy function of an AI speaker with artificial intelligence capabilities. Alternatively, remote control operations can be performed based on gestures. These gestures are not limited to hands or similar objects; all body parts, including the head, can be used. Furthermore, when a user wears smart glasses, gaze can be recognized, and selection and confirmation commands can be executed based on gestures using the eyes (e.g., blinking). Additionally, wireless headphones equipped with motion sensors and capable of sensing head movements based on AI functions can be used. For wireless devices such as wireless headphones, touch sensors can be used to execute selection and confirmation commands.

[0154] Figure 14 This example shows a screen configuration that changes after a specific representative image is selected. In the example shown, the screen zooms in and displays the image selected by the user from the list of representative images (see [link to example]). Figure 13The system selects a representative image and displays partial or complete metadata (MD#1, MD#2, ...) of the scene as information estimated by the scene information estimation system 400. When displaying the representative image, the system can reproduce and output the video and audio streams of the portion that the user has viewed, instead of still images.

[0155] Users can recall memories of scenes they have previously viewed by viewing representative or moving images of the scene they have viewed. Furthermore, users can view metadata and representative images of scenes they have viewed to confirm the reasons or basis for viewing those scenes. Additionally, if a user wishes to change the metadata representing the scene, which is automatically assigned to the scene by the scene information estimation unit 408 based on artificial intelligence or the like, the metadata can be edited (changed, added, deleted, etc.) via remote control operation, audio proxy function, etc., to customize the scene-related metadata for personal use.

[0156] In addition, users can use a remote control or audio agent to indicate the presentation of recommended similar or related content based on the selected gaze scene. Figure 15 An example configuration of a screen for presenting similar or related content recommended based on metadata of a selected scene is shown. The user's gaze scene and its metadata are displayed in the left half of the screen. Additionally, a list of similar or related content is displayed in the right half of the screen. Representative images and metadata of the similar or related content are displayed in the list. The user can determine whether he / she expects to view each segment of similar or related content based on the representative images and metadata. Then, when the user finds similar or related content he / she expects to watch, the user can instruct the content to be reproduced to begin watching via remote control operation, audio proxy function, etc.

[0157] Furthermore, users can view feedback on the scene they are watching on a small screen, such as a smartphone, or on a large screen, such as a television receiver 100. For example, a user can launch an accompanying application on their smartphone that links to the scene acquisition and scene information estimation system 400 to view the scene they have personally watched and its metadata.

[0158] Figure 16 This illustrates an example of how the scene a user is looking at is reflected on the smartphone screen. Smartphone screens are small and therefore cannot be displayed like on a traditional screen. Figure 13 As shown on the television screen, representative images of multiple scenes are displayed in a list. Therefore, only one representative image of a scene is displayed on a single screen. However, for example, representative images of multiple scenes can be virtually arranged in a rotated manner (e.g., Figure 17 (as shown) or virtually arranged in a matrix (e.g.) Figure 18(as shown), and in response to a user tapping a representative image displayed on the smartphone screen in the horizontal or vertical direction, switching to a representative image adjacent in the direction in which the user tapped the image in rotation or matrix.

[0159] Return to reference Figure 16 This describes the configuration of the feedback screen on a smartphone. When a representative image is displayed on the screen, a video and audio stream of the portion the user has been viewing can be reproduced and output, instead of a still image. Furthermore, some or all of the scene's metadata, estimated by the scene information estimation system 400, is displayed below the representative image.

[0160] According to the scene information estimation system 400 disclosed in this specification, a user can recall memories of previously viewed scenes by viewing representative or moving images of the scene being viewed. Furthermore, according to the scene information estimation system 400 disclosed in this specification, a user can view metadata and representative images of scenes he / she has viewed to confirm the reason or basis for viewing those scenes. Moreover, according to the scene information estimation system 400 disclosed in this specification, if a user wishes to change the metadata assigned to scenes, etc., a user can use a touch panel to edit (change, add, delete, etc.) metadata using the editing functions of a smartphone. Furthermore, according to the scene information estimation system 400 disclosed in this specification, by integrating and learning scenes with high user attention, trends in video content with high attention, i.e., user "satisfaction," can be learned. Therefore, the possibility of providing users with video content with high "satisfaction" through artificial intelligence learning capabilities can be improved.

[0161] E. Updates of neural networks

[0162] The document describes the gaze estimation neural network 500, scene information estimation neural network 600, and scene information estimation neural network 900 used in the processing of metadata about the scene that a user has gazed at in video content, based on artificial intelligence.

[0163] These neural networks are used as a function of artificial intelligence and operate in devices that can be directly manipulated by the user, such as a television receiver 100 installed in each home, or in an operating environment (hereinafter also referred to as the "local environment") such as a home in which the device is installed. One effect of operating the neural network as a function of artificial intelligence in the local environment is that, for example, learning can be easily achieved in real time using algorithms such as backpropagation for these neural networks, using feedback from the user as teacher data. Feedback from the user includes, for example, the user's gaze level estimated by the gaze estimation neural network 500, and the user's evaluation of metadata, which is information characterizing the scene estimated by the scene information estimation neural networks 600 or 900. For example, user feedback might be simple, such as OK and NG. User feedback is input to the television receiver 100, for example, via an operation input unit 222 or a remote control, an audio agent as a form of artificial intelligence, a smartphone, etc. Therefore, another aspect of the effect of operating such a neural network as a function of artificial intelligence in the local environment is that the neural network can be customized or personalized for a specific user based on learning using user feedback.

[0164] On the other hand, a method can be envisioned for collecting data from a large number of users to accumulate the learning of neural networks as a function of artificial intelligence, and using the learning results from one or more server devices operating on a cloud (hereinafter also referred to as the "cloud"), which is a collection of server devices on the Internet, to update the neural network in the television receiver 100 in each household. One of the effects of updating the neural network, which acts as artificial intelligence, through the cloud is that a more accurate neural network can be constructed by performing learning using a large amount of data.

[0165] Figure 20 A schematic example of a cloud-based artificial intelligence system 2000 configuration is shown. The cloud-based artificial intelligence system 2000 shown consists of a local environment 2010 and a cloud environment 2020.

[0166] The local environment 2010 corresponds to the operating environment (home) in which the television receiver 100 is installed or installed in a home. Although for simplicity... Figure 20Only a single local environment 2010 is shown, but in practice, it is conceivable to connect a large number of local environments to a single cloud 2020. Furthermore, while in this embodiment, the television receiving device 100 or an operating environment such as the home in which the television receiving device 100 operates is primarily exemplified as local environment 2010, local environment 2010 can be any device that can be operated by the user, such as a smartphone or wearable device, or an environment in which the device operates (including public facilities such as train stations, bus stations, airports, and shopping malls, and workplace facilities such as factories and offices).

[0167] As described above, the television receiver 100 is equipped with a gaze estimation neural network 500 and a scene information estimation neural network 600 or 900 as artificial intelligence. Here, it is assumed that these neural networks installed in and actually used in the television receiver 100 are collectively referred to as computational neural networks 2011. It is assumed that the computational neural network 2011 performs pre-learning using an expert teaching database that includes a large amount of sample data.

[0168] Meanwhile, Cloud 2020 is equipped with an AI server (as described above) that provides AI functionality (including one or more server devices). The AI ​​server has a computational neural network 2021 and an evaluation neural network 2022 that evaluates the computational neural network 2022. The computational neural network 2021 has the same configuration as the computational neural network 2011 provided in the local environment 2010 and is assumed to have been pre-learned using an expert teaching database containing a large amount of sample data. Furthermore, the evaluation neural network 2022 is a neural network used to evaluate the learning progress of the computational neural network 2021.

[0169] On the local environment 2010 side, the computational neural network 2011 receives sensor information such as captured images from camera 311 and a user profile, and outputs a gaze level suitable for the user profile and metadata for each scene. However, when the computational neural network 2011 is a scene information estimation neural network 600, the input is a video stream of the scene the user has gazed at and metadata of the original content. Here, for simplicity, the input of the computational neural network 2011 is simply referred to as the "input value," and the output from the computational neural network 2012 is simply referred to as the "output value."

[0170] A user in the local environment 2010 (e.g., a viewer of the television receiver 100) evaluates the output value of the computational neural network 2011 and feeds back the evaluation result to the television receiver 100, for example, via the operation input unit 222, a remote control, an audio agent, a linked smartphone, etc. Here, for the sake of simplicity, it is assumed that the user feedback is either OK (0) or NG (1).

[0171] Feedback data includes a combination of the input and output values ​​of the computational neural network 2011 and user feedback, which is transmitted from the local environment 2010 to the cloud 2020. In the cloud 2020, the feedback data transmitted from the local environment is accumulated in a feedback database 2023. The feedback database 2023 accumulates a large amount of feedback data describing the correspondence between the input and output values ​​of the computational neural network 2011 and the user.

[0172] Furthermore, Cloud2020 can possess or utilize the expert teaching database 2024, which includes a large amount of sample data used for prior learning of the computational neural network 2011. Individual sample data consists of teacher data describing the correspondence between sensor information and user profiles and the output values ​​of the computational neural network 2011 (or 2021).

[0173] When feedback data is extracted from the feedback database 2023, the input values ​​included in the feedback data (e.g., a combination of sensor information and user profile) are input into the computational neural network 2021. Furthermore, the output value of the computational neural network 2021 and the input values ​​included in the corresponding feedback data (e.g., a combination of sensor information and user profile) are input into the evaluation neural network 2022, and the evaluation neural network 2022 outputs user feedback.

[0174] In Cloud 2020, the learning of the evaluation neural network 2022, which is the first step, and the learning of the computation neural network 2021, which is the second step, are performed alternately.

[0175] The evaluation neural network 2022 is a network that learns the correspondence between the input values ​​of the computational neural network 2021 and the user feedback based on the output of the computational neural network 2021 and the user feedback. Therefore, in the first step, the evaluation neural network 2022 receives the output value of the computational neural network 2021 and the user feedback contained in the corresponding feedback data, and performs learning such that the user feedback output by the evaluation neural network 2022 for the output value of the computational neural network 2021 is consistent with the real user feedback for the output value of the computational neural network 2021. As a result, the evaluation neural network 2022 is trained to output the same user feedback (OK or NG) as the real user feedback for the output of the computational neural network 2021.

[0176] Subsequently, in the second step, the evaluation neural network 2022 is fixed, and the computational neural network 2021 is trained. As described above, when feedback data is extracted from the feedback database 2023, the input values ​​contained in the feedback data are input into the computational neural network 201, and the output values ​​of the computational neural network 2021 and the user feedback data contained in the corresponding feedback data are input into the evaluation neural network 2022. The evaluation neural network 2022 outputs user feedback that is identical to the user feedback from the real user.

[0177] At this point, the computational neural network 2021 applies an evaluation function (e.g., a loss function) to the output from the output layer of the neural network and uses backpropagation to perform learning that minimizes this value. For example, when user feedback is used as teacher data, the computational neural network 2021 performs learning such that the output of the evaluation neural network 2022 becomes OK (0) for all input values. By performing learning in this way, the computational neural network 2021 can output the user feedback OK value (gaze level, metadata about the scene, etc.) relative to any input value (sensor information, user profile, etc.).

[0178] Furthermore, the expert teaching database 2024 can be used as teacher data during the learning of the computational neural network 2021. Additionally, learning can be performed using two or more teacher data sets, such as user feedback and the expert teaching database 2024. In this case, a weighted addition can be performed on the loss function calculated for each teacher data set, and the computational neural network 2021 can be trained to minimize the weighted addition result.

[0179] As described above, the learning of the evaluation neural network 2022 as the first step and the learning of the operational neural network 2021 as the second step are performed alternately to improve the accuracy of the operational neural network 2021. Then, the inference coefficients in the operational neural network 2021, whose accuracy has been improved according to the learning, are provided to the operational neural network 2011 in the local environment 2010, and thus the user can obtain a further trained operational neural network 2011.

[0180] For example, the bitstream of inference coefficients from an operational neural network 2011 can be compressed and downloaded from the cloud 2020 to a local environment. When the size of the compressed bitstream is still large, the inference coefficients can be divided into layers or regions, and the compressed bitstream can be downloaded multiple times.

[0181] [Industrial Applicability]

[0182] The techniques disclosed in this specification have been described in detail with reference to specific embodiments. However, it will be apparent to those skilled in the art that modifications and substitutions can be made to the embodiments without departing from the key technical points disclosed in this specification.

[0183] Although this specification focuses on embodiments in which the techniques disclosed herein are applied to television receivers, the key points of the techniques disclosed herein are not limited thereto. The techniques disclosed herein can also be applied equally to various types of content playback devices that present various types of reproduced content, such as video and audio content, to users.

[0184] In summary, the technology disclosed in this specification has been described in an illustrative form, but the content of this specification should not be interpreted restrictively. The essential elements of the technology disclosed in this specification should be determined by considering the claims.

[0185] Meanwhile, the technology disclosed in this specification can also be configured as follows.

[0186] (1) An artificial intelligence information processing device, comprising: a gaze estimation unit configured to estimate the gaze of a user viewing content based on sensor information and artificial intelligence;

[0187] The acquisition unit is configured to acquire, based on the estimation results of the gaze estimation unit, video of the scene in the content where the user is gazing and information about that content; and

[0188] The scene information estimation unit is configured to estimate scene information about the user's gaze based on the video of the scene being viewed by the user and information about that content, using artificial intelligence.

[0189] (2) According to the artificial intelligence information processing device of (1) above, the scene information estimation unit uses a neural network that has learned the correlation between scene video and information about content and information about scene to estimate information that is relevant to the scene being viewed by the user as an estimate implemented based on artificial intelligence.

[0190] (3) According to the artificial intelligence information processing device of (1) or (2) above, the gaze estimation unit uses a neural network that has learned the correlation between sensor information and the user's gaze to estimate the gaze that is correlated with the sensor information about the user who is watching the content, as an estimate implemented according to artificial intelligence.

[0191] (4) According to the artificial intelligence information processing device of (3) above, wherein the sensor information includes at least a captured image of a user viewing content captured by a camera, and

[0192] The gaze estimation unit uses a neural network that has learned the correlation between facial recognition results and user gaze to estimate the gaze that is correlated with the facial recognition results in the captured image, as an estimate implemented based on artificial intelligence.

[0193] (5) An artificial intelligence information processing device according to any one of (1) to (4) above, wherein the content includes at least one of broadcast content and streaming content.

[0194] (6) An artificial intelligence information processing device according to any one of (1) to (5) above, wherein information about the scene being gazed at by the user is output to the outside.

[0195] (7) An artificial intelligence information processing device according to any one of (1) to (6) above, wherein the user is presented with the scene being looked at and information about the scene.

[0196] (8) An artificial intelligence information processing device, comprising: an input unit configured to receive sensor information about a user viewing content; and

[0197] The scene information estimation unit is configured to use a neural network that has learned the correlation between sensor information, content and information about that content, and information about the scene being viewed by the user to estimate information relevant to the scene being viewed by the user.

[0198] (9) The artificial intelligence information processing device according to (8) above, wherein the sensor information includes at least a captured image of a user viewing content captured by a camera, and

[0199] The scene information estimation unit uses a neural network that has learned the correlation between facial images, content, and information about that content and information about the scene the user is looking at to estimate information that is relevant to the scene the user is looking at.

[0200] (10) An artificial intelligence information processing method, comprising: a gaze estimation step, used to estimate the gaze of a user viewing content based on sensor information and artificial intelligence;

[0201] The acquisition step is used to acquire video footage of the scene in the content where the user is gazing, and information about the content, based on the estimation results in the gaze estimation step; and

[0202] The scene information estimation step is used to estimate scene information about the user's gaze based on the video of the scene and information about the content, using artificial intelligence.

[0203] List of reference numerals

[0204] 100 TV receiver

[0205] 201 Control Unit

[0206] 202 bus

[0207] 203 memory cells

[0208] 204 Communication Interface (IF) Unit

[0209] 205 Expansion Interface (IF) Unit

[0210] 206 tuner / demodulation unit

[0211] 207 Demultiplexer

[0212] 208 video decoder

[0213] 209 audio decoder

[0214] 210 Subtitle Overlay Decoder

[0215] 211 Subtitle Decoder

[0216] 212 Subtitle Synthesis Unit

[0217] 213 Data Decoder

[0218] 214 cache units

[0219] 215 Application Programming (AP) Control Unit

[0220] 216 browser units

[0221] 217 sound source units

[0222] 218 video synthesis units

[0223] 219 display units

[0224] 220 audio synthesis unit

[0225] 221 audio output unit

[0226] 222 Operation Input Unit

[0227] 300 sensor group

[0228] 310 camera unit

[0229] Cameras 311-313

[0230] 320 Status Sensor Unit

[0231] 330 Environmental Sensor Unit

[0232] 340 Equipment Status Sensor Unit

[0233] 350 User Profile Sensor Unit

[0234] 400 Scene Acquisition and Scene Information Estimation System

[0235] 401 Receiving Unit

[0236] 402 Signal Processing Unit

[0237] 403 Buffer Unit

[0238] 404 Output Unit

[0239] 405 sensor unit

[0240] 406 gaze estimation units

[0241] 407 Scene Acquisition Unit

[0242] 408 Scene Information Estimation Unit

[0243] 500-Gaze ​​Estimation Neural Network

[0244] 510 Input Layer

[0245] 520 intermediate layer

[0246] 530 Output Layer

[0247] 600 Scene Information Estimation Neural Network

[0248] 610 Input Layer

[0249] 620 intermediate layer

[0250] 630 output layer

[0251] 800 Scene Acquisition and Scene Information Estimation System

[0252] 801 Receiver Unit

[0253] 802 Signal Processing Unit

[0254] 803 Buffer Unit

[0255] 804 Output Unit

[0256] 805 sensor unit

[0257] 806 Scene Information Estimation Unit

[0258] 900 Scene Information Estimation Neural Network

[0259] 910 Input Layer

[0260] 920 intermediate layer

[0261] 930 output layer

[0262] 1900 monitor

[0263] 1901 speaker unit

[0264] 1902 stent

[0265] 1901-1, 1902-2 exciters

[0266] 2000 cloud-based artificial intelligence systems

[0267] 2010 local environment

[0268] 2011 Computational Neural Networks

[0269] 2020 Cloud

[0270] 2021 Computational Neural Networks

[0271] 2022 Evaluation Neural Networks

[0272] 2023 Feedback Database

[0273] 2024 Expert Teaching Database.

Claims

1. An artificial intelligence information processing apparatus comprising: a gaze degree estimation unit configured to estimate, based on received sensor information about a user who is watching a first content, a gaze degree of the user according to artificial intelligence; an acquisition unit configured to acquire, based on an estimation result of the gaze degree estimation unit, a video of a scene at which the user gazes in the first content and first information about the first content; and a scene information estimation unit configured to estimate, based on the video of the scene at which the user gazes and the first information about the first content, second information about the scene at which the user gazes according to artificial intelligence, wherein the second information about the scene includes first metadata that characterizes a scene, wherein the artificial intelligence information processing apparatus is further configured to: customize the first metadata based on an input of the user, acquire a plurality of second contents and second metadata corresponding to each of the plurality of second contents based on the scene and the customized first metadata; and control a display device to display the second metadata and representative images of the plurality of second contents. The scene information estimation unit estimates information having a correlation with the scene at which the user gazes, using a neural network that has learned a correlation between a video of a scene and information about a content and information about the scene, as an estimation according to artificial intelligence. 2.The artificial intelligence information processing apparatus according to claim 1, wherein The gaze degree estimation unit estimates a gaze degree having a correlation with the sensor information about the user who is watching the content, using a neural network that has learned a correlation between sensor information and a gaze degree of the user, as an estimation according to artificial intelligence. 3.The artificial intelligence information processing apparatus according to claim 1, wherein 4.The artificial intelligence information processing apparatus according to claim 3, wherein the sensor information includes at least a captured image of the user who is watching the first content captured through a camera, and the gaze degree estimation unit estimates a gaze degree having a correlation with a face recognition result in the captured image, using a neural network that has learned a correlation between a face recognition result and a gaze degree of the user, as an estimation according to artificial intelligence. The first content includes at least one of broadcast content and streaming distribution content. 5.The artificial intelligence information processing apparatus according to claim 1, wherein The second information about the scene at which the user gazes is output to the outside. 6.The artificial intelligence information processing apparatus according to claim 1, wherein The scene at which the user gazes and the second information about the scene are presented to the user. 7.The artificial intelligence information processing apparatus according to claim 1, wherein 8.An artificial intelligence information processing apparatus comprising: an input unit configured to receive sensor information about a user who is watching a first content; a gaze degree estimation unit configured to estimate, based on the received sensor information about a user who is watching a first content, a gaze degree of the user according to artificial intelligence; an acquisition unit configured to acquire, based on an estimation result of the gaze degree estimation unit, a video of a scene at which the user gazes in the first content and first information about the first content; and a scene information estimation unit configured to estimate, based on the video of the scene at which the user gazes and the first information about the first content, second information about the scene at which the user gazes according to artificial intelligence, wherein the second information about the scene includes first metadata that characterizes a scene. ​ a scene information estimating unit configured to estimate, based on a video of a scene at which a user gazes and first information about a first content, second information about the scene at which the user gazes using a neural network that has learned a correlation between sensor information, a content, and information about the content and information about a scene at which a user gazes, wherein the second information about the scene includes first metadata that characterizes a scene, wherein the artificial intelligence information processing apparatus is further configured to: customize the first metadata based on an input of the user, acquire a plurality of second contents and second metadata corresponding to each of the plurality of second contents based on the scene and the customized first metadata; and control a display device to display the second metadata and representative images of the plurality of second contents.

9. The artificial intelligence information processing apparatus according to claim 8, wherein the sensor information includes at least a captured image of the user who is watching the content captured by a camera, and the scene information estimating unit estimates the second information about the scene at which the user gazes using a neural network that has learned a correlation between a facial image, a content, and information about the content and information about a scene at which the user gazes.

10. An artificial intelligence information processing method comprising: a gaze degree estimating step of estimating, based on received sensor information about a user who is watching a first content, a gaze degree of the user according to artificial intelligence; an acquiring step of acquiring, based on a result of the estimating in the gaze degree estimating step, a video of a scene at which the user gazes in the first content and first information about the first content; and a scene information estimating step of estimating, based on the video of the scene at which the user gazes and the first information about the first content, second information about the scene at which the user gazes according to artificial intelligence, wherein the second information about the scene includes first metadata that characterizes a scene, wherein the method further comprises: customizing the first metadata based on an input of the user, acquiring a plurality of second contents and second metadata corresponding to each of the plurality of second contents based on the scene and the customized first metadata; and controlling a display device to display the second metadata and representative images of the plurality of second contents. ​

Citation Information

Patent Citations

  • Information providing apparatus and information providing method, and computer program

    JP2008236779A

  • Information processing apparatus, information processing method, and program

    CN102655576A

  • Image signal processing method, device and equipment

    CN109688351A

  • Apparatus and method for identifying gazing direction of human eyes and its use

    CN1423228A