Construction and realization of interaction between digital media and observers

A system that dynamically adjusts digital images based on viewer interactions addresses under/overexposure and focus issues, enhancing the viewing experience by replicating natural perception and guiding viewers through content.

JP7734088B2Active Publication Date: 2025-09-04RICHMOND ROBERT ELLE
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022015381
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2016-03-14
Filing Date
2022-02-03
Publication Date
2025-09-04
Estimated Expiration
2037-03-14

AI Technical Summary

Technical Problem

Existing digital imaging technologies struggle to dynamically adjust image parameters based on viewer interaction, resulting in unsatisfactory viewing experiences due to issues like under/overexposure and focus discrepancies.

Method used

Implement a system that tracks viewer interactions, such as gaze, touch, and ambient light, to dynamically modify image parameters like brightness, focus, and content in real-time, using algorithms to swap or adjust image pixels based on viewer actions.

Benefits of technology

Enhances the viewing experience by replicating natural human perception, providing detailed and focused views of images, and allowing artists to guide viewers through the content, thereby improving engagement and fidelity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007734088000001
    Figure 0007734088000001
  • Figure 0007734088000002
    Figure 0007734088000002
  • Figure 0007734088000003
    Figure 0007734088000003
Patent Text Reader

Abstract

To provide a method for improving the image viewing experience by tracking the location of an image that a viewer is looking at. The system includes a display screen (14) for displaying an image, an image capture device (gaze detection hardware device 16) for capturing an image of the face of an observer (13) viewing the image, and a computer (17). The computer determines, in response to the image of the observer's face captured by the image capture device, a location on the image at which the observer is gazing, searches for an image portion having a different clarity from the clarity of the portion of the image at the determined location, and corrects the clarity by replacing the portion of the image at the location with the searched image portion. Furthermore, in response to another image of the observer's face captured by the image capture device, the computer determines whether the observer is still gazing at the same location on the image, and if it is determined that the observer is no longer gazing, again corrects the clarity by replacing the portion of the image at the location with the searched image portion.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Today's digital cameras and smartphones use computing power to Enhances the image on your screen for immediate or later viewing. One example is HDR or High Dynamic Range. The camera quickly takes multiple photos at different exposures, ensuring that every part is captured in the most accurate way. It creates an image where even the brightest and darkest areas are exposed to bring out all the details.

[0002] Also, to make it more appealing to the eye, or for how long (e.g., average time To measure the value of an advertisement in terms of how many viewers spend their time on it (total time), A viewing test was conducted to determine how long a viewer looked at a given advertisement on a web page. There are also existing systems that use line detection, and stop video playback to conserve battery power. It is also possible to save money or to make the viewer smarter so that they do not miss any part of the video. Eye gaze detection to ensure you are not looking at a display such as a phone display There are also systems that use Summary of the Invention

[0003] Embodiments enhance the image viewing experience by tracking where the viewer is looking The result is a visual experience that feels like watching the original scene. Identification of one or more images that provide an experience or that were captured automatically or by a photographer The photographer's pre-specified intention about what should happen when an observer looks at the part It goes beyond the original experience and elevates it in new ways.

[0004] An embodiment may be used to allow an artist to create an original image or multiple images for a display screen. It allows you to create multiple original images for display screens, which The image changes when the viewer looks at it depending on who the viewer is or what the viewer does. The observer's actions that change the image depend on where the observer looks or what screen they use. touch the screen, or what the viewer says, or how far the viewer moves from the display. It may be the distance, or how many observers there are, or the identity of a specific expected observer, etc. stomach.

[0005] An embodiment may include one or more cameras positioned in front of one or more display screens. One or more devices are used in response to data collected through capturing an image of the observer's face. Method for modifying one or more original images displayed on a display screen - Patents.com The data collected is the location of the gaze of one or more observers in the original image. or changes in the facial features of one or more observers, or the amount of ambient light, or one or more of the distance of the observer from the display screen, the number of observers, or the identity of the observers may be.

[0006] [brightness] The modification may be to change the luminance of part or all of the original image. The change can be effected by replacing the original image with an alternative image. , the content of the alternative image is the same as the original image but at least some of the content of the alternative image is Some or all of the images have a different luminance than the luminance of at least the corresponding part of the original image. For example, a pair of photographic images may be taken with a camera, and the pixels of the images are essentially the same except for different exposures. Part or all of one image may be Alternatively, the algorithm may use one or more In response to actions by the observer, some but not all image pixels The numeric brightness can be adjusted.

[0007] [focus] The modification may be to change the focus of part or all of the original image. can be obtained by replacing the original image with an alternative image, The content is the same as the original image, but at least some of the focus of the alternative image is different from the original. Some or all of the pixels are out of focus in at least the corresponding portion of the null image. Corrected. Correction is performed for closely-viewed objects in some or all of the image. The modification may be to change the focus of the image. The modification may be to change the focus on a part of the image or on an object. The correction may be to change the apparent depth of field of the original image. The correction may be enlargement or reduction by zooming around a part. , may include replacing the original image with an alternative image. The zooming may be synchronized with the change in the measured distance from the object to the observer.

[0008] [color] Correction changes the color balance, saturation, color intensity, or contrast of part or all of an image. This may be achieved by replacing the original image with an alternative image. The content of the alternate image can be the same as the original image but with the At least some of the colors (e.g., color balance or saturation) match at least some of the colors in the original image. Some or all of the pixels are modified to be a different color than the corresponding portion.

[0009] [Clarity] The modification may be to change the sharpness of part or all of the image. The content of the alternate image can be achieved by replacing the null image with an alternate image. is the same as the original image, but at least some of the clarity of the substitute image is different from the original image. Some or all of the pixels have been modified to differ in sharpness from the corresponding portion of the image.

[0010] [Audio Output] The modification involves causing or changing the playback of the sound that accompanies the original image. That's fine.

[0011] [Sprite] The modification was to cause the sprites in the original image to move or stop moving. That's fine.

[0012] Animated GIF (Graphics Interchange Format )] Modify and replace all or part of an image with a few other images that together form an animated GIF. It may be possible to do so.

[0013] [Video Branching] The modification may be to select a branch of a multi-branching video.

[0014] In any of the above embodiments, the one or more original images may be still or video images. , or a still image with moving sprites. The original image is 2D Or it may be three-dimensional.

[0015] In any of the above embodiments, an algorithm for modifying one or more original images is provided. The algorithm is the artist or composer or content creator who selected one or more images. The algorithm for modification may be custom determined by one or more options. Original imagery selected and used by the selected artist or content creator The algorithm may be predetermined by the company that supplied the software. The algorithm may be predetermined by the company that supplied the image, and may be used to generate a The algorithm may be adapted to modify the step change (e.g., brightness). Gradual change in brightness, gradual change in focus, gradual change in color, small (or partial) image to full image The speed at which the transition (enlarging a new image onto a new image, or other transition method) is performed is determined by the artist or may be specified by the content creator.

[0016] [Touch or mouse input] Another embodiment is, for example, when a viewer touches or points at a portion of the screen within the image. It operates in response to data collected from touch, mouse, or other input devices by One or more images displayed on one or more display screens according to an algorithm that generates a method of modifying an original image of a subject, the modification comprising modifying one or more substitute images and the original image; Image replacement, where the content of the substitute image is the same as the original image but in part or in full. All pixels are corrected. Data collected by the mouse is collected by clicking or It may contain data identified by one or more of over, click, drag and move. The data collected from touches includes the touch by one or more fingers, the force of the touch, the duration of the touch, movement of the touch, and when the user's hand / finger contacts the image display screen. One of the gestures that are close together (e.g., pinching, waving, pointing) The modification may include data identified by one or more of the above modifications. One or more original images may be still or animated or comprise moving sprites. The algorithm can be any of the above elements. may include:

[0017] [Voice Input] Another embodiment involves performing a step of the process according to an algorithm that operates in response to data collected from the vocal sounds. Modify one or more original images displayed on one or more display screens 10. A method, comprising: a) modifying an original image by replacing one or more alternative images with the original image; The content is the same as the original image but some or all of the pixels have been modified The modification may be any of the modifications described above. One or more of the original images may be still images. Or it may be animated or animated with moving sprites, or 2D or 3D. The algorithm may include any of the above elements.

[0018] [Accelerometer Input] Another embodiment involves one or more accelerometers embedded in the housing of the display screen. handheld displays according to algorithms that operate on data collected from 1. A method of modifying an original image displayed on a screen, the modification comprising: The original can be resized by expanding or contracting the left, right, top or bottom boundaries. The goal is to change the field of view of the image of the camera. Data collected from one or more accelerometers is The inclination of a first edge of the display that is away from the viewer relative to the opposite edge, The bevel brings more of the image closer to the first edge into view. The correction Enlarging or reducing by zooming around a part (not necessarily the center) It may be as follows.

[0019] [Authoring Tool] Another embodiment receives guidance from the author and, based on such guidance, determines the actions of the observer. a server computer for generating a data set showing an image that changes based on the A method in a system having a client computer, the method comprising: (a) a server; A computer receiving a series of image specifications from a client computer. (b) at the server computer, selecting an original image in the sequence of images; Specification of the data to be collected to trigger the transition from one image to a second image in the sequence. (c) receiving from the client computer an original image in the series of images; The specification of the speed of the transition from the null image to the second image in the sequence is given to the client computer. (d) receiving from the observer's computer a data set that can be transmitted to the observer's computer. and assembling the dataset on a server computer, the dataset being can be observed on the observer's computer and received from the observer's actions. The data input is a sequence of original images that are replaced with a second image at a specified transition rate. This causes a modification of the observed image.

[0020] Data inputs include, but are not limited to, the gaze positions of one or more observers within the original image. or changes in the facial features of one or more observers, or ambient light, or the observer's display screen distance from the camera, or the number of observers, or the identity of the observer, or voice input, or touch input It may include one or more.

[0021] Modifications include, but are not limited to, changing the brightness of part or all of the original image; Changing the focus of a portion of the original image, enlarging or reducing that portion of the original image the color balance or saturation or color density or contrast of part or all of the original image. Changing the last element causes the sprite to move or stop moving within the original image. This may include selecting a branch of a multi-branching video.

[0022] The original image may include, but is not limited to, a still image or a moving image or a moving sprite. It may be a moving image, or two-dimensional or three-dimensional.

[0023] The client computer and the server computer are each housed in a single computer enclosure. Each software program contained within may be accessed by a single user. may be activated. [Brief explanation of the drawings]

[0024] [Figure 1] 1 illustrates a photographer taking multiple images of a single subject, according to one embodiment. [Figure 2] 1 is a flowchart for creating an exposure map for a set of images according to one embodiment. [Figure 3] 1 is a flowchart for creating a focus map for a set of images according to one embodiment. [Figure 4]1 illustrates how an exposure map for a set of images is stored according to one embodiment. [Figure 5] 1 illustrates how a composer determines how an image changes when a viewer looks at different parts of the image, according to one embodiment. [Figure 6] 1 shows a system that changes an image when a viewer looks at different portions of the image, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0025] One or more embodiments of the present invention may be implemented using a method that is either automatic or predetermined by the photographer / composer. and, for example, as later driven by where the viewer is looking in the image. Improve your photos and adjust the brightness, focus, depth of field, and other quality dimensions for both observers and still images. and making videos interactive.

[0026] In its simplest and most limited form, the embodiment assumes that the original scene is How can someone describe the subject of a scene captured in a photograph that has a huge variation in appearance? The following example will clarify this concept.

[0027] 1. Brightness Imagine a grove of trees with the sun setting behind it. As the observer looks towards the sun, the iris contracts, This causes the pupil to decrease in size, allowing less light in and reducing brightness, sees the other half of the red sun (the part of the sun that is not blocked by the tree), but the tree is black. However, when looking at a single specific tree, the pupils open again. , the characteristics and texture of the tree bark become visible.

[0028] A typical conventional photograph of such a scene would be dark enough to see the shape of the setting sun, If the tree is black or light enough to see the bark, the background of the sunset is the sun. It will be limited to one exposure with no white "blown out" areas.

[0029] There is an existing photographic method that attempts to compensate for this problem, and this method is called "HDR" or High Dynamic Range. HDR requires multiple exposures to avoid motion. The images are taken continuously and automatically blended so that all areas of the image have the correct exposure. However, HDR-compensated images can appear "fake" or "saccharine" or "unrealistic." and therefore often appears unsatisfactory to the observer.

[0030] Instead, the embodiment replicates the action of the observer's own pupil exposure system. Like DR, it uses a set of photographs that differ only in the exposure used when they were taken. However, when any one of the photos is displayed, it looks completely different (compared to HDR technology). Things happen.

[0031] For example, a photo with a properly exposed sun will first appear bright on the monitor. When the observer looks at a darker part of the image that was initially underexposed, such as the tree in this example, the system The gaze detection unit detects the viewer's gaze shifting to the darker parts of the image, and the system Find and display another photo from the set in which the area is more exposed.

[0032] In this way, if the observer's pupils dilate because they are looking at a very dark area, The viewer is rewarded with more image detail being developed. It does the opposite: areas are very bright and blown out, but silhouettes are The sun quickly settles into a red sunset among the sparse trees. The "sensation" of viewing a scene is regained and recreated, enhancing the viewing experience.

[0033] Embodiments may be used with a wide range of cameras and sensors beyond simply determining brightness when taking a still picture. It gives the photographer / composer a way to record changes in the display, as well as to The observer can do this naturally by detecting the line of sight on the screen, or by touching or using the mouse. It allows triggering and searching for changes in the data. Triggers can be manual, automatic or pre-set. It may be in accordance with the intention.

[0034] Alternatively, instead of swapping images with multiple images available for swapping, the algorithm A camera captures not all image pixels of an image in response to actions by one or more observers. The brightness of some image pixels may be adjusted (eg, numerically).

[0035] 2.[Focus] The same method may be used for focus, e.g., if the scene contains objects at various distances. When you have a normal photo, it is not possible to get all objects in perfect focus. Given a camera that takes a photograph with at least a finite depth of field and a finite focal length, Thus, the embodiment tracks which object in the photograph the viewer's eye is looking at. Then, from many different sets of photographs, each taken with a different focus, the object You can select a photo that is accurately focused on the subject.

[0036] 3. "Walking to the window" When a viewer walks towards a photograph or other image that does not fill the entire display, The increase in the angle to the eyes is calculated from the change in the distance between the user and the display, and the display A camera installed in the room measures the change in distance and displays more of the photos simultaneously. The outer frame of the photo can be increased as the viewer approaches the actual window. The content appears, but the viewer must "look around" from left to right through the window to see the additional content. The technology simulates walking up to a real window in that it requires a person to see the window.

[0037] 4. Post-processing effects: Cropping / Scrolling After the camera settings are changed and the photo is taken, the photographer has control over how the image will appear. For example, in cropping, the photographer determines the amount of image to be presented to the viewer. Select a more artistic subset of images. However, the photographer may choose to It may be intended that more of the image be seen if the viewer is interested. If the viewer's eyes look down or up, the image can be scrolled up or down. When done, this can provide the experience of peering into the content of an image. For example, In a photograph, the viewer looks down and sees a path leading to the viewpoint. The horizon is cut off from the distant mountains so that the part of the trail in front of the observer takes up most of the frame. feet forward, i.e., the image will scroll up as the viewer looks down. , the image can be scrolled down when the viewer looks up. Similarly, when the viewer looks left or right, In this case, the image can be scrolled to the right or left.

[0038] Acceleration to scroll the image when the display is moved up and down or sideways The same effect can be achieved with a handheld display by using thermometer data. That is, the handheld display allows the viewer to see the image when it is larger than the window. be made to simulate a portable window that can be moved to observe a desired portion of the is possible.

[0039] 5. Shutter time and movement Imagine a still photograph of several maple samaras descending. In one embodiment, a short exposure Multiple photographs of time can be taken with the photograph having a longer duration. When you first look at an individual seed in a photograph with an exposure time, you see a blurry circle around the center. However, when gazing at it, the observer sees a succession of short exposures. View multiple frames formed by a photo (or a portion of a photo containing a seed) with time The seeds were then observed to actually rotate quickly and drop slightly before becoming stationary in place. (e.g., the last frame formed by a photograph with a short exposure time is displayed) While the observer looks away from the part of the picture containing the seed, the seed remains in its initial position and The blurry appearance will return (e.g., if the photo or part of it has a longer exposure time, (Then the above cycle is repeated when the observer looks at the seed again.) It will be repeated.

[0040] Alternatively, an observer may see something that appears to be moving, such as the maple seeds above. When viewing an object, some of the static images are animated to show that the object is moving. It can be replaced with an image, sprite or animated GIF.

[0041] This also means that the gymnast or skier or jumper is facing a single background. "Sequential Photography" is a method of capturing images at multiple positions. In one embodiment, only one position will be shown. As the viewer looks at the display, other shots are shown, each position across the frame. A single still image will guide the eye through the entire image and back from the last one. This is because the subject of the capture is typically less represented than the series of images it depicts. It is valuable for someone who wants to see how you are moving (e.g., a coach or doctor). It would be.

[0042] 6. Facial features Determine whether a person is smiling, frowning, or showing other specific emotions. Image recognition can be used to identify facial changes. Such facial changes can occur in parts of an image or in the entire image. Image recognition can be used to determine the identity of the observer, The determined identification may alter the image.

[0043] 7. Distance from the display to the viewer A camera placed above the display is used to calculate the distance from the display to the observer. may be used, and the distance may be used to cause changes to part or all of the image. (Methods for measuring distance are optical methods using infrared pulses at a focal point or sonic pulses. Any complex method used by a camera with autofocus, such as an ultrasonic method using a laser. There are several ways to do this: a distance measurement device that the camera can use for autofocusing. The setup is based on the display screen on which the camera is mounted or otherwise associated. This can be used to determine the distance of the observer.

[0044] 8. Touch by observer A touch-sensitive input layer on the surface of the display to receive touch input from the viewer. A touch input may be used to cause a change to part or all of the image. It can be used for.

[0045] 9. Voices from the observer A device mounted on or otherwise associated with a display to receive audio input from an observer. An attached microphone may be used and this audio input may be used to modify some or all of the image. The sound may be used to identify the observer, The image may be modified based on the identity of the observer.

[0046] 10. Display Acceleration If the display is handheld, input is received from the observer who is moving the display. Input from an accelerometer may be used to receive a signal, which may be applied to some or all of the image. can be used to cause changes in

[0047] 11. Ambient Light A camera or other light sensor located on the display is used to determine the amount of ambient light. may be used to cause changes to part or all of the image, such as the brightness of the image. It can be used.

[0048] 12. Further features The previous example shows how the original scene appears with higher fidelity than existing imaging techniques. By attempting to recreate the original scene, we enhance the viewing experience. However, there is no reason to limit these processing techniques to accurate scene reconstruction. The technique developed is an art tool that allows the photographer / composer to guide the viewer through the intention of the work. For example, the composer may act as a guide when the viewer first looks at a part of the photograph. The second time, the brightness and focus should be different, or in any subsequent times, the observer may Alternatively, the constructor may intend for the viewer to focus on a portion of the image. the part of the image near the viewer's point of view to emphasize the concept or meaning implied in It may be intentionally blurred or out of focus, or intentionally distract the viewer.

[0049] [Detailed example] In this section, we discuss how embodiments enhance the experience of viewing images and how they can be used to protect the image from unauthorized access. The overall goal is to give photographers the tools to more fully express their observational intent. It also describes the components involved and how they achieve the objectives. This description details a subset of the possible variations of the photo. Although shown, the same basic process can be easily extended to other variations as described above. Therefore, the description is intended to be a non-limiting example of an embodiment.

[0050] Figure 1 shows the image and pre-processing steps in a simple embodiment. Photograph 1 is shown using camera 2 taking a set of images of a scene using a Various photographic parameters such as brightness, focus and shutter time are varied for each image. .

[0051] Photos can be transferred to a computer via a Secure Digital (SD) card or any other transfer mechanism. The capture mechanism and circuitry are already in the computer or are already computer), and may be labeled and stored in database 5. The images are automatically analyzed and each is associated with the expected action of the viewer. For example, underexposed or overexposed in one image and normally exposed in another image. The parts of each image produced are automatically annotated. "Underexposed" means that the "Overexposed" means that the details of the original are indistinguishable from black or zero brightness. This means that the part of the image that is not distinguishable from white or 100% brightness is indistinguishable from the bark of the tree in front of the sunset in the previous example. For reference, the bark of the tree is much smaller than the normal exposed area in the same location on the other images. Images with dark areas in these areas are annotated and the information is saved. When an observer looks at an underexposed area of ​​an image, they see a photo (or The resulting image will be underexposed, dark, and devoid of distinguishable detail. An observer looking at the dark section will instead notice a cracked and uneven appearance. Conversely, the viewer will see something very white, such as the sky, and in photographic terms, "blonde." Instead, when you look at the part of the photo that is "out," you see all the shapes and details of the clouds. is visible and you will see bright white clouds against a slightly less bright blue.

[0052] FIG. 2 illustrates a method for finding areas of an image that are underexposed or overexposed, according to one embodiment. , we show a simple example of how uncompressed image pixel information can be analyzed. To find the area, the average darkness of the pixels in that area is measured and they are Averages within smaller regions within that area to see if they differ by a predetermined amount For example, if the average brightness Bavg for the entire area is ≦5% of the maximum brightness Bmax, In some cases, at least 75% of the 10 square pixel area within the area is greater than the average brightness of the area. If the exposure varies by only ±2% from Bavg to Bmax, the entire area is considered underexposed. A similar decision can be made for overexposure, e.g., for the whole area If the average brightness Bavg is ≥ 95% of the maximum brightness Bmax, At least 75% of the square pixel area is within ±0.5% of the area's average brightness Bavg to Bmax. If there is only a 2% change, the entire area is considered and labeled as overexposed. The x and y coordinates of the observed area are stored in a database 5 associated with the image. If the gaze of the person is directed towards them, these will later be the same parts as defined above. It is the part that is replaced with another image (or part of it) that is not under- or over-exposed. Although example definitions of "underexposed" and "overexposed" are given above, The object or software application may be "underexposed" and "overexposed" in any suitable manner. It should be noted that it is conceivable to define the average brightness of an area, B If avg is less than x% of the maximum brightness Bmax that the area can have, At least t% of s% of the m n pixel area in the Areas of the image may be "underexposed" if ax varies by only ±v%. where x may have any range, such as 5-15, and t may have any range, such as 50-80. Well, s may have any range such as 50-75, and m·n may be 4-500 pixels. 2 etc. and v may have any range such as 1-10. When the average brightness of the rear Bavg is greater than y% of the maximum brightness that the area can have Bmax and at least t% of s% of the area of ​​m n pixels in the area is greater than the average luminance of the area. An area of ​​the image is "overexposed" if it varies from Bavg by only ±v% of Bmax. where y may have any range, such as 85-95, and t may have any range, such as 50-80. may have any range, s may have any range, such as 50-75, and m·n may have any range, such as 4- 500 pixels 2 and v may have any range such as 1-10. and Bmax is the maximum digital luminance value that a pixel can have (e.g., 8 bits). 255 in our system), and Bmin is the minimum digital brightness a pixel can have. It may be a value (e.g., 0 on an 8-bit system).

[0053] Figure 3 shows the range of images from those with low spatial variation (out of focus) to those with clear focus. How to differentiate between sub-portions of an image to distinguish areas with high spatial variability (e.g., Indicates whether an incremental Fast Fourier Transform (FFT) is performed (to prevent out-of-focus, underexposed (e.g. not overexposed (e.g., black) or overexposed (e.g., white) high frequency, i.e., low spatial variation This area is replaced by another image in the display time that has high frequencies in the same area. The "high spatial variability H" may be predetermined, e.g., m n The vertical axis of the m n pixel block is the luminance of the brightest pixel in the pixel block. It may also mean that there is a ≥ ± a% change in brightness in the horizontal or vertical dimension, where a may have any suitable range, such as 70-95, and m n is 4-500 pixels 2 etc. It may have any suitable range. Similarly, the "low spatial variance L" may be predetermined, For example, the brightness of the brightest pixel in an m n pixel block is m n pixels. This means that there is a ≤ ±c% variation in luminance in the vertical or horizontal dimension of the block. where c may have any suitable range, such as 0-30, and m·n may be 4-500 ppi. Xel 2 Examples of "high spatial variability" and "low spatial variability" The definition of the photograph is defined earlier, but the photographer or another person or software application It is understood that an organization may define "high spatial variability" and "low spatial variability" in any suitable way. It should be noted that

[0054] Figure 4 shows how the photo of the sun behind the tree is displayed at different x,y coordinates in the database. The following example illustrates how a tree can be decomposed into rectangles. The rectangles are divided into rectangles, each containing a part of the tree and each rectangle covering a part of the tree. and small enough so as not to cover the space between the trees or to cover only a small part of it. In Figure 4, any system of numbering rectangles to reference their location within the photograph may be used. In this case, the rectangle numbering starts from the bottom left. Each rectangle is a number of pixels. A rectangle (e.g., 100-500 pixels) with x equal to 1 and y equal to 3 ( The rectangle (1,3) is mostly underexposed since it contains part of a tree, and the rectangle (4,4) contains part of the sun. However, rectangle (3,4) contains some trees and the sun. Since it contains both parts of the image, it is effectively half underexposed and half overexposed. When a person looks at this rectangle, the algorithm determines whether it is more or less exposed. You may wonder if a photo should be used instead, but a smaller rectangle is used. If a photo is more interesting to the human eye, it will have more pixels than the default size rectangle. For example, it contains an object that occupies more than 10 pixels vertically and 10 pixels horizontally (photo But the possibility is that a checkerboard of black and white squares that are exactly the same size as a negligibly small rectangle This problem is unlikely to occur unless you have , the computer scans the rectangles (3,4) until each rectangle is primarily overexposed or underexposed. The resolution of the algorithm can be increased by decomposing into smaller rectangles. The overhead of analyzing a smaller rectangle with pixels is less than that of a microprocessor or microprocessor. This is trivial for many of today's computer integrated circuits (ICs), such as microcontrollers. There is no eye gaze detector or anything that can indicate where the mouse or observer is looking. Do not use smaller rectangles, except that the resolution can be less than would otherwise be possible There is no reason to do this. This resolution is determined by the maximum resolution of the hardware and display used. It is defined by the observer's distance. Therefore, the photographer must determine the hardware resolution and distance in advance. You can specify it or use the default value.

[0055] Other photographic parameters may be varied to relate to the xy area that the observer subsequently views. Furthermore, if the identity of a particular potential observer is specified in advance and identified at the time of presentation, The system accepts as input which individual observer is looking at the x and y coordinates, and The response can be varied based on the identity of the

[0056] [Steps taken by the image composer] As shown in Figure 5, a simple user interface on a computer display 12 Using this database, a photographer or other image composer 11 can capture the image as it is viewed by a viewer. Which parts of the image in the source trigger which actions (if any) and in what order? You can quickly indicate and specify what should be done as follows:

[0057] The constructor 11 has the following related information when it is observed (hereinafter also referred to as "display time"). A list of areas of the image that are likely to have action taken is presented. For example, The list can be divided into underexposed areas, overexposed areas, areas with high spatial resolution, or areas with low spatial resolution. The configurator then takes action for each area listed. Without being limited to, an action can be specified to be part of the following: The whole image or just that part of the image can be exposed to different exposure values ​​that are neither underexposed nor overexposed. or, for example, maximum detail or "focus" relative to the resolution of the original camera. "matching" means replacing an image with another image having a different variation in spatial frequency, replacing image number x with If it is the nth substitution, it is replaced with image number y. If only area b is seen, If image z is found, replace it with image z. A different action is performed on each image in the database. Alternatively, a set of actions may be associated with a primary image (see below). Actions may include viewing images pulled from the internet or other systems. Extends to selecting images or videos available in real time as the observer looks It can be done.

[0058] A single photograph is the image to be observed and all for when looking at the xy rectangle within it. Any photo can be designated as the primary image with the orientation of In this example, if the constructor does not specify an action, In this case, the default is "best automatic" as shown below.

[0059] All actions that the observer might take and all or part of the resulting image Once the composer has completed all the permutations, all this information is combined into one "show" or is "Eyetinerary TM This show is all saved in the cloud. It may be stored online or in any other suitable electronic storage device or location.

[0060] As shown in FIG. 6, when an observer 13 views an image, a display 14 It displays photos from the database 5 stored in and loaded from the data 17. For example, The gaze detection hardware device 16, consisting of one or more cameras mounted on the sprayer 14, It gives the x and y coordinates of where the observer's eyes are looking in the photograph. The system calculates these coordinates The system searches for the action specified by the user in the database 5 using , by referencing either the primary photo or the current photo entry for that xy rectangle. So take action.

[0061] If the action is the default "best automatic", the system will The photo being viewed is compared to the best exposed photo for the xy rectangle the viewer is currently viewing. As mentioned above, "best exposed" here means the best exposure of a set of photographers. agree, showing the most detail and the most "blown out" brightness. This replacement is a quick dissolve. The nature and duration of the image may be predetermined for use by the observer. may be specified or a default value will be used.

[0062] The database mentioned above contains an initial image and a set of links or portions of that image. It is stored online in the cloud as an HTML page consisting of JavaScript calls that are sent to the It is important to note that the gaze detection is performed as if the observer As if your eyes were a mouse, they tell the browser where they are looking, and the browser uses the link. By displaying a new image that is the target of the query, or by changing the aspect ratio of a part of the image, for example Automate actions by calling JavaScript functions to change the object. In this way, the configurator can see that everything is in the server's directory. It may be sufficient to send the viewer a link pointing to a .tml page and a set of images. .

[0063] The user has previously used a user interface with a display and a pointer. Instead of specifying a specific While viewing the xy area, the system via voice, or mouse, or any other method You can specify which other photo (or part of it) should be used by commands to These other correlations with other photos (or portions thereof) are stored in a database. For example, if a photographer is looking at the rectangle (1,3) in Figure 4, You can actually say to the system, "When an observer looks at the area I'm looking at now, Replace the photo with a photo named "tree-exposed.jpg". You can say, "Please do so."

[0064] [Additional Use] The embodiments can be used for more than expressing the artist's intent. For example, a vision therapist may improve a patient's ability to coordinate their eyes by drawing attention in front of and behind a screen. Pilots may want to increase the power of their characters based on an existing set of images or sprites. by making dangerous aircraft appear in any location the pilot was not looking at. and flight systems with displays to quickly locate such dangerous air traffic. This can be trained in a simulator system. If the camera is allowed to see further away, embodiments may substitute a photograph taken at a further focus. Thus, the observer can focus in and out of the 3D features and You can train yourself to "travel" outside.

[0065] [Gaze Detection Image Tool] Embodiments are directed to a user interface (UI) based on gaze detection for an image tool. For example, if an observer tilts their head to the right, the tilted head will return to the center. You can tell the system to rotate the image to the right. , which can be controlled by gazing at a part of the image. It reaches a maximum each time the observer blinks their left eye. The brightness may be increased until the right eye blinks and then begins to decrease again. The process can be stopped at the desired selection by closing both eyes for one second. A similar method can be used for crops.

[0066] [Additional Details] In one embodiment, if the viewer's eyes are focused on a portion of the image, that portion of the image is more visible. After the image with the greatest detail or magnification has been substituted, When the observer's eyes remain fixed on that area, the subject will be photographed in a way that is predetermined by the photographer. The links associated with that part of the image lead to other images or videos that can be downloaded. If the photographer does not include a specific link, that part of the image may be substituted. The initial image seen in the section determines what the image is and provides further information or download other images of similar content captured by other photographers. For example, the image may be transmitted to a cloud-based analyzer where it can be displayed. If you are watching a scene in Washington, D.C., and you stare at the Washington Monument, fills half the display with a series of alternate images for that section of the image If the observer continues to gaze at the monument, the image may expand in size to show more detail. A video downloaded from the internet is shown on the screen describing the monument's history and architecture. may play that section of the (It may have been selected automatically by a cloud-based image recognition device.)

[0067] "Flashlight" The photographer knows that any part of the image the viewer sees can be instantly corrected to a higher brightness. The configurator can simply specify "flashlight" for the xy area. Therefore, the flashlight is actually directed at the scene by the observer's eyes. It will appear to be working against the human interface. In applications related to this, existing systems record where people tend to look in an image. In one embodiment, the system may be configured to measure how often a viewer sees a particular illuminated or You can imagine a log that is output based on whether the image returns to the focused area. When an image composer specifies a "path" for an image alternative, i.e., when a sub-portion of image A is viewed, This results in the substitution of image B, and the viewer seeing part of image B returns to A, and the viewer is then able to see the ABA If this path is repeated, the log will store the information of the path and that it was repeated. Such information may be useful in further developing the interaction between the image(s) and the viewer. It will be beneficial for those

[0068] [Networking and games with other observers] More than one gaze detection device is used for games with other co-located observers. The system can also be used to connect with the internet or local networks for various social purposes. For example, more than one observer at different locations can use the same Eye etinerary TM When viewing a scene, the part seen by most observers is magnified. It could be the part that most people saw, or vice versa. If so, EyetineraryTM forces all observers to go to that image. The first person must look in the right place or activate the weapon in the right place. Other games, such as games where two or more gaze detection systems award points based on A type of social interaction can be implemented. Finally, by viewing the part, Portions of the image can be shared with others over the internet.

[0069] [Feedback to the creator] The embodiment may be based on a predetermined viewer view based on where the viewer looks at the time of display. However, the observer can communicate with the constructor in real time via the network. The system can be used to send useful feedback, e.g., images part of the image, for example, a specific part in a photograph of several people that the viewer wants the composer to replace. There is a person, and the observer blinks three times while looking at the person, drawing a circle around them with their eyes. The observer can then say, over the network connection, "When I look at this person, You can also say, "I want you to substitute another image." This part of the face is the audio file. The constructor will then be able to determine whether the image is good or bad based on the observer's input. Eyetinerary TM can be developed.

[0070] Accordingly, while specific embodiments have been described herein for purposes of illustration, it is to be understood that the scope of the present disclosure is not to be limited to the specific embodiments described herein. It is understood that various modifications may be made without departing from the spirit and scope of the present invention. Where alternatives are disclosed for an embodiment, these alternatives are included unless otherwise stated. Furthermore, any component or operation described may be implemented in hardware. Hardware, software, firmware or hardware, software and hardware The present invention may be implemented / performed in any combination of two or more of the above. Or one or more components of the system may be omitted from the description for clarity or other reasons. Furthermore, one or more of the described devices or systems may be included in the description. may be omitted from the device or system.

Claims

1. at least one display screen viewable by an observer; an eye-gaze detection camera coupled to the at least one display screen and configured to capture data related to the eye gaze of the observer when the observer gazes at an image of the scene displayed on the at least one display screen; at least one processor coupled to the at least one display screen and the gaze detection camera; wherein the at least one processor determining from the captured data a gaze position on the at least one display screen at which the viewer's gaze is directed; selecting a display position on the at least one display screen in response to the determined gaze position; and causing a sprite to move upward or a sprite to stop moving upward within the image of the scene displayed on the at least one display screen at a selected display position. configured to: A system wherein none of said at least one display screens has a touch screen interface.

2. The system of claim 1 , wherein the display position corresponds to the gaze position, and the at least one processor is configured to cause upward movement of the sprite at the gaze position.

3. The system of claim 2 , wherein the at least one processor is configured to cause the sprite to move upward at another location different from the gaze location.

4. The system of claim 1 , wherein the display position corresponds to the gaze position, and the at least one processor is configured to cause the sprite to stop moving upward at the gaze position.

5. The system of claim 1 , wherein the image of the scene is a motion image.

6. The system of claim 1 , wherein the at least one processor is configured to cause the at least one display screen to render an image of the scene as a three-dimensional image.

7. capturing data relating to a line of sight of an observer while the observer gazes at an image of a scene displayed on at least one display screen; determining a gaze position on at least one display screen at which the viewer's gaze is directed in response to the captured data; selecting a display location on the at least one display screen in response to the determined gaze position; causing a sprite to move upward or a sprite to stop moving upward within the image of the scene displayed on the at least one display screen at a selected display position; Including, The method, wherein none of the at least one display screens has a touch screen interface.

8. The method of claim 7 , wherein the display position corresponds to the gaze position, the method including causing an upward movement of the sprite at the gaze position.

9. The method of claim 7 , further comprising causing the sprite to move upward at a location different from the gaze location.

10. The method of claim 7 , wherein the display position corresponds to the gaze position, the method including causing the sprite at the gaze position to stop moving upward.

11. A non-transitory processor-readable medium having embodied thereon program instructions configured to be executed by at least one processor, the program instructions, when executed by the at least one processor, cause the at least one processor to: capturing data relating to a gaze of an observer when the observer gazes at an image of a scene displayed on at least one display screen; determining a gaze position on the at least one display screen to which the viewer's gaze is directed in response to the captured data; selecting a display position on the at least one display screen in response to the determined gaze position; and causing a sprite to move upward or a sprite to stop moving upward within an image of the scene displayed at a selected display position on the at least one display screen. Let them do this, A non-transitory processor-readable medium, wherein none of the at least one display screens is a touch screen interface.

12. 12. The non-transitory processor-readable medium of claim 11, wherein the display position corresponds to the gaze position, and the program instructions cause the at least one processor to cause upward movement of the sprite at the gaze position.

13. 12. The non-transitory processor-readable medium of claim 11, wherein the program instructions cause the at least one processor to cause upward movement of the sprite at a location different from the gaze location.

14. The non-transitory processor-readable medium of claim 11 , wherein the display position corresponds to the gaze position and includes causing the sprite at the gaze position to stop moving upward.

Citation Information

Patent Citations

  • Alternative screen display game device and alternative screen display game program

    JP2014158641A

  • Eye-tracking-based audiovisual playback position selection

    JP2014526725A

  • Display device and digital camera

    JP2015232811A

  • Sound-enhanced ebook with sound events triggered by reader progress

    US20120001923A1

  • User interface apparatus and input method

    WO2011074198A1