Image providing apparatus and image providing method

The image providing device generates stereoscopic images from a single image and video by detecting facial regions and adjusting head angles, addressing the inefficiency of existing depth perception methods and offering a seamless 3D experience.

JP2025079329APending Publication Date: 2025-05-21SK HYNIX INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024194100
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-09
Filing Date
2024-11-06
Publication Date
2025-05-21

AI Technical Summary

Technical Problem

Existing image providing technologies require multiple images or videos to create a sense of depth, which is cumbersome and inefficient.

Method used

An image providing device and method that generates stereoscopic images using a single image and video by detecting facial regions, adjusting head angles, and generating similar images based on gaze positions, allowing for real-time depth perception without additional devices.

Benefits of technology

Enables stereoscopic image display using a single image and video, providing a seamless and efficient 3D experience without the need for multiple images or videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025079329000001_ABST
    Figure 2025079329000001_ABST
Patent Text Reader

Abstract

To provide a three-dimensional image.SOLUTION: An image providing device 10 according to an embodiment of the present invention may include a first pre-processing unit 100 that detects a first face area FC1 from each frame of an original video Video and generates a first image IMG1, a second pre-processing unit 200 that detects a second face area FC1 from an original image OI and generates a second image IMG2, and a similar image generating unit 400 that generates similar images SI1 to SIn corresponding to each of the viewer's gaze positions on the basis of the first image IMG1 and the second image IMG2.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to an image providing device and an image providing method capable of providing an image with a three-dimensional effect. [Background technology]

[0002] Image sensing devices are devices that capture optical images using the properties of light-sensitive semiconductor materials that react to light. With the development of industries such as automobiles, medicine, computers, and communications, there is an increasing demand for high-performance image sensing devices in various fields such as smartphones, digital cameras, game consoles, the Internet of Things, robots, security cameras, and medical micro cameras.

[0003] In addition, technology to provide viewers with 3D (3-dimensional) images that can express a sense of depth is being continuously developed, and in recent years, research into images that can produce a sense of depth without the need for separate devices (e.g., 3D glasses) has been ongoing. In particular, technology that realizes parallax by displaying different images according to the viewer's line of sight has been attracting attention. However, this technology requires the preparation of multiple images or videos according to the viewer's line of sight, which is a considerable hassle. Summary of the Invention [Problem to be solved by the invention]

[0004] The technical idea of ​​the present invention is to provide an image providing apparatus and method capable of providing a stereoscopic image without the need to prepare multiple images or moving images.

[0005] The technical problems of the present invention are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art from the following description. [Means for solving the problem]

[0006] An image providing device according to one embodiment of the present invention disclosed in this document may include a first pre-processing unit that detects a first facial area from each frame of an original video and generates a first image, a second pre-processing unit that detects a second facial area from the original image and generates a second image, and a similar image generating unit that generates similar images corresponding to each of the viewer's gaze positions based on the first image and the second image.

[0007] An image providing device according to another embodiment of the present invention may include a similar image generating unit that generates similar images corresponding to each gaze position of a viewer based on a first image generated using each frame of an original video and a second image generated using the original image, and an image display unit that selects a similar image corresponding to a current gaze position from among the similar images generated by the similar image generating unit and outputs it to a screen.

[0008] An image providing method according to one embodiment of the present invention may include the steps of generating a first image using each frame of an original video, generating a second image using the original image, generating similar images corresponding to each of the viewer's gaze positions based on the first image and the second image, and selecting a similar image corresponding to a current gaze position from the generated similar images and outputting it to a screen. Effect of the Invention

[0009] According to the embodiments disclosed herein, a stereoscopic image can be provided using only one image and one video without the need to prepare multiple images or videos. In addition, various other benefits may be provided by this document, either directly or indirectly. [Brief description of the drawings]

[0010] [Figure 1] FIG. 2 is a block diagram illustrating an image providing device according to an embodiment of the present disclosure. [Diagram 2] 2 is a flowchart for explaining an operation of a first pre-processing unit shown in FIG. 1; [Diagram 3] 4 is a flowchart for explaining an operation of a second pre-processing unit shown in FIG. 1; [Figure 4] 2 is a diagram for explaining a detailed operation of the similar image generating unit shown in FIG. 1; [Diagram 5] 5 is a diagram showing an example of a method for training a first feature extraction unit, a second feature extraction unit, and an image restoration unit shown in FIG. 4. FIG. [Figure 6] 2 is a block diagram of an example of a computing device corresponding to the image providing device of FIG. 1. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] Various embodiments will be described below with reference to the accompanying drawings. However, it should be understood that the present disclosure is not limited to specific embodiments, and includes various modifications, equivalents, and / or alternatives of the embodiments. The embodiments of the present disclosure can provide various effects that are directly or indirectly recognizable by the present disclosure.

[0012] FIG. 1 is a block diagram showing an image providing device according to an embodiment of the present disclosure. 1, an image providing device 10 is a device for providing a viewer with a stereoscopic image, and can provide different images depending on the viewer's gaze position. According to an embodiment, the image providing device 10 may be a head mounted display or a stereoscopic display.

[0013] The image providing device 10 may include a first pre-processing unit 100, a second pre-processing unit 200, a head angle information storage unit 300, a similar image generator 400, and an image display unit 500.

[0014] The first pre-processing unit 100 may detect a face region from each frame of an original video (Video) to generate a first image (IMG1). The original video may be an image including at least one person's face. In this disclosure, it is assumed that the original video includes one person's face, but the scope of the present invention is not limited thereto, and the technical idea of ​​the present invention may be substantially similarly applied even when the original video includes multiple person's faces (however, in this case, an operation of selecting one of the multiple faces may be added). The original video may be composed of multiple frames, and the operation of the first pre-processing unit 100 may be performed sequentially on the multiple frames.

[0015] Also, the face region of the first image (IMG1) may be a region cropped from each frame of the original video by the first pre-processing unit 100 so as to include the person's face. For example, the face region may have a rectangular shape, but the scope of the present invention is not limited thereto, and may have any shape (e.g., circular, hexagonal, etc.). According to one embodiment, if a frame of the original video itself corresponds to a face region, the first pre-processing unit 100 may generate a first image (IMG1) by omitting an operation of cropping the face region from the frame of the original video.

[0016] Additionally, each frame of the original video can include a color image composed of RGB (Red, Green, Blue) which is color information about the scene, and a depth image which contains information about the distance to the scene.

[0017] The second pre-processing unit 200 may detect a face region from an original image (Image) to generate a second image (IMG2). The original image may be an image including at least one human face. In this disclosure, it is assumed that the original image includes one human face, but the scope of the present invention is not limited thereto, and the technical idea of ​​the present invention may be substantially similarly applied even when the original image includes multiple human faces (however, in this case, an operation of selecting one of the multiple faces may be added). Also, the face region of the second image (IMG2) may be a region cropped from the original image by the second pre-processing unit 200 to include the person's face. For example, the face region may have a rectangular shape, but the scope of the present invention is not limited thereto, and may have any shape (e.g., circular, hexagonal, etc.). According to an embodiment, when the original image itself corresponds to a face region, the second pre-processing unit 200 may generate the second image (IMG2) without cropping the face region from the original image.

[0018] The original image may also include a color image composed of RGB (Red, Green, Blue), which is color information about the scene, and a depth image, which includes information about the distance to the scene.

[0019] In the present disclosure, a face region included in each frame of the original video can be defined as a first face region, and a face region included in the original image can be defined as a second face region. On the other hand, the person corresponding to the first face area may be different from the person corresponding to the second face area, but the scope of the present invention is not limited to this, and they may be the same person.

[0020] The head angle information storage unit 300 may store head angle information (HA) including head angles corresponding to each of the viewer's gaze positions. The viewer's gaze position may refer to the relative position of the viewer's head with respect to the screen of the image display unit 500. The head angle may include angles rotated around a first axis (e.g., x-axis), a second axis (e.g., y-axis), and a third axis (e.g., z-axis) that are perpendicular to each other. According to an embodiment, the angle around the first axis may be defined as a pitch angle, the angle around the second axis as a rotation axis as a yaw angle, and the angle around the third axis as a rotation axis as a roll angle. That is, the head angle information (HA) may include pitch angle, yaw angle, and roll angle, which are head angles corresponding to each of the viewer's gaze positions.

[0021] The similar image generating unit 400 can generate similar images (SI1 to SIn; n is an integer equal to or greater than 2) corresponding to each of the viewer's gaze positions based on the first image (IMG1), the second image (IMG2), and the head angle information (HA). The similar images (SI1 to SIn) may be images in which the face in the second image (IMG2) has an expression corresponding to the face in the first image (IMG1) and has a head angle of the head angle information (HA). Here, the number of similar images (SI1 to SIn) can be predetermined based on the range that the viewer's gaze position may have and the number of images to be expressed differently within that range. A method for generating the similar images (SI1 to SIn) by the similar image generating unit 400 will be described in detail later with reference to FIG.

[0022] The image display unit 500 can select a similar image corresponding to the current gaze position of the viewer from among the similar images (SI1 to SIn) generated by the similar image generation unit 400 and output the selected similar image to the screen. In this case, the image display unit 500 can output an image obtained by synthesizing the face area of ​​the original video with the similar image, rather than outputting only the similar image to the screen. Here, synthesis may mean replacing the face of a person included in the face area of ​​the original video with the face of a person included in the similar image. Such a synthesis operation may be performed by the image display unit 500 or a separate image synthesis unit (not shown).

[0023] Meanwhile, the image display unit 500 can use another sensor (e.g., a gyro sensor, a pupil sensor, an ultrasonic sensor, an image sensor, etc.) included in the image providing device 10 to grasp the relative position of the viewer's gaze position with respect to the screen of the image display unit 500, and obtain the viewer's current gaze position.

[0024] In the example of FIG. 1, when the viewer's current gaze position is 510, the image display unit 500 can select a first similar image (SI1) that is a similar image corresponding to the current gaze position 510 and output it onto the screen.

[0025] Alternatively, when the viewer's current gaze position is 520, the image display unit 500 may select and output to the screen a third similar image (SI3) that is a similar image corresponding to the current gaze position 520. As illustrated in the example of FIG. 1, when the viewer's current gaze position is 510, a face close to the front may be displayed on the screen, and when the viewer's current gaze position is 520, a face close to the left side, not the front, may be displayed on the screen.

[0026] According to the image providing device 10 of the present disclosure, it is possible to provide a stereoscopic image corresponding to the current gaze position of the viewer by using only one original video and one image.

[0027] FIG. 2 is a flowchart for explaining the operation of the first pre-processing unit shown in FIG. Referring to FIG. 2, the first pre-processing unit 100 may receive an input of an original video (OV) (S110). The original video (OV) may be stored in a memory (not shown) inside the image providing device 10 and provided to the first pre-processing unit 100. The original video (OV) may include a plurality of frames that proceed sequentially in chronological order, and each of the plurality of frames may be classified by a frame index (i). For example, the first frame of the original video (OV) may have a frame index of 0 (i=0), and from the second frame onwards, the frame index may increase by 1 from 0. According to one embodiment, the initial value of the frame index (i) may be set to 0.

[0028] The first pre-processing unit 100 can extract a frame (FRi) corresponding to a current frame index (i) from an original video (OV) (S120). The first pre-processing unit 100 can detect a face region (FC1) in the frame (FRi) (S130). In one embodiment, the method of detecting the face region (FC1) may be a method of using a face detection algorithm (e.g., a Haar-like feature algorithm). In another embodiment, the method of detecting the face region (FC1) may be a method of manually selecting the face region (FC1) by input by a user of the image providing device 10 (e.g., dragging a mouse).

[0029] If a face region (FC1) is not detected from the frame (FRi) (No in S140), the first preprocessing unit 100 can increment the frame index (i) by 1 (i=i+1) and perform step S120 again on the next frame.

[0030] If a face region (FC1) is detected from the frame (FRi) (Yes in S140), the first preprocessing unit 100 can extract the face region (FC1) from the frame (FRi) (S150).

[0031] The first pre-processing unit 100 may adjust the size of the face region (FC1) to a predetermined reference size (S160). Here, the reference size may be set to any value within a range smaller than the size of the frame (FRi). For example, if the size of the frame (FRi) is 1920×1080, the reference size may be 100×100. In the example of FIG. 2, an example is shown in which the size of the face region (FC1) is enlarged to the reference size, but according to another example, the size of the face region (FC1) may be reduced to the reference size. Here, the enlargement may be performed using an upscaling technique (e.g., bilinear, bicubic, nearest neighbor, lanczos, spline, etc.), and the reduction may be performed using a downscaling technique (e.g., median, mean, etc.). The face area adjusted to the reference size can be defined as the first image (IMG1).

[0032] If the frame index (i) does not correspond to the last frame of the original video (OV) (No in S170), step S145 is performed and step S120 can be performed again for the next frame.

[0033] If the frame index (i) corresponds to the last frame of the original video (OV) (Yes in S170), the operation of the first pre-processing unit 100 can be completed.

[0034] FIG. 3 is a flowchart illustrating the operation of the second pre-processing unit shown in FIG. 3, the second pre-processing unit 200 may receive an original image (OI) (S210). The original image (OI) may be stored in a memory (not shown) in the image providing device 10 and provided to the second pre-processing unit 200.

[0035] The second pre-processing unit 200 may detect a face region (FC2) in the original image (OI) (S220). In one embodiment, the method of detecting the face region (FC2) may be a method of using a face detection algorithm (e.g., a Haar-like feature algorithm). In another embodiment, the method of detecting the face region (FC2) may be a method of manually selecting the face region (FC2) by input by a user of the image providing device 10 (e.g., dragging a mouse).

[0036] If the face region (FC2) is not detected from the original image (OI) (No in S230), the second pre-processing unit 200 can receive an input of another original image (OI) from a memory (not shown).

[0037] If a face region (FC2) is detected from the original image (OI) (Yes in S230), the second pre-processing unit 200 can extract the face region (FC2) from the original image (OI) (S240).

[0038] The second pre-processing unit 200 may adjust the position of the face included in the face region (FC2) to a predetermined reference position (S250). The reference position may be a position where a straight line connecting the centers of two eyes of the face coincides with the horizontal direction and a straight line connecting the centers of the nose and mouth of the face coincides with the vertical direction. According to an embodiment, the second pre-processing unit 200 may use a contour analysis technique (e.g., wavelet transform) to detect the eyes, nose, and mouth of the face. A face region including a face having a position adjusted to correspond to the reference position may be defined as an adjusted face region (FC2').

[0039] The second pre-processing unit 200 can adjust the size of the adjusted face region (FC2') to a predetermined reference size (S260). Here, the reference size may be the same as the reference size described in Fig. 2. In the example of Fig. 3, an example is shown in which the size of the adjusted face region (FC2') is enlarged to the reference size, but according to another example, the size of the adjusted face region (FC2') may be reduced to the reference size.

[0040] The adjusted face region adjusted to the reference size may be defined as a second image (IMG2). When the second image (IMG2) is generated, the operation of the second pre-processing unit 200 may end. By adjusting the face region to the reference position and the reference size, the speed and accuracy of subsequent image processing by the similar image generating unit may be further improved.

[0041] FIG. 4 is a diagram for explaining a detailed operation of the similar image generating unit shown in FIG. Referring to FIG. 4, the similar image generator 400 may include a first feature extractor 410, a second feature extractor 420, a head rotating unit 430, and an image restoration unit 440.

[0042] The first feature extraction unit 410 can extract first feature information (FI1) from a first image (IMG1) including a first depth image (IMG1_D) and a first color image (IMG1_C). The first depth image (IMG1_D) can include information about the distance to the scene for each pixel. The first color image (IMG1_C) can include color information (red, green, blue) about the scene for each pixel.

[0043] Meanwhile, the first feature information (FI1) may be information indicating features related to a face included in the first image (IMG1). According to an embodiment, the first feature information (FI1) may include landmark information and rotation information.

[0044] The landmark information may include three-dimensional coordinates for a facial landmark (e.g., at least one of the eyes, nose, mouth, chin, or ears). Here, the landmark may mean a representative point of each of the eyes, nose, mouth, chin, and ears, and for example, the representative point may be the center of each of the eyes, nose, mouth, chin, and ears, although the scope of the present invention is not limited thereto.

[0045] The rotation information may include the rotation angles (i.e., pitch angle, yaw angle, and roll angle) of the face rotated about the first axis (x-axis), second axis (y-axis), and third axis (z-axis), respectively.

[0046] According to one embodiment, the first feature extraction unit 410 may generate first feature information (FI1) from the first image (IMG1) using a deep learning model (e.g., a U-net encoder, a self-attention GAN (Generative Adversarial Networks) encoder, etc.) for analyzing image features.

[0047] The second feature extraction unit 420 can extract second feature information (FI2) from the second image (IMG2) including a second depth image (IMG2_D) and a second color image (IMG2_C). The second depth image (IMG2_D) can include information about the distance to the scene for each pixel. The second color image (IMG2_C) can include color information (red, green, blue) about the scene for each pixel.

[0048] Meanwhile, the second feature information (FI2) may be information indicating facial features included in the second image (IMG2). According to an embodiment, the second feature information (FI2) may include landmark information and rotation information.

[0049] The landmark information may include three-dimensional coordinates of landmarks for the eyes, nose, mouth, chin, and ears of the face, respectively. The rotation information may include rotation angles (i.e., pitch angle, yaw angle, and roll angle) of the face rotated about a first axis (x-axis), a second axis (y-axis), and a third axis (z-axis), respectively.

[0050] According to one embodiment, the second feature extraction unit 420 may generate second feature information (FI2) from the second image (IMG2) using a deep learning model (e.g., a U-net encoder, a self-attention GAN (Generative Adversarial Networks) encoder, etc.) for analyzing image features.

[0051] The head rotation unit 430 may convert the first feature information (FI1) according to the head angle information (HA) to generate corrected feature information (FI1'). According to an embodiment, the head rotation unit 430 may calculate the three-dimensional coordinates of each landmark changed according to the rotation when rotating the face such that the pitch angle, yaw angle, and roll angle included in the rotation information of the first feature information (FI1) are equal to the pitch angle, yaw angle, and roll angle included in the head angle information (HA), respectively. For example, as illustrated in FIG. 4, the pitch angle and roll angle included in the rotation information of the first feature information (FI1) are equal to the pitch angle and roll angle included in the head angle information (HA), respectively, but the yaw angle of the rotation information is equal to the yaw angle of the head angle information (HA), so that the face rotates counterclockwise around the second axis (y) by α. m If the head must be rotated by a certain amount, head rotation unit 430 can use 3D modeling techniques to calculate the three-dimensional coordinates of each landmark that will change in response to such rotation.

[0052] The image restoration unit 440 may generate similar images (SI1-SIn) based on the corrected feature information (FI1'), the second feature information (FI2), and the second image (IMG2). Each of the similar images (SI1-SIn) may correspond to one head angle information (HA), and the number of similar images (SI1-SIn) may be the same as the number of head angles included in the head angle information (HA). That is, the operation of the image restoration unit 440 to generate similar images may be repeated n times for each head angle information (HA).

[0053] The image restoration unit 440 can generate a similar image by reflecting the facial expression (e.g., the mouth pattern as illustrated in FIG. 4) and head angle appearing in the corrected feature information (FI1') in the second feature information (FI2) and restoring the second feature information (FI2) to a color image by referring to the second image (IMG2).

[0054] According to one embodiment, the image restoration unit 440 may generate similar images (SI1 to SIn) based on the corrected feature information (FI1'), the second feature information (FI2), and the second image (IMG2) using a deep learning model (e.g., a U-net decoder, a self-attention GAN decoder, etc.) for restoring an image from image features.

[0055] Meanwhile, the image restoration unit 440 may generate a current similar image by referring to a previously generated similar image, thereby improving the consistency between similar images.

[0056] FIG. 5 is a diagram showing an example of a method for training the first feature extraction unit, the second feature extraction unit, and the image restoration unit shown in FIG. 5, when the first feature extraction unit 410, the second feature extraction unit 420, and the image restoration unit 440 are realized by a deep learning model, learning for the deep learning model can be performed using a first learning image (TIMG1) and a second learning image (TIMG2) extracted from one video. The first learning image (TIMG1) may be an image obtained by cropping the face of a specific person included in one frame of the video, and the second learning image (TIMG2) may be an image obtained by cropping the face of the specific person included in another frame of the video. The frames corresponding to the first learning image (TIMG1) and the second learning image (TIMG2) can be selected randomly.

[0057] The first feature extraction unit 410 may extract first training feature information (TFI1) from a first training image (TIMG1) including a first training depth image (TIMG1_D) and a first training color image (TIMG1_C).

[0058] The second feature extraction unit 420 can extract second training feature information (TFI2) from the second training images (TIMG2) including a second training depth image (TIMG2_D) and a second training color image (TIMG2_C).

[0059] The image restoration unit 440 may generate a training similar image (TSI) based on the first training feature information (TFI1), the second training feature information (TFI2), and the second training image (TIMG2).

[0060] The image restoration unit 440 can generate a training similar image (TSI) by reflecting the facial expressions and head angles appearing in the first training feature information (TFI1) in the second training feature information (TFI2) and restoring the second training feature information (TFI2) to a color image by referring to the second training image (TIMG2). Also, the image restoration unit 440 can generate a training similar image (TSI) by referring to a previously generated training similar image.

[0061] The learning of the deep learning model can be progressed by changing the weights constituting the first feature extraction unit 410, the second feature extraction unit 420, and the image restoration unit 440 so that the difference value (loss) between the correct image (CAI) which is the first learning color image (TIMG1_C) and the learning similar image (TSI) is minimized. It is more effective to carry out such learning by using as many learning images and as many different videos as possible as learning data.

[0062] FIG. 6 is a block diagram of an example of a computing device corresponding to the image providing device of FIG. Referring to FIG. 6, a computing device 1000 may represent one embodiment of a hardware configuration for performing the operations of the image providing device 10 of FIG.

[0063] Computing device 1000 may include a processor 1010, a memory 1020, an input / output interface 1030, and a communication interface 1040.

[0064] The processor 1010 is capable of processing data and / or instructions required to perform the operations of the components 100 to 500 of the image providing device 10 described in FIG.

[0065] The memory 1020 can store data and / or instructions necessary to perform operations of the components 100 to 500 of the image providing device 10, and can be accessed by the processor 1010. For example, the memory 1020 can be realized with a volatile memory (e.g., Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), etc.) or a non-volatile memory (e.g., Programmable Read Only Memory (PROM), Erasable PROM (EPROM), Electrically Erasable PROM (EEPROM), flash memory, etc.).

[0066] That is, a computer program for performing the operation of the image providing device 10 disclosed in this document can be recorded in the memory 1020 and executed and processed by the processor 1010 to realize the operation of the image providing device 10.

[0067] The input / output interface 1030 can provide an interface to connect an external input device (e.g., a keyboard, a mouse, a touch panel, etc.) and / or an external output device (e.g., a display) to the processor 1010 and enable data to be transmitted and received.

[0068] The communication interface 1040 is configured to be capable of transmitting and receiving various data to and from an external device (eg, an application processor, an external memory, etc.), and may be a device capable of supporting wired or wireless communication. [Explanation of symbols]

[0069] 10 Image Providing Device 100 First pre-processing section 200 Second pre-processing section 300 Head angle information storage unit 400 Similar Image Generation Unit 410 First feature extraction unit 420 Second feature extraction unit 430 Head Rotation Unit 440 Image Restoration Unit 500 Image display unit 510 Gaze position 520 Gaze position 1000 Computing Devices 1010 Processor 1020 Memory 1030 Input / Output Interface 1040 Communication Interface

Claims

1. a first pre-processing unit that detects a first face area from each frame of the original video and generates a first image; a second pre-processing unit for detecting a second face region from the original image and generating a second image; a similar image generating unit configured to generate similar images corresponding to respective gaze positions of a viewer based on the first image and the second image; 2. An image providing device comprising:

2. The similar image generating unit includes: a first feature extraction unit for generating first feature information indicative of facial features included in the first image; a second feature extraction unit for generating second feature information indicative of facial features included in the second image; a head rotation unit that converts the first feature information according to head angle information including head angles corresponding to the respective gaze positions of the viewer to generate corrected feature information; The image providing apparatus of claim 1 , further comprising: an image restoration unit configured to generate the similar image based on the corrected feature information, the second feature information, and the second image.

3. The image providing device of claim 2, wherein each of the first feature information and the second feature information includes landmark information including three-dimensional coordinates of facial landmarks, and rotation information including a rotation angle of the face rotated around each of first to third axes perpendicular to each other.

4. The image providing device of claim 3 , wherein the head rotation unit calculates the three-dimensional coordinates of the first feature information that are changed according to a rotation so that the rotation angle of the first feature information becomes the same as the head angle.

5. The image providing apparatus of claim 3 , wherein the landmarks include at least one of the eyes, nose, mouth, chin, or ears of the face.

6. The image providing device of claim 3 , wherein the rotation angles include a pitch angle, a yaw angle, and a roll angle.

7. The image providing apparatus of claim 2 , wherein the similar image is an image in which a face in the second image has the facial expression and head angle contained in the first image.

8. The image providing apparatus of claim 2 , wherein each of the first feature extracting unit and the second feature extracting unit is a U-net encoder or a self-attention Generative Adversarial Networks (GAN) encoder.

9. The image providing device of claim 2 , wherein the image restoration unit is a U-net decoder or a self-attention GAN decoder.

10. The image providing apparatus of claim 2 , wherein the image restoration unit generates the similar image by referring to a previously generated similar image.

11. The image providing apparatus of claim 1 , further comprising an image display unit for selecting a similar image corresponding to a current gaze position from among the similar images generated by the similar image generating unit and outputting the selected similar image on a screen.

12. The image providing device of claim 11 , wherein the image display unit outputs, to the screen, a composite image in which a face included in the first face region of the original video is replaced with a face included in the selected similar image.

13. The image providing device of claim 1 , wherein the first pre-processing unit generates the first image by cropping the first face region from each frame of the original video.

14. The image providing apparatus of claim 1 , wherein the second pre-processing unit generates the second image by cropping the second face region from the original image.

15. The image providing apparatus of claim 1 , wherein the first preprocessing unit generates the first image by adjusting the first face area to a predetermined reference size.

16. 16. The image providing apparatus of claim 15, wherein the second pre-processing unit adjusts a position of a face included in the second face region to a predetermined reference position, and adjusts the adjusted second face region to the reference size to generate the second image.

17. a similar image generating unit that generates similar images corresponding to each of the viewer's gaze positions based on a first image generated using each frame of the original video and a second image generated using the original image; an image display unit that selects a similar image corresponding to a current gaze position from among the similar images generated by the similar image generating unit and outputs the selected similar image to a screen; 2. An image providing device comprising:

18. The similar image generating unit includes: a first feature extraction unit for generating first feature information indicative of facial features included in the first image; a second feature extraction unit for generating second feature information indicative of facial features included in the second image; a head rotation unit that converts the first feature information according to head angle information including head angles corresponding to the respective gaze positions of the viewer to generate corrected feature information; The image providing apparatus of claim 17 , further comprising: an image restoration unit configured to generate the similar image based on the corrected feature information, the second feature information, and the second image.

19. 20. The image providing device of claim 18, wherein each of the first feature information and the second feature information includes landmark information including three-dimensional coordinates of facial landmarks, and rotation information including a rotation angle of the face around each of first to third axes perpendicular to each other.

20. The image providing apparatus of claim 19 , wherein the head rotation unit calculates the three-dimensional coordinates of the first feature information that are changed according to a rotation so that the rotation angle of the first feature information becomes equal to the head angle.

21. 20. The image providing apparatus of claim 18, wherein the similar image is an image in which a face in the second image has the facial expression and head angle contained in the first image.

22. generating a first image using each frame of the original video; generating a second image using the original image; generating similar images corresponding to respective gaze positions of a viewer based on the first image and the second image; selecting a similar image corresponding to a current gaze position from among the generated similar images and outputting the selected similar image to a screen; How to provide images, including:

23. The step of generating a first image comprises:

23. A method for providing an image according to claim 22, further comprising the step of extracting a next frame if no face region is detected from any one of the frames.

24. The step of generating a second image comprises:

23. The image providing method of claim 22, further comprising the step of receiving an input of another original image if a face region is not detected from the original image.