Image processing device, image processing method, and program

By adaptively applying viewpoint transformation only to areas of high interest in MR systems, the image processing apparatus addresses the inefficiencies of existing viewpoint conversion processes, improving processing time and reducing computational costs.

JP2026078874APending Publication Date: 2026-05-15CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
CANON KK
Filing Date
2024-10-29
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

The existing viewpoint conversion process for converting a captured camera image into an image of a different viewpoint requires significant computational processing time and cost, leading to inefficiencies in Mixed Reality (MR) systems.

Method used

An image processing apparatus that determines the degree of attention in a captured image and adaptively applies viewpoint transformation only to areas of high interest, using depth, gaze, motion, or object detection to reduce processing time and computational resources.

Benefits of technology

This approach reduces processing time and computational costs by selectively applying viewpoint transformation, thereby enhancing the efficiency and speed of image transformation in MR systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026078874000001_ABST
    Figure 2026078874000001_ABST
Patent Text Reader

Abstract

To reduce the processing time and computational cost required for image transformation through viewpoint transformation processing. [Solution] The image processing device includes an image acquisition means for acquiring a first image captured from a first viewpoint in real space, a determination means for determining the degree of attention in the first image, and a generation means for performing a viewpoint transformation process on the first image according to the degree of attention in the first image to generate a second image corresponding to a second viewpoint different from the first viewpoint.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus, an image processing method, and a program.

Background Art

[0002] As a technology for seamlessly and real - time fusing the real world and the virtual world, a so - called MR (Mixed Reality) technology of composite reality is known. One of the MR technologies is a method in which an HMD (Head Mounted Display) user observes an image in which CG (Computer Graphics) is superimposed on a captured image of the real space using a video see - through type HMD. At this time, if the arrangement of the camera for imaging the real space is different from the viewpoint position where the HMD user observes the display, it will cause a sense of discomfort to the HMD user. Patent Document 1 proposes a technology for converting a camera image into an image of a different viewpoint using distance information (depth information) to a subject in the camera image.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, the viewpoint conversion process for converting a captured camera image into an image of a different viewpoint requires enormous computational processing, and the processing time and computational cost required for image conversion by the viewpoint conversion process are large. The object of the present invention is to reduce the processing time and computational cost required for image conversion by the viewpoint conversion process.

Means for Solving the Problems

[0005] The image processing apparatus according to the present invention is characterized by comprising: an image acquisition means for acquiring a first image captured from a first viewpoint in real space; a determination means for determining the degree of attention in the first image; and a generation means for generating a second image corresponding to a second viewpoint different from the first viewpoint by performing a viewpoint transformation process on the first image according to the degree of attention in the first image. [Effects of the Invention]

[0006] According to the present invention, it is possible to reduce the processing time and computational cost required for image transformation through viewpoint transformation processing. [Brief explanation of the drawing]

[0007] [Figure 1] This figure shows an example configuration of a system to which an image processing device is applied. [Figure 2] This figure shows an example of the functional configuration of an image processing device. [Figure 3] This figure shows an example of the hardware configuration of an image processing device. [Figure 4] This is a flowchart showing an example of processing performed by an image processing device. [Figure 5] This is a schematic diagram showing an overhead view of what is seen through the HMD (Head-Mounted Display). [Figure 6] This diagram illustrates the images processed by the HMD (Head-Mounted Display). [Figure 7] This is a flowchart showing an example of the attention level determination process in the first embodiment. [Figure 8] This is a flowchart showing an example of the attention level determination process in the second embodiment. [Figure 9] This is a flowchart showing an example of the attention level determination process in the third embodiment. [Figure 10] This is a flowchart showing an example of the attention level determination process in the fourth embodiment. [Figure 11] This is a flowchart showing an example of the attention level determination process in the fifth embodiment. [Modes for carrying out the invention]

[0008] Embodiments of the present invention will be described below with reference to the drawings.

[0009] (First embodiment) Figure 1 shows an example of the configuration of a system to which the image processing device in this embodiment is applied. As shown in Figure 1, the system in this embodiment has a head-mounted display device (HMD) 100 and a virtual object presentation device 104. The virtual object presentation device 104 is implemented, for example, by an independent PC (personal computer) and connected to the HMD 100 by an interface cable or a wireless network. Alternatively, the virtual object presentation device 104 may be a computer device built into the HMD 100. The virtual object presentation device 104 can be replaced by a general personal computer (information processing device) that operates by a computer program.

[0010] The HMD100 includes a camera 101, a depth sensor 102, and a display 103. These may be built into the HMD100 or they may be separate devices. The camera 101 is a camera (imaging device) for capturing images of the real world. The depth sensor 102 is a sensor for acquiring depth information (depth information, distance information) to objects in the real world (hereinafter also referred to as "real objects"). The display 103 is used to present images to the HMD user wearing the HMD100.

[0011] The virtual object display device 104 reads content information stored internally or externally and renders the virtual object as Computer Graphics (CG) based on the content information. Although there are multiple algorithms for CG calculation, here we assume that the polygon-based calculation method (polygon rendering), which is widely used in the field known as real-time rendering, is used. The details of this process are widely known and are performed by the rendering engine, so a detailed explanation is omitted. In this embodiment, an image containing CG is represented as a virtual image, and the virtual image may contain color information, transparency information, depth information, etc.

[0012] Figure 2 shows an example of the functional configuration of the image processing device in this embodiment. The image processing device in this embodiment includes an imaging unit 201, a depth acquisition unit 202, a display unit 203, a memory 204, a focus determination unit 205, a viewpoint transformation unit 206, a virtual object acquisition unit 207, and a synthesis unit 208. The image processing device in this embodiment also includes a gaze acquisition unit 209, a motion acquisition unit 210, and an object detection unit 211. Some or all of the functions shown in Figure 2 are implemented by, for example, the HMD 100. Alternatively, some or all of the functions shown in Figure 2 may be implemented by the virtual object presentation device 104.

[0013] The imaging unit 201 acquires a camera image captured by the camera 101 in real space. The depth acquisition unit 202 performs measurements in real space using the depth sensor 102 and acquires a depth image showing depth information (depth information, distance information) to real objects. The imaging unit 201 is an example of an image acquisition means, and the depth acquisition unit 202 is an example of a depth acquisition means. Here, the depth sensor 102 is positioned so that the acquired depth image covers the shooting range of the camera image, and can measure the distance to real objects shown in the camera image. Alternatively, the depth image may be generated based on the camera image. In this case, for example, the distance in real space can be measured using stereo matching between the camera image captured by the camera 101 and an image captured from a different viewpoint by another camera (not shown) separate from the camera 101. That is, it is sufficient, but not limited to, that the distance to real objects shown in the image captured by the camera 101 can be measured.

[0014] The attention level determination unit 205 determines areas of high attention (attention areas) and areas of low attention (non-attention areas) in the camera image. In this embodiment, viewpoint transformation processing is adaptively performed on the camera image according to the attention level determined by the attention level determination unit 205. That is, the attention level determination unit 205 determines the areas in the camera image to which viewpoint transformation processing will be applied. Details of the attention level determination unit 205 will be described later.

[0015] The viewpoint conversion unit 206 converts the camera image captured by the camera 101 into an image viewed from a different viewpoint. The viewpoint conversion unit 206 is an example of a generation means. Since the display 103 that displays the camera image captured by the camera 101 is arranged so as to be seen from a viewpoint different from that of the camera 101, a sense of incongruity in appearance occurs when the camera image is displayed as it is. Therefore, the appearance of the image is corrected by the viewpoint conversion process to generate an image from a different viewpoint corresponding to the camera image. As a method, it is common to use the depth information (distance information) of the real object shown in the camera image and the relative positional relationship between the camera and the display to convert the real object in the camera image into an image from a new viewpoint. Since these viewpoint conversion processes are well-known, details thereof are omitted. At this time, it is necessary to accurately measure the depth information of the real object in the camera image. In this embodiment, it is assumed that the depth information can be accurately measured, and it is explained that the disocclusion caused by the difference between viewpoints can also be interpolated by another means. That is, this embodiment can be implemented in combination with a technique of mapping the depth image to match the camera image, a technique of interpolating the result of the viewpoint conversion process at locations where depth information or the image is missing, and the like.

[0016] The virtual object acquisition unit 207 acquires, as a virtual image, an image in which the virtual object generated by the virtual object presentation device 104 is arranged. The acquired virtual image is supplied to the synthesis unit 208 via the memory 204. The synthesis unit 208 synthesizes the real image and the virtual image to generate a synthesized image. That is, the synthesized image is a mixed reality (MR) image in which the real space and the virtual space are fused. By presenting the synthesized image to the display unit 203, the MR image is displayed on the display 103, and the user of the HMD 100 can experience MR.

[0017] The line-of-sight acquisition unit 209 acquires the line-of-sight position of the user of the HMD 100. The motion acquisition unit 210 acquires the motion of the HMD 100. By the line-of-sight acquisition unit 209 and the motion acquisition unit 210 acquiring various information, the position and orientation of the HMD 100 can be acquired. The object detection unit 211 detects a specific object from the camera image acquired by the imaging unit 201. This function will be described in another embodiment.

[0018] Note that the configuration shown in FIG. 2 is an example and is not limited thereto. For example, depending on the implementation mode of the viewpoint conversion processing in the image processing device, a configuration that does not have some of the functional units shown in FIG. 2 may be adopted.

[0019] FIG. 3 is a diagram showing an example of the hardware configuration of the image processing device in the present embodiment. The image processing device 300 in the present embodiment includes a CPU 301, a ROM 302, a RAM 303, a storage device 304, an input unit 305, a display unit 306, a communication unit 307, and a system bus 308. The CPU 301, the ROM 302, the RAM 303, the storage device 304, the input unit 305, the display unit 306, and the communication unit 307 are communicably connected via the system bus 308. Note that the image processing device 300 in the present embodiment may further have other configurations.

[0020] The CPU (Central Processing Unit) 301 controls the entire image processing device 300 using computer programs and data stored in the ROM 302 and the RAM 303, thereby realizing, for example, each function of the image processing device described above. Note that the image processing device 300 may have one or more dedicated hardware different from the CPU 301, and at least a part of the processing by the CPU 301 may be executed by the dedicated hardware. Examples of the dedicated hardware include an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), a DSP (Digital Signal Processor), and a GPU (Graphics Processing Unit).

[0021] ROM (Read Only Memory) 302 stores programs and other data that do not require modification. RAM (Random Access Memory) 303 temporarily stores programs and data supplied from the storage device 304, as well as data supplied from external sources via the communication unit 307. RAM 303 also functions as the main memory and work area of ​​the CPU 301. Storage device 304 is composed of, for example, an HDD (Hard Disk Drive) or an SSD (Solid State Drive) and stores various types of data. Storage device 304 stores, for example, various data necessary for the CPU 301 to perform program-related processing, and various data obtained as a result of the CPU 301 performing program-related processing.

[0022] The input unit 305 consists of, for example, a keyboard, mouse, joystick, touch panel, etc., and receives various instructions from the user and inputs them to the CPU 301. The display unit 306 consists of, for example, an LCD display or LED, and presents display images to the user. The display unit 306 displays, for example, a GUI (Graphical User Interface) for the user to operate the image processing device 300, or the processing results from the CPU 301. Note that the input unit 305 and the display unit 306 may exist as separate devices outside the image processing device 300. The communication unit 307 connects the image processing device 300 to a network and controls communication with other devices, etc.

[0023] The operation of the image processing apparatus in this embodiment will be described with reference to Figure 4. Figure 4 is a flowchart showing an example of processing by the image processing apparatus in this embodiment. In step S401, the imaging unit 201 acquires the camera image captured by the camera 101.

[0024] Next, in step S402, the attention determination unit 205 determines the level of attention within the camera image acquired in step S401. The attention determination unit 205 determines, for example, areas with a high level of attention (areas of interest) within the acquired camera image. In the first embodiment, the attention determination unit 205 determines areas where virtual objects are placed as areas with a low level of attention (areas of non-attention), and other areas as areas with a high level of attention (areas of interest).

[0025] In step S403, the viewpoint transformation unit 206 determines the degree of attention within the camera image in order to adaptively perform viewpoint transformation processing according to the degree of attention within the image. Since areas with a high degree of attention are areas where viewpoint transformation processing is performed, the viewpoint transformation unit 206 performs viewpoint transformation processing on areas that it determines to have a high degree of attention, and the process proceeds to step S404. On the other hand, the viewpoint transformation unit 206 does not perform viewpoint transformation processing on areas that do not have a high degree of attention, i.e., areas that it determines to have a low degree of attention, and the process proceeds to step S406. In this way, the viewpoint transformation unit 206 applies viewpoint transformation processing to areas with a high degree of attention (areas of interest) in the camera image, and does not apply viewpoint transformation processing to areas with a low degree of attention (areas of non-attention).

[0026] In step S404, the depth acquisition unit 202 acquires a depth image for use in the viewpoint transformation process. In step S405, the viewpoint conversion unit 206 performs viewpoint conversion processing in areas of high interest and converts the camera image to match a viewpoint that feels natural when the user views the display 103, which is the display destination. The viewpoint conversion unit 206 performs viewpoint conversion processing on areas of high interest in the camera image based on the depth information in the camera image and the difference between the viewpoint of the camera 101 and the viewpoint of the user viewing the display.

[0027] In step S406, the synthesis unit 208 generates a composite image by combining the real image, which has undergone adaptive viewpoint transformation processing according to its level of attention as described above, with the virtual image in which the virtual object acquired by the virtual object acquisition unit 207 is placed. In other words, for areas of high attention in the camera image, the transformed image transformed by the viewpoint transformation unit 206 is used as the real image, and for other areas, the viewpoint transformation processing is skipped, meaning the captured camera image is used as the real image, and a composite image is generated. In step S407, the display unit 203 displays the composite image generated in step S406. Then, the process shown in Figure 4 is terminated.

[0028] The image processing (viewpoint transformation processing) in this embodiment will be explained with reference to Figures 5 and 6. Figure 5 is a schematic diagram showing a top-down view of a real space as seen from the HMD 100, and Figure 6 is a diagram illustrating the image processed within the HMD 100.

[0029] As shown in Figure 5, the camera 101 of the HMD 100 captures an image of the field of view within the range indicated by the dotted line, and the real objects A501, B502, and C503 are included within the shooting range. The camera image at this time is shown in Figure 6(a). In addition, the HMD 100 can view the displayed image using the display 103, and as shown in Figure 5, the image of the field of view within the range indicated by the solid line can be viewed as the displayed image. As mentioned above, if the image captured by the camera 101 (camera image) is displayed directly on the display 103, it will be an unnatural image, so the viewpoint conversion unit 206 performs viewpoint conversion processing on the camera image. At this time, the image after viewpoint conversion processing of the camera image shown in Figure 6(a) is shown in Figure 6(b). Comparing the changes in the images in Figure 6(a) and Figure 6(b), the real object C, which is close to the HMD 100, changes in appearance significantly in the image, while the change in appearance is smaller for real objects that are farther away from the HMD 100. This is because the degree to which the viewpoint change affects the user differs depending on the distance from the HMD100. The viewpoint conversion unit 206 uses a depth image that shows depth information (depth information, distance information) from the HMD100 to perform correction that takes these effects into account.

[0030] Furthermore, as shown in Figure 5, a virtual object 504 is placed between real objects A501 and B502 and real object C503. Figure 6(c) shows a virtual image in which the virtual object acquired by the virtual object acquisition unit 207 is placed. The virtual object 504 also includes depth information from the HMD 100, and the synthesis unit 208 generates a composite image from the camera image and the depth image. Figure 6(d) shows a composite image obtained by combining the virtual image in Figure 6(c) with the image after viewpoint transformation processing shown in Figure 6(b). By displaying the image in Figure 6(d) on the display 103, the user of the HMD 100 can experience MR.

[0031] Next, the attention determination process by the attention determination unit 205 in the first embodiment will be described. Figure 7 is a flowchart showing an example of the attention determination process in the first embodiment. The process in the flowchart shown in Figure 7 is executed in step S402 of the flowchart shown in Figure 4. When the attention determination process begins, in step S701, the virtual object acquisition unit 207 acquires a virtual image containing the virtual object, as shown in Figure 6(c) as an example.

[0032] Next, in step S702, the attention determination unit 205 detects the position of the virtual object in the real image from the virtual image acquired in step S701, and determines the level of attention in the camera image based on the detection result. In this example, the attention determination unit 205 determines the level of attention so that the level of attention is low in the area where the virtual object is placed, and high in the area other than the area where the virtual object is placed. That is, the attention determination unit 205 determines the area where the virtual object is placed as a non-attention area, and the other areas as attention areas. Then, the process shown in Figure 7 is completed, and the process proceeds to step S403 in Figure 4.

[0033] Figure 6(e) shows the determination of which areas of the image in Figure 6(a) are of high importance using this embodiment. The area where the virtual object is placed is set as area 602 with low importance, and the other areas are set as area 601 with high importance. As mentioned above, it can be seen that the area where the virtual object is placed is determined to have low importance. This is because the area where the virtual object is placed is composited using a virtual image, and the real image behind the virtual object is hidden and not visible, so there is no sense of incongruity, and there is no need to perform the viewpoint transformation processing performed by the viewpoint transformation unit 206. Note that if semi-transparency information is added to the virtual image, a part of the real image hidden by the virtual image will become transparent. In this case, the area containing semi-transparency information may be configured to be set as an area with high importance.

[0034] As explained above, in this embodiment, the attention determination unit 205 determines the level of attention in the camera image, and performs viewpoint transformation processing on areas with high attention (areas of interest), while skipping viewpoint transformation processing on areas other than those areas (areas of interest). This reduces the processing time and computational resources required for image transformation by viewpoint transformation processing, making it possible to speed up and improve the efficiency of processing.

[0035] (Second embodiment) A second embodiment will now be described. In the second embodiment described below, the attention level determination unit 205 determines the level of attention in the camera image based on the depth information (depth information, distance information) contained in the depth image. The process is the same as in the first embodiment described above, except for the attention level determination process, so the explanation will be omitted, and the attention level determination process by the attention level determination unit 205 in the second embodiment will be described below.

[0036] Figure 8 is a flowchart showing an example of the attention determination process in the second embodiment. The process shown in the flowchart in Figure 8 is executed in step S402 of the flowchart shown in Figure 4. When the attention determination process begins, in step S801, the depth acquisition unit 202 acquires depth information associated with the camera image to be processed for viewpoint transformation.

[0037] Next, in step S802, the attention determination unit 205 determines the level of attention in the camera image from the depth information acquired in step S801. In this example, the attention determination unit 205 determines the level of attention based on the acquired depth information, such that the closer the depth direction is to the HMD 100, the higher the level of attention, and the further away the depth direction is, the lower the level of attention. That is, the attention determination unit 205 determines the area where the depth direction is close to the HMD 100 as the area of ​​attention, and the area where the depth direction is far as the area of ​​non-attention. For example, the attention determination unit 205 determines the area where the depth of the real object is below a predetermined threshold as the area of ​​attention. Then, the process shown in Figure 8 is completed, and the process proceeds to step S403 in Figure 4.

[0038] This is because the closer a real-world object is to the HMD100 in the depth direction, the greater the parallax effect caused by viewpoint transformation. Conversely, the further a real-world object is from the HMD100 in the depth direction, the smaller the parallax effect caused by viewpoint transformation. Therefore, in this embodiment, since a small parallax effect reduces the likelihood of discomfort before and after transformation, real-world objects that are far away in the depth direction within the image are assigned a lower level of attention, thereby omitting viewpoint transformation processing. This reduces the processing time and computational resources required for image transformation due to viewpoint transformation processing, making it possible to speed up and improve the efficiency of processing.

[0039] (Third embodiment) A third embodiment will now be described. In the third embodiment described below, the attention level determination unit 205 determines the level of attention in the camera image based on the gaze information acquired by the gaze acquisition unit 209. Since the process is the same as in the first embodiment described above, the explanation will be omitted, and the attention level determination process by the attention level determination unit 205 in the third embodiment will be described below.

[0040] Figure 9 is a flowchart showing an example of the attention determination process in the third embodiment. The process shown in the flowchart in Figure 9 is executed in step S402 of the flowchart shown in Figure 4. When the attention determination process begins, in step S901, the gaze acquisition unit 209 acquires gaze information indicating the point of focus (gaze position) of the display 103 that the user of the HMD 100 is looking at.

[0041] Next, in step S902, the attention level determination unit 205 determines the level of attention in the camera image from the gaze information acquired in step S901. In this example, the attention level determination unit 205 determines the level of attention based on the acquired gaze information such that the closer to the user's point of gaze, that is, the closer to the area the user's gaze is directed, the higher the level of attention. In other words, the attention level determination unit 205 determines the area close to the user's point of gaze as the area of ​​attention. For example, the attention level determination unit 205 determines the area within a predetermined range from the user's point of gaze (gaze position) (the area corresponding to the gaze position) as the area of ​​attention. Then, the process shown in Figure 9 is completed, and the process proceeds to step S403 in Figure 4.

[0042] This is because, within the display image shown on display 103, areas closer to the user's point of focus are regions where the user's visual sensitivity is higher, and therefore require viewpoint transformation processing. On the other hand, areas further from the user's point of focus are regions where the user's visual sensitivity is lower, and for these areas, the level of attention is not determined to be high, thus omitting viewpoint transformation processing. This reduces the processing time and computational resources required for image transformation due to viewpoint transformation processing, making it possible to speed up and improve the efficiency of processing.

[0043] (Fourth embodiment) A fourth embodiment will now be described. In the fourth embodiment described below, the attention level determination unit 205 determines the attention level in the camera image based on the motion information acquired by the motion acquisition unit 210. Since the process is the same as in the first embodiment described above, the explanation will be omitted, and the attention level determination process by the attention level determination unit 205 in the fourth embodiment will be described below.

[0044] Figure 10 is a flowchart showing an example of the attention determination process in the fourth embodiment. The process shown in the flowchart in Figure 10 is executed in step S402 of the flowchart shown in Figure 4. When the attention determination process begins, in step S1001, the motion acquisition unit 210 acquires motion information of the HMD 100. Here, motion information refers to motion information relating to the movement of the HMD 100 when the user wearing the HMD 100 moves. The motion acquisition unit 210 obtains motion information of the HMD 100 by acquiring angular velocity information, acceleration information, etc., obtained from an angular velocity sensor, an acceleration sensor, etc. In this embodiment, motion information is described as velocity information obtained by integrating the angular velocity.

[0045] Next, in step S1002, the attention level determination unit 205 determines the level of attention in the camera image based on the motion information (speed information) acquired in step S1001. The attention level determination unit 205 determines whether the motion information (speed information) acquired in step S1001 is below a predetermined threshold. If the attention level determination unit 205 determines that the acquired motion information (speed information) is below a predetermined threshold, it determines the level of attention so that the overall level of attention in the camera image is high. On the other hand, if the attention level determination unit 205 determines that the acquired motion information (speed information) is not below a predetermined threshold, i.e., is greater than a predetermined threshold, it determines the level of attention so that the overall level of attention in the camera image is low. In other words, if the attention level determination unit 205 determines that the acquired motion information (speed information) is below a predetermined threshold, it designates the entire camera image as a region of attention, and if it determines that it is greater than a predetermined threshold, it designates the entire camera image as a region of non-attention. Then, the process shown in Figure 10 is completed, and the process proceeds to step S403 in Figure 4.

[0046] By determining the degree of attention to the camera image in this way, if the user wearing the HMD100 moves their head quickly, it is determined that the user is not paying attention to the displayed image, and the viewpoint transformation process for the entire image is omitted. This reduces the processing time and computational resources required for image transformation due to viewpoint transformation, making it possible to speed up and improve the efficiency of processing.

[0047] (Fifth embodiment) A fifth embodiment will now be described. In the fifth embodiment described below, the attention level determination unit 205 determines the level of attention in the camera image based on the object information detected by the object detection unit 211. The third embodiment is the same as the first embodiment described above except for the attention level determination process, so the explanation will be omitted, and the attention level determination process by the attention level determination unit 205 in the third embodiment will be described below.

[0048] Figure 11 is a flowchart showing an example of the attention determination process in the fifth embodiment. The process shown in the flowchart in Figure 11 is executed in step S402 of the flowchart shown in Figure 4. When the attention determination process begins, in step S1101, the object detection unit 211 acquires the camera image obtained from the imaging unit 201 and stored in the memory 204.

[0049] In step S1102, the object detection unit 211 analyzes the camera image acquired in step S1101 and detects a specific object in the camera image. This specific object could be, for example, the hand of the HMD100 user. Possible means for detecting an object from a camera image include, for example, an object estimation method using machine learning, or a method for separating the target object from registered color information. In other words, this embodiment can be implemented without particularly limiting the object detection method or the object to be detected.

[0050] In step S1103, the attention determination unit 205 determines the level of attention in the camera image based on the object detection result in step S1102. In this example, the attention determination unit 205 determines the level of attention so that the area of ​​the specific object detected in step S1102 is highly attentional, and the level of attention for the other areas is low. That is, the attention determination unit 205 determines the detected object area as an area of ​​attention and the other areas as areas of non-attention. Then, the process shown in Figure 11 is completed, and the process proceeds to step S403 in Figure 4.

[0051] By determining the level of attention in the camera image in this way, objects that are highly unnatural to the HMD100 user (such as the user's hands) are given priority for viewpoint transformation processing, while other objects are deemed not to be of interest and are omitted without viewpoint transformation. This reduces the processing time and computational resources required for image transformation due to viewpoint transformation processing, enabling faster and more efficient processing.

[0052] Furthermore, the first to fifth embodiments described above are not limited to being configured independently, but may also be configured by appropriately combining the embodiments described above.

[0053] (Other embodiments of the present invention) The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.

[0054] It should be noted that the embodiments described above are merely examples of how the present invention can be implemented, and the technical scope of the present invention should not be interpreted as being limited by them. In other words, the present invention can be implemented in various forms without departing from its technical concept or its main features.

[0055] The disclosure of this embodiment includes the following configurations and methods, etc. (Composition 1) Image acquisition means for acquiring a first image captured from a first viewpoint of real space, A determination means for determining the level of attention in the first image, An image processing apparatus characterized by comprising: a generation means that performs a viewpoint transformation process on the first image according to the degree of attention in the first image to generate a second image corresponding to a second viewpoint different from the first viewpoint. (Configuration 2) The image processing apparatus according to configuration 1, characterized in that the generation means performs viewpoint transformation processing based on depth information relating to a real object included in the first image and the difference between the first viewpoint and the second viewpoint. (Composition 3) The image processing apparatus according to configuration 1 or 2, characterized by having a synthesis means for generating a composite image in which a virtual object is placed on the second image. (Composition 4) It has a display means for displaying an image corresponding to the second viewpoint, The image processing apparatus according to configuration 3, characterized in that the second image or the composite image is displayed on the display means. (Composition 5) The determination means determines the region of interest within the first image, The image processing apparatus according to any one of configurations 1 to 4, characterized in that the generation means performs the viewpoint transformation process on the region of interest in the first image to generate the second image. (Composition 6) The image processing apparatus according to configuration 5, characterized in that the determination means determines an area other than the area where the virtual object is placed as the area of ​​interest. (Composition 7) The image processing apparatus according to configuration 6, characterized in that the determination means determines that a region in the virtual image relating to the virtual object that contains semi-transparent information is the region of interest. (Composition 8) Depth acquisition means for acquiring depth information relating to real objects included in the first image. The image processing apparatus according to any one of configurations 5 to 7, characterized in that the determination means determines the region of interest to be a region in which the depth of the real object is less than or equal to a predetermined threshold. (Composition 9) The system includes a gaze acquisition means for acquiring the user's gaze position relative to a display means that displays an image corresponding to the second viewpoint, The image processing apparatus according to any one of configurations 5 to 8, characterized in that the determination means determines the area corresponding to the user's line of sight as the area of ​​interest. (Composition 10) The system includes a detection means for detecting a predetermined object included in the first image, The image processing apparatus according to any one of configurations 5 to 9, characterized in that the determination means determines the region in which the predetermined object is detected as the region of interest. (Composition 11) The image processing device has an operation acquisition means for acquiring speed information related to the movement of the image processing device, The image processing apparatus according to any one of configurations 5 to 10, characterized in that the determination means determines the entire first image as the region of interest if the speed information is below a predetermined threshold. (Method 1) The image acquisition process involves obtaining a first image by capturing the real space from a first viewpoint, A determination step for determining the level of attention in the first image, An image processing method characterized by comprising: a generation step of performing a viewpoint transformation process on the first image according to the degree of attention in the first image to generate a second image corresponding to a second viewpoint different from the first viewpoint. (Program 1) In the computer of the image processing device, An image acquisition step in which a first image is obtained by capturing the real space from a first viewpoint, A decision step to determine the level of attention in the first image, A program for performing a generation step of performing a viewpoint transformation process on the first image according to the degree of attention in the first image to generate a second image corresponding to a second viewpoint different from the first viewpoint. [Explanation of Symbols]

[0056] 100: HMD 101: Camera 102: Depth sensor 103: Display 104: Virtual object presentation device 201: Imaging unit 202: Depth acquisition unit 203: Display unit 204: Memory 205: Attention level determination unit 206: Viewpoint transformation unit 207: Virtual object acquisition unit 208: Synthesis unit 209: Gaze acquisition unit 210: Motion acquisition unit 211 Object detection unit

Claims

1. Image acquisition means for acquiring a first image captured from a first viewpoint of real space, A determination means for determining the degree of attention in the first image, An image processing apparatus characterized by comprising: generation means, which performs a viewpoint transformation process on the first image according to the degree of attention in the first image to generate a second image corresponding to a second viewpoint different from the first viewpoint.

2. The image processing apparatus according to claim 1, characterized in that the generation means performs viewpoint transformation processing based on depth information relating to a real object included in the first image and the difference between the first viewpoint and the second viewpoint.

3. The image processing apparatus according to claim 1, characterized by having a synthesis means for generating a composite image in which a virtual object is placed on the second image.

4. It has a display means for displaying an image corresponding to the second viewpoint, The image processing apparatus according to claim 3, characterized in that the second image or the composite image is displayed on the display means.

5. The determination means determines the region of interest within the first image, The image processing apparatus according to any one of claims 1 to 4, characterized in that the generation means performs the viewpoint transformation process on the region of interest in the first image to generate the second image.

6. The image processing apparatus according to claim 5, characterized in that the determination means determines an area other than the area in which the virtual object is placed as the area of ​​interest.

7. The image processing apparatus according to claim 6, characterized in that the determination means determines that a region in the virtual image relating to the virtual object that contains semi-transparent information is the region of interest.

8. Depth acquisition means for acquiring depth information relating to real objects included in the first image. The image processing apparatus according to claim 5, characterized in that the determination means determines the region of interest to be a region in which the depth of the real object is less than or equal to a predetermined threshold.

9. The system includes a gaze acquisition means for acquiring the user's gaze position relative to a display means that displays an image corresponding to the second viewpoint, The image processing apparatus according to claim 5, characterized in that the determination means determines the area corresponding to the user's line of sight as the area of ​​interest.

10. The system includes a detection means for detecting a predetermined object contained in the first image, The image processing apparatus according to claim 5, characterized in that the determination means determines the region in which the predetermined object is detected as the region of interest.

11. The image processing device has an operation acquisition means for acquiring speed information related to the movement of the image processing device, The image processing apparatus according to claim 5, characterized in that the determination means determines the entire first image as the region of interest if the velocity information is less than or equal to a predetermined threshold.

12. The image acquisition process involves obtaining a first image by capturing the real space from a first viewpoint, A determination step for determining the degree of attention in the first image, An image processing method characterized by comprising: a generation step of performing a viewpoint transformation process on the first image according to the degree of attention in the first image to generate a second image corresponding to a second viewpoint different from the first viewpoint.

13. In the computer of the image processing device, An image acquisition step in which a first image is obtained by capturing the real space from a first viewpoint, A decision step for determining the level of attention in the first image, A program for performing a generation step of performing a viewpoint transformation process on the first image according to the degree of attention in the first image, thereby generating a second image corresponding to a second viewpoint different from the first viewpoint.