Image processing device and method, and program
The image processing method stabilizes markerless AR image quality by correcting superimposed positions using past frame references, addressing accuracy issues without increasing hardware performance, thus maintaining image quality and reducing costs.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2026-04-09
AI Technical Summary
Markerless AR image quality is compromised due to insufficient camera tracking and feature point detection accuracy, leading to unnatural object positioning, and enhancing hardware performance to improve accuracy results in increased costs.
An image processing method that sequentially acquires and corrects the superimposition position of digital content in current frames based on past frame positions identified by feature points, using 2D coordinates to maintain image quality without increasing hardware performance.
This method suppresses the reduction in markerless AR image quality while preventing increased costs by correcting superimposed positions using 2D coordinates, thereby stabilizing virtual object placement and reducing processing time.
Smart Images

Figure JP2025032817_09042026_PF_FP_ABST
Abstract
Description
Image Processing Apparatus and Method, and Program
[0001] The present disclosure relates to an image processing apparatus and method, and a program, and particularly to an image processing apparatus and method, and a program capable of suppressing a reduction in the quality of markerless AR images while suppressing an increase in cost.
[0002] Conventionally, there has been markerless AR (Augmented Reality), a technology that recognizes the real-world environment and superimposes digital content without the need for specific markers or tags. This technology performs 3D mapping of the surrounding environment by using a camera to recognize the shapes and positions of walls, floors, objects, etc., and estimates the movement of the camera by detecting and tracking characteristic points (edges, corners, etc.) in the camera image, thereby providing a natural and seamless augmented reality experience.
[0003] Also, in such markerless AR, a method has been considered in which an IMU (Inertial Measurement Unit) including an acceleration sensor and a gyroscope is used to detect changes in the movement and orientation of the device, and this is used for complementing camera tracking or tracking high-speed movement, or a depth sensor is used to more accurately capture the 3D structure of the environment (see, for example, Patent Document 1).
[0004] Japanese Unexamined Patent Application Publication No. 2024-45724
[0005] However, if the accuracy of camera tracking or the detection accuracy of feature points such as corners is not sufficiently high, there is a risk that the quality of a markerless AR image, which is an image of a real space with a virtual object (an object of digital content) superimposed by applying this markerless AR technology, will be reduced. For example, in such a markerless AR image, there is a risk that the virtual object may be superimposed at an incorrect position, causing the position of the virtual object to change unnecessarily in the time direction and appear to behave unnaturally.
[0006] However, in order to suppress the occurrence of such phenomena and ensure that the accuracy of camera tracking and feature point detection is sufficiently high, it would be necessary to improve hardware performance, such as by applying a processor with very high image processing capabilities or a very high-precision sensor, which could lead to increased costs.
[0007] This disclosure is made in light of these circumstances and aims to suppress the reduction in the quality of markerless AR images while suppressing the increase in costs.
[0008] One aspect of this technology is an image processing device comprising: an image acquisition unit that sequentially acquires the current frame of an image captured from real space; an overlay unit that sequentially overlays objects of digital content onto the acquired current frame; and a position correction unit that sequentially corrects the current frame overlay position, which is the overlay position of the object in the current frame, based on the past frame overlay position, which is the overlay position of the object in the past frame, identified based on the feature points of the past frame of the image.
[0009] One aspect of this technology is an image processing method that includes sequentially acquiring the current frame of an image captured from real space, sequentially superimposing digital content objects onto the acquired current frame, and sequentially correcting the current frame superimposition position, which is the superimposition position of the object in the current frame, based on the past frame superimposition position, which is the superimposition position of the object in the past frame, identified based on the feature points of the past frames of the image.
[0010] One aspect of this technology is a program that causes a computer to perform the following processes: sequentially acquiring the current frame of an image captured from real space; sequentially superimposing digital content objects onto the acquired current frame; and sequentially correcting the current frame superposition position, which is the superposition position of the objects in the current frame, based on the past frame superposition position, which is the superposition position of the objects in the past frame, identified based on the feature points of past frames of the captured image.
[0011] In one aspect of this technology, the image processing apparatus, method, and program involve sequentially acquiring the current frame of an image captured from real space, sequentially superimposing digital content objects onto the acquired current frame, and sequentially correcting the current frame superimposition position, which is the superimposition position of objects in the current frame, based on the past frame superimposition position, which is the superimposition position of objects in the past frame, identified based on the feature points of the past frames of the image.
[0012] This is a diagram illustrating an example of markerless AR. This is a diagram illustrating an example of markerless AR. This is a diagram illustrating an example of a markerless AR image. This is a diagram illustrating an example of a method for correcting the superposition position. This is a diagram illustrating an example of the process of correcting the superposition position. This is a diagram illustrating an example of feature points for deriving the relative position of the superposition position. This is a diagram illustrating an example of the process of correcting the superposition position. This is a diagram illustrating an example of the relative position of the superposition position based on feature points. This is a block diagram illustrating an example of the main configuration of an imaging device. This is a flowchart illustrating an example of the flow of markerless AR processing. This is a diagram illustrating an example of the process of deriving the relative position of the superposition position. This is a diagram illustrating an example of the process of correcting the superposition position. This is a diagram illustrating an example of the result of correcting the superposition position. This is a block diagram illustrating an example of the main configuration of an image processing system. This is a flowchart illustrating an example of the flow of markerless AR processing. This is a flowchart following Figure 16, showing an example of the flow of markerless AR processing. This is a block diagram illustrating an example of the main configuration of a computer.
[0013] The following describes the embodiments for implementing this disclosure. The description will be in the following order: 1. Supporting literature, etc., for technical content and technical terminology 2. Markerless AR 3. Correction of superposition position using feature points of captured images 4. First embodiment (imaging device) 5. Second embodiment (image processing system) 6. Appendix
[0014] <1. Supporting Documents for Technical Content and Terminology> The scope disclosed in this technology includes not only the contents described in the embodiments, but also the contents described in the following patent and non-patent documents that were publicly known at the time of filing, as well as the contents of other documents referenced in the following patent and non-patent documents.
[0015] Patent Document 1: (as described above) Non-Patent Document 1: Samuele Salti and Luigi Di Stefano, "SVR-Based Jitter Reduction for Markerless Augmented Reality", DEIS, University of Bologna, Bologna 40136, IT, Image Analysis and Processing-ICIAP 2009, https: / / link.springer.com / chapter / 10.1007 / 978-3-642-04146-4_5 Non-Patent Document 2: Ke Xu, Kar Wee Chia, Adrian David Cheok, "Real-time camera tracking for marker-less and unprepared augmented reality environments", Image and Vision Computing Volume 26, Issue 5, 1 May 2008, Pages 673-689, https: / / www.sciencedirect.com / science / article / abs / pii / S0262885607001266
[0016] In other words, the contents described in the aforementioned patent and non-patent documents, as well as the contents of other documents referenced in those patent and non-patent documents, will also serve as a basis for determining the support requirements.
[0017] <2. Markerless AR> Traditionally, there has been a technology called markerless AR (Augmented Reality) that recognizes the real-world environment and overlays digital content without requiring specific markers or tags. This technology uses a camera to recognize the shape and position of walls, floors, objects, etc., thereby creating a 3D map of the surrounding environment. It also estimates camera movement by detecting and tracking characteristic points (edges, corners, etc.) in the camera image, providing a natural and seamless augmented reality experience.
[0018] For example, Patent Document 1 proposes a method for detecting changes in device movement and orientation using an IMU (Inertial Measurement Unit) that includes an accelerometer and a gyroscope in such markerless AR, and using this to complement camera tracking or track high-speed movements, or to capture the 3D structure of the environment more accurately using a depth sensor.
[0019] For example, as shown in the upper part of Figure 1, in markerless AR, an environment map of the real space 10 is generated based on the camera's captured image and IMU sensor information, and feature points (A, B, C) in the real space 10 are determined using this environment map. Then, the position and angle of the camera 11 are detected based on these feature points. Then, based on the position and angle of the camera 11, as shown in the lower part of Figure 1, the digital content object 12 (virtual object) is virtually placed in the real space 10. In other words, a position for compositing the object 12 is set, and the object 12 is superimposed on the image of the real space 10 captured by the camera 11 so that it appears as if the object 12 exists at that position. In this specification, an image in which digital content is superimposed on an image of the real space using markerless AR technology is also referred to as a "markerless AR image".
[0020] However, if the accuracy of camera tracking or the detection of feature points such as corners is not sufficiently high, the quality of markerless AR images may be reduced. For example, as shown in Figure 2, if camera 11 is actually located at the position of camera 11A in real space 10 but is mistakenly detected as being at the position of camera 11B, then object 12, which should be placed at the position of object 12A, may be placed at the position of object 12B. When object 12 is placed in the wrong position in this way, as shown in Figure 3, object 22 should be superimposed at the position of object 22A in the markerless AR image 20, but may be superimposed at the position of object 22B. When objects are superimposed in the wrong position in markerless AR images, for example, the position of the object may change unnecessarily in the time direction, appearing to vibrate erratically, or exhibiting other unnatural behavior. In other words, the quality of markerless AR images may be reduced.
[0021] To suppress the occurrence of such phenomena, and to ensure that camera tracking and feature point detection accuracy is sufficiently high, it would be necessary to enhance hardware performance, such as by applying a processor with very high image processing capabilities or a very high-precision sensor, which could lead to increased costs.
[0022] Furthermore, Non-Patent Documents 1 and 2, for example, proposed methods for correcting and improving the accuracy of camera orientation on a three-dimensional environment map. For instance, Non-Patent Document 1 proposed a method using Support Vector Regression (SVR). Non-Patent Document 2 proposed a method for correction using prior measurement data and a Kalman filter. However, in these methods, position correction was performed in three-dimensional space. In other words, the calculation of position correction was performed using three-dimensional coordinates, which increased the amount of computation. Consequently, such position correction could increase the processing load. This could also lead to increased costs, such as the need for a more powerful processor.
[0023] <3. Correction of Superposition Position Using Feature Points of Captured Images> <Method 1> As shown in the top row of the table in Figure 4, the current frame of the captured image is acquired sequentially, digital content is superimposed on the current frame, and the superposition position of the current frame is corrected based on the superposition position of the past frame, which is expressed as the relative position to the feature points in the past frame (Method 1). In other words, when generating a markerless AR image by applying markerless AR technology, the superposition position of the digital content object (virtual object) relative to the current frame is corrected using the superposition position in the past frame (expressed as the relative position from the feature points). Such corrections are performed sequentially.
[0024] For example, an image processing device may include an image acquisition unit that sequentially acquires the current frame of an image captured from real space, an overlay unit that sequentially overlays digital content objects onto the acquired current frame, and a position correction unit that sequentially corrects the current frame overlay position, which is the overlay position of objects in the current frame, based on the past frame overlay position, which is the overlay position of objects in the past frame, identified based on the feature points of the past frames of the image.
[0025] Furthermore, the image processing method performed by the image processing device may include sequentially acquiring the current frame of an image captured in real space, sequentially superimposing digital content objects onto the acquired current frame, and sequentially correcting the current frame superimposition position, which is the superimposition position of the objects in the current frame, based on the past frame superimposition position, which is the superimposition position of the objects in the past frame, identified based on the feature points of the past frames of the image.
[0026] Furthermore, the program may cause the computer to perform a process that includes sequentially acquiring the current frame of an image captured from real space, sequentially superimposing digital content objects onto the acquired current frame, and sequentially correcting the current frame superimposition position, which is the superimposition position of the objects in the current frame, based on the past frame superimposition position, which is the superimposition position of the objects in the past frame, identified based on the feature points of the past frames of the captured image.
[0027] In the following explanation, the captured image on which digital content objects are superimposed will be described as a captured image of real space and as a moving image (composed of multiple frames).
[0028] Furthermore, the digital content superimposed on the captured image can be any virtual object that does not exist in the real space contained in the captured image. For example, it may be a so-called CG image generated by CG (Computer Graphics), an animated image, or a captured image of real space taken in a different space or at a different time. This virtual object may be 3D data or 2D data. Also, this virtual object may be static data that does not change in the direction of time (e.g., a still image) or dynamic data that changes in the direction of time (e.g., a moving image). In the following explanation, the location in real space where the virtual object is placed is assumed to be fixed (invariant) in the direction of time. That is, even as time progresses, the virtual object is assumed not to move in real space. However, the virtual object may move in real space. In that case, it is necessary to set the superimposed position of the virtual object while taking into account the movement of the virtual object.
[0029] Furthermore, the current frame refers to the frame being processed. Past frames, on the other hand, refer to frames that have already been processed (i.e., frames that were previously the target of processing). For example, this could be a frame processed immediately before the current frame, or a frame processed two or more frames before the current frame. It could also be a keyframe (a processed keyframe) set according to a predetermined interval or condition.
[0030] Figure 5 shows an example of a markerless AR image. In the past frame 110 shown at the top of Figure 5, a digital content object 111 is superimposed at the superimposition position K. In this specification, this superimposition position K is also referred to as the "past frame superimposition position". For a feature point A detected using an environment map or the like, its position in this past frame 110 is derived, and as shown by the double arrow 112, the relative position (Δx1, Δy1) of the superimposition position K with respect to the position of feature point A in that past frame 110 is derived and stored. This superimposition position K is assumed to be the correct superimposition position.
[0031] In contrast, in the current frame 120 shown at the bottom of Figure 5, the digital content object 111 is superimposed at the superimposition position L. In this specification, this superimposition position L is also referred to as the "current frame superimposition position". As shown in Figure 5, this superimposition position L is different from the superimposition position K, for example, due to low accuracy of camera tracking. In other words, the object 111 is superimposed at the wrong position.
[0032] Therefore, as described above, Method 1 is applied to correct the superposition position L of the current frame 120 based on the superposition position K of the past frame 110, which is the correct superposition position. For example, the superposition position L is corrected to the superposition position K as shown by the arrow 121 in the current frame 120.
[0033] By suppressing the shift in superimposed positions in this way, unnecessary temporal changes in the position of virtual objects can be suppressed, preventing them from appearing to behave unnaturally. In other words, the degradation of markerless AR image quality can be suppressed. Furthermore, since there is no need to increase the accuracy of camera tracking or the detection of feature points such as corners, there is no need to increase the performance of the hardware, thus suppressing the increase in costs.
[0034] Furthermore, the superimposed position K of the past frame 110 and the superimposed position L of the current frame 120 are expressed as relative positions (2D coordinates) from feature points on the markerless AR image, as described above. For example, the superimposed position K is expressed as a relative position (Δx1, Δy1) from feature point A, as described above. The superimposed position L is expressed as a relative position (Δx2, Δy2) from feature point A, as shown by the arrow 122 in the current frame 120. In this way, the correction of the superimposed position is performed using 2D coordinates without using 3D coordinates, so a higher-performance processor is not required, and the increase in cost can be suppressed. In other words, by applying Method 1, it is possible to suppress the reduction in the quality of markerless AR images while suppressing the increase in cost.
[0035] Furthermore, as mentioned above, since it is possible to suppress the reduction in quality of markerless AR images while suppressing the increase in costs, the increase in processing time required for correcting the superposition position can be suppressed. Therefore, this correction of the superposition position can be performed sequentially with respect to image acquisition (for example, in parallel with the imaging process).
[0036] Furthermore, the relative position from a feature point can also be described as the Euclidean distance from that feature point.
[0037] For example, the current frame superposition position (superposition position L) may be corrected to approach the past frame superposition position (superposition position K). For example, in an image processing device, the position correction unit may correct the current frame superposition position to approach the past frame superposition position. In the example of frame 120 in Figure 5, the superposition position L is corrected to move toward superposition position K (i.e., to approach superposition position K) as shown by arrow 121. By doing so, errors in the superposition position can be suppressed, and the reduction in the quality of markerless AR images can be suppressed.
[0038] Furthermore, the current frame superposition position (superposition position L) may be corrected using the difference (dx, dy) for each coordinate component between the past frame superposition position and the current frame superposition position. For example, in an image processing device, the position correction unit may correct the current frame superposition position using the difference between the past frame superposition position and the current frame superposition position. For example, in the case of the current frame 120 in Figure 5, the difference (dx, dy) between superposition position K and superposition position L is derived, and the superposition position L is corrected using this difference as shown by arrow 121. By applying this difference, the current frame superposition position (superposition position L) can be easily corrected to approach the past frame superposition position (superposition position K).
[0039] Alternatively, the difference (dx, dy) between the superimposed position K and the superimposed position L may be derived using a first relative position indicating the relative position of the current frame superimposed position with respect to the feature point, and a second relative position indicating the superimposed position of the past frame with respect to that feature point. For example, in an image processing device, the position correction unit may derive the difference (dx, dy) between the superimposed position of the past frame and the superimposed position of the current frame using a first relative position indicating the relative position of the current frame superimposed position with respect to the feature point, and a second relative position indicating the superimposed position of the past frame with respect to the feature point. For example, in the case of the current frame 120 in Figure 5, the difference (dx, dy) between the superimposed position K and the superimposed position L is derived using the relative positions from the feature point A described above ((Δx1, Δy1) and (Δx2, Δy2)). In this way, by using the relative position (2D coordinates) from the feature points, the difference (dx, dy) between the superimposed position K and the superimposed position L can be easily derived.
[0040] Furthermore, the captured image (moving image) overlaid with the digital content object may be generated by capturing images of the real world (the subject). For example, in an image processing device, the image acquisition unit may sequentially acquire the current frame by capturing images of the real world and generating moving images. In other words, this "acquisition" of the captured image includes the "generation" of the captured image through imaging. By doing so, the superposition position can be corrected in parallel with imaging (generation of captured images). Also, the captured image (moving image) overlaid with the digital content object may be supplied from another device. For example, in an image processing device, the image acquisition unit may sequentially acquire the current frame supplied from another device. By doing so, the superposition position can be corrected in parallel with the acquisition of captured images supplied from another device.
[0041] Also, the position where the digital content object is to be superimposed may be determined based on an environmental map generated from a captured image and the object may be superimposed. For example, in an image processing apparatus, a superimposing unit may generate an environmental map of the current frame, derive the superimposing position of the current frame based on the generated environmental map, and superimpose the digital content object at the derived superimposing position of the current frame.
[0042] Also, the superimposing position of the current frame may be derived using the coordinates (X, Y, Z) of the object in the camera coordinate system, the focal lengths (fx, fy), and the center position (cx, cy) of the captured image. For example, in an image processing apparatus, a superimposing unit may derive the superimposing position of the current frame using the coordinates of the object in the coordinate system of an imaging unit that captures the real space, the focal length of the captured image, and the center position of the captured image.
[0043] For example, the superimposing position (x, y) of the current frame may be derived as in the following equations (1) and (2).
[0044] x = (X × fx) / Z + Cx... (1) y = (Y × fy) / Z + Cy... (2)
[0045] Incidentally, the number of feature points (and relative positions) serving as the reference for the relative position for specifying the superimposition position may be any number. It may be one as in the example of FIG. 5, or it may be plural as in the example of FIG. 6. In the case of the past frame 130 shown in FIG. 6, the feature point A, the feature point B, and the feature point C are detected, and as indicated by the double arrows 131 to the double arrows 133, the relative position (Δx1a, Δy1a) based on the feature point A, the relative position (Δx1b, Δy1b) based on the feature point B, and the relative position (Δx1c, Δy1c) based on the feature point C are respectively derived. That is, the superimposition position K (past frame superimposition position) is expressed by these relative positions. Of course, the number of these feature points and relative positions may be four or more. This relative position may be derived using all of the feature points detected in each frame, or may be derived only for some of the feature points. The number of feature points detected in each frame is arbitrary and variable between frames. Therefore, the number of relative positions is also variable between frames.
[0046] Incidentally, the digital content to be superimposed on the captured image may be stored in advance in the image processing apparatus that corrects the superimposition position. For example, the image processing apparatus may further include a digital content storage unit that stores the digital content, and the superimposing unit may superimpose the object of the digital content read from the digital content storage unit on the current frame.
[0047] Further, the image processing apparatus that corrects the superimposition position may display the current frame (markerless AR image) on which the object is superimposed. For example, the image processing apparatus may further include an image display unit that displays the current frame (that is, the markerless AR image) in which the current frame superimposition position is corrected.
[0048] Further, the image processing apparatus that corrects the superimposition position may output the current frame (markerless AR image) on which the object is superimposed. For example, the image processing apparatus may further include an image output unit that outputs the current frame (that is, the markerless AR image) in which the current frame superimposition position is corrected.
[0049] Furthermore, an image processing device that corrects the superposition position may store the current frame (markerless AR image) with the object superimposed. For example, the image processing device may further include an image storage unit that stores the current frame (i.e., a markerless AR image) with the superposition position corrected.
[0050] <Method 1-1> Alternatively, as shown in the second row from the top of the table in Figure 4, the amount of correction for the superposition position may be adjusted using a correction coefficient k (Method 1-1). For example, the correction amount may be the difference (dx, dy) multiplied by the correction coefficient k. For example, in an image processing device, the position correction unit may correct the current frame superposition position using the difference between the past frame superposition position and the current frame superposition position multiplied by the correction coefficient.
[0051] For example, in the past frame 140 shown at the top of Figure 7, let's assume that the digital content object 141 is superimposed at superposition position K. In this case, let's assume that the relative positions from feature points A, B, and C are the values shown in table A of Figure 8. Also, in the current frame 150 shown at the bottom of Figure 7, let's assume that the object 141 is superimposed at superposition position L. In this case, let's assume that the relative positions from feature points A, B, and C are the values shown in table B of Figure 8.
[0052] In this case, if the superposition position L is corrected (shifted) by the difference (dx, dy), the superposition position L will coincide with the superposition position K. However, the correct superposition positions of objects are not necessarily the same in each frame. The (correct) superposition position of an object can change between frames depending on changes in, for example, the position of feature points, the camera position and angle, etc. Therefore, it is not always correct to match the superposition position in the current frame to the superposition position in past frames.
[0053] For example, in the case of Figure 7, the camera position and angle are different in past frame 140 and current frame 150. Therefore, the correct superposition position of object 141 in current frame 150 is different from that in past frame 140. In current frame 150, let's assume that superposition position K' is the correct superposition position of object 141. In this case, matching superposition position L, which is the superposition position in the current frame, to superposition position K is not necessarily the optimal correction (it does not reduce the deviation in superposition position the most).
[0054] Therefore, as described above, the correction amount of the current frame superposition position can be adjusted using a correction coefficient k. For example, the current frame superposition position is corrected by multiplying the difference (dx, dy) between the past frame superposition position and the current frame superposition position by the correction coefficient. For example, the current frame superposition position (x, y) may be corrected as shown in equations (3) and (4) below.
[0055] x = (X × fx) / Z + Cx + kdx ・・・(3) y = (Y × fy) / Z + Cy + kdy ・・・(4)
[0056] The variables kdx and kdy are the same as in equations (1) and (2). This allows the current frame superposition position to be corrected to a more favorable position, as shown by the arrow 151 in the current frame 150 of Figure 7.
[0057] The value of this correction coefficient k must satisfy 0 ≤ k ≤ 1. When k = 0, the current frame superposition position is not corrected (the correction is skipped). When k = 1, the current frame superposition position is corrected to the past frame superposition position.
[0058] The larger the value of the correction coefficient k, the more accurately the current frame's superimposed position can be corrected, for example, when camera tracking is significantly off (when the error in superimposed position is large). Conversely, the smaller the value of the correction coefficient k, the more accurately the current frame's superimposed position can be corrected to a less jarring position when the correct superimposed position changes between frames.
[0059] This correction coefficient may be set by the user or other party. In other words, this correction coefficient k may be input to the image processing device that performs the correction of the superposition position. For example, the image processing device may further include a correction coefficient input unit that receives the input of the correction coefficient. The position correction unit may then use the value obtained by multiplying the received correction coefficient by the difference to correct the superposition position of the current frame. In this way, it is possible to correct the superposition position according to the user's intentions.
[0060] Furthermore, this correction coefficient may be set by the device designer or manufacturer. In other words, this correction coefficient k may be stored in advance in the image processing device that performs the superposition position correction. For example, the image processing device may further include a correction coefficient storage unit that stores the correction coefficient. The position correction unit may then use the value obtained by multiplying the difference by the correction coefficient read from the correction coefficient storage unit to correct the current frame superposition position. In this way, it is possible to correct the superposition position according to the intentions of the designer or manufacturer.
[0061] Furthermore, the image processing device that performs the superposition position correction may set the correction coefficient k based on some information. Also, this correction coefficient k may be set independently for each coordinate axis (x coordinate, y coordinate). That is, the correction coefficient kx for the x coordinate and the correction coefficient ky for the y coordinate may be set independently of each other. In other words, the values of the correction coefficient kx and the correction coefficient ky may be different from each other.
[0062] <Method 1-2> If the correction of the superimposed position is performed too finely, the object may appear to vibrate unnaturally in the moving image. In other words, the quality of the markerless AR image may be reduced. Therefore, in order to allow a certain degree of error (to suppress the correction), a threshold TH may be set for the correction. That is, as shown in the bottom row of the table in Figure 4, the execution of the superimposed position correction may be controlled using a threshold TH (Method 1-2). For example, the superimposed position correction may be performed when the difference is greater than or equal to the threshold TH, and the correction may be skipped when it is less than the threshold TH. For example, in an image processing device, the position correction unit may correct the superimposed position of the current frame when the difference between the superimposed position of the past frame and the superimposed position of the current frame is greater than or equal to a threshold. By doing so, unnecessary corrections can be suppressed and the reduction in the quality of the markerless AR image can be suppressed.
[0063] The value of this threshold TH can be any value. Furthermore, this threshold TH may be set by the user or other means. In other words, this threshold TH may be input to the image processing device that performs the superposition position correction. For example, the image processing device may further include a threshold input unit that receives the threshold input. The position correction unit may then correct the current frame's superposition position if the difference in superposition positions is greater than or equal to the accepted threshold. In this way, superposition position correction according to the user's intentions can be achieved.
[0064] Furthermore, this threshold TH may be set by the device designer or manufacturer. In other words, this threshold TH may be stored in advance in the image processing device that performs the superposition position correction. For example, the image processing device may further include a threshold storage unit that stores the threshold. The position correction unit may then correct the current frame's superposition position if the difference in superposition position is greater than or equal to the threshold read from the threshold storage unit. In this way, it is possible to correct the superposition position according to the intentions of the designer or manufacturer.
[0065] For example, if the camera's movement speed or rotation is large, the variation in the superimposed position between frames tends to be large as well. In such cases, the shift in the superimposed position due to errors becomes less noticeable, and the importance of correcting the superimposed position decreases. Therefore, when such conditions are anticipated, setting the threshold TH to a larger value can suppress unnecessary corrections and prevent a decrease in the quality of markerless AR images.
[0066] Furthermore, when the camera's movement speed or rotation is small, the superimposed position shift due to errors becomes more noticeable, increasing the importance of correcting the superimposed position. Therefore, when such conditions are anticipated, setting the threshold TH to a smaller value can suppress unnecessary skipping of the superimposed position correction, thereby suppressing the degradation of markerless AR image quality.
[0067] Furthermore, the image processing device that corrects the superposition position may set a threshold TH based on some information. In that case, for example, the camera's movement may be detected, and the threshold TH may be set as described above according to the magnitude of the movement. By doing so, the degradation of markerless AR image quality can be suppressed in a wider variety of situations.
[0068] <4. First Embodiment> <Imaging Device> This technology can be applied to any device. For example, this technology can be applied to an imaging device. Figure 9 is a block diagram showing an example of the configuration of an imaging device, which is one embodiment of an image processing device to which this technology is applied. The imaging device 300 shown in Figure 9 is a device that captures images of real space, superimposes digital content objects onto the captured images using markerless AR technology, and generates markerless AR images. Therefore, the imaging device 300 can also be called a markerless AR processing device that performs markerless AR processing. Furthermore, the imaging device 300 can also be called a markerless AR image generation device that generates markerless AR images.
[0069] Figure 9 shows the main components such as the processing unit and data flow, but it does not necessarily represent an exhaustive list. In other words, the imaging device 300 may have processing units that are not shown as blocks in Figure 9, or processes and data flows that are not shown as arrows or other symbols in Figure 9.
[0070] As shown in Figure 9, the imaging device 300 includes an imaging processing unit 310 and an information processing unit 320. The imaging processing unit 310 performs processing related to imaging. The imaging processing unit 310 includes, for example, an imaging unit 311, a sensor unit 312, and a supply unit 313.
[0071] The imaging unit 311 has an imaging function and performs processing related to imaging. For example, the imaging unit 311 may image a subject, generate an image, and supply it to the supply unit 313 frame by frame. This image may be, for example, an RGB image (or an image generated by converting or compressing such an RGB image). This image may also be a moving image. The imaging unit 311 can also be called an image acquisition unit.
[0072] The sensor unit 312 includes an IMU, a depth sensor, etc., and performs processing related to sensing external information. For example, the sensor unit 312 may use its sensors to detect information regarding the position and orientation of the imaging unit 311, and supply the detected sensor information to the supply unit 313.
[0073] The supply unit 313 performs processing related to supplying the information obtained by the imaging processing unit 310 to the information processing unit 320. For example, the supply unit 313 may acquire the captured image supplied from the imaging unit 311 and the sensor information supplied from the sensor unit 312, and supply them to the information processing unit 320 (specifically, its acquisition unit 321).
[0074] The information processing unit 320 performs processing on the information obtained by the imaging processing unit 310. The information processing unit 320 includes, for example, an acquisition unit 321, a digital content storage unit 322, an overlay unit 323, a correction parameter storage unit 324, a correction parameter input unit 325, a position correction unit 326, a display unit 327, an output unit 328, and a storage unit 329.
[0075] The acquisition unit 321 performs processing related to the acquisition of information supplied from the imaging processing unit 310. For example, the acquisition unit 321 may acquire the captured image and sensor information supplied from the supply unit 313 and supply them to the superimposition unit 323.
[0076] The digital content storage unit 322 has a storage medium and performs processing related to the storage of digital content. For example, the digital content storage unit 322 may store digital content in its storage medium. Alternatively, the digital content storage unit 322 may read the digital content stored in its storage medium and supply it to the superimposition unit 323.
[0077] The superposition unit 323 performs processing related to the superposition of digital content onto the captured image. For example, the superposition unit 323 may acquire the captured image and sensor information supplied from the acquisition unit 321. Alternatively, the superposition unit 323 may acquire digital content read from the digital content storage unit 322. Furthermore, the superposition unit 323 may use markerless AR technology to superimpose digital content onto the captured image and generate a markerless AR image. The superposition unit 323 may also supply the generated markerless AR image, etc., to the position correction unit 326.
[0078] The correction parameter storage unit 324 has a storage medium and performs processing related to the storage of correction parameters. The correction parameters are information used for correcting the superimposed position and may include the correction coefficient k and threshold TH mentioned above. The correction parameter storage unit 324 may store the correction parameters in its storage medium. Alternatively, the correction parameter storage unit 324 may read the correction parameters stored in its storage medium and supply them to the position correction unit 326. The correction parameter storage unit 324 can also be called a correction coefficient storage unit that stores the correction coefficient k. The correction parameter storage unit 324 can also be called a threshold value storage unit that stores the threshold TH.
[0079] The correction parameter input unit 325 has an input device and performs processing related to the input of correction parameters. The correction parameter input unit 325 may receive correction parameters input via its input device and supply them to the position correction unit 326. The correction parameter input unit 325 can also be called a correction coefficient input unit that inputs a correction coefficient k. The correction parameter input unit 325 can also be called a threshold input unit that inputs a threshold value TH.
[0080] The position correction unit 326 performs processing related to the correction of the superimposed position. For example, the position correction unit 326 may acquire a markerless AR image supplied from the superimposition unit 323. The position correction unit 326 may also acquire correction parameters supplied from the correction parameter storage unit 324. The position correction unit 326 may also acquire correction parameters supplied from the correction parameter input unit 325. The position correction unit 326 may apply this technology to the acquired markerless AR image to correct the superimposed position. In this case, the position correction unit 326 may perform the correction using the acquired correction parameters (correction coefficient k, threshold TH, etc.). The position correction unit 326 may supply the markerless AR image (current frame) with the corrected superimposed position to the display unit 327, the output unit 328, and the storage unit 329.
[0081] The display unit 327 has a display device and performs processing related to the display of markerless AR images. For example, the display unit 327 may acquire a markerless AR image (the current frame in which the superposition position of digital content objects has been appropriately corrected) supplied from the position correction unit 326 and display it on its display device. The display unit 327 can also be called an image display unit.
[0082] The output unit 328 has output devices such as output terminals and communication devices, and performs processing related to the output of markerless AR images. For example, the output unit 328 may acquire a markerless AR image (the current frame in which the superposition position of digital content objects has been appropriately corrected) supplied from the position correction unit 326 and output it to another device via its output device. The output unit 328 can also be called an image output unit.
[0083] The storage unit 329 has a storage medium and performs processing related to the storage of markerless AR images. For example, the storage unit 329 may acquire a markerless AR image (the current frame in which the superimposed position of objects in the digital content has been appropriately corrected) supplied from the position correction unit 326 and store it in its storage medium. The storage unit 329 can also be called an image storage unit.
[0084] <Flow of Markerless AR Processing> An example of the flow of markerless AR processing performed by the imaging device 300 with this configuration will be explained with reference to the flowchart in Figure 10.
[0085] When markerless AR processing is started, the position correction unit 326 sets correction parameters (such as the correction coefficient k and threshold TH) in step S301. These correction parameters may be read from the correction parameter storage unit 324 or received by the correction parameter input unit 325.
[0086] In step S302, the imaging unit 311 captures an image of the subject and generates an image of the current frame. The sensor unit 312 also detects external information using the sensor and generates sensor information.
[0087] In step S303, the superimposition unit 323 acquires this information, generates an environment map from the captured image (current frame) using markerless AR technology, and superimposes digital content objects onto the captured image (current frame) based on the environment map. These digital content objects may be read from the digital content storage unit 322.
[0088] In step S304, the superimposing unit 323 detects a feature point in its current frame and derives the relative position of the superimposing position in the current frame based on that feature point.
[0089] In step S305, the superimposing unit 323 stores the derived relative position as the superimposing position (relative position) of the past frame.
[0090] For example, as shown in Figure 11A, when object 411 is superimposed on the superimposition position K of frame 410, feature points A, B, and C are detected as shown in Figure 11B. Then, the relative positions of the superimposition position K from each feature point are derived from the double-headed arrows 412 and 414 shown in Figure 11C.
[0091] In this way, the relative positions of the superimposed positions of past frames (n frames) are stored, as shown in the table in Figure 12A. Similarly, the relative positions of the current frame (n+1 frames) are derived and stored as shown in the table in Figure 12B.
[0092] In step S306, the position correction unit 326 determines whether the current frame is the first frame. If it is determined that it is not the first frame (i.e., a past frame exists), the process proceeds to step S307.
[0093] In step S307, the position correction unit 326 determines whether the difference between the current frame superimposed position and the past frame superimposed position is greater than or equal to the threshold TH. This threshold TH may be supplied from the correction parameter storage unit 324 or the correction parameter input unit 325. If it is determined that the difference between the current frame superimposed position and the past frame superimposed position is greater than or equal to the threshold TH, the process proceeds to step S308.
[0094] In step S308, the position correction unit 326 applies the above-described technology and corrects the current frame superposition position using the past frame superposition position. At this time, the position correction unit 326 may use a correction coefficient k. This correction coefficient k may be supplied from the correction parameter storage unit 324 or the correction parameter input unit 325.
[0095] As a result of this correction, the superposition position L of object 411 is corrected based on the superposition position K, as shown in Figure 12C.
[0096] When the processing in step S308 is completed, the process proceeds to step S309. Also, if it is determined in step S306 that the current frame is the first frame (no past frames exist), the process proceeds to step S309. Also, if it is determined in step S307 that the difference between the superimposed position of the current frame and the superimposed position of the past frame is less than the threshold TH, the process proceeds to step S309.
[0097] In step S309, the display unit 327 displays the current frame (markerless AR image).
[0098] In step S310, the output unit 328 outputs the current frame (markerless AR image).
[0099] In step S311, the memory unit 329 stores the current frame (markerless AR image).
[0100] In step S312, the imaging unit 311 determines whether or not to terminate imaging. If it is determined not to terminate imaging, the process returns to step S302. If it is determined in step S312 to terminate imaging, the markerless AR process is terminated.
[0101] By performing each process as described above, the imaging device 300 can correct the superposition position of digital content objects superimposed on the captured image using markerless AR technology by applying the aforementioned technology. Therefore, the imaging device 300 can suppress the reduction in quality of markerless AR images while suppressing an increase in costs.
[0102] <Use Cases for Controlling Correction Execution> Use cases for such corrections will be explained using Figures 13 and 14. In the example in Figure 13, assume that object 421 is superimposed at superposition position K in frame 420 (n frames) shown in Figure 13A. In contrast, assume that object 421 is superimposed at superposition position L in frame 430 (n+1 frames) shown in Figure 13B. Since superposition position L is close to superposition position K and the difference is smaller than the threshold TH, the correction is skipped. Assume that object 421 is superimposed at superposition position M in frame 440 (n+2 frames) shown in Figure 13C. Since superposition position M is far from superposition position K and the difference is greater than or equal to the threshold TH, the correction is performed as shown by arrow 441.
[0103] In the example shown in Figure 14, let's assume that in frame 450 (n frames) shown in Figure 14A, object 451 is superimposed at superposition position K. In contrast, let's assume that in frame 460 (n+1 frames) shown in Figure 14B, object 451 is superimposed at superposition position L. Since superposition position L is close to superposition position K, and the difference is smaller than the threshold TH, the correction is skipped. Similarly, let's assume that in frame 470 (n+2 frames) shown in Figure 14C, object 451 is superimposed at superposition position M. Since superposition position M is close to superposition position K, and the difference is smaller than the threshold TH, the correction is skipped.
[0104] <5. Second Embodiment> <Image Processing System> This technology can also be applied to a system composed of multiple devices. For example, the configuration of the imaging device 300 described above may be divided into multiple devices. Figure 15 is a block diagram showing an example of the main configuration of an image processing system to which this technology is applied. The image processing system 500 shown in Figure 15 is a system that, similar to the imaging device 300 described above, captures images of real space, superimposes digital content objects onto the captured images using markerless AR technology, and generates markerless AR images. Therefore, the image processing system 500 can also be called a markerless AR processing system that performs markerless AR processing. Furthermore, the image processing system 500 can also be called a markerless AR image generation system that generates markerless AR images.
[0105] Figure 15 shows the main components such as processing units and data flows, but it does not necessarily represent an exhaustive list of all components. In other words, each device constituting the image processing system 500 may have processing units that are not shown as blocks in Figure 15, or processing and data flows that are not shown as arrows or other symbols in Figure 15.
[0106] As shown in Figure 15, the image processing system 500 includes an imaging device 510 and an information processing device 520. The imaging device 510 and the information processing device 520 are connected to each other via a network 550 so that they can communicate with one another.
[0107] This network 550 is a communication network composed of any communication medium. Communication conducted through network 550 may be wired communication, wireless communication, or both. In other words, network 550 may be a communication network for wired communication, a communication network for wireless communication, or a communication network composed of both. Furthermore, network 550 may be composed of a single communication network or of multiple communication networks.
[0108] For example, the internet may be included in this network 550. Public telephone network may also be included in this network 550. Furthermore, wide-area communication networks for wireless mobile devices, such as so-called 3G and 4G networks, may also be included in this network 550. For example, LPWA (Low Power Wide Area) communication networks such as LTE-M, which enable long-distance data communication and have low power consumption, may also be included in this network 550. Furthermore, WAN (Wide Area Network) and LAN (Local Area Network) may also be included in this network 550. Furthermore, wireless communication networks that perform communication compliant with the Bluetooth (registered trademark) standard may also be included in this network 550. Furthermore, communication channels for short-range wireless communication such as NFC (Near Field Communication) may also be included in this network 550. Furthermore, communication channels for infrared communication may also be included in this network 550. Furthermore, wired communication networks compliant with standards such as HDMI (High-Definition Multimedia Interface) (registered trademark) and USB (Universal Serial Bus) (registered trademark) may also be included in this network 550. Thus, the network 550 may include communication networks and communication channels of any communication standard. Furthermore, these communication networks and communication channels may include not only communication media such as cables, but also devices and circuits necessary for communication, such as communication equipment and relay equipment.
[0109] The imaging device 510 is a device that performs imaging of real space. The imaging device 510 has basically the same configuration as the imaging processing unit 310 in Figure 9 and performs the same processing. That is, the imaging unit 511 is the same processing unit as the imaging unit 311 and performs the same processing. The sensor unit 512 is the same processing unit as the sensor unit 312 and performs the same processing. The communication unit 513 is basically the same processing unit as the supply unit 313 and performs the same processing. However, the communication unit 513 can communicate with the information processing device 520 (specifically its communication unit 521) via the network 550. Through this communication, the communication unit 513 supplies the captured image and sensor information to the information processing device 520 (specifically its communication unit 521).
[0110] The information processing device 520 is a device that performs processing on the information obtained by the imaging device 510. The information processing device 520 has basically the same configuration as the information processing device 320 in Figure 9 and performs the same processing. In other words, the communication unit 521 is basically the same processing unit as the acquisition unit 321 and performs the same processing. However, the communication unit 521 can communicate with the imaging device 510 (its communication unit 513) via the network 550. Through this communication, the communication unit 521 acquires the captured image and sensor information supplied from the imaging device 510 (its communication unit 513) and supplies it to the superimposition unit 523. The communication unit 521 can also be called an image acquisition unit that acquires the captured image.
[0111] The digital content storage unit 522 is a processing unit similar to the digital content storage unit 322 and performs the same processing. The superimposition unit 523 is a processing unit similar to the superimposition unit 323 and performs the same processing.
[0112] The correction parameter storage unit 524 is a processing unit similar to the correction parameter storage unit 324 and performs the same processing. The correction parameter storage unit 524 can also be called a correction coefficient storage unit that stores the correction coefficient k. The correction parameter storage unit 524 can also be called a threshold value storage unit that stores the threshold value TH. The correction parameter input unit 525 is a processing unit similar to the correction parameter input unit 325 and performs the same processing. The correction parameter input unit 525 can also be called a correction coefficient input unit that inputs the correction coefficient k. The correction parameter input unit 525 can also be called a threshold value input unit that inputs the threshold value TH. The position correction unit 526 is a processing unit similar to the position correction unit 326 and performs the same processing.
[0113] The display unit 527 is a processing unit similar to the display unit 327 and performs the same processing. The display unit 527 can also be called an image display unit. The output unit 528 is a processing unit similar to the output unit 328 and performs the same processing. The output unit 528 can also be called an image output unit. The storage unit 529 is a processing unit similar to the storage unit 329 and performs the same processing. The storage unit 529 can also be called an image storage unit.
[0114] <Flow of Markerless AR Processing> An example of the flow of markerless AR processing performed by the image processing system 500 with the above configuration will be explained with reference to the flowcharts in Figures 16 and 17.
[0115] When markerless AR processing is started, the position correction unit 526 of the information processing device 520 sets correction parameters (such as the correction coefficient k and threshold TH) in step S521.
[0116] In step S511, the imaging unit 511 of the imaging device 510 captures an image of the subject and generates an image of the current frame. The sensor unit 512 also detects external information using the sensor and generates sensor information.
[0117] In step S512, the communication unit 513 of the imaging device 510 supplies that information to the information processing device 520. In step S522, the communication unit 521 of the information processing device 520 acquires that information.
[0118] The processes in steps S523 to S525 in Figure 16 and steps S531 to S537 in Figure 17 are executed in the same manner as in Figure 10.
[0119] If it is determined in step S537 that imaging should not be terminated, the process returns to step S511 in Figure 16.
[0120] By performing each process as described above, the image processing system 500 (information processing device 520) can correct the superposition position of digital content objects superimposed on the captured image using markerless AR technology by applying the aforementioned technology. Therefore, the image processing system 500 (information processing device 520) can suppress the reduction in quality of markerless AR images while suppressing an increase in costs.
[0121] <6. Addendum> <Computer> The series of processes described above can be executed by hardware or by software. When the series of processes are executed by software, the programs that make up the software are installed on a computer. Here, a computer includes computers built into dedicated hardware, as well as general-purpose personal computers, for example, that can perform various functions by installing various programs.
[0122] Figure 18 is a block diagram showing an example of the hardware configuration of a computer that executes the series of processes described above using a program.
[0123] In the computer 900 shown in Figure 18, the CPU (Central Processing Unit) 901, ROM (Read Only Memory) 902, and RAM (Random Access Memory) 903 are interconnected via a bus 904.
[0124] An input / output interface 910 is also connected to the bus 904. An input / output interface 910 is connected to an input unit 911, an output unit 912, a storage unit 913, a communication unit 914, and a drive 915.
[0125] The input unit 911 consists of, for example, a keyboard, mouse, microphone, touch panel, and input terminals. The output unit 912 consists of, for example, a display, speaker, and output terminals. The storage unit 913 consists of, for example, a hard disk, RAM disk, and non-volatile memory. The communication unit 914 consists of, for example, a network interface. The drive 915 drives removable media 921 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory.
[0126] In a computer configured as described above, the CPU 901 loads, for example, a program stored in the memory unit 913 into the RAM 903 via the input / output interface 910 and the bus 904, and executes it, thereby performing the series of processes described above. The RAM 903 also appropriately stores data necessary for the CPU 901 to perform various processes.
[0127] The program executed by the computer can be recorded and applied, for example, on removable media 921 such as a package medium. In this case, the program can be installed in the storage unit 913 via the input / output interface 910 by inserting the removable media 921 into the drive 915.
[0128] Furthermore, this program can also be provided via wired or wireless transmission media such as a local area network, the internet, or digital satellite broadcasting. In that case, the program can be received by the communication unit 914 and installed in the storage unit 913.
[0129] In addition, this program can be pre-installed in ROM 902 or memory unit 913.
[0130] <Applications of this technology> This technology can be applied to any configuration. For example, this technology can be applied to various electronic devices.
[0131] Furthermore, this technology can also be implemented as part of a device, such as a processor as a system LSI (Large Scale Integration) (e.g., a video processor), a module using multiple processors (e.g., a video module), a unit using multiple modules (e.g., a video unit), or a set with additional functions added to a unit (e.g., a video set).
[0132] Furthermore, this technology can also be applied to network systems composed of multiple devices. For example, this technology may be implemented as cloud computing, where multiple devices share and collaborate on processing via a network. For example, this technology may be implemented in a cloud service that provides image (video) related services to any terminal such as computers, AV (Audio Visual) equipment, portable information processing terminals, and IoT (Internet of Things) devices.
[0133] In this specification, a system refers to a collection of multiple components (devices, modules (parts), etc.), regardless of whether all components are located in the same enclosure. Therefore, multiple devices housed in separate enclosures and connected via a network, and a single device containing multiple modules within a single enclosure, are both considered systems.
[0134] <Applicable Fields and Applications of This Technology> Systems, devices, and processing units incorporating this technology can be used in any field, such as transportation, medical care, security, agriculture, livestock farming, mining, beauty, factories, home appliances, weather, and nature monitoring. Furthermore, the applications are entirely arbitrary.
[0135] <Other> In this specification, "flag" refers to information used to identify multiple states, and includes not only information used to identify two states, true (1) or false (0), but also information capable of identifying three or more states. Therefore, the values that this "flag" can take are, for example, two values, 1 / 0, or three or more values. In other words, the number of bits that constitute this "flag" is arbitrary, and can be 1 bit or multiple bits. Furthermore, identification information (including flags) is envisioned not only in the form of including the identification information itself in the bitstream, but also in the form of including difference information of the identification information relative to a certain reference information in the bitstream. Therefore, in this specification, "flag" and "identification information" include not only the information itself, but also difference information relative to the reference information.
[0136] Furthermore, various types of information (metadata, etc.) related to encoded data (bitstream) may be transmitted or recorded in any form as long as they are associated with the encoded data. Here, the term "associate" means, for example, making it possible to use (link) one data when processing the other. In other words, associated data may be combined into a single data, or they may be individual data. For example, information associated with encoded data (image) may be transmitted on a different transmission path than the encoded data (image). Also, for example, information associated with encoded data (image) may be recorded on a different recording medium (or a different recording area on the same recording medium) than the encoded data (image). Note that this "association" may not apply to the entire data, but only to a part of it. For example, an image and the information corresponding to that image may be associated with each other in any unit, such as multiple frames, one frame, or a part within a frame.
[0137] In this specification, terms such as "combine," "multiplex," "add," "integrate," "include," "store," "insert," "insert," and "place" mean combining multiple things into one, such as combining encoded data and metadata into a single data, and represent one method of "associating" as described above.
[0138] Furthermore, the embodiments of this technology are not limited to those described above, and various modifications are possible without departing from the gist of this technology.
[0139] For example, the configuration described as a single device (or processing unit) may be divided and configured as multiple devices (or processing units). Conversely, the configurations described above as multiple devices (or processing units) may be combined and configured as a single device (or processing unit). Furthermore, it is also possible to add configurations other than those described above to the configuration of each device (or each processing unit). In addition, if the overall system configuration and operation are substantially the same, a part of the configuration of one device (or processing unit) may be included in the configuration of another device (or other processing unit).
[0140] Furthermore, for example, the program described above may be executed on any device. In that case, the device should have the necessary functions (such as functional blocks) and be able to obtain the necessary information.
[0141] Furthermore, for example, each step of a flowchart may be executed by one device, or it may be divided among multiple devices. Additionally, if a single step includes multiple processes, these processes may be executed by one device, or they may be divided among multiple devices. In other words, multiple processes included in a single step can be executed as multiple steps. Conversely, processes described as multiple steps can be combined and executed as a single step.
[0142] Furthermore, for example, a program executed by a computer may be structured so that the steps of the program are executed chronologically in the order described herein, or they may be executed in parallel or individually at necessary times, such as when a call is made. In other words, the steps may be executed in an order different from the order described above, as long as no inconsistencies arise. Moreover, the steps of this program may be executed in parallel with the processing of other programs, or in combination with the processing of other programs.
[0143] Furthermore, for example, multiple technologies relating to this technology can be implemented independently, as long as they do not create a contradiction. Of course, any multiple technologies can also be implemented in combination. For example, some or all of the technologies described in one embodiment can be implemented in combination with some or all of the technologies described in another embodiment. Also, some or all of the above-mentioned technologies can be implemented in combination with other technologies not mentioned above.
[0144] Furthermore, this technology can also be configured as follows: (1) An image processing device comprising: an image acquisition unit that sequentially acquires the current frame of an image captured from real space; an overlay unit that sequentially overlays objects of digital content onto the acquired current frame; and a position correction unit that sequentially corrects the current frame overlay position, which is the overlay position of the object in the current frame, based on the past frame overlay position, which is the overlay position of the object in the past frame, which is identified based on the feature points of the past frame of the image. (2) The image processing device according to (1), wherein the position correction unit corrects the current frame overlay position so that it approaches the past frame overlay position. (3) The image processing device according to (2), wherein the position correction unit corrects the current frame overlay position using the difference between the past frame overlay position and the current frame overlay position. (4) The image processing apparatus according to (3), wherein the position correction unit derives the difference using a first relative position indicating the relative position of the current frame superimposed position with respect to the feature point and a second relative position indicating the past frame superimposed position with respect to the feature point. (5) The image processing apparatus according to (3) or (4), wherein the position correction unit corrects the current frame superimposed position using a value obtained by multiplying the difference by a correction coefficient. (6) The image processing apparatus according to (5), further comprising a correction coefficient input unit that receives input of the correction coefficient, wherein the position correction unit corrects the current frame superimposed position using a value obtained by multiplying the difference by the received correction coefficient. (7) The image processing apparatus according to (5) or (6), further comprising a correction coefficient storage unit that stores the correction coefficient, wherein the position correction unit corrects the current frame superimposed position using a value obtained by multiplying the difference by the correction coefficient read from the correction coefficient storage unit. (8) The image processing apparatus according to any one of (3) to (7), wherein the position correction unit corrects the current frame superimposed position when the difference is greater than or equal to a threshold. (9) The image processing apparatus according to (8), further comprising a threshold input unit that receives input of the threshold, wherein the position correction unit corrects the current frame superimposed position when the difference is greater than or equal to the received threshold.(10) The image processing apparatus according to (8) or (9), further comprising a threshold storage unit for storing the threshold, wherein the position correction unit corrects the current frame superposition position when the difference is greater than or equal to the threshold read from the threshold storage unit. (11) The image processing apparatus according to any one of (1) to (10), wherein the image acquisition unit sequentially acquires the current frame by capturing the real space and generating the captured image of a moving image. (12) The image processing apparatus according to any one of (1) to (11), wherein the image acquisition unit sequentially acquires the current frame supplied from another device. (13) The image processing apparatus according to any one of (1) to (12), wherein the superposition unit generates an environment map of the current frame, derives the current frame superposition position based on the generated environment map, and superimposes the object on the derived current frame superposition position. (14) The image processing apparatus according to (13), wherein the superposition unit derives the current frame superposition position using the coordinates of the object in the coordinate system of the imaging unit that images the real space, the focal length of the captured image, and the center position of the captured image. (15) The image processing apparatus according to any one of (1) to (14), further comprising a digital content storage unit for storing the digital content, wherein the superposition unit superimposes the object of the digital content read from the digital content storage unit onto the current frame. (16) The image processing apparatus according to any one of (1) to (15), further comprising an image display unit for displaying the current frame on which the current frame superposition position has been corrected. (17) The image processing apparatus according to any one of (1) to (16), further comprising an image output unit for outputting the current frame on which the current frame superposition position has been corrected. (18) The image processing apparatus according to any one of (1) to (17), further comprising an image storage unit for storing the current frame on which the current frame superposition position has been corrected.(19) An image processing method comprising: sequentially acquiring the current frame of an image captured in real space; sequentially superimposing objects of digital content onto the acquired current frame; and sequentially correcting the current frame superimposition position, which is the superimposition position of the object in the current frame, based on the past frame superimposition position, which is the superimposition position of the object in the past frame, identified based on the feature points of past frames of the image. (20) A program for causing a computer to perform a process comprising: sequentially acquiring the current frame of an image captured in real space; sequentially superimposing objects of digital content onto the acquired current frame; and sequentially correcting the current frame superimposition position, which is the superimposition position of the object in the current frame, based on the past frame superimposition position, which is the superimposition position of the object in the past frame, identified based on the feature points of past frames of the image.
[0145] 300 Imaging device, 310 Imaging processing unit, 311 Imaging unit, 312 Sensor unit, 313 Supply unit, 320 Information processing unit, 321 Acquisition unit, 322 Digital content storage unit, 323 Overlay unit, 324 Correction parameter storage unit, 325 Correction parameter input unit, 326 Position correction unit, 327 Display unit, 328 Output unit, 329 Storage unit, 500 Image processing system, 510 Imaging device, 511 Imaging unit, 512 Sensor unit, 513 Communication unit, 520 Information processing unit, 521 Communication unit, 522 Digital content storage unit, 523 Overlay unit, 524 Correction parameter storage unit, 525 Correction parameter input unit, 526 Position correction unit, 527 Display unit, 528 Output unit, 529 Storage unit, 550 Network, 900 Computer
Claims
1. An image processing apparatus comprising: an image acquisition unit that sequentially acquires the current frame of an image captured from real space; an overlay unit that sequentially overlays objects of digital content onto the acquired current frame; and a position correction unit that sequentially corrects the current frame overlay position, which is the overlay position of the object in the current frame, based on the past frame overlay position, which is the overlay position of the object in the past frame, identified based on the feature points of the past frame of the image.
2. The image processing apparatus according to claim 1, wherein the position correction unit corrects the current frame superposition position so that it approaches the past frame superposition position.
3. The image processing apparatus according to claim 2, wherein the position correction unit corrects the current frame superimposition position using the difference between the past frame superimposition position and the current frame superimposition position.
4. The image processing apparatus according to claim 3, wherein the position correction unit derives the difference using a first relative position indicating the relative position of the current frame superimposed position with respect to the feature point and a second relative position indicating the past frame superimposed position with respect to the feature point.
5. The image processing apparatus according to claim 3, wherein the position correction unit corrects the current frame superimposition position using a value obtained by multiplying the difference by a correction coefficient.
6. The image processing apparatus according to claim 5, further comprising a correction coefficient input unit that receives input of the correction coefficient, wherein the position correction unit corrects the current frame superimposition position using a value obtained by multiplying the received correction coefficient by the difference.
7. The image processing apparatus according to claim 5, further comprising a correction coefficient storage unit for storing the correction coefficient, wherein the position correction unit corrects the current frame superimposition position using a value obtained by multiplying the difference by the correction coefficient read from the correction coefficient storage unit.
8. The image processing apparatus according to claim 3, wherein the position correction unit corrects the current frame superimposition position when the difference is greater than or equal to a threshold.
9. The image processing apparatus according to claim 8, further comprising a threshold input unit for receiving the input of the threshold, wherein the position correction unit corrects the current frame superimposed position when the difference is greater than or equal to the received threshold.
10. The image processing apparatus according to claim 8, further comprising a threshold storage unit for storing the threshold, wherein the position correction unit corrects the current frame superimposed position when the difference is greater than or equal to the threshold read from the threshold storage unit.
11. The image processing apparatus according to claim 1, wherein the image acquisition unit acquires the current frame sequentially by capturing images of the real space and generating the captured images of a moving image.
12. The image processing apparatus according to claim 1, wherein the image acquisition unit sequentially acquires the current frame supplied from another device.
13. The image processing apparatus according to claim 1, wherein the superposition unit generates an environment map of the current frame, derives the current frame superposition position based on the generated environment map, and superimposes the object on the derived current frame superposition position.
14. The image processing apparatus according to claim 13, wherein the superimposing section derives the current frame superimposing position using the coordinates of the object in the coordinate system of the imaging unit that images the real space, the focal length of the captured image, and the center position of the captured image.
15. The image processing apparatus according to claim 1, further comprising a digital content storage unit for storing the digital content, wherein the superimposing unit superimposes the objects of the digital content read from the digital content storage unit onto the current frame.
16. The image processing apparatus according to claim 1, further comprising an image display unit that displays the current frame whose current frame superimposition position has been corrected.
17. The image processing apparatus according to claim 1, further comprising an image output unit that outputs the current frame whose current frame superposition position has been corrected.
18. The image processing apparatus according to claim 1, further comprising an image storage unit that stores the current frame whose current frame superimposition position has been corrected.
19. An image processing method comprising: sequentially acquiring the current frame of an image captured from real space; sequentially superimposing objects of digital content onto the acquired current frame; and sequentially correcting the current frame superimposition position, which is the superimposition position of the object in the current frame, based on the past frame superimposition position, which is the superimposition position of the object in the past frame, identified based on the feature points of past frames of the image.
20. A program for causing a computer to perform a process that includes: sequentially acquiring the current frame of an image captured from real space; sequentially superimposing objects of digital content onto the acquired current frame; and sequentially correcting the current frame superimposition position, which is the superimposition position of the object in the current frame, based on the past frame superimposition position, which is the superimposition position of the object in the past frame, identified based on the feature points of past frames of the image.