Homography image conversion system
The homography image conversion system uses machine learning to automatically generate transformation matrices, addressing the complexity of manual coordinate acquisition in conventional systems by enabling efficient and precise image conversion from one perspective to another.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- MEIJO UNIVERSITY
- Filing Date
- 2024-11-08
- Publication Date
- 2026-05-20
AI Technical Summary
Conventional image conversion systems require manual coordinate acquisition, making the process complicated.
A homography image conversion system utilizing machine learning to automatically generate a transformation matrix and remove distortion from input images, enabling automatic image conversion without manual coordinate determination.
Enables accurate conversion of images from one perspective to another without manual coordinate input, allowing for efficient and precise image reconstruction.
Smart Images

Figure 2026083792000001_ABST
Abstract
Description
Technical Field
[0006]
[0001] The disclosed technology relates to a homography image conversion system.
Background Art
[0002] Conventionally, converting a planar image into another planar image has been performed. As a result, based on an image of an object viewed from a certain line of sight, an image of the object viewed from another line of sight can be obtained. As a prior art related to such image conversion, the one described in Patent Document 1 can be cited.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the conventional technology, coordinates had to be manually obtained during conversion. Therefore, the execution of image conversion was complicated. For example, in the technology of Patent Document 1, it is stated that information such as "facility position coordinates" is stored in the "storage unit 20". Coordinates for storing in this "storage unit 20" had to be manually obtained in advance.
[0005] An object of the disclosed technology is to provide a homography image conversion system that can obtain a converted image based on an input image without the need to manually obtain coordinates.
Means for Solving the Problems
[0006] A homography image conversion system in one aspect of the disclosed technology involves adding distortion to a reference image. An encoder that receives an input image, which is a homography image of the input image, and an encoder that processes the input image into a homography image. A system having a decoder that outputs a converted output image, wherein the encoder is Machine learning based on input images generates a transformation matrix, which is an indicator of the distortion of the input image from a reference image. The extraction of the features and the distribution of the position vectors corresponding to the features of the input image were normalized to remove distortion. This process obtains features from a reference image and represents positional information within the region of the output image. A coordinate sample output unit that outputs a coordinate sample which is an information matrix, and the coordinate sample, and mechanical engineering The decoder further includes a coupling calculation unit that calculates the coupling with the position vector after learning, and the decoder is coupling This process receives input and constructs the output image.
[0007] In the homography image conversion system described above, the encoder converts the input image to a Deep learning is used to extract the transformation matrix, which is an indicator of distortion. Furthermore, it is normalized. The position vector is obtained. The coordinate sample output from the coordinate sample output unit is and the normalized distribution. The coupling calculation unit calculates the coupling with the given position vector. Based on this coupling, the decoder outputs The output image is a reference image reconstructed from the input image. In other words, the output image is a reference image reconstructed from the input image. This is an image from which the distortion defined by the transformation matrix has been removed.
[0008] In the homography image conversion system described above, the decoder uses a conversion matrix It is preferable that the output image is constructed based solely on stitching, without distortion. The characteristic can be eliminated.
[0009] In the homography image conversion system according to any of the above aspects, the input image is an image of the target object viewed from a certain perspective, and the output image can be a planar image of the target object viewed from a different perspective from that of the input image. The input image is an image of the target object viewed from a certain perspective, and the output image can be a planar image of the target object viewed from a different perspective from that of the input image. It can be a planar image of the target object viewed from a different perspective from that of the input image.
[0010] In the homography image conversion system according to any of the above aspects, the encoder can be configured to receive inputs of a plurality of input images. In that case, the plurality of input images are images of a common target object viewed from different perspectives, and the plurality of output images can all be planar images of the target object viewed from the same perspective different from any of the different perspectives. The encoder can be configured to receive inputs of a plurality of input images. In that case, the plurality of input images are images of a common target object viewed from different perspectives, and the plurality of output images can all be planar images of the target object viewed from the same perspective different from any of the different perspectives. The plurality of input images are images of a common target object viewed from different perspectives, and the plurality of output images can all be planar images of the target object viewed from the same perspective different from any of the different perspectives. It can be a planar image of the target object viewed from the same perspective different from any of the different perspectives. The plurality of planar images thus obtained are easy to integrate. For example, an image obtained by imaging a parking lot can be used as the input image. For example, an image obtained by imaging a parking lot can be used as the input image.
Advantages of the Invention
[0011] According to the disclosed technology, there is provided a homography image conversion system capable of obtaining a converted image based on an input image without the need to manually obtain coordinates. There is provided a homography image conversion system capable of obtaining a converted image based on an input image without the need to manually obtain coordinates.
Brief Description of the Drawings
[0012] [Figure 1] It is a block diagram showing the configuration of the homography image conversion system according to the embodiment. [Figure 2] It is a schematic diagram showing the encoding of an input image by an encoder. <000009th>It is a schematic diagram showing the conversion of position information by a conversion matrix. [Figure 4] It is a schematic diagram (Part 1) showing the reconstruction of an image by a decoder. [Figure 5] It is a schematic diagram (Part 2) showing the reconstruction of an image by a decoder. [Figure 6]This is a diagram showing a first example of image conversion by the homography image conversion system according to an embodiment. [Figure 7] This is a diagram showing a second example of image conversion by the homography image conversion system according to an embodiment.
[0013] Embodiments embodying the disclosed technology will be described. The homography image conversion system according to this embodiment is configured as shown in FIG. 1. The homography image conversion system 1 in FIG. 1 has an encoder 2 and a decoder 3. The encoder 2 receives the input of the input image GT and the decoder 3 outputs output images rec and rec*. The output images rec and rec* are reconstructed through homography image conversion based on the input image GT and are. The difference between the output image rec and the output image rec* will be described later. The input image GT is an image with some distortion added to the reference image.
[0014] The encoder ② is a variational auto-encoder (Va riational Auto-Encoder) that converts the input image GT into encoded information. The decoder 3 decodes the encoded information and outputs the image re c to reconstruct. The homography image conversion system 1 in FIG. 1 further has a coordinate sample output unit 4 and a concatenation calculation unit 5 between the encoder 2 and the decoder 3. As a result, in the homography image conversion system 1, the output image rec is reconstructed by the decoder 3 based only on the part of the encoded information that removes the distortion characteristics of the input image GT. be able to.
[0015] As shown in FIG. 2, the encoded information output from the encoder 2 is the position vector z'. This corresponds to multiplying by the transformation matrix H. The position vector z' is the image of the input image GT. This corresponds to a feature. The transformation matrix H is a matrix that represents the distortion of the input image GT from the reference image. The input image GT in Figure 2 is a grid-like figure, which is the reference image, that has been appropriately distorted. The overall shape of the input image GT is square. The transformation matrix H is given by the input image GT in its undistorted state. This is an indicator that shows how distorted the original grid figure is.
[0016] However, from the encoded information output from encoder 2, the position vector z' and the transformation matrix are obtained. H cannot be read directly. Therefore, the feature removal unit 4 processes the entire input image GT. From the encoded information z, the position vector z' and transformation matrix H are extracted using machine learning. Machine learning ensures that the position vector z' for the entire input image GT follows a normal distribution. This is how the representation is obtained. Furthermore, the transformation matrix H is extracted from the encoded information z after machine learning. (Mechanics) The post-training position vector z' is derived from the original reference image features of the input image GT after distortion has been removed. This information reflects the fact that, in this way, the homography image conversion system 1, Using learning, the transformation matrix H and the position vector in the world coordinate system are obtained from encoded information z. Separate the regular representation of z'.
[0017]
number
[0018] The relationship between the transformation matrix H and the coordinate sample uν can be expressed as shown in Equation 1. The coordinate sample uν on the right side corresponds to the position before the transformation, and the transformed coordinate sample uν* on the left side This corresponds to the position before the transformation. Equation 1 is the transformation performed by multiplying the coordinate sample uν by the transformation matrix H from the left. This is the form in which the coordinate sample uν* is obtained. The coordinate sample uν and the transformed coordinate sample uν* are It has three elements, but since the third component is fixed at "1", it is essentially two-dimensional. The transformation matrix H is a 3x3 matrix, but the bottom right component is fixed at "1". Therefore, it effectively has 8 degrees of freedom.
[0019] The transformation matrix H can be used to reconstruct the image from the encoded information. This is shown in Figure 3. This will be explained using Figure 4. Figure 3 shows the region S, the coordinate sample uν, the transformation matrix H, and the transformation The reconstituted coordinate sample uν* appears. Region S represents the region of the image to be reconstructed. In Figure 3, region S is a square region on the xy coordinate system. The coordinate sample uν is, This is a position information matrix representing the position information within region S. Each component of the coordinate sample uν is -1. The value is within the range of ~1. If the length of each segment of region S is 64 pixels, then the coordinate sample Each component of uν is represented by a 6-digit binary value. The generation of the coordinate sample uν based on region S is This is the function of the coordinate sample output unit 4 in Figure 1. The transformed coordinate sample uν* is a coordinate sample This is the transformed position information matrix obtained by multiplying ν by the transformation matrix H.
[0020] The coordinate sample uν is defined as shown in Equation 2. In the definition in Equation 2, the range of region S is x Both the axis and the y-axis are defined as ranging from "-1" to "1".
number
[0021] Figure 4 shows the position vector z', the transformed coordinate sample uν*, and the decoder 3. It is. z' in Figure 4 is equivalent to z' in Figure 2. However, in Figure 4, the position Multiple z' values are displayed depending on the information. Also, the z' values in Figure 4 are those after machine learning. The transformed coordinate sample uν* in Figure 4 is the same as the transformed coordinate sample uν* in Figure 3. Yes. The position vector z' and the transformed coordinate sample uν* are input to decoder 3.
[0022]
number
[0023] Specifically, the input to decoder 3 is the position vector z' and the transformation coordinate sample uν. *This is a connection shown in equation 3. The calculation of the connection is a function of the connection calculation unit 5 in Figure 1. This allows decoder 3 to reconstruct the output image rec* and output it. The output image rec* in Figure 4 is almost identical to the input image GT in Figure 2. This indicates that the image reconstruction by decoder 3 was successful.
[0024] The loss function in the image reconstruction described above is, in the case of a typical variational autoencoder, Similarly, it can be expressed by the number 4.
number
[0025] In this embodiment of homography image conversion system 1, the conversion matrix H is not used, and the decoder Image reconstruction can be performed using method -3. This is explained in Figure 5. In Figure 4, the transformed coordinate sample uν* is replaced with the coordinate sample uν. As explained in Figure 3, the target sample uν is a positional information defined by equation 2 for the region S. It is a matrix. In other words, in Figure 5, the coordinate sample uν is not multiplied by the transformation matrix H, and the coordinate The sample uν is used as is. This function is also included in the functions of the concatenation calculation unit 5.
[0026]
number
[0027] In this case, the connection between the position vector z' and the coordinate sample uν, as shown in equation 5, is deco It is input to decoder 3. As a result, the output image rec is reconstructed by decoder 3. The output is generated. The output image rec in Figure 5 is slightly different from the input image GT in Figure 2. This is because the transformation matrix H is not reflected. Therefore, the output image rec is the same as the input. This image is obtained by removing the distortion indicated by the transformation matrix H from the original image GT.
[0028] In this embodiment of homography image conversion system 1, the input image GT is homographed as described above. The image is converted to an output image rec through graphic image conversion. Thus, the output image rec contains: The difference between this image and the output image rec* is that the distortion characteristics due to the transformation matrix H are not reflected. The output image rec is different from the input image GT in that the distortion features have been removed. Yes. For example, if the input image GT is an image of the object viewed from a first viewpoint, then the output image The image rec is an image of the object viewed from a different, second viewpoint. Output image rec and output image Of the image rec*, the output image rec is a distinctive feature of the technology disclosed here.
[0029] If the image of the object viewed from a second perspective is the reference image, then the input image GT has another This introduces distortion due to viewing from the first viewpoint. The output image rec is, This image can be described as a reproduction of the reference image by removing distortion from the input GT image. There is no need to know what shape the reference image is. The quality of the output image rec is important. The evaluation will involve comparing the image with a reference image.
[0030] (Example 1) An example of image conversion using the homography image conversion system 1 of this embodiment will be described. An example is shown in Figure 6. Figure 6 is an example of the conversion of handwritten numbers into images. The upper part of Figure 6 shows "GT " is the actual handwritten character. This is the input image GT to encoder 2 and The middle row "rec*" and the bottom row "rec" are output images from decoder 3. "rec*" is the case when the transformation matrix H is used, and "rec" is the case when the transformation matrix H is used. This is the case where they were not present.
[0031] The image "rec*" is an image with a high degree of similarity to the input image GT. This indicates that the Deco This shows that the image reconstruction accuracy by Order 3 is high. In contrast, "rec" The image differs slightly from the input image GT, but within a range where the numbers are still legible as the same. Overall, in "rec," the distortion in the numbers in the input image GT is removed. For example, the second from the left and the second from the right in the image are both "1", but While there are considerable differences in shape in the "GT" image, the differences are small in the "rec" image. It's blooming.
[0032] (Example 2) A second embodiment is shown in Figure 7. Figure 7 is an example of the transformation of a grid-like image. The upper part of Figure 7 "GT" is a shape created by artificial intelligence to appropriately distort the original grid pattern. There are eight variations shown here. A distorted grid pattern is placed on it. This will be the input image GT to encoder 2. The lower section is the output image from decoder 3, as in the case of Figure 6.
[0033] The image labeled "rec*" in the middle section, as in Figure 6, is an image with a high degree of similarity to the input image GT. This is an image. In the second embodiment as well, the accuracy of image reconstruction by decoder 3 is high. This is shown. In the lower part of Figure 7, under "rec", all eight grid-like figures are almost distorted. It has a shape that lacks distortion. This is because the transformation matrix H was not used, resulting in the original shape before distortion. This is thought to be a reproduction of the original grid pattern.
[0034] As described in detail above, in the homography image conversion system 1 according to this embodiment, By performing machine learning with encoder 2, the distortion features are removed from the input image GT. The goal is to obtain the vector z'. Furthermore, the coordinate sample output unit 4 outputs the coordinate sample uν The goal is to obtain the connection between the vector z' and the coordinate sample uν. The goal is to obtain a result. This allows the decoder 3 to convert the input image GT into an image. This is how we obtain the output image rec. Thus, there is no need to manually determine the coordinates. A homography image conversion system has been realized that can obtain a converted image based on the input image. It is.
[0035] Thus, the disclosed technology uses deep learning to automatically generate a transformation matrix from an input image. It learns in a specific way. This allows it to convert a planar image into a planar image from a different viewpoint. Therefore, for example, when using multiple photos taken while changing the camera position as input images... However, it eliminates the need to manually determine the coordinates for each camera position. Also, the projection transformation matrix and By separating the normalized representation of the world coordinate system and performing image reconstruction learning, a more accurate visual representation can be achieved. Implement point transformation.
[0036] For example, in a parking lot, based on images captured by the onboard cameras of individual vehicles, this disclosure The technology allows us to obtain a plan view of the parking lot from above. This can be done by taking photos with an in-vehicle camera. Instead of using existing images, images taken by surveillance cameras installed in or around the parking lot were used. It is also possible to do so. Or, conversely, when there are a certain number of vehicles parked in the parking lot... Based on images taken from above by a camera mounted on a drone, the parking lot in question An image of the situation can be obtained from an entrance / exit on the ground. Also, a certain person can be viewed from a certain direction. Based on the image captured, it is also possible to obtain an image of the person viewed from a different direction.
[0037] Furthermore, images of the object taken from an appropriate viewpoint (for example, from diagonally above) can be obtained using the disclosed technology. It can be converted into a planar image viewed from a different viewpoint (for example, vertically upward). Furthermore, Coder 2 may be made capable of receiving input images GT from multiple sources. Furthermore, using the disclosed technology, multiple input images of a common object viewed from different viewpoints can all be processed in the same manner. The image is converted into a planar image viewed from the same viewpoint (for example, vertically upward) that is different from any of the other viewpoints. By doing so, an integrated plan view can be easily obtained by combining them. This is because, as mentioned above, It can be used as a parking lot, and it can also be used for processing using object tracking technology, etc.
[0038] This embodiment is merely illustrative and does not limit the disclosed technology in any way. Therefore, the disclosed technology can naturally be improved and modified in various ways, without departing from its essence. That is the case. [Explanation of Symbols]
[0039] 1. Homography Image Conversion System 2 Encoders 3 Decoder 4. Coordinate Sample Output Unit 5. Consolidation Calculation Unit
Claims
1. An encoder that receives input of an input image which is a reference image with distortion added to it, and the input A homograph having a decoder that outputs an output image obtained by homographing a force image. An image conversion system, The encoder, based on machine learning of the input image, Extraction of a transformation matrix, which is an indicator of the distortion of the input image from the reference image, The distribution of position vectors corresponding to the features of the input image is normalized to remove distortion. This involves obtaining features from a quasi-image. The coordinate sample is a position information matrix that represents the position information within the region of the output image. Sample output unit, Concatenation calculation: This calculates the concatenation between the coordinate sample and the position vector after machine learning. It further has a department, The decoder receives the input of the concatenation and constructs the output image. Raffi image conversion system.
2. A homography image conversion system according to claim 1, The decoder constructs the output image based solely on the concatenation without using the transformation matrix. A homography image conversion system that accomplishes this task.
3. A homography image conversion system according to claim 1 or claim 2, The aforementioned input image is an image of the object viewed from a certain viewpoint. The output image is a planar image of the object viewed from a different viewpoint than the viewpoint of the input image. This is a homography image conversion system.
4. A homography image conversion system according to claim 1 or claim 2, The encoder is a homography image transformer that accepts multiple input images. system.
5. A homography image conversion system according to claim 4, The aforementioned multiple input images are images of a common object viewed from different viewpoints. All of the output images are identical to the object from any of the separate viewpoints. A homography image conversion system that displays a planar image from a specific perspective.
6. A homography image conversion system according to claim 5, Homography image conversion system where all of the above input images are images of a parking lot. Stem.