Convex mirror reflected image generation, semantic segmentation method and device
By constructing coordinate system relationships to generate high-quality convex mirror reflection images and performing adversarial learning, the geometric morphological differences between convex mirror reflection images and normal images are resolved, thereby improving the accuracy of the semantic segmentation model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PEKING UNIV
- Filing Date
- 2022-02-23
- Publication Date
- 2026-05-01
AI Technical Summary
The scarcity of existing convex mirror reflection image data leads to significant domain differences between different domains. Traditional methods struggle to effectively address the geometrical differences between convex mirror reflection images and normal images, thus affecting the accuracy of semantic segmentation models.
By constructing the relationship between the world coordinate system and the camera coordinate system, high-quality convex mirror reflection image data is generated. Adversarial learning is then performed using a pose estimator and a convolutional network to reduce differences in the geometric morphology domain and improve semantic segmentation accuracy.
It effectively reduces the geometric morphological domain differences between normal images and convex mirror reflection images, and improves the semantic segmentation accuracy of convex mirror reflection images.
Smart Images

Figure CN116681708B_ABST
Abstract
Description
Methods and apparatus for generating and semantically segmenting images reflected by convex mirrors. Technical Field
[0001] This invention relates to the field of convex mirror reflection images, and more particularly to a method and apparatus for generating and semantically segmenting convex mirror reflection images. Background Technology
[0002] Convex mirrors are commonly found at street corners, reflecting scenes in blind spots to provide safety for pedestrians and drivers. In recent years, deep learning-based scene segmentation techniques have developed rapidly, but training a segmentation network in a supervised manner requires a large amount of labeled data. However, existing convex mirror reflection images are scarce, making it impossible to obtain comprehensive training data. A primitive approach is to train on data in one domain and test on data in another; however, due to domain differences, such models often have low accuracy.
[0003] To address this issue, some methods have proposed unsupervised domain adaptation approaches, such as those based on adversarial learning or self-training. However, these methods primarily address domain differences caused by traditional style variations, neglecting the fact that the difference between convex mirror reflections and normal images mainly stems from geometric morphology; that is, convex mirror reflections exhibit severe distortion, while normal images do not. Therefore, traditional methods for addressing style variations cannot be directly applied to resolving geometric morphological differences between convex mirror reflections and normal images. Summary of the Invention
[0004] To overcome the shortcomings of existing technologies, this invention provides a method and apparatus for generating and semantically segmenting convex mirror reflection images. By generating high-quality convex mirror reflection image data, the semantic segmentation model achieves higher accuracy on convex mirror reflection images.
[0005] The technical solution of the present invention is as follows:
[0006] A method for generating a reflection image from a convex mirror, comprising the following steps:
[0007] Construct a world coordinate system and a camera coordinate system, and base the image on the reflection from the convex mirror I. q The attitude parameters are used to establish the relationship between the world coordinate system and the camera coordinate system. The relationship includes the distance between the origin of the world coordinate system and the camera coordinate system. The attitude parameters include the tilt angle parameter α, the tilt angle parameter β, the radial twist parameter k, and the square ratio of the distance between the origin to the focal length of the camera coordinate system. q is the number of the image reflected by the convex mirror.
[0008] Based on the radial distortion parameter k, for the planar image I placed in the camera coordinate system sRadial distortion is performed to obtain the distorted image in the world coordinate system, where s is the number of the planar image;
[0009] Based on the tilt angle parameters α and β, the distorted image is rotated around the X and Y axes of the world coordinate system to obtain the planar image I. s Image I′ of convex mirror reflection s .
[0010] Furthermore, the method for obtaining the attitude parameters includes: reflecting the image I from the convex mirror. q The input pose estimator, and the network structure for training the pose estimator includes a ResNet18 network.
[0011] Furthermore, the method for training the pose estimator includes:
[0012] Extracting the reflection image from the convex mirror I q Image I′ reflected by a convex mirror s The geometric edges, and based on the obtained target boundary With source boundary Adversarial learning is performed to obtain the first loss of the pose estimator;
[0013] Based on the reflection image of a convex mirror I q Image I′ reflected by a convex mirror s Adversarial learning is performed on the semantic segmentation results to obtain the second loss of the pose estimator;
[0014] The parameters of the attitude estimator are adjusted by backpropagating based on the first loss and the second loss.
[0015] Furthermore, the planar image I placed in the camera coordinate system s Radial twisting includes:
[0016] 1) For the planar image I s Each coordinate point in [x] is [x] o y o ] T Calculate parameters
[0017] 2) Based on parameter r o Using the radial distortion parameter k, the coordinates of the distorted image are calculated as [x...]. b y b ] T .
[0018] A semantic segmentation method, the steps of which include:
[0019] Based on the reflection image I from the convex mirror qCompared with the convex mirror reflection image I′ obtained using any of the above methods s Construct a training set;
[0020] A segmentation network is obtained by training a convolutional network using the training set.
[0021] The target convex mirror reflection image is input into the segmentation network to obtain the semantic segmentation result of the target convex mirror reflection image.
[0022] A storage medium storing a computer program, wherein the computer program is configured to execute any of the methods described above when run.
[0023] An electronic device is characterized by comprising a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform any of the methods described above.
[0024] By taking the above steps, the domain differences in geometric shape between normal images and convex mirror reflection images can be effectively reduced, thereby improving the semantic segmentation results of real convex mirror reflection images.
[0025] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0026] This invention relies on a convex mirror simulation module to simulate convex mirror imaging, and uses adversarial learning to make the simulated convex mirror image and the real image geometrically similar, thereby improving the semantic segmentation accuracy of the real convex mirror reflection image. Attached Figure Description
[0027] Figure 1 shows the framework provided by the present invention for reducing the difference between normal images and convex mirror reflection image domains.
[0028] Figures 2, 3 and 4 illustrate the simulated convex mirror imaging process in an embodiment of the present invention.
[0029] Figure 5 shows a real convex mirror image and a simulated convex mirror image provided by the invention.
[0030] Figures 6 and 7 show the input images and corresponding semantic segmentation results provided by the present invention. Detailed Implementation
[0031] The present invention will be further described below with reference to the accompanying drawings and embodiments, but the scope of the invention is not limited in any way.
[0032] Figure 1 is a framework diagram of the unsupervised semantic segmentation of convex mirror reflection images provided by the present invention. As can be seen from the flowchart in Figure 1, the entire system includes a designed convex mirror simulation layer, a pose estimator to estimate the parameters of the target domain image, adversarial processing on the edges of the input image, and adversarial processing on the semantic boundaries of the output result.
[0033] First stage: Establishing a convex mirror simulation layer; Figures 2, 3 and 4 show the convex mirror simulation layer provided by the present invention, which includes establishing the relationship between the camera coordinate system and the world coordinate system, placing the radially distorted image in the world coordinate system, rotating the placed image and imaging it.
[0034] Specifically, we first establish the relationship between the camera coordinate system and the world coordinate system; both are right-handed coordinate systems. As shown in Figure 2, let P be a point in the world coordinate system and P' be a point in the camera coordinate system. c The rotation transformation between the world coordinate system and the camera coordinate system is represented by the identity matrix I, and the translation vector t = [0 0 d]. T Where d is the distance between the two coordinate systems. The transformation relationship between the camera coordinate system and the world coordinate system is P. c = P + t, where P c Let be a point in the camera coordinate system. That is, the X-axis, Y-axis, and Z-axis of the two coordinate systems are parallel to each other, and the distance between the origins of the two coordinate axes is d.
[0035] Assume a point on the imaging plane is Intrinsic parameter matrix
[0036]
[0037] Then you can get Where f is the focal length.
[0038] Secondly, the image is radially distorted, and the coordinates of the normal image are set to [x]. o y o ] T The coordinates of the barrel after distortion are [x b yb] T The distortion process is as follows:
[0039]
[0040] in k is the radial torsion parameter.
[0041] Finally, place the radially warped image onto the X coordinate system of the world coordinate system. w OY w Plane, and around OX w Axis and OY w The axis is rotated by angles α and β to simulate the perspective distortion of a convex mirror.
[0042] The final transformation can be expressed by the following formula.
[0043]
[0044]
[0045] in Let be the homogeneous coordinates of a point on the final synthesized convex mirror image. Let be the homogeneous coordinates of a point on the barrel-shaped deformed image.
[0046]
[0047] Where s1 = sinα, s2 = sinβ, c1 = cosα, c2 = cosβ. When hour, when hour, In practical applications, we set the focal length f = 1 and use the ResNet18 pose estimator to estimate the tilt angle parameters α and β, as well as the distance d between the two coordinate systems and the radial twist parameter k. To make the simulated convex mirror image more similar to the real convex mirror image, this invention performs adversarial learning at the edges of the input space and the semantic boundaries of the output space. Figure 5 shows the real convex mirror image and the synthesized convex mirror image.
[0048] The second stage involves adversarial learning of edges in the input space; to make the synthesized image more realistic, the source domain image I′ is processed separately. s and target domain image I q Extract geometric edges and perform adversarial learning. Assume the source boundary of the synthesized source domain image after edge extraction is... The target boundary after edge extraction from the target domain image is Then the adversarial loss of the training discriminant network in the input space can be obtained as follows:
[0049]
[0050] Among them, D geo For the discriminator of the input space, L bce This is the binary cross-entropy loss. Similarly, the loss for training the pose estimator can be obtained as follows:
[0051]
[0052] The third stage: Adversarial learning of the semantic boundaries of the output space. Assume the semantic boundaries of the source segmentation result are... The semantic boundary of the target segmentation result is Then the discriminator D can be obtained.sem The loss is
[0053]
[0054] Similarly, the loss function for training the pose estimator can be obtained as follows:
[0055]
[0056] The loss L in the overall adversarial input space geo Then the loss function for training the pose estimator is L. geo +L sem The adversarial loss of the discriminative network in the input space is... The adversarial loss of the discriminative network in the output space is Figure 6 shows the input target, and Figure 7 shows the target segmentation result.
[0057] It should be noted that the purpose of disclosing the embodiments is to help further understand the present invention. However, those skilled in the art will understand that various substitutions and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the embodiments, and the scope of protection of the present invention is defined by the scope of the claims.
Claims
1. A method for generating a reflection image from a convex mirror, comprising the following steps: Construct a world coordinate system and a camera coordinate system, and base the image on the reflection of the convex mirror. The attitude parameters are used to establish the relationship between the world coordinate system and the camera coordinate system. This relationship includes the distance between the origins of the world coordinate system and the camera coordinate system. The attitude parameters include tilt angle parameters. Tilt angle parameters Radial torsion parameters The ratio of the distance to the origin to the square of the focal length in the camera coordinate system. The image is numbered by the convex mirror reflection image; the method for obtaining the attitude parameters includes: displaying the convex mirror reflection image... The input pose estimator, the network structure for training the pose estimator includes a ResNet18 network, and the method for training the pose estimator includes: extracting images reflected from a convex mirror. Image reflected by a convex mirror The geometric edges, and based on the obtained target boundary With source boundary Adversarial learning is performed to obtain the first loss of the pose estimator; based on the convex mirror reflection image. Image reflected by a convex mirror Adversarial learning is performed on the semantic segmentation results to obtain the second loss of the pose estimator; backpropagation is performed based on the first loss and the second loss to adjust the parameters of the pose estimator; based on the radial twist parameter... For a planar image placed in the camera coordinate system Radial distortion is performed to obtain a distorted image in the world coordinate system. Numbering of planar images; based on tilt angle parameters With tilt angle parameters The distorted image is rotated around the world coordinate system. shaft and Rotate the axis to obtain the planar image. Convex mirror reflection image 。 2. The method as described in claim 1, characterized in that, The planar image positioned in the camera coordinate system Performing radial distortion includes: 1) to the planar image Each coordinate point in the array is Calculate parameters 2) Based on parameters With radial torsion parameters The coordinates of the distorted image are calculated as follows: 。 3. A semantic segmentation method, comprising the following steps: Based on the image reflected by the convex mirror The convex mirror reflection image obtained using any of the methods in claims 1-2 above. A training set is constructed; a convolutional network is trained using the training set to obtain a segmentation network; the reflection image of the target convex mirror is input into the segmentation network to obtain the semantic segmentation result of the reflection image of the target convex mirror.
4. A storage medium storing a computer program, wherein, The computer program is configured to execute any of the methods described in claims 1-3 at runtime.
5. An electronic device comprising a memory and a processor, the memory storing a computer program, the processor being configured to run the computer program to perform the method as claimed in any one of claims 1-3.