Method for generating physical consistency polarization image data based on semantic segmentation
By generating polarization image data through semantic segmentation and physical consistency verification, the problems of high cost and single scene were solved, and high-quality and diversified data generation was achieved, thereby improving the model's generalization ability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-03
- Publication Date
- 2026-03-13
AI Technical Summary
Existing polarization image data acquisition is costly, small in scale, and limited in scenarios. The generation methods lack physical consistency and fail to fully utilize semantic information, resulting in poor generalization ability of deep learning models.
The semantic information of the image is extracted by a semantic segmentation model, and polarization parameters are assigned by combining the material and polarization characteristics of the object. Physical consistency is checked to generate large-scale polarization image data that conforms to physical laws.
The low-cost generation of high-quality, diverse polarization image data enhances the generalization ability and practical application performance of deep learning models.
Smart Images

Figure CN121661342A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for generating physically consistent polarization image data based on semantic segmentation. It is suitable for providing training data for deep learning models in fields such as polarization image demodulation, target detection and recognition, and polarization 3D reconstruction. It can be widely applied in multiple technical fields such as computational imaging, machine vision, autonomous driving, remote sensing monitoring, and medical imaging, and belongs to the interdisciplinary field of computational imaging and artificial intelligence. Background Technology
[0002] Polarized images contain rich information such as light intensity, polarization direction, and degree of polarization. Compared to traditional RGB images, they provide more information about the physical properties of object surfaces, giving them unique advantages in tasks such as target recognition and scene understanding. With the development of deep learning technology, data-driven polarized image processing models have become a research hotspot. However, the performance of these models is highly dependent on large-scale, high-quality polarized image datasets.
[0003] Currently, acquiring polarization image data primarily relies on specialized polarization imaging equipment, such as polarization cameras and polarization sensors. These devices acquire images in different polarization directions and demodulate them to obtain polarization image data containing four parameters: S0, S1, S2, and S3. In fields such as remote sensing monitoring and autonomous driving, polarization images have begun to be used to improve the robustness of target detection, for example, by distinguishing objects of different materials on a road based on their polarization characteristics. Meanwhile, semantic segmentation technology has made significant progress in the field of image understanding, accurately identifying object categories and region ranges in images, providing a technological foundation for generating targeted data based on object characteristics.
[0004] Current methods for acquiring and generating polarization image data have many limitations:
[0005] Challenges in acquiring real-world datasets: Professional polarization imaging equipment is expensive, with a single unit typically costing hundreds of thousands of yuan. Furthermore, the data acquisition process is complex, requiring strict control over lighting conditions and shooting angles. This results in real-world polarization image datasets being characterized by high acquisition costs, small scale, and limited scene representation, making it difficult to meet the needs of deep learning models for large-scale, diverse data.
[0006] Traditional generation methods lack physical consistency: Existing simulation-based polarization image generation methods mostly generate polarization parameter maps by simply assigning parameter values, without considering the correspondence between object material and polarization characteristics, resulting in polarization images lacking physical plausibility. While some methods attempt to incorporate object surface characteristics, they lack rigorous physical consistency verification, often resulting in polarization degrees of polarization (DOP) greater than 1, which violates physical laws and affects the reliability of training data.
[0007] Insufficient utilization of semantic information: Existing technologies do not fully integrate the semantic information of images for polarization parameter allocation. The polarization parameters of the same object region in the generated polarized image lack consistency, and the differences in polarization characteristics between different objects are not obvious, failing to truly reflect the polarization reflection patterns of objects of different materials in real-world scenes.
[0008] Insufficient data diversity: Traditional generation methods generate datasets with limited scenarios, making it difficult to cover real-world application scenarios such as complex weather, different lighting conditions, and diverse materials. This results in poor generalization ability of deep learning models trained on these datasets, leading to a significant performance drop in real-world scenarios.
[0009] These problems severely restrict the development of deep learning technology in the field of polarization image processing, and there is an urgent need for a method that can generate polarization image data that conforms to physical laws at low cost and on a large scale.
[0010] This invention proposes a method for generating physically consistent polarized image data based on semantic segmentation, which makes a breakthrough improvement on the core defects of prior art: it extracts semantic information of images through an artificial intelligence panoramic segmentation model, allocates polarization parameters by combining the correspondence between object material and polarization characteristics, and ensures the rationality of the generated data through strict physical consistency verification, thus solving the technical problems of high cost, small scale, and single scene in obtaining real polarized image datasets. Compared with traditional generation methods, this invention makes full use of semantic information to achieve accurate allocation of polarization parameters, ensuring the consistency of polarization parameters within the same semantic region and the difference in polarization characteristics between different semantic regions, and the generated polarized images are closer to real scenes. Through parameter normalization processing and data augmentation operations, it not only ensures the physical consistency of the data, but also improves the diversity of the dataset, which can provide high-quality and diversified training data for deep learning models, significantly improving the generalization ability and practical application performance of the models. Summary of the Invention
[0011] The purpose of this invention is to provide a low-cost, large-scale, and high-quality method for generating physically consistent polarization image data based on semantic segmentation, thereby solving the problems of difficulty in obtaining existing polarization image datasets, poor physical consistency, and insufficient diversity.
[0012] The objective of this invention is achieved as follows:
[0013] Image reception and conversion: Receive a standard digital image and convert it into a grayscale image using a weighted average method. The converted grayscale image is directly used as the total polarization intensity S0 map to ensure that the S0 parameter can accurately reflect the total light intensity information.
[0014] Semantic Region Segmentation: The high-performance AI-powered panoptic segmentation model mask2former-swin-large-coco-panoptic is used to analyze and process standard digital images. This model leverages massive image feature knowledge learned during pre-training to perform panoptic segmentation end-to-end, accurately dividing independent object regions within the image and labeling them with semantic categories, resulting in multiple regions with semantic labels. Simultaneously, edge smoothing and hole filling are performed on the segmented semantic regions, effectively eliminating potential errors such as region discontinuities and jagged edges during segmentation, ensuring the spatial integrity and continuity of the semantic regions, and laying the foundation for accurate subsequent polarization parameter allocation.
[0015] Polarization parameter allocation: Based on the segmented semantic region features, a uniform allocation strategy is adopted to map a portion of the semantic region to the channels corresponding to polarization parameters S1, S2, and S3, respectively. The allocated semantic regions retain the original image's grayscale brightness for display, while the background portion of the unallocated semantic regions is displayed according to the original image. Figure 1 The grayscale brightness of / 5 is darkened. This allocation method not only ensures the semantic distinction of each polarization channel, but also improves the authenticity and rationality of polarization image data by darkening the background to match the physical characteristics of low polarization of background light in real scenes.
[0016] Physical consistency verification: through the degree of polarization (DOP) calculation formula The four polarization parameters are verified. When DOP > 1, the values of S1, S2, and S3 are adjusted through parameter normalization. The normalization formula is as follows: α is an adjustment coefficient between 0.8 and 0.95, which ensures the physical constraint of DOP≤1 while preserving the differences in polarization characteristics between different objects, and finally obtains the desired target image.
[0017] The beneficial effects of this invention are: low-cost, large-scale generation; no reliance on specialized polarization imaging equipment; polarization image data can be generated using only ordinary standard digital images, significantly reducing dataset acquisition costs and enabling rapid generation of large-scale training data; rigorous DOP verification and parameter normalization ensure that the generated polarization images conform to the physical laws of polarization, solving the problem of insufficient physical rationality in traditional generation methods; combining image semantic information to allocate polarization parameters makes polarization characteristics highly matched with object materials, resulting in images that are closer to real-world scenes and improving the effectiveness of training data; through semantic segmentation to adapt images to different scenes, combined with data augmentation operations, polarization datasets covering various materials, lighting, and scenes can be generated, improving the generalization ability of deep learning models. Attached Figure Description
[0018] Figure 1This is a schematic diagram of the overall process of generating physically consistent polarized image data based on semantic segmentation according to the present invention. It shows the complete process from standard digital image input to polarized image dataset output, including four core links: image conversion, semantic segmentation, parameter allocation and consistency verification. It clearly presents the logical relationship and data flow of each step.
[0019] Figure 2 The diagram shows the semantic segmentation results. (a) is the input RGB standard digital image, which in this example contains scene elements such as motorcyclists, dirt roads, mountain vegetation, and sky. (b) is the semantic label map after panoramic segmentation. Different colors represent different semantic regions, clearly separating categories such as sky, mountains, roads, vegetation, and motorcyclists. (c) is the refined semantic region map after edge smoothing and hole filling, eliminating edge jaggedness and internal hole problems in the original segmentation.
[0020] Figure 3 The diagram shows the results of polarization parameter allocation, illustrating the generation of four polarization parameters: S0, S1, S2, and S3. (a) is the converted S0 total intensity map, which retains the brightness distribution characteristics of the original image; (b) is the S1 parameter map; (c) is the S2 parameter map; and (d) is the S3 parameter map.
[0021] Figure 4 The results of the physical consistency verification show the light intensity of the four Stokes parameters and the intensity curve of the entire DOP. This ensures the physical constraint that DOP ≤ 1. Detailed Implementation
[0022] The present invention will be further illustrated below with reference to specific embodiments.
[0023] like Figure 1 As shown, a method for generating physically consistent polarization image data based on semantic segmentation includes:
[0024] A1, convert a standard digital image into a grayscale image as the total intensity (S0) map;
[0025] A2 uses a panoramic segmentation model to segment a standard digital image into multiple regions with semantic labels;
[0026] A3, according to the preset allocation rules, the semantic region is allocated to S1, S2 and S3 in the spatial dimension corresponding to S0, resulting in the initial four Stokes parameters;
[0027] A4 calculates the degree of polarization based on the initial Stokes parameters and generates a temporary polarization image;
[0028] A5 verifies the temporary polarization image, corrects the illegal areas where DOP>1, and finally outputs the target polarization image data.
[0029] Specifically, in step A1, a standard digital image is received. Taking an RGB image as an example, it is converted into a grayscale image using a weighted average method. The conversion formula is Gray = 0.299R + 0.587G + 0.114B (where R, G, and B are the pixel values of the red, green, and blue channels of the RGB image, respectively). The converted grayscale value range is [0, 255]. This grayscale image is directly used as the total polarization intensity (S0) map to ensure that S0 accurately reflects the total light intensity information.
[0030] Specifically, in step A2, as follows: Figure 2 As shown, a pre-trained mask2former-swin-large-coco-panoptic AI panoramic segmentation model is used to perform end-to-end analysis and processing on the received standard digital images. This model relies on the massive image feature knowledge learned during the pre-training stage to accurately identify the boundaries and extents of different objects in the image, segmenting the image into multiple independent regions and labeling each region with a corresponding semantic category label (such as motorcycle, road, grass, mountain, sky, etc.). Finally, it outputs multiple independent regions with semantic labels, providing a foundation for subsequent polarization parameter allocation.
[0031] Specifically, in step A3, as follows: Figure 3 As shown, polarization parameter allocation is performed according to a preset uniform allocation rule: First, the total number of semantic regions output in step A2 is counted. Following the principle of "average allocation + remainder priority" (e.g., when the total number of regions is 10, S1 allocates 4, S2 allocates 3, and S3 allocates 3), each semantic region is mapped to the three polarization parameter channels S1, S2, and S3 respectively in the spatial dimension corresponding to the S0 image. Pixels allocated to semantic regions retain their original grayscale brightness, while background pixels not allocated to semantic regions retain their original grayscale brightness. Figure 1 The grayscale brightness is displayed at 5 / 5, and the initial S1, S2 and S3 images are generated. Combined with the S0 image, the initial four Stokes parameters (S0, S1, S2 and S3) are obtained.
[0032] Specifically, in step A4, based on the initial four Stokes parameters, the polarization degree is calculated pixel-by-pixel using the degree of polarization (DOP) calculation formula to generate a temporary polarization degree image. The DOP calculation formula is as follows: S0-S3 are the four Stokes parameters obtained in step A3. Each pixel value of this temporary polarization image directly corresponds to the polarization degree at the same location in the original image, intuitively presenting the polarization degree distribution of each region.
[0033] Specifically, in step A5, as follows: Figure 4As shown, a physical consistency check is performed on the generated temporary polarization image: First, all pixels in the image are traversed to identify all non-compliant pixel regions with a DOP value greater than 1 (these regions violate the physical laws of polarization and need to be corrected); for these non-compliant regions, a normalization formula is used. The corresponding S1, S2, and S3 values were adjusted, with the adjustment coefficient α set to 0.9 (within a reasonable range of 0.8 to 0.95). After correction, the DOP value of all pixels was ≤1, which not only met the physical consistency requirements but also preserved the differences in polarization characteristics between different objects, ultimately yielding target polarization image data that met the requirements.
[0034] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to specific embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications and substitutions should be covered within the scope of the claims of the present invention. Technical, shape, and structural parts not described in detail in this invention are all well-known technologies.
Claims
1. A method for generating physically consistent polarization image data based on semantic segmentation, characterized in that, Includes the following steps: A1. Receive a standard digital image and convert it into a grayscale image. The grayscale image is directly used as the total polarization intensity (S0) map. A2. Analyze the standard digital image using an artificial intelligence panoramic segmentation model, outputting multiple independent regions with semantic labels. A3. Based on preset semantic rules, assign each semantic region obtained in step A2 to polarization parameters S1, S2, and S3 in the spatial dimension corresponding to the S0 map, generating initial S1, S2, and S3 maps, thus obtaining the initial four Stokes parameters (S0, S1, S2, S3). A4. Based on the initial four Stokes parameters, calculate the polarization degree pixel-by-pixel using the polarization degree calculation formula, generating a temporary polarization degree image. A5. Perform physical consistency verification on the temporary polarization degree image, identifying all pixel regions with a DOP value greater than 1. For these non-compliant pixel regions, normalize and correct their corresponding S1, S2, and S3 values to ensure that the DOP of all pixels after correction is ≤ 1, ultimately obtaining target polarization image data that meets the physical consistency requirements.
2. The image panoramic segmentation method according to claim 1, characterized in that: The AI panoramic segmentation model mentioned in step A2 is the mask2former-swin-large-coco-panoptic model. This model is based on the Transformer encoder-decoder architecture and mask attention mechanism. Through the massive image feature knowledge learned in the pre-training stage, it performs end-to-end panoramic segmentation processing on standard digital images. It can simultaneously achieve object instance segmentation and semantic classification, and has high regional boundary positioning accuracy and strong internal pixel consistency. It provides reliable spatial region division and semantic support for the targeted allocation of subsequent polarization parameters S1, S2, and S3.
3. The polarization parameter allocation method according to claim 1, characterized in that: Step A3 includes semantic region refinement, which involves edge smoothing and hole filling of the segmented semantic regions to eliminate regional discontinuities caused by segmentation errors and ensure the spatial continuity of polarization parameter allocation.
4. The parameter-satisfying physical consistency method according to claim 1, characterized in that: In step A3, preset high polarization values are set for each semantic region assigned to polarization parameters S1, S2, and S3, and preset low polarization values are set for the background portion of the unassigned semantic regions. Initial S1, S2, and S3 images and corresponding temporary polarization degree (DOP) images are generated by combining the S0 image. In step A5, by identifying all non-compliant pixel regions with DOP values greater than 1 in the temporary polarization degree image, the S1, S2, and S3 values of these regions are normalized and corrected to ensure that the DOP of all pixels is ≤1 after correction. Finally, target polarization image data that meets the physical consistency requirements is output.