Standard-view white model generation method based on digital twinning

CN122617656APending Publication Date: 2026-08-21BEIJING YUANDIAN FUTURE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610755687.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-28
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

然而,由于采集视角通常带有明显的倾斜关系,原始图像中普遍存在透视畸变、局部尺度不一致、边缘模糊以及背景纹理干扰较强等问题,导致所得结果难以直接满足标准视角白模生成对于结构一致性和空间对应性的要求

Benefits of technology

本发明提出的基于数字孪生的标准视角白模生成方法,围绕原始图像进入数字孪生建模前的标准化处理需求,构建了由图像标准化、特征点识别、透视校正、结构增强、主体分割以及基准数据封装组成的连续处理链路,使采集端获得的投影机镜头视角图像能够被转换为适用于后续白模生成的规范化输入数据。与现有技术相比,本发明通过对原始图像的图像解码、色彩空间转换、白平衡处理、噪声抑制以及尺寸归一化处理,提高了输入图像在亮度、颜色和数据格式上的一致性,减弱了设备差异、光照波动和局部噪声对后续处理结果的影响;通过在候选区域内采用AKAZE图像特征匹配进行多尺度特征点检测,并结合主体角点、主体标记点与目标正面视图角点坐标之间的对应关系求解单应性矩阵,实现了从倾斜透视视图到正向平行视图的稳定映射,能够在校正过程中兼顾主体轮廓准确恢复和有效像素保留,提升标准视角构建的几何可靠性;通过对校正彩图进行灰度化、局部对比度增强、边缘检测和边缘强化处理,使目标主体的外轮廓信息、内部结构边界和几何走向特征得到集中表达,弱化了色彩纹理和背景干扰对主体识别的影响。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122617656A_ABST
    Figure CN122617656A_ABST
Patent Text Reader

Abstract

The application discloses a standard visual angle white model generation method based on digital twinning, which comprises the following steps: step one, image decoding, color space conversion, white balance and noise suppression processing are performed on an original image; step two, multi-scale feature point detection is performed; step three, a homography matrix is calculated and perspective transformation is performed to obtain a corrected color image; step four, the corrected color image is subjected to grayscale, contrast enhancement and edge enhancement to generate a structure-enhanced image; step five, the structure-enhanced image is input into an improved Mask2Former model, and a binary mask is obtained through a structure coding module, a main body query module, a mask decoding module and a graph cut refinement module; and step six, the corrected color image, the structure-enhanced image and the binary mask are subjected to pixel-level alignment and encapsulation to generate a reference image data package. The improved Mask2Former model is used to realize standard visual angle correction and white model reference data integrated generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision processing technology, and in particular to a method for generating standard-view white models based on digital twins. Background Technology

[0002] As digital twin technology is increasingly applied in scenarios such as virtual display, device mapping, spatial reconstruction, and digital content generation, the stable acquisition of standard perspective base data suitable for subsequent modeling processing from actual acquired images has gradually become a crucial factor affecting the quality of white model generation. Current technologies mostly use raw images from the camera or projector lens as input for image acquisition of the target subject, followed by general image enhancement, contour extraction, or semantic segmentation methods to identify the subject region. However, due to the often significant tilt of the acquisition perspective, the raw images commonly suffer from perspective distortion, inconsistent local scales, blurred edges, and strong background texture interference, making it difficult for the results to directly meet the requirements of structural consistency and spatial correspondence for standard perspective white model generation.

[0003] Existing solutions for these types of problems typically focus on a single step, such as image correction, edge detection, or foreground segmentation, lacking a continuous processing mechanism from original image standardization, feature point selection, frontal view mapping, structural enhancement to subject mask generation and data encapsulation. On one hand, some methods fail to uniformly process the channel order, color balance, and noise perturbations of the acquired images, leading to insufficient stability in subsequent feature detection. On the other hand, conventional perspective correction methods have limited accuracy in recognizing subject corners or markers, easily resulting in mapping offsets, boundary folding, or insufficient retention of effective pixels. Furthermore, existing segmentation methods are prone to subject boundary breaks, local misclassification, and mask discontinuities in complex boundaries, narrow structures, and weakly textured regions. Moreover, the corrected color image, structural enhancement image, and segmentation results often lack pixel-level alignment under a unified coordinate reference, making it difficult to directly form standardized benchmark data suitable for generating digital twin white models.

[0004] Therefore, how to provide a standard perspective for white model generation based on digital twins is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] One objective of this invention is to propose a standard-view white model generation method based on digital twins. This invention achieves integrated generation of standard-view white model benchmark data through image standardization, perspective correction, structural enhancement, subject segmentation, and pixel-level alignment encapsulation, thereby improving the accuracy of subject contour extraction, boundary continuity, and consistency of modeling data.

[0006] The standard-perspective white model generation method based on digital twins according to embodiments of the present invention includes the following steps: Step 1: Receive the original image captured from the projector lens, and perform image decoding, color space conversion, white balance and noise suppression on the original image to obtain a standardized image; Step 2: Use AKAZE image feature matching to perform multi-scale feature point detection on the standardized image, identify the main corner points or main marker points of the main body contour in the standardized image, and obtain the corrected feature point set; Step 3: Calculate the homography matrix based on the feature points in the correction feature point set and the preset target front view corner coordinates, and perform perspective transformation based on the homography matrix to obtain the corrected color image; Step 4: Convert the corrected color image to grayscale, enhance its contrast, and use the Canny edge detection operator to extract contours for edge enhancement, generating a structure enhancement image; Step 5: Input the structure enhancement graph into the improved Mask2Former model, and obtain the binary mask through the structure encoding module, the main body query module, the mask decoding module, and the graph cutting and refining module; Step 6: Perform pixel-level alignment on the corrected color image, structure enhancement image, and binary mask, and encapsulate metadata to generate the final baseline image data package.

[0007] Optionally, step one specifically includes: Receive the original image captured from the perspective of the projector lens, read the encoding header information, resolution information and channel information of the original image, and decode it according to the corresponding encoding format to obtain the decoded image; Check the channel arrangement order and bit depth format of the decoded image. If the channel arrangement order of the decoded image is the same as the channel arrangement order of the acquisition device, convert the decoded image into a preset unified channel arrangement order to obtain a unified channel image. A color space conversion is performed on the channel unified image, which converts the channel unified image from the device output color space to a preset processed color space to obtain a color-converted image. The pixel distribution of each color channel in the color-converted image is statistically analyzed. Based on the mean relationship of each color channel, the gain of each color channel is adjusted until the adjusted color channels reach the preset color balance condition, thus obtaining the white balance image. The white balance image is subjected to noise suppression processing, which involves eliminating isolated noise points, brightness jitter areas, and local particle areas using a neighborhood smoothing method, while preserving pixel changes at the subject edge position to obtain a noise-reduced image. The denoised image is scaled or padded to a preset input size, and the pixel value range is normalized to obtain a standardized image.

[0008] Optionally, step two specifically involves: Edge response detection and local texture scanning are performed on the standardized image to determine the candidate region where the target subject is located, and the search range of feature points is limited within the candidate region; Within the candidate region, AKAZE image feature matching is used to perform multi-scale feature point detection, extracting corner features and marker features that characterize the contour changes of the target subject, and removing invalid feature points located in the background region, low contrast region and repeated texture region to obtain an initial feature point set. The low contrast region is a region with a contrast less than a preset threshold. The initial feature point set is subjected to positional consistency screening and orientation consistency screening, specifically including: classifying the initial feature point set according to the boundary distribution relationship of the target body contour, identifying the body corner points corresponding to the outer contour of the target body or the body marker points corresponding to the preset marker area, and obtaining the corrected feature point set.

[0009] Optionally, step three specifically includes: Read the pre-established target front view corner point coordinate template, match each main corner point or main marker point in the correction feature point set with the corresponding target point in the target front view corner point coordinate template, and generate a feature correspondence set; For each set of matching samples in the feature correspondence set, two linear constraint equations are established according to the homography mapping relationship. All linear constraint equations are combined into a matrix equation and singular value decomposition is performed. The eigenvector corresponding to the smallest singular value is taken as the solution of the parameter vector in the matrix equation. The parameter vector is reconstructed into a matrix of a preset size to obtain the homography matrix from the standardized image coordinate system to the target front view coordinate system; The homography matrix is ​​subjected to mapping rationality verification to obtain a homography matrix that meets the perspective correction conditions. The mapping rationality verification is to eliminate mapping relationships with mapping offsets greater than a preset value, boundary flips, or local compression. A perspective transformation is performed on the standardized image based on the homography matrix that satisfies the perspective correction conditions. The perspective transformation maps the oblique perspective shape of the target subject from the perspective of the projector lens to the normal parallel view shape, and determines the effective pixel retention range after transformation according to the outer region of the target subject to obtain the corrected color image.

[0010] Optionally, step four specifically includes: The color channel information of each pixel in the corrected color image is converted into single-channel grayscale information to obtain a grayscale image; The grayscale image is divided into multiple adjacent local sub-regions according to a preset window size. The grayscale distribution in each local sub-region is truncated according to a preset contrast limit threshold. The truncated grayscale distribution is then uniformly mapped to obtain the enhancement result of each local sub-region. The enhancement results of each local sub-region are stitched together and smoothed to obtain a contrast-enhanced image; Gradient change analysis is performed on the contrast-enhanced image to determine the location of gray-level abrupt changes in the image. Canny edge detection processing is then performed based on preset high and low thresholds to extract the outer contour edges and internal structure edges of the target subject, thus obtaining an initial edge image. The initial edge image is subjected to edge enhancement processing, which involves connecting broken edges, removing isolated noise edges, and retaining continuous edge regions that are consistent with the contour direction of the target subject, to obtain an enhanced edge image. The enhanced edge image and the contrast-enhanced image are fused according to their positions to obtain a structure-enhanced image.

[0011] Optionally, the improved Mask2Former model is specifically as follows: The structure enhancement map is input into the structure representation module, and multi-layer convolutional scanning and downsampling processing are performed on the structure enhancement map to extract edge texture information, contour direction information and region morphology information under different receptive ranges, so as to obtain feature maps at different levels. Size alignment processing is performed on the feature maps obtained at different levels. The size alignment processing involves mapping the feature maps at each level to a unified feature representation scale, and then fusing the mapped feature maps at each level layer by layer to obtain a structural representation feature map. The structural representation feature map is input into the subject guidance module, and multiple subject query vectors are established in the structural representation feature map. Each subject query vector corresponds to a different candidate region response center of the target subject. Each subject query vector is paired with the positional features corresponding to each spatial location in the structural representation feature map to obtain the response intensity distribution of each subject query vector to different spatial locations. Based on the response intensity distribution of each spatial location, the location features with response intensity greater than a preset threshold are aggregated. The aggregation is to eliminate the location features with response intensity less than the preset threshold, and to separate the regional features related to the creative subject from the background interference features to obtain the subject guidance features. The subject guidance features obtained from the previous aggregation are used as the new subject query vector, and relevance calculation and aggregation are performed to obtain the updated subject guidance features; The updated subject guidance features are input into the probability generation module. The updated subject guidance features and the structural representation feature map are mapped position by position to obtain the response intensity corresponding to each pixel position. The ratio of the response intensity of each pixel belonging to the creative subject to the sum of the response intensities of all pixels in the current creative subject is used as the probability value corresponding to the current pixel to generate a pixel probability map. Regions with increased probability values ​​in the pixel probability map are identified as candidate regions for the main body, and regions with discrete and abrupt changes in probability are identified as regions with unstable boundaries, thus forming a preliminary segmentation result. The initial segmentation results are input into the boundary optimization module, and pixel association constraints in the conditional random field are constructed by combining the gray-level relationship between adjacent pixels, the edge continuity relationship and the spatial adjacency relationship in the structure enhancement map. Based on the pixel association constraint relationship, the pixels at the main body boundary in the preliminary segmentation result are relabeled to obtain the optimized pixel probability map; The optimized pixel probability map is subjected to threshold segmentation processing, which involves determining pixels with a pixel probability value greater than or equal to a preset threshold as the main subject pixels and pixels with a pixel probability value less than the preset threshold as non-main subject pixels, thereby generating a binary mask.

[0012] Optionally, the step of establishing multiple subject query vectors in the structural representation feature map, where each subject query vector corresponds to a different candidate region response center of the target subject, specifically involves: The structural characterization feature map is uniformly divided into multiple spatial grid regions, and the edge density, contour continuity, region closure and local response intensity of each spatial grid region are statistically analyzed. The edge density is the number of edge pixels per unit area within the current spatial grid region; The contour continuity is obtained by detecting the connection relationship between adjacent edge pixels along the edge direction within the current spatial grid region, counting the length of continuously connected edge segments and the number of break points, and weighting them according to a preset weight based on the length of continuously connected edge segments and the number of break points to obtain the contour continuity value of the current spatial grid region. The region closure degree is the percentage of the area within the current spatial grid region that is enclosed by the edges of the closed outline; The local response intensity is the average of the characteristic response values ​​of each region in the structural characterization feature map of the current spatial grid region; Candidate regions that meet the preset subject determination criteria are selected based on the edge density, contour continuity, region closure, and local response intensity of each spatial grid region. Extract the center position of each candidate region and map the center position of each candidate region to the feature coordinate space of the structural representation feature map to obtain the set of candidate center points; The positional features of each candidate center point in the structural representation feature map are used as initial seed features, and a set of local features within the neighborhood of each candidate center point is extracted. The local feature sets corresponding to each candidate center point are aggregated and calculated. The aggregated local features are then concatenated with the location encoding information of the candidate center points to generate the initial query representation corresponding to each candidate region. The initial query representations are subjected to dimensional unification and numerical normalization to obtain multiple subject query vectors.

[0013] Optionally, step six specifically includes: The image size information, pixel coordinate range, channel information, and storage format information of the corrected color image, structure enhancement image, and binary mask are read respectively to generate the corresponding image attribute set; Based on the image attribute set, coordinate reference consistency checks are performed on the corrected color image, the structure enhancement image, and the binary mask. If the coordinate reference is inconsistent, the pixel coordinate system of the corrected color image is used as the reference coordinate system, and coordinate remapping is performed on the structure enhancement image and the binary mask to obtain the aligned structure enhancement image and the aligned binary mask that maintain the pixel-level correspondence with the corrected color image. The corrected color image, the aligned structure enhancement image, and the aligned binary mask are encapsulated according to a preset data packet format to obtain a reference image data packet.

[0014] The beneficial effects of this invention are: The proposed method for generating a standard-perspective white model based on digital twins addresses the standardization requirements of the original image before it enters the digital twin modeling process. It constructs a continuous processing chain consisting of image standardization, feature point recognition, perspective correction, structural enhancement, subject segmentation, and reference data encapsulation, enabling the projector lens perspective image obtained at the acquisition end to be converted into standardized input data suitable for subsequent white model generation. Compared with existing technologies, this invention improves the consistency of the input image in terms of brightness, color, and data format by performing image decoding, color space conversion, white balance processing, noise suppression, and size normalization on the original image, thereby reducing the impact of device differences, illumination fluctuations, and local noise on subsequent processing results. By using AKAZE image feature matching in the candidate region for multi-scale feature point detection and solving the homography matrix by combining the correspondence between the coordinates of the subject corner points, subject marker points, and target front view corner points, a stable mapping from the oblique perspective view to the frontal parallel view is achieved. This ensures that the correction process can balance accurate restoration of the subject contour and preservation of effective pixels, improving the geometric reliability of the standard viewpoint construction. By performing grayscale conversion, local contrast enhancement, edge detection, and edge strengthening on the corrected color image, the outer contour information, internal structural boundaries, and geometric orientation features of the target subject are centrally expressed, weakening the impact of color texture and background interference on subject recognition.

[0015] By employing an improved Mask2Former model, structural representation, subject guidance, probabilistic generation, and boundary optimization are seamlessly integrated. This enables the segmentation process not only to identify the subject region but also to perform constraint optimization for unstable boundary regions, improving the performance of binary masks in terms of edge continuity, regional integrity, and preservation of local details. Pixel-level alignment is performed on the corrected color image, structural enhancement image, and binary mask, and the processing results are uniformly encapsulated with metadata into a reference image data package. This establishes a clear correspondence between various result data under the same coordinate reference, facilitating subsequent white model generation, digital twin mapping, and structural reuse. This invention effectively solves the problems of difficult perspective distortion elimination, unstable subject contour extraction, easy misalignment of segmentation boundaries, and inconsistent modeling reference data in existing technologies, improving the stability, accuracy, and usability of standard viewpoint white model generation. Attached Figure Description

[0016] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is an overall flowchart of the standard perspective white model generation method based on digital twin proposed in this invention; Figure 2 This is a schematic diagram of the improved Mask2Former model processing steps of the standard perspective white model generation method based on digital twin proposed in this invention; Figure 3 This is a flowchart illustrating the main query vector generation process of the standard perspective white model generation method based on digital twins proposed in this invention. Detailed Implementation

[0017] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0018] refer to Figures 1-3 The standard white model generation method based on digital twins includes the following steps: Step 1: Receive the original image captured from the projector lens, and perform image decoding, color space conversion, white balance and noise suppression on the original image to obtain a standardized image; Step 2: Use AKAZE image feature matching to perform multi-scale feature point detection on the standardized image, identify the main corner points or main marker points of the main body contour in the standardized image, and obtain the corrected feature point set; Step 3: Calculate the homography matrix based on the feature points in the correction feature point set and the preset corner coordinates of the target front view, and perform perspective transformation based on the homography matrix to obtain the corrected color image; Step 4: Convert the corrected color image to grayscale, enhance its contrast, and use the Canny edge detection operator to extract contours for edge enhancement, generating a structure enhancement image; Step 5: Input the structure enhancement graph into the improved Mask2Former model, and obtain the binary mask through the structure encoding module, the main body query module, the mask decoding module, and the graph cutting and refining module; Step 6: Perform pixel-level alignment on the corrected color image, structure enhancement image, and binary mask, and encapsulate metadata to generate the final baseline image data package.

[0019] This step addresses the requirements of data standardization, structural accuracy, and result reusability in generating standard-view white models within digital twin scenarios. It establishes a complete and tightly integrated benchmark image construction mechanism, significantly improving the stability and consistency of the conversion from actual captured images to digital spatial base models. By standardizing the input images, it effectively mitigates image representation differences caused by different acquisition devices, lighting conditions, and local noise disturbances, ensuring subsequent recognition and correction are based on more stable data, thus improving the reliability of the overall processing chain from the source. Targeted identification of key features of the subject's contour and establishment of a stable mapping relationship with the target's frontal view effectively corrects perspective shifts and morphological distortions caused by shooting angles. This results in stronger geometric regularity and perspective uniformity in the output, facilitating accurate representation of the subject's size, boundaries, and relative spatial relationships during subsequent white model construction. By enhancing structural information in the image and suppressing irrelevant textures and color interference, the subject's contour boundaries, transition areas, and internal structural layers are more clearly highlighted, improving the distinguishability between the subject and background areas and reducing the impact of complex backgrounds on result quality. By employing an improved segmentation mechanism to discriminate and refine the main region, the masking results demonstrate enhanced performance in terms of boundary continuity, local integrity, and detail preservation, reducing the adverse effects of edge breaks, local missing areas, and misclassified regions on the modeling results. By unifying the correction color image, structural enhancement image, and binary mask under the same pixel coordinate reference and encapsulating them with metadata, a clear, stable, and traceable correspondence can be established between various result data, facilitating direct retrieval in subsequent digital twin white model generation, contour reuse, model mapping, and data management. Therefore, this invention not only improves the accuracy and stability of standard-view white model generation but also enhances the standardization, data integrity, and engineering application value of the generated results, significantly contributing to improving the usability of front-end basic data for digital twin modeling.

[0020] In this embodiment, step one specifically includes: Receive the original image captured from the perspective of the projector lens, read the encoding header information, resolution information and channel information of the original image, and decode it according to the corresponding encoding format to obtain the decoded image; Check the channel arrangement order and bit depth format of the decoded image. If the channel arrangement order of the decoded image is the same as the channel arrangement order of the acquisition device, convert the decoded image into a preset unified channel arrangement order to obtain a unified channel image. A color space conversion is performed on the channel unified image, which converts the channel unified image from the device output color space to a preset processed color space to obtain a color-converted image. The pixel distribution of each color channel in the color-converted image is statistically analyzed. Based on the mean relationship of each color channel, the gain of each color channel is adjusted until the adjusted color channels reach the preset color balance condition, thus obtaining the white balance image. The white balance image is subjected to noise suppression processing, which involves eliminating isolated noise points, brightness jitter areas, and local particle areas using a neighborhood smoothing method, while preserving pixel changes at the subject edge position to obtain a noise-reduced image. The denoised image is scaled or padded to a preset input size, and the pixel value range is normalized to obtain a standardized image.

[0021] This step standardizes the preprocessing of the original image before it enters the subsequent correction and segmentation processes, significantly improving the uniformity, stability, and processability of the input data, providing a more reliable image foundation for the subsequent standard-view white model generation. By reading the encoded header information, resolution information, and channel information of the original image and decoding it according to the corresponding encoding format, data parsing errors caused by directly mixing different acquisition formats can be avoided, ensuring that the image content can enter the subsequent processing flow in the correct pixel organization manner. By uniformly converting the channel arrangement order and bit depth format, the differences in output format between different acquisition devices can be eliminated, ensuring that subsequent processing algorithms face a consistent image representation, thereby improving the compatibility and execution stability of the entire processing chain. By converting the device's output color space to a preset processing color space, the image color expression can fall under a unified standard, reducing the interference caused by device color deviations to subject recognition and edge judgment.

[0022] Based on this, by adjusting the gain according to the mean relationship of each color channel, the color imbalance caused by differences in ambient lighting, lens conditions, and projected light sources during the shooting process can be effectively corrected, making the overall brightness and color distribution of the image more coordinated, which is conducive to enhancing the distinguishability between the subject area and the background area. By using a neighborhood smoothing method to suppress isolated noise points, brightness jitter areas, and local grainy areas, while preserving the pixel changes at the subject edge position, the clarity of the subject outline and local structure can be maintained as much as possible while reducing noise interference, avoiding boundary blurring and detail loss caused by excessive smoothing. By performing size unification and pixel value normalization processing on the denoised image, images from different sources and resolutions can achieve a consistent state in spatial scale and numerical range, which is convenient for subsequent feature point detection, homography matrix solution, and structure enhancement processing. It can be seen that this invention not only improves the visual quality of the original image, but more importantly, it improves the input reliability of subsequent geometric correction, edge extraction, and subject segmentation, reduces false detections, false negatives, and processing deviations caused by data inconsistency, and has a significant effect on improving the accuracy, robustness, and engineering applicability of the entire standard view white model generation method.

[0023] In this embodiment, step two specifically includes: Edge response detection and local texture scanning are performed on the standardized image to determine the candidate region where the target subject is located, and the search range of feature points is limited within the candidate region; Within the candidate region, AKAZE image feature matching is used to perform multi-scale feature point detection, extracting corner features and marker features that characterize the contour changes of the target subject, and removing invalid feature points located in the background region, low contrast region and repeated texture region to obtain an initial feature point set. The low contrast region is a region with a contrast less than a preset threshold. The initial feature point set is subjected to positional consistency screening and orientation consistency screening, specifically including: classifying the initial feature point set according to the boundary distribution relationship of the target body contour, identifying the body corner points corresponding to the outer contour of the target body or the body marker points corresponding to the preset marker area, and obtaining the corrected feature point set.

[0024] This step addresses the issues of unstable subject positioning, susceptibility of feature points to background interference, and unreliable perspective correction during the standard viewpoint white model generation process. It specifically optimizes the target subject feature extraction stage, significantly improving the accuracy of subsequent geometric correction and the robustness of the overall processing chain. By pre-defining more representative candidate regions in the image, the scope of subsequent feature analysis is effectively narrowed, reducing interference from complex backgrounds, irrelevant textures, and edge noise on keypoint extraction. This allows the feature search process to focus more on areas related to the true contour of the target subject, thereby improving the relevance and effectiveness of feature detection results. Combining multi-scale extraction of corner and marker features corresponding to changes in the subject contour helps to simultaneously consider both the overall shape and local detail differences, resulting in more comprehensive feature information in terms of scale adaptability and spatial representation, avoiding the problems of detail omission or structural distortion that occur at a single scale.

[0025] During feature selection, excluding invalid feature points from background, low-contrast, and repetitive texture areas significantly reduces the adverse effects of false, weak, and confusing features on subsequent registration relationships, making the retained feature points more representative of contours and of geometric discriminative value. This approach not only improves the purity of the feature point set but also helps reduce mapping deviations, boundary folding, and local correction anomalies caused by incorrect matching. Furthermore, by imposing positional and directional consistency constraints on the initial feature point set, the spatial distribution of the finally identified subject corner points and subject marker points better matches the true boundary relationship of the target subject, and the directional changes are closer to the continuous direction of the subject contour, thereby enhancing the consistency between the feature point set and the geometric shape of the target subject. The resulting corrected feature point set has stronger stability, accuracy, and structural representativeness, providing a more reliable constraint basis for solving the homography matrix and improving the accuracy and repeatability of frontal view mapping. Overall, this invention enhances the anti-interference capability, structural expression capability, and feature effectiveness in the key feature extraction stage, and has significant technical effects on improving the subsequent perspective correction effect, enhancing the quality of standard viewpoint white model generation, and increasing the credibility of digital twin basic image data.

[0026] In this embodiment, step three specifically includes: Read the pre-established target front view corner point coordinate template, match each main corner point or main marker point in the correction feature point set with the corresponding target point in the target front view corner point coordinate template, and generate a feature correspondence set; For each set of matching samples in the feature correspondence set, two linear constraint equations are established according to the homography mapping relationship. All linear constraint equations are combined into a matrix equation and singular value decomposition is performed. The eigenvector corresponding to the smallest singular value is taken as the solution of the parameter vector in the matrix equation. The parameter vector is reconstructed into a matrix of a preset size to obtain the homography matrix from the standardized image coordinate system to the target front view coordinate system; The homography matrix is ​​subjected to mapping rationality verification to obtain a homography matrix that meets the perspective correction conditions. The mapping rationality verification is to eliminate mapping relationships with mapping offsets greater than a preset value, boundary flips, or local compression. A perspective transformation is performed on the standardized image based on the homography matrix that satisfies the perspective correction conditions. The perspective transformation maps the oblique perspective shape of the target subject from the perspective of the projector lens to the normal parallel view shape, and determines the effective pixel retention range after transformation according to the outer region of the target subject to obtain the corrected color image.

[0027] The corrected color image is subjected to boundary integrity checks and subject proportion checks. When the corrected color image meets the conditions of complete subject outline, continuous main structure and effective pixel area reaching the preset retention conditions, the corrected color image is subjected to structural enhancement processing.

[0028] This step involves matching the set of corrected feature points with a pre-established target frontal view corner coordinate template, and then solving the homography matrix accordingly. This constructs a geometric correction mechanism that stably maps the captured viewpoint image to a standard frontal view, effectively improving the spatial regularity and structural consistency of the front-end image generated from the digital twin white model. Compared to existing technologies that rely on coarse perspective adjustments or empirical corrections, this invention utilizes the correspondence between feature points to establish explicit geometric constraints, giving the image mapping process stronger mathematical certainty and higher correction accuracy. This helps reduce positional shifts, contour stretching, and proportional imbalances in the target subject during viewpoint transitions. By solving the matrix equations formed by all linear constraint equations, the generation of the homography matrix is ​​based on a global matching relationship, rather than local approximations, thereby improving the global stability of the mapping relationship and its adaptability to complex shooting postures.

[0029] By performing mapping rationality checks on the homography matrix and eliminating abnormal relationships such as excessive mapping offsets, boundary folding, or local compression, problems such as subject shape distortion, boundary misalignment, and local area distortion in the perspective correction results can be further avoided, making the retained mapping results more consistent with the real spatial structure of the target subject. This processing method not only improves the accuracy of the corrected color image in contour representation but also enhances its reliability in terms of geometric dimensions, boundary continuity, and subject shape integrity. By converting the tilted perspective form into a normal parallel view form based on the homography matrix that meets the conditions, the perspective error caused by the projector lens angle can be significantly reduced, allowing the target subject to obtain a more standardized perspective representation in the output image that is more suitable as a white model base. This is beneficial for maintaining a unified reference benchmark in subsequent structural enhancement, subject segmentation, and model construction operations. Furthermore, by checking the boundary integrity and subject proportion of the corrected color image, the result quality can be reconfirmed before entering subsequent processing, preventing incomplete images or images with insufficient effective subject areas from continuing into subsequent processes, thereby improving the reliability and stability of the entire processing chain. Overall, this invention achieves a complete closed loop from feature constraints and mapping solutions to result verification in the geometric correction stage, significantly improving the problems of unstable standard viewpoint recovery, distorted correction results, and insufficient retention of effective pixels in the prior art. It plays a significant role in improving the accuracy, continuity, and engineering practical value of standard viewpoint white model generation.

[0030] In this embodiment, step four specifically includes: The color channel information of each pixel in the corrected color image is converted into single-channel grayscale information to obtain a grayscale image; The grayscale image is divided into multiple adjacent local sub-regions according to a preset window size. The grayscale distribution in each local sub-region is truncated according to a preset contrast limit threshold. The truncated grayscale distribution is then uniformly mapped to obtain the enhancement result of each local sub-region. The enhancement results of each local sub-region are stitched together and smoothed to obtain a contrast-enhanced image; Gradient change analysis is performed on the contrast-enhanced image to determine the location of gray-level abrupt changes in the image. Canny edge detection processing is then performed based on preset high and low thresholds to extract the outer contour edges and internal structure edges of the target subject, thus obtaining an initial edge image. The initial edge image is subjected to edge enhancement processing, which involves connecting broken edges, removing isolated noise edges, and retaining continuous edge regions that are consistent with the contour direction of the target subject, to obtain an enhanced edge image. The enhanced edge image and the contrast-enhanced image are fused according to their positions, so that the geometric contour information and local structural hierarchy information of the target subject are consistently expressed in the same image, resulting in a structural enhancement image.

[0031] This step establishes an enhancement mechanism oriented towards the geometric structure representation of the target subject by performing grayscale conversion, local contrast enhancement, edge extraction, edge strengthening, and structural fusion processing on the corrected color image. This significantly improves the recognition quality of the subject's outline and internal structure during the standard-view white model generation process. Compared to directly using the corrected color image for subsequent segmentation or modeling, this invention converts color information into a grayscale representation that is more conducive to structural analysis. This effectively reduces the interference caused by color differences, surface textures, and lighting changes on the judgment of the subject's outline, allowing subsequent processing to focus more on the morphological boundaries and structural levels of the target subject itself. By locally partitioning the grayscale image according to a preset window size and implementing controlled enhancement of each local region in conjunction with a contrast limiting threshold, the problems of uneven brightness distribution, insufficient detail contrast, and buried edge information in local areas of the image can be improved. This results in clearer grayscale level representation at subject outline transitions, boundary intersections, and areas of local structural change, thereby improving the discernibility of the target subject in complex backgrounds or low-contrast environments.

[0032] In the edge information extraction stage, gradient change analysis is performed on the enhanced image to extract the outer contour edges and internal structural edges. This allows for the concentrated expression of key locations reflecting the geometric features of the subject, helping to further distinguish the subject boundary from the background region. By enhancing the initial edge image by connecting broken edges, removing isolated noise edges, and retaining continuous edge regions with consistent orientation, the adverse effects of edge breaks, false edges, and scattered noise on subsequent recognition results can be reduced. This makes the final retained edge information more consistent with the true structural features of the target subject in terms of continuity, integrity, and directional consistency. This approach not only improves the stable representation of the outer contour but also enhances the presentation of internal structural levels and local geometric relationships. Furthermore, by positionally fusing the enhanced edge image with the contrast-enhanced image, contour information and regional hierarchical information are consistently expressed in the same image. This simultaneously considers the overall shape and local structural details of the target subject, avoiding the one-sided expression problem caused by relying solely on edge information or solely on grayscale enhancement information. The resulting enhanced structure map shows significant improvements in the clarity of the subject boundary, the distinguishability of the internal structure, and the ability to resist background interference. It provides a more stable, accurate, and structurally representative input basis for subsequent subject segmentation, mask generation, and white model benchmark data construction, and has a significant effect on improving the accuracy and usability of the entire digital twin standard perspective white model generation method.

[0033] In this embodiment, the improved Mask2Former model is specifically as follows: The structure enhancement map is input into the structure representation module, and multi-layer convolutional scanning and downsampling processing are performed on the structure enhancement map to extract edge texture information, contour direction information and region morphology information under different receptive ranges, so as to obtain feature maps at different levels. Size alignment processing is performed on the feature maps obtained at different levels. The size alignment processing is to map the feature maps at each level to a unified feature expression scale, and then fuse the mapped feature maps at each level layer by layer, so that the outer contour features, local turning features and internal connectivity features of the target subject are aggregated in a unified feature space to obtain a structural representation feature map. The structural representation feature map is input into the subject guidance module, and multiple subject query vectors are established in the structural representation feature map. Each subject query vector corresponds to a different candidate region response center of the target subject. Each subject query vector is paired with the positional features corresponding to each spatial location in the structural representation feature map to obtain the response intensity distribution of each subject query vector to different spatial locations. Based on the response intensity distribution of each spatial location, the location features with response intensity greater than a preset threshold are aggregated. The aggregation is to eliminate the location features with response intensity less than the preset threshold, and to separate the regional features related to the creative subject from the background interference features to obtain the subject guidance features. The subject guidance features obtained from the previous aggregation are used as the new subject query vector, and relevance calculation and aggregation are performed to obtain the updated subject guidance features; The updated subject guidance features are input into the probability generation module. The updated subject guidance features and the structural representation feature map are mapped position by position to obtain the response intensity corresponding to each pixel position. The ratio of the response intensity of each pixel belonging to the creative subject to the sum of the response intensities of all pixels in the current creative subject is used as the probability value corresponding to the current pixel to generate a pixel probability map. Regions with increased probability values ​​in the pixel probability map are identified as candidate regions for the main body, and regions with discrete and abrupt changes in probability are identified as regions with unstable boundaries, thus forming a preliminary segmentation result. The initial segmentation results are input into the boundary optimization module, and pixel association constraints in the conditional random field are constructed by combining the gray-level relationship between adjacent pixels, the edge continuity relationship and the spatial adjacency relationship in the structure enhancement map. Based on the pixel association constraint relationship, the pixels at the main body boundary in the preliminary segmentation result are relabeled to obtain the optimized pixel probability map; Based on the pixel association constraint relationship, the blurred pixels at the main boundary, the broken pixels at the narrow connection and the abnormal pixels in the local misclassified area in the preliminary segmentation result are relabeled so that pixels with the same attributes in adjacent areas tend to maintain the same label, and the optimized pixel probability map is obtained. The optimized pixel probability map is subjected to threshold segmentation processing, which involves determining pixels with pixel probability values ​​greater than or equal to a preset threshold as the main subject pixels and pixels with pixel probability values ​​less than the preset threshold as non-main subject pixels, thereby generating a binary mask. The binary mask is subjected to hole inspection and isolated area inspection. The hole inspection and isolated area inspection involve filling the hollow areas located inside the main body and removing isolated small areas that are detached from the main body area.

[0034] The improved Mask2Former model proposed in this step is similar to the traditional Mask2Former model in that both aim at image segmentation and follow the overall processing approach of "feature extraction—subject representation—mask generation—result refinement". Both require multi-level feature extraction of the input image, followed by the establishment of a subject representation in the feature space. Then, pixel-level segmentation results are generated based on the correspondence between the subject representation and image features. Finally, the segmentation boundaries are optimized to obtain a more stable mask output. The improved model also retains the core task attributes of pixel-level segmentation in the traditional Mask2Former, still focusing on subject region recognition, boundary localization, and mask output. It also exhibits the same basic functions: multi-scale feature processing, subject query guidance, pixel probability generation, and post-processing refinement. Therefore, this improved solution does not deviate from the traditional Mask2Former technology route to reconstruct a completely independent segmentation framework. Instead, it continues the original segmentation task chain and basic functional structure, and makes targeted enhancements to meet the needs of main structure expression, candidate region selection and boundary correction in the scenario of generating white models from the standard perspective of digital twins. This maintains consistency with the traditional Mask2Former in terms of overall purpose, basic processing direction and core segmentation objectives.

[0035] The difference lies in the fact that the improved Mask2Former model proposed in this step does not simply perform conventional segmentation on the input image. Instead, it constructs a more explicit modular processing chain around the specific input form of the structure enhancement map, and refines and constrains the process of establishing the subject query vector. Traditional Mask2Former usually focuses more on query-driven mask prediction under a general segmentation framework, while this scheme emphasizes the joint extraction and unified scale fusion of edge texture information, contour direction information, and regional morphological information in the structure representation module, making the input features more biased towards the geometric structure of the subject. In the subject guidance module, instead of directly using a preset query, the structure representation feature map is first divided into spatial grid regions, and then candidate regions are screened, adjacent regions are merged, and abnormal regions are removed by combining edge density, contour continuity, region closure, and local response intensity. Only then is the subject query vector formed, making the query establishment process directly related to the geometric features of the subject, regional connectivity, and response stability. In subsequent processing, structural information, spatial relationships, and label correction are further coupled through position-wise pairing of response intensity distribution, iterative updating of subject guidance features, and boundary optimization under conditional random field constraints. In other words, this improved model is more specific and more geared towards white model generation tasks than the traditional Mask2Former in terms of input feature organization, source of main query vector, query update mechanism, and boundary optimization basis.

[0036] The improvements enhance the segmentation model's ability to recognize the structural contours, local transitions, and regional connectivity of the target subject, thereby improving the subject extraction quality during the standard-perspective white model generation process of digital twins. By first performing multi-level structural representation on the structure enhancement map, and then filtering candidate regions based on edge density, contour continuity, region closure, and local response intensity, the establishment of the subject query vector no longer relies on a broad global response but is based on regions that better match the true geometric distribution of the subject. This helps reduce the impact of background interference, repetitive textures, and irrelevant regions on the segmentation results. By aggregating positional features with response intensities greater than a preset threshold and continuously updating the query vector with the aggregated subject-guided features, the model's ability to focus on the core region and boundary transition regions of the subject is enhanced, improving the retention of complex contours, narrow connections, and locally weak structural regions. Furthermore, by combining conditional random fields to construct pixel association constraints and relabeling blurred boundaries, broken pixels, and locally misclassified regions, the continuity of mask boundaries, regional integrity, and pixel attribution consistency can be further improved. Therefore, this improved scheme can effectively improve the problems of rough boundaries, local breaks and incomplete main areas that are prone to occur in the traditional segmentation model in the white model generation scenario, and improve the accuracy, stability and subsequent usability of binary masking.

[0037] In this embodiment, the step of establishing multiple subject query vectors in the structural representation feature map, where each subject query vector corresponds to a different candidate region response center of the target subject, specifically involves: The structural characterization feature map is uniformly divided into multiple spatial grid regions, and the edge density, contour continuity, region closure and local response intensity of each spatial grid region are statistically analyzed. The edge density is the number of edge pixels per unit area within the current spatial grid region; The contour continuity is obtained by detecting the connection relationship between adjacent edge pixels along the edge direction within the current spatial grid region, counting the length of continuously connected edge segments and the number of break points, and weighting them according to a preset weight based on the length of continuously connected edge segments and the number of break points to obtain the contour continuity value of the current spatial grid region. The region closure degree is the percentage of the area within the current spatial grid region that is enclosed by the edges of the closed outline; The local response intensity is the average of the characteristic response values ​​of each region in the structural characterization feature map of the current spatial grid region; The edge density, contour continuity, region closure, and local response intensity of each spatial grid region are used as the region determination feature group for the corresponding spatial grid region; Each spatial grid region's region determination feature group is compared with the preset subject determination conditions. Spatial grid regions with edge density values ​​greater than preset edge thresholds, contour continuity values ​​greater than preset continuity thresholds, region closure values ​​greater than preset closure thresholds, and local response intensity values ​​greater than preset response thresholds are retained to obtain the initial screening region set. Adjacency detection is performed on adjacent spatial grid regions in the initial screening region set. Adjacent spatial grid regions with common boundaries and whose region determination feature group changes do not exceed a preset difference threshold are merged to obtain a connected candidate region set. For each connected candidate region in the connected candidate region set, perform area and shape checks, remove abnormal regions with an area smaller than a preset area threshold, an aspect ratio exceeding a preset range, or an overly discrete internal edge distribution, and retain regions that meet the subject determination requirements to obtain candidate regions; For each candidate region, extract the region center location, region range information, and feature aggregation value within the region, and map the region center location corresponding to each candidate region to the feature coordinate space of the structural representation feature map to obtain the candidate center point set; The positional features of each candidate center point in the structural representation feature map are used as initial seed features, and a set of local features within the neighborhood of each candidate center point is extracted. The local feature sets corresponding to each candidate center point are aggregated and calculated. The aggregated local features are then concatenated with the location encoding information of the candidate center points to generate the initial query representation corresponding to each candidate region. The initial query representations are subjected to dimensional unification and numerical normalization to obtain multiple subject query vectors.

[0038] In this embodiment, step six specifically includes: The image size information, pixel coordinate range, channel information, and storage format information of the corrected color image, structure enhancement image, and binary mask are read respectively to generate the corresponding image attribute set; Based on the image attribute set, coordinate reference consistency checks are performed on the corrected color image, the structure enhancement image, and the binary mask to determine whether the image width, image height, pixel start coordinates, and pixel arrangement order of the three are consistent, and the consistency check results are obtained. When the consistency check results show that there are differences in size, coordinate offset or pixel arrangement among the three, the pixel coordinate system of the corrected color image is used as the reference coordinate system, and coordinate remapping is performed on the structure enhancement image and the binary mask to obtain the first alignment image and the second alignment image. The boundary positions of the first alignment image and the second alignment image are checked. The overlap between the structural edge position in the first alignment image and the main body outline position in the corrected color image is detected, and the overlap between the mask boundary position in the second alignment image and the main body area boundary in the corrected color image is detected, so as to obtain the boundary check result. Based on the boundary verification results, local offset correction processing is performed on the first alignment map and the second alignment map. The local areas where the boundary deviation exceeds the preset tolerance range are finely adjusted in pixel position according to the corresponding boundary direction to obtain an alignment structure enhancement map and an alignment binary mask that maintain a pixel-level correspondence with the corrected color image. The corrected color image, the aligned structure enhancement image, and the aligned binary mask are associated and encapsulated according to a unified file organization rule, and respectively assigned corresponding data identifiers, image type identifiers, and association index identifiers to obtain the aligned image group; Extract the processing metadata corresponding to the aligned image group. The processing metadata includes the original image identifier, normalization processing parameters, geometric correction parameters, structural enhancement parameters, segmentation processing parameters, pixel coordinate reference information, and image size information to obtain a metadata set. The aligned image group and metadata set are encapsulated according to a preset data packet format, and unified data packet header information, index information and content description information are written to generate a baseline image data packet.

[0039] This step ensures that the corrected color image, structural enhancement image, and binary mask establish a stable correspondence under a unified pixel coordinate benchmark, reducing the impact of size differences, coordinate offsets, and boundary misalignments on subsequent white model generation. Through boundary verification, local fine-tuning, and unified metadata encapsulation, the consistency, traceability, and reusability of multi-source image results can be improved, making the generated benchmark image data package more suitable for standard perspective modeling in digital twin scenarios.

[0040] Example 1: To verify the feasibility of this invention in practice, it was applied to the digital twin content production scenario of a digital exhibition production center in a certain city. Staff needed to quickly convert a set of three-dimensional display subjects with regular outer contours and clear boundary features in the exhibition hall into standard-view white models for subsequent digital space mapping, virtual scene pre-placement, and interactive projection debugging. The original workflow of this production center mainly relied on manual image cropping, manual perspective correction, manual outlining, and subsequent re-segmentation to generate the basic white model image. This method often resulted in incomplete subject edge recognition, localized proportional distortion after perspective correction, significant mask boundary jitter, and inconsistencies in coordinates between different processing results when there were large changes in shooting angle, uneven lighting, localized reflections on the subject surface, or decorative textures in the background. Especially when the image results needed to be directly fed into the digital twin engine for white model mapping, if a strict pixel-level correspondence was not established between the correction color image, structural enhancement image, and mask image, subsequent mesh generation, contour fitting, and spatial positioning would accumulate errors, leading to a significant deviation between the virtual model and the real target.

[0041] In this scenario, the on-site image acquisition device was fixed near the projector, with the lens facing the main subject. The acquired raw image resolution was 1920×1080, containing the target subject, ground reflections, background wall textures, and side lighting. After the raw image entered the processing flow, the image's encoding header, resolution, and channel information were first read and decoded. Then, the channel arrangement order was unified, and the device's output color space was converted to a preset processing color space. Due to the simultaneous presence of overhead lighting and side lighting, the mean values ​​of different color channels had significant deviations. Therefore, gain adjustments were made to each color channel during the white balance stage to correct the color imbalance between bright and dark areas on the subject's surface. Subsequently, neighborhood smoothing noise reduction was performed on the white balance image to eliminate isolated noise and localized grain noise, while preserving as much of the grayscale variation characteristics as possible at the subject's edges. After processing, the image was uniformly scaled to 1536×1024, and the pixel value range was normalized to generate a standardized image. Through this process, the input image achieves uniformity in color representation, spatial size, and noise level, providing a stable input for subsequent geometric correction.

[0042] Subsequently, edge response detection and local texture scanning were performed on the standardized image. Based on the approximate distribution of the subject within the image, candidate subject regions were first determined. Then, AKAZE image feature matching was used within these candidate regions for multi-scale feature point detection. The displayed subject has a relatively clear outer contour transition and a small number of auxiliary marker areas, thus allowing for the extraction of relatively stable subject corner points and subject marker points at different scales. Invalid feature points generated in the background wall with repetitive textures, low-contrast shadow areas, and ground reflection areas were removed using filtering rules. The remaining initial feature point set was then filtered for positional and directional consistency to identify the subject corner points truly corresponding to the target subject's outer contour and the subject marker points corresponding to the preset marker areas, forming a corrected feature point set. This set was then matched one-to-one with a pre-established target front view corner point coordinate template to establish a feature correspondence set. A matrix equation was constructed using linear constraint equations to complete singular value decomposition, yielding a homography matrix from the standardized image coordinate system to the target front view coordinate system. The homography matrix is ​​then used to perform a perspective transformation on the image, mapping the originally tilted subject image to a normal, parallel view. Simultaneously, the effective pixel range of the subject's outer region is preserved to generate a corrected color image. For cases with excessive mapping offset or abnormal local compression, the system directly removes the corresponding mapping relationship, ensuring the final correction result remains stable in terms of boundaries and scale.

[0043] After obtaining the corrected color image, it is converted to grayscale and divided into multiple adjacent local sub-regions according to a preset window size. For each local sub-region, a restricted contrast enhancement method is used to truncate and equalize the grayscale distribution. Then, the enhancement results are stitched together and a smooth transition is applied to obtain a contrast-enhanced image. Because there are shadow transition areas on one side of the subject and bright reflection areas on the other side in the on-site acquisition environment, ordinary global enhancement methods easily cause overexposure in bright areas and loss of detail in dark areas. However, local restricted enhancement can improve the visibility of dark area contours while suppressing excessive brightness amplification. Then, gradient change analysis and Canny edge detection are used to extract the outer contour edges and internal structural edges to generate an initial edge image. Further edge enhancement processing is then performed to connect broken edges, remove isolated noise edges, and retain only continuous edge regions consistent with the subject contour direction. Finally, the enhanced edge image and the contrast-enhanced image are fused according to their positions to form a structural enhancement map. This structural enhancement map exhibits stronger contour recognizability in practical applications, providing clearer structural representations of subject boundaries, corners, connections, and local concave areas.

[0044] In the main body segmentation stage, the structure enhancement map is input into the improved Mask2Former model. The model first performs convolutional scanning and downsampling on the structure enhancement map to extract edge texture information, contour direction information, and region morphology information at different levels. Then, it performs size alignment and layer-by-layer fusion on the features at different levels to form a structure representation feature map. Subsequently, multiple main body query vectors are established in the feature map, corresponding to the response centers of different candidate regions of the target main body. The response intensity distribution of each main body query vector to spatial position is obtained by positional pairing. Position features above a preset threshold are aggregated to obtain main body guiding features, which are then iteratively updated and input into the probability generation process to form a pixel probability map. For unstable boundary regions in the probability map, a conditional random field constraint is constructed by combining the gray-level relationship of adjacent pixels, edge continuity relationship, and spatial adjacency relationship in the structure enhancement map. Labels are reassigned to boundary pixels, and finally, a binary mask is obtained through threshold segmentation. This binary mask can better preserve the integrity of the main body region and significantly reduce edge burrs, local breaks, and internal holes when used for subsequent white model contour generation.

[0045] After segmentation, the image size, pixel coordinate range, channel information, and storage format information of the corrected color image, structural enhancement image, and binary mask are uniformly read, and a coordinate reference consistency check is performed. If local deviations exist, the pixel coordinate system of the corrected color image is used as the reference to perform coordinate remapping on the structural enhancement image and binary mask, resulting in an alignment result that strictly maintains the pixel-level correspondence with the corrected color image. Finally, the three types of images and the corresponding metadata are encapsulated into a reference image data package and sent directly to the digital twin white model generation engine. After the engine calls this data package, manual retouching and alignment are no longer required, and the white model contour extraction, base surface generation, and scene mapping can be completed. On-site technicians reported that after adopting the method of this invention, the processing flow that originally required switching between multiple software programs has been compressed into a unified flow, and the consistency and reusability of the generated results have been significantly improved, especially in multi-subject continuous processing tasks, where the stability between batches of results has been significantly improved.

[0046] To verify the practical effect of this invention, 60 sets of projector lens perspective images were selected as test samples within the same production center. These samples varied in subject size, boundary complexity, background texture intensity, and lighting balance. The method of this invention was compared with the center's previous manual perspective correction and conventional edge segmentation process. Evaluation indicators included standard perspective correction error, subject outline integrity rate, mask cross-union ratio, boundary continuity rate, number of isolated misclassified regions, processing time, effective pixel retention rate, subsequent white model bonding deviation, and the success rate of direct data package retrieval. The test locations were the center's digital modeling experimental area and joint debugging demonstration area. All samples were processed under the same hardware environment, with the computing device being a workstation equipped with an independent graphics processing unit. The results are shown in the table below.

[0047] Table 1. Comparison of Standard View White Model Baseline Image Generation Results ; As shown in Table 1, the method of this invention outperforms the original process in terms of geometric correction accuracy, contour extraction quality, and data availability. The average error of standard viewpoint correction decreased from 5.8 pixels to 1.9 pixels, indicating that after solving the homography matrix through feature correspondence and performing rationality verification, the geometric shape of the target subject in the frontal view can be recovered more accurately. The subject contour integrity rate increased from 88.6% to 97.4%, and the boundary continuity rate increased from 81.3% to 95.1%, indicating that after structural enhancement and conditional random field optimization, the continuity and integrity of the subject boundary are significantly improved, especially the preservation of narrow connecting areas and edge transition areas is more stable. The number of isolated misclassified areas and internal holes decreased significantly, also indicating that the present invention has a stronger suppression effect on background interference and local misclassification.

[0048] From the practical application results of white model generation, the method of this invention shows superior performance in terms of effective pixel retention rate and white model contour fitting deviation. The effective pixel retention rate is increased to 93.8%, indicating that the effective area corresponding to the subject can be more fully preserved during perspective transformation, reducing information loss caused by local compression and boundary clipping. The average deviation of white model contour fitting is reduced from 7.6mm to 2.8mm, indicating that the corrected color image, structural enhancement image, and binary mask after pixel-level alignment can provide a more accurate spatial reference basis for subsequent white model construction. At the same time, the success rate of direct call to the reference image data package reaches 98.3%, and the rework rate is reduced to 5.0%, indicating that the data generated by this invention not only has higher result quality, but also stronger compatibility, stability, and engineering implementation capabilities in subsequent system calls.

[0049] In practical applications at this digital exhibition production center, this invention effectively solves the problems of inconsistent input image standards, unstable perspective correction, easily broken boundary extraction, numerous misclassifications of subject masking, and lack of pixel-level correspondence between multiple results in the original methods. The baseline image data package generated by the method of this invention can directly serve the construction of digital twin white models, virtual model mapping, and subsequent spatial configuration processing, which not only improves processing efficiency but also enhances the accuracy of white model generation and engineering reusability, meeting the requirements of stability, accuracy, and batch processing capabilities in actual production environments.

[0050] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A standard-perspective white model generation method based on digital twins, characterized in that, The steps include the following: Step 1: Receive the original image captured from the projector lens, and perform image decoding, color space conversion, white balance and noise suppression on the original image to obtain a standardized image; Step 2: Use AKAZE image feature matching to perform multi-scale feature point detection on the standardized image, identify the main corner points or main marker points of the main body contour in the standardized image, and obtain the corrected feature point set; Step 3: Calculate the homography matrix based on the feature points in the correction feature point set and the preset target front view corner coordinates, and perform perspective transformation based on the homography matrix to obtain the corrected color image; Step 4: Convert the corrected color image to grayscale, enhance its contrast, and use the Canny edge detection operator to extract contours for edge enhancement, generating a structure enhancement image; Step 5: Input the structure enhancement graph into the improved Mask2Former model, and obtain the binary mask through the structure encoding module, the main body query module, the mask decoding module, and the graph cutting and refining module; Step 6: Perform pixel-level alignment on the corrected color image, structure enhancement image, and binary mask, and encapsulate metadata to generate the final baseline image data package.

2. The standard perspective white model generation method based on digital twins according to claim 1, characterized in that, Step one specifically involves: Receive the original image captured from the perspective of the projector lens, read the encoding header information, resolution information and channel information of the original image, and decode it according to the corresponding encoding format to obtain the decoded image; Check the channel arrangement order and bit depth format of the decoded image. If the channel arrangement order of the decoded image is the same as the channel arrangement order of the acquisition device, convert the decoded image into a preset unified channel arrangement order to obtain a unified channel image. A color space conversion is performed on the channel unified image, which converts the channel unified image from the device output color space to a preset processed color space to obtain a color-converted image. The pixel distribution of each color channel in the color-converted image is statistically analyzed. Based on the mean relationship of each color channel, the gain of each color channel is adjusted until the adjusted color channels reach the preset color balance condition, thus obtaining the white balance image. The white balance image is subjected to noise suppression processing, which involves eliminating isolated noise points, brightness jitter areas, and local particle areas using a neighborhood smoothing method, while preserving pixel changes at the subject edge position to obtain a noise-reduced image. The denoised image is scaled or padded to a preset input size, and the pixel value range is normalized to obtain a standardized image.

3. The standard perspective white model generation method based on digital twins according to claim 1, characterized in that, Step two specifically involves: Edge response detection and local texture scanning are performed on the standardized image to determine the candidate region where the target subject is located, and the search range of feature points is limited within the candidate region; Within the candidate region, AKAZE image feature matching is used to perform multi-scale feature point detection, extracting corner features and marker features that characterize the contour changes of the target subject, and removing invalid feature points located in the background region, low contrast region and repeated texture region to obtain an initial feature point set. The low contrast region is a region with a contrast less than a preset threshold. The initial feature point set is subjected to positional consistency screening and orientation consistency screening, specifically including: classifying the initial feature point set according to the boundary distribution relationship of the target body contour, identifying the body corner points corresponding to the outer contour of the target body or the body marker points corresponding to the preset marker area, and obtaining the corrected feature point set.

4. The standard perspective white model generation method based on digital twins according to claim 1, characterized in that, Step three specifically involves: Read the pre-established target front view corner point coordinate template, match each main corner point or main marker point in the correction feature point set with the corresponding target point in the target front view corner point coordinate template, and generate a feature correspondence set; For each set of matching samples in the feature correspondence set, two linear constraint equations are established according to the homography mapping relationship. All linear constraint equations are combined into a matrix equation and singular value decomposition is performed. The eigenvector corresponding to the smallest singular value is taken as the solution of the parameter vector in the matrix equation. The parameter vector is reconstructed into a matrix of a preset size to obtain the homography matrix from the standardized image coordinate system to the target front view coordinate system; The homography matrix is ​​subjected to mapping rationality verification to obtain a homography matrix that meets the perspective correction conditions. The mapping rationality verification is to eliminate mapping relationships with mapping offsets greater than a preset value, boundary flips, or local compression. A perspective transformation is performed on the standardized image based on the homography matrix that satisfies the perspective correction conditions. The perspective transformation maps the oblique perspective shape of the target subject from the perspective of the projector lens to the normal parallel view shape, and determines the effective pixel retention range after transformation according to the outer region of the target subject to obtain the corrected color image.

5. The standard perspective white model generation method based on digital twins according to claim 1, characterized in that, Step four specifically involves: The color channel information of each pixel in the corrected color image is converted into single-channel grayscale information to obtain a grayscale image; The grayscale image is divided into multiple adjacent local sub-regions according to a preset window size. The grayscale distribution in each local sub-region is truncated according to a preset contrast limit threshold. The truncated grayscale distribution is then uniformly mapped to obtain the enhancement result of each local sub-region. The enhancement results of each local sub-region are stitched together and smoothed to obtain a contrast-enhanced image; Gradient change analysis is performed on the contrast-enhanced image to determine the location of gray-level abrupt changes in the image. Canny edge detection processing is then performed based on preset high and low thresholds to extract the outer contour edges and internal structure edges of the target subject, thus obtaining an initial edge image. The initial edge image is subjected to edge enhancement processing, which involves connecting broken edges, removing isolated noise edges, and retaining continuous edge regions that are consistent with the contour direction of the target subject, to obtain an enhanced edge image. The enhanced edge image and the contrast-enhanced image are fused according to their positions to obtain a structure enhancement map.

6. The standard perspective white model generation method based on digital twins according to claim 1, characterized in that, The improved Mask2Former model is specifically as follows: The structure enhancement map is input into the structure representation module, and multi-layer convolutional scanning and downsampling processing are performed on the structure enhancement map to extract edge texture information, contour direction information and region morphology information under different receptive ranges, so as to obtain feature maps at different levels. Size alignment processing is performed on the feature maps obtained at different levels. The size alignment processing involves mapping the feature maps at each level to a unified feature representation scale, and then fusing the mapped feature maps at each level layer by layer to obtain a structural representation feature map. The structural representation feature map is input into the subject guidance module, and multiple subject query vectors are established in the structural representation feature map. Each subject query vector corresponds to a different candidate region response center of the target subject. Each subject query vector is paired positionally with the positional features corresponding to each spatial location in the structural representation feature map to obtain the response intensity distribution of each subject query vector to different spatial locations. Based on the response intensity distribution of each spatial location, the location features with response intensity greater than a preset threshold are aggregated. The aggregation is to eliminate the location features with response intensity less than the preset threshold, and to separate the regional features related to the creative subject from the background interference features to obtain the subject guidance features. The subject guidance features obtained from the previous aggregation are used as the new subject query vector, and relevance calculation and aggregation are performed to obtain the updated subject guidance features; The updated subject guidance features are input into the probability generation module. The updated subject guidance features and the structural representation feature map are mapped position by position to obtain the response intensity corresponding to each pixel position. The ratio of the response intensity of each pixel belonging to the creative subject to the sum of the response intensities of all pixels in the current creative subject is used as the probability value corresponding to the current pixel to generate a pixel probability map. Regions with increased probability values ​​in the pixel probability map are identified as candidate regions for the main body, and regions with discrete and abrupt changes in probability are identified as regions with unstable boundaries, thus forming a preliminary segmentation result. The initial segmentation results are input into the boundary optimization module, and pixel association constraints in the conditional random field are constructed by combining the gray-level relationship between adjacent pixels, the edge continuity relationship and the spatial adjacency relationship in the structure enhancement map. Based on the pixel association constraint relationship, the pixels at the main body boundary in the preliminary segmentation result are relabeled to obtain the optimized pixel probability map; The optimized pixel probability map is subjected to threshold segmentation processing, which involves determining pixels with a pixel probability value greater than or equal to a preset threshold as the main subject pixels and pixels with a pixel probability value less than the preset threshold as non-main subject pixels, thereby generating a binary mask.

7. The standard perspective white model generation method based on digital twins according to claim 6, characterized in that, The process of establishing multiple subject query vectors in the structural representation feature map, where each subject query vector corresponds to a different candidate region response center of the target subject, is as follows: The structural characterization feature map is uniformly divided into multiple spatial grid regions, and the edge density, contour continuity, region closure and local response intensity of each spatial grid region are statistically analyzed. The edge density is the number of edge pixels per unit area within the current spatial grid region; The contour continuity is obtained by detecting the connection relationship between adjacent edge pixels along the edge direction within the current spatial grid region, counting the length of continuously connected edge segments and the number of break points, and weighting them according to a preset weight based on the length of continuously connected edge segments and the number of break points to obtain the contour continuity value of the current spatial grid region. The region closure degree is the percentage of the area within the current spatial grid region that forms a closed outline by its edges; The local response intensity is the average of the characteristic response values ​​of each region in the structural characterization feature map of the current spatial grid region; Candidate regions that meet the preset subject determination criteria are selected based on the edge density, contour continuity, region closure, and local response intensity of each spatial grid region. Extract the center position of each candidate region and map the center position of each candidate region to the feature coordinate space of the structural representation feature map to obtain the set of candidate center points; The positional features of each candidate center point in the structural representation feature map are used as initial seed features, and a set of local features within the neighborhood of each candidate center point is extracted. The local feature sets corresponding to each candidate center point are aggregated and calculated. The aggregated local features are then concatenated with the location encoding information of the candidate center points to generate the initial query representation corresponding to each candidate region. The initial query representations are subjected to dimensional unification and numerical normalization to obtain multiple subject query vectors.

8. The standard perspective white model generation method based on digital twins according to claim 1, characterized in that, Step six specifically involves: The image size information, pixel coordinate range, channel information, and storage format information of the corrected color image, structure enhancement image, and binary mask are read respectively to generate the corresponding image attribute set; Based on the image attribute set, coordinate reference consistency checks are performed on the corrected color image, the structure enhancement image, and the binary mask. If the coordinate reference is inconsistent, the pixel coordinate system of the corrected color image is used as the reference coordinate system, and coordinate remapping is performed on the structure enhancement image and the binary mask to obtain the aligned structure enhancement image and the aligned binary mask that maintain the pixel-level correspondence with the corrected color image. The corrected color image, the aligned structure enhancement image, and the aligned binary mask are encapsulated according to a preset data packet format to obtain a reference image data packet.