Three-dimensional cultural relic 3D scanning ultra-high-definition enhancement method, device and equipment and medium
By using local time-series exposure consistency analysis and texture completion processing, the problems of uneven exposure and insufficient texture fusion accuracy in the acquisition of surface texture details of cultural relics were solved, thereby improving the realism and refinement of the 3D model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XIAN UNVERSITY OF ARTS & SCI
- Filing Date
- 2026-01-15
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies suffer from uneven local exposure and insufficient texture fusion accuracy when acquiring surface texture details of cultural relics under complex lighting conditions, which affects the realism and refinement of 3D models.
By analyzing the exposure consistency of local time series and enhancing the lighting, structurally consistent image pairs are generated. Minimum coverage rectangle region extraction and texture completion processing are then performed. Combined with deep semantic segmentation and semantic consistency verification, the final texture completion map is obtained. Finally, the three-dimensional cultural relic model is reconstructed through triangulation.
It achieves balanced lighting adjustment in areas with uneven exposure, improves the fidelity and detail of texture information, and enhances the realism and refinement of the 3D model.
Smart Images

Figure CN121961963A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of 3D modeling technology, and in particular to a method, apparatus, equipment and medium for ultra-high-definition enhancement of 3D scanning of three-dimensional cultural relics. Background Technology
[0002] In the digital preservation and research of three-dimensional cultural relics, 3D scanning technology has become an important means to realize the digitization of the appearance, shape and texture information of cultural relics. The conventional 3D modeling process of cultural relics usually includes multi-view image acquisition, image preprocessing, feature extraction and matching, 3D reconstruction and texture mapping. In recent years, the development of high-resolution imaging and computer vision algorithms has enabled the acquisition of high-precision geometric structure and surface texture information of cultural relics through structured light scanning, photogrammetry and multi-scale feature matching. These methods have been widely used in archaeological research, digital display and virtual restoration, and can present the morphological features and surface details of cultural relics in a relatively comprehensive way.
[0003] However, conventional methods have limitations in handling local exposure variations and restoring detailed textures when acquiring surface textures of cultural relics under complex lighting conditions. On the one hand, traditional image enhancement methods often rely on global exposure adjustment, making it difficult to provide targeted compensation for local dark or bright areas while maintaining overall brightness balance. On the other hand, during multi-view texture fusion, texture consistency and detail fidelity may be insufficient due to differences in lighting, camera angle changes, and feature matching accuracy. These problems accumulate and are amplified during high-precision modeling and texture mapping, affecting the realism and detail of the final 3D model. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides an ultra-high-definition enhancement method for 3D scanning of three-dimensional cultural relics to solve the problems of uneven local exposure and insufficient texture fusion accuracy.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides a method for ultra-high-definition enhancement of 3D scanning of three-dimensional cultural relics, comprising, Collect original images of three-dimensional cultural relics and construct an original image set; Local time-series exposure consistency analysis was performed on the original image set, and supplementary lighting enhancement processing was applied to obtain a supplementary lighting enhancement image set; A dense matching strategy is applied to the illuminated and enhanced atlas to generate image pairs with consistent structure. Minimum coverage rectangle region extraction is performed on image pairs with consistent structure to generate redundant completion region pairs; Texture completion processing is performed on the receptor region image in the redundant completion region pair to obtain a texture completion map set; Deep semantic segmentation is performed on the donor region image in the redundant completion region pair to generate a semantically consistent mask atlas. The semantic consistency of the texture completion map is checked based on the semantic consistency mask map to obtain the final texture completion map. Key image pairs are extracted based on the final texture completion map, and a set of matching point pairs is obtained through scale-invariant feature transformation. The three-dimensional coordinate set of the matching points is obtained by performing triangulation on each matching point pair. The three-dimensional coordinate set of matching points is used to reconstruct a three-dimensional network model, and the semantic consistency mask atlas is mapped onto the surface of the three-dimensional mesh model to generate a three-dimensional cultural relic model.
[0007] As a preferred embodiment of the 3D scanning ultra-high-definition enhancement method for three-dimensional cultural relics described in this invention, the following steps are taken: Local time-series exposure consistency analysis is performed on the original image set, followed by supplementary lighting enhancement processing to obtain a supplementary lighting enhanced image set. Extract the pixel coordinates of each image in the original image set, and collect the original brightness values of the pixel coordinate images to generate an original brightness matrix set. Sort the original brightness values with the same pixel coordinates in the same group in chronological order to obtain the brightness dataset of the local time series. The brightness difference between adjacent frames is calculated on the brightness dataset of the local time series to obtain the brightness difference dataset of the local time series, and the time-weighted exposure consistency index is calculated. Based on the time-weighted exposure consistency index, the exposure consistency of the brightness dataset is determined, the exposure consistency discrimination result set is obtained, and the original image set is processed by pixel-by-pixel brightness compensation and color equalization to obtain the supplementary lighting enhancement image set.
[0008] As a preferred embodiment of the 3D scanning ultra-high-definition enhancement method for three-dimensional cultural relics described in this invention, the method involves: applying a dense matching strategy to the supplementary lighting enhancement image set to generate image pairs with consistent structure, as follows: Based on the supplementary lighting enhancement atlas, an initial set of matching point pairs is obtained through dense keypoint feature extraction and matching algorithms. Then, based on the matching distance sorting results and the distribution of the number of matching points in the initial set of matching point pairs, image pairs in the supplementary lighting enhancement atlas are filtered to obtain image pairs with consistent structures.
[0009] As a preferred embodiment of the 3D scanning ultra-high-definition enhancement method for three-dimensional cultural relics described in this invention, the following steps are taken: Minimum coverage rectangle region extraction is performed on image pairs with consistent structure to generate redundant completion region pairs. Based on structurally consistent image pairs, the coordinate set of all matching points in the structurally consistent image pairs is extracted by the region boundary localization method, and the minimum coverage rectangle region covering all matching points is calculated. In a pair of images with consistent structure, the smallest coverage rectangle is cropped to obtain a pair of redundant completion regions. The pair of redundant completion regions is defined based on the difference in shooting angle and the number of matching points. The redundant completion region pair includes the recipient region image and the donor region image.
[0010] As a preferred embodiment of the 3D scanning ultra-high-definition enhancement method for three-dimensional cultural relics described in this invention, the following steps are taken: Texture completion processing is performed on the recipient region image in the redundant completion region pair to obtain a texture completion atlas. A multi-scale decomposition structure of the receptor region image is constructed using the multi-scale Laplacian pyramid algorithm. In each scale layer of the multi-scale decomposition structure, the texture details of the donor region image are fused into the same scale layer of the recipient region image according to spatial relationships, and the fused multi-scale decomposition structure is reconstructed layer by layer to generate a recipient region image with fused texture. Based on the arrangement order of redundant completion region pairs, the receptor region images of the fused texture are integrated to obtain a texture completion atlas.
[0011] As a preferred embodiment of the 3D scanning ultra-high-definition enhancement method for three-dimensional cultural relics described in this invention, the following steps are taken: Semantic consistency is checked on the texture completion map based on a semantic consistency mask map to obtain the final texture completion map. Based on the deep semantic segmentation network, the pixels in the donor region image are semantically classified to generate a semantic label map of the donor region image. Then, pixel-level connectivity analysis is performed on the semantic label map of the donor region image to generate a semantic segmentation mask of the donor region image. Based on the order of the redundant completion regions, a semantically consistent mask atlas is generated. The semantic consistency mask atlas and the texture completion atlas are aligned pixel-level according to the order of redundant completion region pairs. The semantic consistency of each pixel semantic label in the texture completion atlas and the pixel semantic labels in the semantic consistency mask atlas are checked by consistency comparison, and the final texture completion atlas is selected.
[0012] As a preferred embodiment of the 3D scanning ultra-high-definition enhancement method for three-dimensional cultural relics described in this invention, the steps are as follows: Key image pairs are extracted based on the final texture completion atlas, and a set of matching point pairs is obtained through scale-invariant feature transformation. The three-dimensional coordinate set of the matching points is obtained by performing triangulation on each matching point pair. Based on the principle of maximizing perspective differences, key image pairs are selected from the final texture completion map set; Based on the scale-invariant feature transformation algorithm, feature descriptor subsets are extracted from key image pairs; Feature matching is performed on the feature descriptor subsets of key image pairs to generate a set of matching point pairs; Based on the set of matching point pairs, perform triangulation to obtain the three-dimensional coordinate set of the matching points.
[0013] Secondly, this invention provides a 3D scanning ultra-high-definition enhancement system for three-dimensional cultural relics, comprising, The image acquisition module is used to acquire original images of three-dimensional cultural relics and construct an original image set; The enhancement processing module is used to perform local time-series exposure consistency analysis on the original image set and perform supplementary lighting enhancement processing to obtain a supplementary lighting enhancement image set; The region extraction module is used to perform a dense matching strategy on the supplementary lighting enhancement atlas to generate structurally consistent image pairs, and to extract the minimum coverage rectangle region from the structurally consistent image pairs to generate redundant completion region pairs. The texture restoration module is used to perform texture restoration processing on the recipient region image in the redundant restoration region pair to obtain a texture restoration map set, perform deep semantic segmentation on the donor region image in the redundant restoration region pair to generate a semantic consistency mask map set, perform semantic consistency verification on the texture restoration map set based on the semantic consistency mask map set, and obtain the final texture restoration map set. The 3D reconstruction module is used to extract key image pairs based on the final texture completion map, obtain a set of matching point pairs through scale-invariant feature transformation, obtain the 3D coordinate set of matching points by triangulation of each matching point pair, reconstruct a 3D network model using the 3D coordinate set of matching points, and map the semantic consistency mask map onto the surface of the 3D mesh model to generate a 3D cultural relic model.
[0014] Thirdly, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the ultra-high-definition enhancement method for 3D scanning of three-dimensional cultural relics as described in the first aspect of the present invention.
[0015] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the ultra-high-definition enhancement method for 3D scanning of three-dimensional cultural relics as described in the first aspect of the present invention.
[0016] The beneficial effects of this invention are as follows: by extracting the pixel coordinates and original brightness values of each image, and combining them with the time-weighted exposure consistency index to perform pixel-by-pixel brightness compensation and color balancing, the illumination balance adjustment of local uneven exposure areas is achieved, ensuring the illumination stability of the subsequent feature extraction and matching process, thereby improving the fidelity and detail integrity of the overall texture information. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A flowchart for a method to enhance the 3D scanning of three-dimensional cultural relics to ultra-high definition.
[0019] Figure 2 This is a schematic diagram of a 3D scanning ultra-high-definition enhancement system for three-dimensional cultural relics.
[0020] Figure 3 This is a flowchart for image enhancement and feature matching.
[0021] Figure 4 This is a flowchart for texture restoration and semantic verification. Detailed Implementation
[0022] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0023] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0024] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0025] Reference Figures 1-4 As an embodiment of the present invention, this embodiment provides a method for ultra-high-definition enhancement of 3D scanning of three-dimensional cultural relics, including the following steps: S1. Collect original images of three-dimensional cultural relics and construct an original image set.
[0026] The three-dimensional artifacts are placed in an adjustable lighting environment, and the original images of the three-dimensional artifacts are captured at equal intervals and angles.
[0027] Furthermore, the three-dimensional artifact is placed in the center of the enclosed adjustable lighting acquisition chamber, aligning the artifact with the central axis of the rotating platform of the enclosed adjustable lighting acquisition chamber, so as to ensure that the artifact remains in the center of the camera's field of view during the rotating shooting process.
[0028] Turn on the evenly distributed adjustable color temperature LED lights in the acquisition chamber, and adjust the color temperature and brightness of the lights according to the material and color of the three-dimensional artifact to ensure uniform lighting on the surface of the artifact and minimize shadows and highlights; mount the high-definition camera on the electric turntable and set the fixed focal length, aperture, shutter speed and ISO of the high-definition camera.
[0029] By controlling the rotating platform, the three-dimensional cultural relic is rotated at equal intervals in the horizontal direction. After each rotation of an angle, the image of the three-dimensional cultural relic is captured, and the shooting sequence and rotation angle information are recorded.
[0030] The original images of the three-dimensional cultural relics are preprocessed, including numbering and grouping them according to the shooting time sequence and shooting angle information, resolution detection, distortion correction, and classification according to shooting angle, to obtain the original image set.
[0031] Furthermore, each original image of the three-dimensional cultural relic is assigned a unique number according to the time sequence and angle information, and the images are grouped together according to the same shooting angle to generate an image set with unique numbers and group information; resolution detection is performed on each original image of the three-dimensional cultural relic, and distortion correction is performed on the original images of the three-dimensional cultural relic that pass the resolution detection, using the calibration parameters of the high-definition camera to eliminate radial and tangential distortion; the original images of the three-dimensional cultural relic after distortion correction are classified according to the shooting angle information to generate image subsets distributed by angle; all original images of the three-dimensional cultural relic classified by shooting angle are counted to generate the original image set.
[0032] Furthermore, when capturing the original image of the three-dimensional cultural relic, the external parameter information of the camera is simultaneously acquired, including the three-dimensional coordinate position of the camera's optical center and the direction vector of the camera's principal optical axis.
[0033] S2. Perform local time-series exposure consistency analysis on the original image set and perform supplementary lighting enhancement processing to obtain a supplementary lighting enhancement image set.
[0034] Extract the pixel coordinates of each image in the original image set, and collect the original brightness values of the pixel coordinate images to generate an original brightness matrix set. Sort the original brightness values with the same pixel coordinates in the same group in chronological order to obtain a local time series brightness dataset.
[0035] Furthermore, by traversing each original image in the original image set, the two-dimensional grid positions of the pixels in the original image are read, and the two-dimensional grid positions of the pixels are recorded as pixel horizontal and vertical coordinates respectively; the red, green, and blue channels are extracted from the original images, and the pixel values of the two-dimensional grid positions are read based on the red, green, and blue channels of the original images. The original brightness values are obtained through the original brightness calculation formula to generate the brightness matrix of the original images; the brightness matrix of each original image is stored according to the acquisition order of the original images to generate an original brightness matrix set composed of multiple brightness matrices; based on the unique number and grouping information of the original images, within the same group, for each position with the same pixel coordinates in spatial location, all the original brightness values in the brightness matrix are extracted, and the original brightness values are arranged sequentially according to the image capture time order to generate a time-series brightness record of pixel coordinates within the group. The time-series brightness records of all pixel coordinates are counted to generate a local time-series brightness dataset.
[0036] It should be noted that the original brightness calculation formula, expressed as: ; in, Represents time frame The original image in pixel coordinates The original brightness value at that location, Indicates the time frame number, representing the first frame in the original image set. Image Represents time frame The original image in pixel coordinates The red channel pixel value at that location, Represents time frame Image in pixel coordinates The green channel pixel value at that location, Represents time frame Image in pixel coordinates The blue channel pixel value at that location, This represents the weighting coefficient for the red channel. This represents the weighting coefficient for the green channel. This represents the weighting coefficient for the green channel.
[0037] It should be noted that, , , Based on the existing standard system, values are assigned according to the differences in human visual sensitivity to different colors. The value range of is [0,1]. The value range of is [0,1]. The value range of is [0,1], and satisfies .
[0038] The brightness difference between adjacent frames is calculated on the brightness dataset of the local time series to obtain the brightness difference dataset of the local time series, and the time-weighted exposure consistency index is calculated.
[0039] Furthermore, based on the local time-series brightness difference dataset, the time-weighted exposure consistency index is calculated, expressed as: ; ; ; in, This represents the time-weighted exposure consistency index. This represents the total number of frames in the time series. Represents the first in the luminance difference dataset Frame and the Frame in prime coordinates The difference in brightness, This represents a very small constant to prevent the denominator from being zero. This indicates that time sovereignty is paramount. Indicates time from weight, This indicates the time center, which is the middle frame number of the entire time series. Indicates the attenuation coefficient. This indicates the inter-frame distance corresponding to the half-life.
[0040] To maintain consistency in the exposure calculation process, and considering the uniformity of ambient light changes across different time periods during shooting, and with the time weights of the brightness baseline and brightness changes attenuating around the same center frame, the following will be implemented: The value as The value; It determines the rate at which the time weight decays with time distance.
[0041] Based on the time-weighted exposure consistency index, the exposure consistency of the brightness dataset is determined, the exposure consistency discrimination result set is obtained, and the original image set is processed by pixel-by-pixel brightness compensation and color equalization to obtain the supplementary lighting enhancement image set.
[0042] Furthermore, a global judgment threshold is set. Based on the time-weighted exposure consistency index and the global judgment threshold, the exposure consistency of the brightness dataset is judged. When the time-weighted exposure consistency index of the pixel coordinates... When determining the global threshold, the current pixel coordinates are determined to have stable exposure throughout the entire time series and are marked as having consistent exposure.
[0043] When the time-weighted exposure consistency index of a pixel coordinate is greater than the global judgment threshold, the current pixel coordinate is determined to be unstable in exposure throughout the entire time series and is marked as inconsistent exposure. The results of exposure consistency judgment for all pixel coordinates are statistically analyzed to generate an exposure consistency discrimination result set.
[0044] Based on the exposure consistency discrimination result set, the exposure consistency judgment results of all pixel coordinates in the exposure consistency discrimination result set are mapped according to the two-dimensional coordinate positions in the original image. The judgment state of each pixel coordinate is converted into a binary code. All binary codes construct a binary matrix with the same size as the original image. A value of 1 in the binary matrix indicates consistent exposure, and a value of 0 in the binary matrix indicates inconsistent exposure. For the pixel coordinates with a value of 1 in the binary matrix, the brightness compensation gain is calculated.
[0045] Furthermore, the brightness compensation gain of the pixel coordinates is calculated as follows: First, the median brightness value of the pixel coordinates in the entire time series is obtained as the target brightness value. Then, the mean brightness value of the pixel coordinates in the entire time series is obtained as the current brightness reference value. The ratio of the target brightness value to the current brightness reference value is calculated to obtain the brightness gain factor of the pixel coordinates. The brightness gain factor is then truncated to prevent distortion caused by excessively high or low compensation.
[0046] For each frame in the original image set, all pixel coordinates in the image are traversed. When the exposure consistency judgment result set is marked as consistent at the pixel coordinates, the original brightness value of the pixel coordinates is multiplicatively aggregated with the brightness gain factor. When the exposure consistency judgment result set is marked as inconsistent at the pixel coordinates, the original brightness value remains unchanged. After completing pixel-by-pixel brightness compensation, color equalization processing is performed on each frame image. Specifically, the global mean of the three channels (red, green, and blue) in the current frame image is calculated, and the average of the three channel means is used as the target gray value. The ratio of the target gray value to the three channel means is calculated respectively to obtain the red channel gain, green channel gain, and blue channel gain. The channel gain is multiplicatively aggregated with each pixel value of the channel respectively. All frame images are reorganized according to the numbering order and grouping information of the original image set to generate a consistent supplementary lighting enhancement image set that has completed brightness compensation and color equalization.
[0047] It should be noted that setting a global judgment threshold involves: summarizing the calculated time-weighted exposure consistency indices to generate a consistency index dataset, which includes the exposure stability quantification results of all pixel coordinates across the entire image set without any filtering; removing outliers from the consistency index dataset; performing statistical analysis on the dataset after outlier removal, and determining the overall stability level of the consistency index by calculating the median, quartiles, and skewness of the data distribution; and selecting a quantile that can effectively distinguish between stable and unstable exposure areas as the global judgment threshold based on the distribution of the consistency index.
[0048] It should be noted that, based on the distribution of the consistency index, a quantile that can effectively distinguish between stable and unstable exposure areas is selected as the global judgment threshold. Specifically, a distribution analysis is performed on the consistency index dataset. When the time-weighted exposure consistency index exhibits two local maxima and a local minimum between them, the quantile of the local minimum is selected as the reference quantile. When the distribution of the time-weighted exposure consistency index is relatively smooth, the lower quartile or median is selected as the reference quantile. The value of the time-weighted exposure consistency index corresponding to the reference quantile is set as the global judgment threshold.
[0049] The global decision threshold remains constant throughout the exposure consistency assessment of the entire image set. The value of the global decision threshold effectively eliminates pixels with unstable lighting without mistakenly affecting pixels with stable exposure; its typical range is [range missing]. .
[0050] S3. Apply a dense matching strategy to the supplementary lighting enhancement atlas to generate image pairs with consistent structure.
[0051] Based on the supplementary lighting enhancement atlas, an initial set of matching point pairs is obtained through dense keypoint feature extraction and matching algorithms. Then, based on the matching distance sorting results and the distribution of the number of matching points in the initial set of matching point pairs, image pairs in the supplementary lighting enhancement atlas are filtered to obtain image pairs with consistent structures.
[0052] Furthermore, a scale-invariant feature transform algorithm is used to uniformly sample feature points on the pixel grid of the images in the illumination enhancement dataset, and a scale-space neighborhood is constructed at each feature point location. The gradient orientation histogram within the scale-space neighborhood is statistically analyzed and normalized to generate a feature descriptor subset with rotation, scale, and illumination invariance. For each set of feature descriptors from two different images in the illumination enhancement dataset, a fast approximate nearest neighbor search algorithm is used for matching. The Euclidean distance between each feature point and the candidate matching point is calculated and sorted, and a ratio test strategy is used to obtain the best matching distance sorting result. Image pairs with consistent structure are screened based on the number of matching points threshold and the matching distance sorting result. Image pairs that simultaneously meet both the number of matching points threshold and the matching distance sorting result are defined as image pairs with consistent structure.
[0053] It should be noted that the ratio test strategy is as follows: the matching ratio is calculated according to the matching ratio test formula. When the matching ratio is less than the matching threshold, the matching relationship between the nearest neighbor matching point and the current feature point is determined to be the best matching result of the current feature point. The matching pair between the nearest neighbor matching point and the current feature point is defined as the initial matching point pair, and all initial matching point pairs constitute the initial matching point pair set.
[0054] Matching ratio, the expression is: ; in, Indicates the match ratio; This represents the Euclidean distance between a feature point and its nearest matching neighbor. This represents the Euclidean distance between the feature point and its second nearest neighbor matching point.
[0055] The matching threshold is set by statistically analyzing the matching distance distribution of three-dimensional cultural relic samples under different shooting conditions. The value of the matching threshold needs to balance the matching accuracy and the number of matching points, and the value range is [insert range here]. .
[0056] The matching point number threshold determination is as follows: the total number of matching points in the initial matching point pair set of each image pair is counted, and a matching point number threshold is set based on the image resolution and sampling interval. When the total number of matching points in the initial matching point pair set of the current image pair is less than the number threshold, the current image pair is directly removed.
[0057] The threshold for the number of matching points is set by linearly mapping the image resolution to the angular interval between adjacent sampling viewpoints and statistically calculating historical scaling factors. The threshold value must satisfy the condition that the number of matching points for each image pair in the initial set of matching point pairs reaches a minimum, and its range is [value missing]. .
[0058] The matching distance sorting results are as follows: for each image pair, all matching point pairs are sorted by matching distance from smallest to largest, and the matching point pairs ranked in the top R% are taken as the candidate valid matching set. The number of matching points in the candidate valid matching set is calculated. When the number of matching points in the candidate valid matching set is higher than the total number of matching points in the initial matching point pair set of the image pair, it indicates that the matching quality of the current image pair is concentrated and stable, and the current image pair is retained.
[0059] Based on the matching distance distribution characteristics of image pairs from different viewpoints in the supplementary lighting enhancement dataset, the matching distances of multiple sets of sample image pairs are sorted and the matching stability interval is statistically analyzed. The value of R needs to balance the number of matching points and the matching accuracy, and the range of R is [value missing]. The matching distance threshold is determined by statistical analysis of the matching quality of image pairs under different shooting angles, based on the proportion of effective matching points. The value of the matching distance threshold balances matching stability and retention rate, and its range is [value missing]. .
[0060] Image pairs that simultaneously satisfy both the threshold for the number of matching points and the sorting result of the matching distance are defined as structurally consistent image pairs.
[0061] S4. Extract the minimum coverage rectangle region from the image pairs with consistent structure to generate redundant completion region pairs.
[0062] Based on structurally consistent image pairs, the coordinate set of all matching points in the structurally consistent image pairs is extracted by the region boundary localization method, and the minimum coverage rectangle region covering all matching points is calculated.
[0063] Furthermore, based on structurally consistent image pairs, the set of matching point pairs is first extracted from the structurally consistent image pairs. Each pair of matching points in the set is then split into the coordinates of the matching points in the first image and the coordinates of the matching points in the second image. These coordinates are stored in two independent coordinate lists to obtain the set of matching point coordinates for the first image and the set of matching point coordinates for the second image. Subsequently, a rectangular region boundary division operation is performed on the set of matching point coordinates for the first and second images to obtain the minimum coverage rectangular region of the two images.
[0064] It should be noted that the rectangular region boundary division operation specifically involves obtaining the minimum, maximum, minimum, and maximum values of the horizontal and vertical coordinates of all matching points in the set of matching point coordinates, and determining the rectangular region boundary that can cover all matching points based on these four extreme values.
[0065] In a pair of images with consistent structure, the smallest coverage rectangle is cropped to obtain a pair of redundant completion regions. The redundant completion regions are then defined based on the difference in shooting angle and the number of matching points.
[0066] The redundant completion region pair includes both the recipient region image and the donor region image.
[0067] Furthermore, in the structurally consistent image pair, image cropping operations are performed on the first and second images respectively based on the boundary coordinates of the minimum coverage rectangle region, to obtain rectangular cropped region images from the first and second images. The two rectangular cropped region images together constitute a redundant completion region pair. Subsequently, the shooting parameters of the redundant completion region pair are analyzed, the shooting angle difference between the two rectangular cropped region images is calculated, and the number of matching points between the two rectangular cropped region images is counted. Based on the shooting angle difference and the number of matching points, the recipient region image and the donor region image are defined, thereby obtaining a redundant completion region pair including the recipient region image and the donor region image.
[0068] It should be noted that the recipient region image and donor region image are determined based on the difference in shooting angle and the number of matching points. Specifically, the two rectangular cropped region images are named Image A and Image B, respectively.
[0069] The difference in shooting angle is determined as follows: when the angle between the shooting direction of image A and the surface normal vector of the target area is less than the angle between the shooting direction of image B and the surface normal vector of the target area, image A is determined to be near-frontal view; when the angle between the shooting direction of image B and the surface normal vector of the target area is less than the angle between the shooting direction of image A and the surface normal vector of the target area, image B is determined to be near-frontal view.
[0070] The determination of the number of matching points is as follows: if the number of matching points in image A is less than the number of matching points in image B, then image A is determined to have fewer matching points; if the number of matching points in image B is less than the number of matching points in image A, then image B is determined to have fewer matching points.
[0071] The rectangular cropped region image that is determined to be near-frontal and has a small number of matching points is defined as the recipient region image, and the other rectangular cropped region image is defined as the donor region image.
[0072] S5. Perform texture completion processing on the receptor region image in the redundant completion region pair to obtain a texture completion map set.
[0073] A multi-scale decomposition structure of the receptor region image is constructed using the multi-scale Laplacian pyramid algorithm.
[0074] It should be noted that a multi-scale decomposition structure of the receptor region image is constructed using a multi-scale Laplacian pyramid algorithm. Specifically, the receptor region image is decomposed into a Gaussian pyramid, and Gaussian blurring is applied layer by layer, followed by sampling according to the pyramid scaling factor to obtain a multi-level Gaussian image sequence from the original resolution to lower resolution. The first layer is the original resolution image, images closer to the first layer are high-resolution images, and the last layer is the lowest resolution image. Upsampling and difference operations are performed between adjacent layers of the Gaussian pyramid. The low-resolution layer image is upsampled and Gaussian smoothed, and then pixel-differencing is performed with the high-resolution layer image to obtain Laplacian image layers at different scales. All Laplacian image layers contain image detail information within different spatial frequency ranges. The lowest resolution Gaussian image is retained as the residual layer of the pyramid. Combined with each Laplacian image layer, the multi-scale decomposition structure of the receptor region image is constructed, thereby separating the low-frequency structure from the high-frequency details of the image.
[0075] The pyramid scaling factor is set according to the resolution and detail fidelity requirements of the receptor region image. The value of the pyramid scaling factor balances the number of pyramid layers with computational efficiency, and its range is [range missing]. .
[0076] In each scale layer of the multi-scale decomposition structure, the texture details of the donor region image are fused into the same scale layer of the recipient region image according to spatial relationships, and the fused multi-scale decomposition structure is reconstructed layer by layer to generate a recipient region image with fused texture.
[0077] Based on the arrangement order of redundant completion region pairs, the receptor region images of the fused texture are integrated to obtain a texture completion atlas.
[0078] Furthermore, in each scale layer of the multi-scale decomposition structure, the donor region image is decomposed into Gaussian pyramids and Laplacian pyramids at the same resolution and scale as the recipient region image to obtain the multi-scale decomposition structure of the donor region image. In each scale layer, using the scale layer of the recipient region image as the base, the texture details of the donor region image in the same scale layer are fused pixel-level according to the spatial correspondence of redundant completion region pairs. Specifically, while preserving the low-frequency structure of the recipient region image, the high-frequency detail information of the donor region image in the scale layer is superimposed onto the position of the recipient region image, thereby enhancing and supplementing the texture details. After the fusion of all scale layers is completed, upsampling and additive reconstruction are performed layer by layer starting from the lowest resolution layer, and the high-frequency details are superimposed level by level to finally generate the recipient region image with fused texture. According to the arrangement order of the redundant completion region pairs, all recipient region images with fused texture are stitched together and integrated to generate a complete texture completion atlas.
[0079] S6. Perform deep semantic segmentation on the donor region image in the redundant completion region pair to generate a semantic consistency mask map. Based on the semantic consistency mask map, perform semantic consistency verification on the texture completion map to obtain the final texture completion map.
[0080] Based on the deep semantic segmentation network, the pixels in the donor region image are semantically classified to generate a semantic label map of the donor region image. Then, pixel-level connectivity analysis is performed on the semantic label map of the donor region image to generate a semantic segmentation mask of the donor region image.
[0081] Based on the order of redundant completion regions, a semantically consistent mask atlas is generated.
[0082] Furthermore, pixel-level manually labeled semantic tags are generated for the donor region images to obtain a data training set. The donor region images and their training set are then input into the DeepLabV3+ deep semantic segmentation network. The difference between the DeepLabV3+ network labels and the pixel-level manually labeled semantic tags is calculated using the cross-entropy loss function. The learning rate is adjusted using the Adam optimizer, and data augmentation is performed through random rotation, cropping, and scaling. The network undergoes at least 100 iterations of training, with monitoring of the cross-entropy loss function and the performance of the DeepLabV3+ deep semantic segmentation network during these iterations to prevent overfitting. After multiple iterations, the DeepLabV3+ deep semantic segmentation network infers from the input donor region images, performing pixel-by-pixel semantic category determination for all pixels in each donor region image. Each pixel is assigned a unique semantic label according to its identified semantic category, thus generating a semantic label map of the donor region images.
[0083] Pixel-level connectivity analysis is performed on the semantic label map of the donor region image. Connected regions are divided according to the semantic labels of adjacent pixels, and each connected region is marked as an independent mask region to generate a semantic segmentation mask for the donor region image. The semantic segmentation masks of each donor region image are sequentially integrated according to the arrangement order of redundant completion region pairs to obtain a semantically consistent mask map set arranged according to spatial relationships.
[0084] It should be noted that connected regions are divided based on the semantic labels of adjacent pixels, and each connected region is marked as an independent mask region to generate a semantic segmentation mask for the donor region image. Specifically, after generating the semantic label map of the donor region image, for each pixel in the semantic label map, the set of adjacent pixels is determined according to the 8-neighborhood connectivity rule; it is then determined whether the semantic labels of the adjacent pixels of the current pixel are completely consistent with the semantic label of the current pixel. When the semantic labels are consistent, the current pixel and its adjacent pixels are classified into the same connected region; the current connected region is continuously expanded by traversing a queue until there are no new adjacent pixels with consistent semantic labels in the connected region, thus obtaining a complete connected region; subsequently, all pixels in the connected region are marked as pixels in the same mask region and assigned a unified mask number; the above process is repeated until all pixels in the semantic label map have been classified into the corresponding connected regions, thus obtaining the semantic segmentation masks for the entire donor region image.
[0085] The semantic consistency mask atlas and the texture completion atlas are aligned pixel-level according to the order of redundant completion region pairs. The semantic consistency of each pixel semantic label in the texture completion atlas and the pixel semantic labels in the semantic consistency mask atlas are checked by consistency comparison, and the final texture completion atlas is selected.
[0086] Furthermore, the semantic consistency mask atlas and texture completion atlas are paired one by one according to the order of redundant completion region pairs. Each paired semantic consistency mask atlas and texture completion atlas are spatially aligned in the pixel coordinate system so that pixels at the same position in both have the same pixel coordinate index. Under the spatially aligned pixel coordinate index, the semantic label of each pixel in the texture completion atlas is compared pixel by pixel with the semantic label of the same coordinate pixel in the semantic consistency mask atlas. When the semantic labels are the same, the pixel value in the texture completion atlas is retained; otherwise, the pixel is marked as invalid. Consistency comparison and filtering are performed on all paired images to generate the final texture completion atlas containing only pixels with matching semantic labels.
[0087] S7. Extract key image pairs based on the final texture completion map set, and obtain a set of matching point pairs through scale-invariant feature transformation. Obtain the three-dimensional coordinate set of matching points by performing triangulation on each matching point pair. Reconstruct the three-dimensional network model using the three-dimensional coordinate set of matching points, and map the semantic consistency mask map set onto the surface of the three-dimensional mesh model to generate a three-dimensional cultural relic model.
[0088] Based on the principle of maximizing perspective differences, key image pairs are selected from the final texture completion map set.
[0089] Furthermore, based on the camera's extrinsic parameters and according to the principle of maximizing viewpoint difference, all images in the final texture completion dataset are combined and arranged. Different image pairs are selected from all possible combinations, and the shooting position and pose parameters of each image pair are analyzed. Specifically, the camera extrinsic parameters of each image in the image pair are extracted, and the angle between the optical center line and the principal optical axis of each image is calculated as a measure of viewpoint difference between the images. Among all possible image pair combinations, the viewpoint difference measures of each pair are sorted in descending order, prioritizing the selection of image pairs with higher rankings to ensure that the selected image pairs are spatially... A large parallax at the shooting location is beneficial for improving the depth accuracy of 3D structure reconstruction. For each selected image pair, the overlapping area of the two images is calculated through image registration, and the number of pixels in the overlapping area is counted. When the number of pixels in the overlapping area is less than the lower limit of the overlapping pixel ratio, the current image pair is discarded and the next candidate combination is selected until an image pair that simultaneously meets the requirements of having a large parallax at the spatial shooting location and having an overlapping pixel number not less than the lower limit of the overlapping pixel ratio is selected. This pair is defined as the key image pair to ensure that subsequent feature matching can both utilize the large parallax to improve spatial positioning accuracy and rely on sufficient overlapping area to ensure that the number of matching points is sufficient and evenly distributed.
[0090] It should be noted that the lower limit of the overlap pixel ratio is set based on the ratio of the minimum number of matching pixels required for 3D reconstruction to the image resolution. The value of the lower limit of the overlap pixel ratio satisfies both sufficient parallax and enough matching feature points, and its range is [value missing]. .
[0091] Based on the scale-invariant feature transformation algorithm, feature descriptor subsets are extracted from key image pairs; Furthermore, based on the scale-invariant feature transformation algorithm, multi-scale spatial extremum point detection is performed on each image in the key image pair to determine the location of potentially stable features. At the detected feature point locations, the gradient direction distribution in the scale space neighborhood is calculated using the gradient direction histogram calculation method to determine the principal direction for achieving rotation invariance. In the neighborhood of each feature point, a description vector based on gradient magnitude and direction is constructed to generate feature descriptors with scale invariance and rotation invariance. All feature descriptors with scale invariance and rotation invariance are counted to generate a feature descriptor set, which is then saved as the feature descriptor set of the key image pair.
[0092] It should be noted that a description vector based on gradient magnitude and direction is constructed in the neighborhood of each feature point to generate a feature descriptor with scale invariance and rotation invariance. Specifically, in the neighborhood of each feature point, a sub-region of fixed size is calculated with the feature point as the center. Within each sub-region of fixed size, the gradient components of each pixel in the neighborhood of each feature point in the horizontal and vertical directions are calculated to obtain the gradient magnitude and gradient direction. The gradient direction is divided into direction intervals according to fixed angular intervals. The gradient magnitude in all direction intervals is counted to generate the gradient direction histogram of the current sub-region. The gradient direction histograms of all sub-regions are spliced together in spatial order to form the description vector of the feature point, generating a feature descriptor with rotation invariance and scale invariance.
[0093] Feature matching is performed on the feature descriptor subsets of key image pairs to generate a set of matching point pairs.
[0094] Furthermore, when performing feature matching on the feature descriptor set of the key image pair, for each feature descriptor of the first image in the key image pair, the Euclidean distance with all feature descriptors in the second image is calculated, and the feature descriptor with the smallest Euclidean distance is selected as the first candidate matching point, and the feature descriptor with the second smallest Euclidean distance is selected as the second candidate matching point. The ratio test method is used to calculate the ratio of the Euclidean distance between each pair of first candidate matching points and second candidate matching points. When the ratio is less than the upper limit of the matching confidence ratio, it is determined that the first candidate matching point and the original feature descriptor constitute a valid matching relationship. After completing the ratio test of all feature descriptors, all valid matching relationships are summarized to generate a matching point pair set. The matching point pair set is composed of pairs of pixel coordinates of the two images in the key image pair, which are used for subsequent triangulation calculations.
[0095] It should be noted that the upper limit of the match confidence ratio is set by statistically analyzing the distance ratio distribution between correct and incorrect matches on the feature matching validation set. This ratio is chosen to maximize precision while maintaining recall. The upper limit of the match confidence ratio balances matching accuracy and the number of matches, and its value range is [range missing]. .
[0096] Based on the set of matching point pairs, perform triangulation to obtain the three-dimensional coordinate set of the matching points.
[0097] Furthermore, based on the set of matching point pairs, the pixel coordinates of each matching point in the two images of the key image pair are extracted, and the camera intrinsic matrix and relative pose matrix of the key image pair are called to convert the pixel coordinates into normalized image plane coordinates through the camera projection model. The normalized image plane coordinates of the same matching point in the two images are substituted into the triangulation equation, and the three-dimensional coordinates of the matching point are obtained by solving the minimum distance intersection point of the two light rays in three-dimensional space. The above operation is performed on all matching points in the set of matching point pairs in turn to finally generate the three-dimensional coordinate set of matching points.
[0098] Furthermore, after obtaining the 3D coordinate set of matching points, the 3D coordinate set of matching points is input into the 3D mesh generation algorithm according to the point cloud topology reconstruction principle. The connection structure of vertices, edges and faces is constructed according to the spatial adjacency relationship between matching points to generate a 3D mesh model with a complete geometric surface. The semantic consistency mask atlas is texture mapped to the 3D mesh model according to the arrangement order of redundant completion region pairs to ensure that the semantic label of each pixel in the mask atlas accurately corresponds to the texture coordinates of the mesh surface, thereby generating a semantically consistent texture distribution on the surface of the 3D mesh model, and finally obtaining the 3D cultural relic model.
[0099] This embodiment also provides a 3D scanning ultra-high-definition enhancement system for three-dimensional cultural relics, including: an image acquisition module, an enhancement processing module, a region extraction module, a texture restoration module, and a three-dimensional reconstruction module; The image acquisition module is used to acquire original images of three-dimensional cultural relics and construct an original image set; The enhancement processing module is used to perform local time-series exposure consistency analysis on the original image set and perform supplementary lighting enhancement processing to obtain a supplementary lighting enhancement image set; The region extraction module is used to perform a dense matching strategy on the supplementary lighting enhancement atlas to generate structurally consistent image pairs, and to extract the minimum coverage rectangle region from the structurally consistent image pairs to generate redundant completion region pairs. The texture restoration module is used to perform texture restoration processing on the recipient region image in the redundant restoration region pair to obtain a texture restoration map set, perform deep semantic segmentation on the donor region image in the redundant restoration region pair to generate a semantic consistency mask map set, perform semantic consistency verification on the texture restoration map set based on the semantic consistency mask map set, and obtain the final texture restoration map set. The 3D reconstruction module is used to extract key image pairs based on the final texture completion map, obtain a set of matching point pairs through scale-invariant feature transformation, obtain the 3D coordinate set of matching points by triangulation of each matching point pair, reconstruct a 3D network model using the 3D coordinate set of matching points, and map the semantic consistency mask map onto the surface of the 3D mesh model to generate a 3D cultural relic model.
[0100] This embodiment also provides a computer device applicable to the ultra-high-definition enhancement method for 3D scanning of three-dimensional cultural relics, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to realize the ultra-high-definition enhancement method for 3D scanning of three-dimensional cultural relics as proposed in the above embodiment.
[0101] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0102] This embodiment also provides a storage medium storing a computer program, which, when executed by a processor, implements the method for ultra-high-definition enhancement of 3D scanning of three-dimensional cultural relics as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0103] In summary, this invention achieves balanced illumination adjustment for unevenly exposed areas by extracting the pixel coordinates and original brightness values of each image and combining them with a time-weighted exposure consistency index for pixel-by-pixel brightness compensation and color balancing. This ensures the illumination stability of subsequent feature extraction and matching processes, thereby improving the fidelity and detail integrity of the overall texture information.
[0104] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for ultra-high-definition enhancement of 3D scanning of three-dimensional cultural relics, characterized in that: include, Collect original images of three-dimensional cultural relics and construct an original image set; Local time-series exposure consistency analysis was performed on the original image set, and supplementary lighting enhancement processing was applied to obtain a supplementary lighting enhancement image set; A dense matching strategy is applied to the illuminated and enhanced atlas to generate image pairs with consistent structure. Minimum coverage rectangle region extraction is performed on image pairs with consistent structure to generate redundant completion region pairs; Texture completion processing is performed on the receptor region image in the redundant completion region pair to obtain a texture completion map set; Deep semantic segmentation is performed on the donor region image in the redundant completion region pair to generate a semantically consistent mask atlas. The semantic consistency of the texture completion map is checked based on the semantic consistency mask map to obtain the final texture completion map. Key image pairs are extracted based on the final texture completion map, and a set of matching point pairs is obtained through scale-invariant feature transformation. The three-dimensional coordinate set of the matching points is obtained by performing triangulation on each matching point pair. The three-dimensional coordinate set of matching points is used to reconstruct a three-dimensional network model, and the semantic consistency mask atlas is mapped onto the surface of the three-dimensional mesh model to generate a three-dimensional cultural relic model.
2. The method for ultra-high-definition enhancement of 3D scanning of three-dimensional cultural relics as described in claim 1, characterized in that: The steps for performing local time-series exposure consistency analysis on the original image set and then performing supplementary lighting enhancement processing to obtain a supplementary lighting enhancement image set are as follows: Extract the pixel coordinates of each image in the original image set, and collect the original brightness values of the pixel coordinate images to generate an original brightness matrix set. Sort the original brightness values with the same pixel coordinates in the same group in chronological order to obtain the brightness dataset of the local time series. The brightness difference between adjacent frames is calculated on the brightness dataset of the local time series to obtain the brightness difference dataset of the local time series, and the time-weighted exposure consistency index is calculated. Based on the time-weighted exposure consistency index, the exposure consistency of the brightness dataset is determined, the exposure consistency discrimination result set is obtained, and the original image set is processed by pixel-by-pixel brightness compensation and color equalization to obtain the supplementary lighting enhancement image set.
3. The method for ultra-high-definition enhancement of 3D scanning of three-dimensional cultural relics as described in claim 2, characterized in that: The steps for using a dense matching strategy to generate structurally consistent image pairs from the enhanced illumination atlas are as follows: Based on the supplementary lighting enhancement atlas, an initial set of matching point pairs is obtained through dense keypoint feature extraction and matching algorithms. Then, based on the matching distance sorting results and the distribution of the number of matching points in the initial set of matching point pairs, image pairs in the supplementary lighting enhancement atlas are filtered to obtain image pairs with consistent structures.
4. The method for ultra-high-definition enhancement of 3D scanning of three-dimensional cultural relics as described in claim 3, characterized in that: The steps for extracting the minimum coverage rectangle region from structurally consistent image pairs and generating redundant completion region pairs are as follows: Based on structurally consistent image pairs, the coordinate set of all matching points in the structurally consistent image pairs is extracted by the region boundary localization method, and the minimum coverage rectangle region covering all matching points is calculated. In a pair of images with consistent structure, the smallest coverage rectangle is cropped to obtain a pair of redundant completion regions. The pair of redundant completion regions is defined based on the difference in shooting angle and the number of matching points. The redundant completion region pair includes the recipient region image and the donor region image.
5. The method for ultra-high-definition enhancement of 3D scanning of three-dimensional cultural relics as described in claim 4, characterized in that: The steps for performing texture completion processing on the receptor region image in the redundant completion region pair to obtain a texture completion map set are as follows: A multi-scale decomposition structure of the receptor region image is constructed using the multi-scale Laplacian pyramid algorithm. In each scale layer of the multi-scale decomposition structure, the texture details of the donor region image are fused into the same scale layer of the recipient region image according to spatial relationships, and the fused multi-scale decomposition structure is reconstructed layer by layer to generate a recipient region image with fused texture. Based on the arrangement order of redundant completion region pairs, the receptor region images of the fused texture are integrated to obtain a texture completion atlas.
6. The method for ultra-high-definition enhancement of 3D scanning of three-dimensional cultural relics as described in claim 5, characterized in that: The semantic consistency check of the texture completion map based on the semantic consistency mask map to obtain the final texture completion map is performed as follows: Based on the deep semantic segmentation network, the pixels in the donor region image are semantically classified to generate a semantic label map of the donor region image. Then, pixel-level connectivity analysis is performed on the semantic label map of the donor region image to generate a semantic segmentation mask of the donor region image. Based on the order of the redundant completion regions, a semantically consistent mask atlas is generated. The semantic consistency mask atlas and the texture completion atlas are aligned pixel-level according to the order of redundant completion region pairs. The semantic consistency of each pixel semantic label in the texture completion atlas and the pixel semantic labels in the semantic consistency mask atlas are checked by consistency comparison, and the final texture completion atlas is selected.
7. The method for ultra-high-definition enhancement of 3D scanning of three-dimensional cultural relics as described in claim 6, characterized in that: The steps are as follows: Key image pairs are extracted based on the final texture completion map, and a set of matching point pairs is obtained through scale-invariant feature transformation. Then, the 3D coordinates of the matching points are obtained by triangulation of each matching point pair. Based on the principle of maximizing perspective differences, key image pairs are selected from the final texture completion map set; Based on the scale-invariant feature transformation algorithm, feature descriptor subsets are extracted from key image pairs; Feature matching is performed on the feature descriptor subsets of key image pairs to generate a set of matching point pairs; Based on the set of matching point pairs, perform triangulation to obtain the three-dimensional coordinate set of the matching points.
8. A 3D scanning ultra-high-definition enhancement system for three-dimensional cultural relics, based on the 3D scanning ultra-high-definition enhancement method for three-dimensional cultural relics as described in any one of claims 1 to 7, characterized in that: include, The image acquisition module is used to acquire original images of three-dimensional cultural relics and construct an original image set; The enhancement processing module is used to perform local time-series exposure consistency analysis on the original image set and perform supplementary lighting enhancement processing to obtain a supplementary lighting enhancement image set; The region extraction module is used to perform a dense matching strategy on the supplementary lighting enhancement atlas to generate structurally consistent image pairs, and to extract the minimum coverage rectangle region from the structurally consistent image pairs to generate redundant completion region pairs. The texture restoration module is used to perform texture restoration processing on the recipient region image in the redundant restoration region pair to obtain a texture restoration map set, perform deep semantic segmentation on the donor region image in the redundant restoration region pair to generate a semantic consistency mask map set, perform semantic consistency verification on the texture restoration map set based on the semantic consistency mask map set, and obtain the final texture restoration map set. The 3D reconstruction module is used to extract key image pairs based on the final texture completion map, obtain a set of matching point pairs through scale-invariant feature transformation, obtain the 3D coordinate set of matching points by triangulation of each matching point pair, reconstruct a 3D network model using the 3D coordinate set of matching points, and map the semantic consistency mask map onto the surface of the 3D mesh model to generate a 3D cultural relic model.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the ultra-high-definition enhancement method for 3D scanning of three-dimensional cultural relics as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the ultra-high-definition enhancement method for 3D scanning of three-dimensional cultural relics as described in any one of claims 1 to 7.