Interactive text travel image vectorization method and device
By preprocessing cultural and tourism images and interactively segmenting them using a deep segmentation model, combined with Bezier curve edge fitting, a high-quality SVG standard semantic vector image is generated. This solves the problem of unsatisfactory vectorization effects caused by uneven image quality and improves the structural expressiveness and semantic readability of the image.
Patent Information
- Application Number
- CN202511280548.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-09-09
AI Technical Summary
In the existing technology of cultural and tourism image collection, the image quality is uneven due to the influence of the image collection environment. Directly vectorizing the entire image leads to unsatisfactory results, which reduces the use value of the materials in the material library.
The image is preprocessed through geometric correction, size mapping and multi-stage enhancement. Interactive semantic segmentation is performed by combining the deep segmentation model that introduces the attention mechanism and conditional random field. Connected regions are extracted and the boundary structure is optimized. Bezier curves are used to fit the edges. Color-calibrated images are generated in combination with user interaction. Finally, a semantic vector graph that conforms to the SVG standard is constructed.
It significantly improves the structural fidelity and texture clarity of the image, generates a structurally continuous, semantically clear, and editable vector image, and improves the structural expressiveness and semantic readability of the image.
Smart Images

Figure CN120765796A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and in particular to a method and device for interactive cultural and tourism image vectorization. Background Art
[0002] In the field of culture and tourism, the application of vectorization technology has been upgraded from the early "image digitization tool" to the "infrastructure for the dataization of cultural heritage". Vectorized cultural and tourism images have many advantages over raster images, such as image quality is independent of resolution and can be scaled arbitrarily without distortion; compact storage, small file size, and efficient storage and transmission; support for editing and modification of geometric elements, which can be highly consistent with the needs of scenarios such as cultural relics protection, cultural and creative development, and digital display.
[0003] However, existing technologies still have bottlenecks in natural texture processing, semantic understanding, real-time computing, etc. In cultural and tourism image acquisition, the image quality is uneven due to the influence of the image acquisition environment. For images with unsatisfactory results, the entire image is directly vectorized, and the effect of the obtained vector image is greatly reduced, and its cultural value cannot be fully exerted, thereby reducing the use value of the materials in the material library.
[0004] Therefore, there is an urgent need for an interactive cultural and tourism image vectorization method and device. Summary of the Invention
[0005] The present application provides an interactive cultural and tourism image vectorization method and device, which solves the problem that the effect of the vector image obtained is greatly reduced when the entire image is directly vectorized for an image with unsatisfactory effect.
[0006] In a first aspect of the present application, an interactive cultural and tourism image vectorization method is provided, which includes: acquiring a cultural and tourism image, and preprocessing the cultural and tourism image to obtain a preprocessed image; segmenting the preprocessed image to obtain a segmented image; performing edge extraction and fitting operations on the segmented image to obtain image edge features; performing color attribute assignment operations based on the image edge features to obtain a color calibration image; performing semantic encoding operations based on the color calibration image to obtain a vector diagram corresponding to the cultural and tourism image.
[0007] Optionally, obtaining cultural and tourism images specifically includes: obtaining cultural and tourism images through preset acquisition equipment, the preset acquisition equipment includes a digital camera and a scanner; obtaining the image type of the cultural and tourism images, and setting differentiated acquisition parameters based on the image type, the image types include cultural relics image type, architectural image type and landscape image type; based on the differentiated acquisition parameters, constructing the coordinate mapping relationship corresponding to the cultural and tourism images through geometric correction operations and diffraction mapping operations.
[0008] Optionally, the cultural and tourism image is preprocessed to obtain a preprocessed image, specifically including: converting the cultural and tourism image into a grayscale image by a weighted averaging method; determining the texture entropy of the cultural and tourism image based on the grayscale image, and performing hierarchical noise reduction on the cultural and tourism image by texture entropy; performing image enhancement processing on the denoised cultural and tourism image, and the image enhancement processing includes Retinex processing and CLAHE enhancement processing; based on the peak signal-to-noise ratio and structural similarity of the cultural and tourism image after image enhancement, performing an evaluation operation on the enhanced cultural and tourism image; if the evaluation operation is confirmed as the first evaluation result, returning the original cultural and tourism image and marking the noise area in the cultural and tourism image; if the evaluation operation is confirmed as the second evaluation result, adjusting the filter kernel size in the cultural and tourism image and enhancing the number of iterations; if the evaluation operation is confirmed as the third evaluation result, outputting the image-enhanced cultural and tourism image as a preprocessed image.
[0009] Optionally, the preprocessed image is segmented to obtain a segmented image, specifically including: constructing an image segmentation model based on a neural network structure and conditional random field that introduces an attention mechanism, and interactively labeling the preprocessed image through the image segmentation model; performing a segmentation operation on the labeled preprocessed image based on the image segmentation model to segment the labeled preprocessed image into multiple connected regions; based on the characteristic parameters of each connected region, performing an erosion or expansion operation of a corresponding kernel on the connected region, and obtaining a high-uncertainty region in the multiple connected regions through a preset evaluation operation, wherein the characteristic parameters include area parameters, perimeter parameters, aspect ratio parameters, and edge complexity parameters; optimizing and iterating the high-uncertainty region to complete the image segmentation operation, and outputting the segmented image.
[0010] Optionally, edge extraction and fitting operations are performed on the segmented image to obtain image edge features, specifically including: performing edge extraction operations on each sub-element in the segmented image to obtain sub-element information; and using a cubic Bezier curve to fit the sub-element information to obtain image edge features.
[0011] Optionally, a color attribute assignment operation is performed based on the edge features of the image to obtain a color-calibrated image, specifically including: extracting pixel values for each connected area in the sub-element based on the coordinates of the image edge and the mapping relationship between the coordinates and the color values, and calculating the average color as the filling color; filling the cultural and tourism image with color based on the filling color, and judging whether the color of the image after color filling meets the preset requirements; if the image color does not meet the preset requirements, increasing or decreasing the color through a preset interactive method; if the image color meets the preset requirements, the image after color filling is used as the color-calibrated image.
[0012] Optionally, a semantic encoding operation is performed based on the color calibration image to obtain a vector diagram corresponding to the cultural and tourism image, specifically including: encoding the color attributes corresponding to the color calibration image into XML format, and adding corresponding semantic tags in the XML format; outputting the color calibration image with added semantic tags as an SVG vector diagram file to obtain a vector diagram corresponding to the cultural and tourism image.
[0013] In a second aspect of the present application, an interactive cultural tourism image vectorization device is provided, the device comprising an acquisition module and a processing module, wherein: The acquisition module is used to acquire cultural tourism images and preprocess the cultural tourism images to obtain preprocessed images.
[0014] The processing module is used to segment the preprocessed image to obtain a segmented image; perform edge extraction and fitting operations on the segmented image to obtain image edge features; perform color attribute assignment operations based on the image edge features to obtain a color calibration image; perform semantic encoding operations based on the color calibration image to obtain a vector diagram corresponding to the cultural and tourism image.
[0015] In the third aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device performs any of the methods described above.
[0016] In a fourth aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and the computer program is executed by a processor to perform any of the above methods.
[0017] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: 1. Acquire cultural and tourism images and pre-process them through geometric correction, size mapping and multi-stage enhancement, significantly improving the structural fidelity and texture clarity of the images, and outputting pre-processed images with uniform scale and high contrast; based on the deep segmentation model that introduces the attention mechanism and conditional random field, perform interactive semantic segmentation on the pre-processed images, extract connected regions, and optimize the boundary structure through high-uncertainty screening and active learning mechanisms to obtain segmented images with continuous structure and clear semantics; perform gradient perception and edge extraction of connected domain markers on the segmented images, and use Bezier curves to perform high-precision fitting of the edges to generate image edge features with precise shape and editable shape; perform color statistics and compression on each area based on the edge structure, and implement differentiated processing on the fill color and stroke color, and combine user interaction to generate color-calibrated images with realism and consistency; finally, based on the fusion of the three elements of edge, color and semantics, construct a vector structure that conforms to the SVG standard, embed semantic tags and output as a semantic vector graph, comprehensively improving the image structure expression, semantic readability and later usability.
[0018] 2. By constructing an adaptive preprocessing process based on dynamic feedback of image quality, grayscale conversion, noise identification, hierarchical noise reduction and brightness enhancement are implemented on the acquired cultural and tourism images, and an image quality evaluation mechanism is introduced to realize hierarchical judgment and processing strategy regulation, thereby achieving a balanced improvement in image quality in complex lighting and noise interference environments.
[0019] 3. Build an image segmentation model with attention mechanism and edge optimization capabilities, and complete the labeling of key semantic areas under the guidance of user interaction to generate an initial segmentation map; extract multiple connected regions in the image based on the model prediction results, and implement differentiated structural correction processing based on regional geometry and boundary features; introduce an evaluation mechanism to quantify the uncertainty of each connected region, so as to perform focused iterative optimization on areas with blurred boundaries or semantically unstable regions; finally, output a segmentation image with clear boundaries, complete structure and accurate semantics, thereby improving the adaptability and expression accuracy of complex cultural and tourism images in structural analysis through the combination of full-process semantic drive and local structure regulation, laying a stable foundation for subsequent edge fitting and semantic vector expression. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 This is a flow chart of an interactive cultural tourism image vectorization method provided in an embodiment of the present application; Figure 2 This is a module diagram of an interactive cultural and tourism image vectorization device provided in an embodiment of the present application; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0021] Explanation of the reference numerals: 21, acquisition module; 22, processing module; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. DETAILED DESCRIPTION
[0022] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.
[0023] The terms used in the following examples of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of the present application, the singular expressions "a", "an", "said", "above", "the", and "this" are intended to include plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to and includes any or all possible combinations of one or more of the listed items.
[0024] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of the technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of this application, unless otherwise specified, "plurality" means two or more.
[0025] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.
[0026] Please refer to Figure 1 , which shows a flow chart of an interactive cultural and tourism image vectorization method provided in an embodiment of the present application, and the flow chart mainly includes the following steps: S101 to S105.
[0027] Step S101: Acquire a cultural tourism image and pre-process the cultural tourism image to obtain a pre-processed image.
[0028] Specifically, we acquire multiple types of cultural and tourism images and build a unified processing framework. By configuring acquisition parameters and modeling physical scenes for different image types, combined with image distortion correction, size mapping construction, and color mapping generation, we achieve spatial geometric consistency and color consistency of the image. Subsequently, we implement grayscale conversion, hierarchical adaptive noise reduction driven by texture entropy calculation, multi-scale Retinex enhancement, and local contrast optimization fusion on the image. We establish a multi-stage collaborative mechanism between image structure preservation, noise suppression, and detail enhancement, and finally output a pre-processed image that meets the accuracy requirements of subsequent segmentation and feature extraction. In this possible implementation, step S101 includes the following sub-steps: S111 to S113.
[0029] Step S111: Acquire cultural and tourism images through a preset acquisition device, which includes a digital camera and a scanner.
[0030] Specifically, cultural and tourism images are acquired through preset acquisition devices such as digital cameras and scanners. During the acquisition process, the image sensor in the preset acquisition device can convert the collected light signals into RAW data, which is then processed by the processor into a standard format image such as JPEG, PNG, etc., while retaining the original color information.
[0031] Step S112: obtaining the image type of the cultural tourism image and setting differentiated acquisition parameters based on the image type. The image type includes a cultural relic image type, a building image type, and a landscape image type.
[0032] Specifically, differentiated acquisition parameters are set for different types of images, including but not limited to cultural relic images, architectural images, and landscape images. For example, for cultural relic images, multiple light sources arranged in a ring (intervals ≤ 10°) are used to illuminate the surface of the object being measured from different angles, and a polarizer is installed in front of the ring light source to generate linearly polarized light. Afterwards, the resolution is set to ≥ 600dpi, and polarized images at different incident angles are obtained through circular scanning, which can effectively eliminate surface reflections and improve image contrast and feature detection accuracy. For architectural images, drone oblique photography (to obtain panoramic structures) is combined with close-range scanner scanning (to obtain local details such as doors, windows, and carvings), and the SIFT (Scale-Invariant Feature Transform) algorithm is used to match feature points and stitch them together to generate a complete image, ensuring that the geometric structure is distortion-free.
[0033] Step S113: Based on the differentiated acquisition parameters, a coordinate mapping relationship corresponding to the cultural tourism image is constructed through geometric correction operations and diffraction mapping operations.
[0034] Specifically, the geometric correction operation: For distortion correction, use a standard checkerboard calibration plate (e.g., 10×7 corner points) and ensure that its actual size is known (e.g., each square has a side length of 20mm). Detect the checkerboard corners in the image to obtain pixel coordinates. Call the function (cv2.calibrated) to calculate the camera intrinsic parameter matrix and distortion coefficients (dist_coeffs) to eliminate barrel or pincushion distortion. For size mapping, select the pixel coordinates of the calibration plate's four corner points and the actual physical coordinates (e.g., the physical coordinates of the checkerboard corners). Use the function (cv2.getPerspectiveTransform) to generate a 3×3 transformation matrix M to map the pixel coordinates to the physical coordinate system. Use the function (cv2.warpPerspective) to perform perspective correction on the image, and calculate the pixel-to-physical size ratio (e.g., 1 pixel = 0.1mm) based on the calibration plate's size.
[0035] Diffraction mapping operations: By traversing the pixels, the coordinates (i, j) and corresponding RGB values are recorded for each pixel in the corrected image, generating a color map. This color map is saved as a JSON or CSV file for easy access by subsequent vectorization tools (such as AutoCAD). For example, in 3D modeling, the RGB value of the pixel (i, j) can be reverse-queried based on the physical coordinates (x, y) to achieve accurate material mapping.
[0036] In a possible implementation, step S101 further includes the following sub-steps: S121 to S127.
[0037] Step S121: convert the cultural tourism image into a grayscale image by weighted averaging method.
[0038] Specifically, using the weighted average method, ,Converting the color image into a single-channel grayscale image can reduce color ,interference and facilitate subsequent texture analysis.
[0039] Step S122: Determine the texture entropy of the cultural tourism image based on the grayscale image, and perform hierarchical noise reduction on the cultural tourism image through the texture entropy.
[0040] Specifically, the image texture entropy is calculated by the gray-level co-occurrence matrix (GLCM) to distinguish Gaussian noise from salt and pepper noise. For Gaussian noise, a 5×5 Gaussian kernel is used; for salt and pepper noise, a median filter is used.
[0041] In this embodiment, the entropy value of Gaussian noise is low, the pixel gray scale changes gently, and the entropy value of salt and pepper noise is high, containing discrete black and white noise points. According to the characteristics of Gaussian noise and salt and pepper noise, adaptive filtering can be performed; for Gaussian noise, a 5*5 Gaussian kernel is used, and the sigma value is adaptively adjusted according to the noise intensity (the more serious the noise, the larger the sigma, such as sigma = 1.5-2.5), to suppress the noise in the smooth area; for salt and pepper noise, median filtering is used, and the kernel size is (3*3 or 5*5), which is dynamically selected according to the noise density (the density is high, the kernel size is increased), and the edge details are protected.
[0042] Step S123, image enhancement processing is performed on the travel image after noise reduction, and the image enhancement processing includes Retinex processing and CLAHE enhancement processing.
[0043] Specifically, the multi-scale Retine algorithm (MSRCR) is used to separate the illumination component and the reflection component of the image, and the color constancy is restored; the dark area detection is performed on the image after Retinex processing, and the CLAHE (Contrast Limited Adaptive Histogram Equalization) is applied to the detected dark area, and the original contrast of other areas is preserved; the dark area enhanced by CLAHE is fused with the global image processed by Retinex, and the bilinear interpolation or Poisson fusion is used to eliminate the boundary artifacts.
[0044] In this embodiment, the image after noise reduction is enhanced in multiple stages, which can avoid the loss of image details. Through MSRCR (Multi-Scale Retinal Model), the illumination component (uneven illumination) and the reflection component (real texture) of the image are separated, the color constancy is restored (such as eliminating the influence of glass reflection and shadow on color), and the dark area positioning is performed on the image processed by the MSRCR model through threshold segmentation or gradient analysis. Only the CLAHE (Contrast Limited Adaptive Histogram Equalization) is applied to the area to avoid overexposure caused by global enhancement, preserve the bright part details, and finally use the bilinear interpolation or Poisson fusion algorithm to seamlessly splice the enhanced dark part and the global image to eliminate the edge brightness jump.
[0045] Step S124, based on the peak signal-to-noise ratio and the structural similarity of the travel image after image enhancement, the travel image after enhancement is evaluated.
[0046] Specifically, based on the PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity) of the travel image, the processed image is evaluated.
[0047] Step S125, if the evaluation operation confirms the first evaluation result, the original travel image is returned and the noise area in the travel image is marked.
[0048] Specifically, if the first evaluation result is PSNR < 22dB and SSIM < 0.6, the original cultural and tourism image is returned and the noisy areas are manually marked. The image in the first evaluation result contains severe noise, blur, or compression distortion. Direct denoising and enhancement cannot distinguish between real details and noise. This may lead to the accidental deletion of cultural relic textures (such as historical traces of rust on bronze artifacts) or excessive restoration of architectural lines (such as worn details on ancient architectural carvings), resulting in irreversible damage. Therefore, it is necessary to return to the original cultural and tourism image and manually mark the noisy areas (such as stains on the surface of cultural relics and blurred areas in architectural images). Afterwards, a restoration strategy is selected based on the noise type and marking results.
[0049] In step S126, if the evaluation operation confirms the second evaluation result, the filter kernel size of the cultural tourism image is adjusted and the number of iterations is increased. If the image in the second evaluation result has some noise, blur, or lack of contrast, but the overall structure and details are not severely damaged, adjusting the filter kernel size or increasing the number of iterations (for example, increasing the Gaussian kernel σ value) can optimize local quality while preserving existing details.
[0050] Specifically, the second evaluation result is 22dB≤PSNR<32dB and 0.6≤SSIM<0.8. If the second evaluation result is confirmed, the filter kernel size is adjusted and the number of iterations is increased.
[0051] Step S127: If the evaluation operation is confirmed as the third evaluation result, the cultural tourism image after image enhancement is output as a pre-processed image.
[0052] Specifically, if the third evaluation result is PSNR ≥ 32dB and SSIM ≥ 0.8, the enhanced cultural tourism image is output as the preprocessed image. The image in the third evaluation result has minimal pixel distortion and well-preserved structural details, indicating high quality. At this point, the preprocessed image can be directly output and the image segmentation process can proceed.
[0053] Step S102: segment the pre-processed image to obtain a segmented image.
[0054] Specifically, based on the image segmentation model that introduces the attention mechanism and conditional random field structure, the image is interactively semantically labeled under the guidance of user annotation, and the initial segmentation of multi-category target areas is completed; then the geometric and boundary feature parameters of each connected area are extracted, and the erosion or expansion kernel size is adaptively configured according to the features to achieve structural repair and boundary smoothing; combined with entropy evaluation or edge overlap indicators, high-uncertainty areas are identified, and iterative optimization is performed through pseudo-label generation and consistency regularization mechanism to obtain high-quality segmented images with complete structure, continuous semantics and clear boundaries.
[0055] In a possible implementation, step S102 further includes the following sub-steps: S211 to S214.
[0056] Step S211: construct an image segmentation model based on the neural network structure and conditional random field that introduces the attention mechanism, and interactively mark the preprocessed image through the image segmentation model.
[0057] Specifically, a U-Net++ neural network is used, an attention mechanism (such as SENet channel attention) is added to the encoder, and a conditional random field (CRF) is introduced in the decoder; the key areas of cultural and tourism images (such as cultural relics outlines and building lines) are focused on, and the edge continuity of the segmentation results is optimized; based on sample complexity, different weights w are assigned to each type of image in the training sample; w1=0.5 for cultural relic images, w2=0.3 for building images, and w3=0.2 for landscape images. The total loss function is cross entropy loss + Dice coefficient loss, which improves the segmentation accuracy of images with rich details.
[0058] The encoder consists of four downsampling blocks, each of which includes two 3×3 convolutions (with padding=1), a batch normalization layer, and a ReLU activation, followed by a 2×2 max pooling layer (stride=2). SENet channel attention (compression ratio r=16) is inserted after the last convolutional layer of each encoder block to weight the feature maps per channel, enhancing the channel responses corresponding to artifact outlines and architectural lines. The decoder consists of four upsampling blocks, which upsample using either deconvolution (2×2, stride=2) or bilinear interpolation, and fuse the skip-connected features of the corresponding encoders. The output of the last decoder block is fed into a fully connected CRF (DenseCRF). This is implemented using the PyTorch-CRF library, with relevant parameters. The Gaussian kernel standard deviation is sigma_space=5 (controls spatial proximity); the bilateral kernel standard deviations are sigma_color=30 and sigma_space=10 (balances color and spatial similarity and protects edges). The number of iterations is 5 during inference and 2 during training to reduce computational overhead.
[0059] During the sample weighting phase, complexity is assigned based on user requirements. For cultural relic images, c1 = 5 (most complex details, highest cultural value), for architectural images, c2 = 3 (complex geometric structure), and for landscape images, c3 = 2 (relatively simple contours). The weighting formula is: , that is, cultural relics map ,Cultural Relics Map , Cultural Relics Map The weight of each sample is directly assigned according to its type (for example, if the label contains the category of cultural relic images, the weight is 0.5).
[0060] Among them, the total loss function Consists of two parts: weighted cross entropy loss and weighted Dice loss , weighted cross entropy loss By assigning different weights to different samples, it can deal with the problem of class imbalance and focus on difficult samples; weighted Dice loss By calculating the similarity between the predicted results and the true labels to optimize the boundary continuity and small area segmentation accuracy, the combination of these two loss functions can effectively improve the performance of the image segmentation task. The formula is as follows: ; ; Where N is the total number of samples, is the weight of the nth sample, used to deal with category imbalance, is the true label (one-hot encoding) of the c-th category of the n-th sample at position (i, j), is the predicted probability of the cth category of the nth sample at position (i, j); ;in, is the weight of the nth sample, is the true label of the cth category of the nth sample at position (i, j), is the predicted probability of the cth category for the nth sample at position (i, j).
[0061] The user labels key areas, and the image segmentation model generates initial segmentation results based on the labels; the graph cut algorithm is used to diffuse the labeled areas, generate pseudo labels and add them to training samples; the model parameters are optimized by combining supervision loss and consistency regularization loss; the entropy value of the unlabeled area is calculated, and the area with the highest entropy value is selected for user labeling, forming a closed-loop iteration.
[0062] For key area annotation, for cultural relic images, only the main outlines (such as the outer edge of a bronze vessel or the rim of a ceramic vessel) are annotated, eliminating the need for detailed inscriptions or decorative patterns, thus reducing labor costs. For architectural images, key structural lines (such as beam-column intersections and bracket outlines) are annotated, creating a "skeleton"-style annotation. During annotation, polygon tools or brushes can be used to quickly outline and generate a binary mask (annotated area = 1, unannotated = 0, background = 2). For initial segmentation prediction, the image with the annotated mask is input, and a pre-trained U-Net++ neural network (including SENet and CRF) is used to generate an initial segmentation probability map (covering the entire image). For unannotated areas, the model predicts the category (e.g., cultural relic, building, landscape, background) based on contextual information (e.g., color and texture similarity).
[0063] Build a graph cut model and define image pixels (the labeled region S and the unlabeled region U), the pixels in the labeled region S are forced to belong to the labeled category, and the weight of the pixels in the unlabeled region U is determined by the initial model prediction probability (such as P (cultural relics The color similarity (RGB Euclidean distance) and the position distance of adjacent pixels and are used to construct a bilateral weight, which can balance the spatial and color similarity; the specific formula is as follows: ; Among them, , , indicate the bilateral weight between pixels and ; and respectively indicate the position coordinates of pixels and ; and respectively indicate the RGB color values of pixels and ; is the spatial difference, which can control the influence degree of the position distance on the weight; indicates the color standard deviation, which can control the influence degree of the color similarity on the weight.
[0064] Generate pseudo-labels to minimize the energy function, and the formula is as follows: , wherein, indicates the energy function, indicating the total energy or cost of the entire model; f indicates the pseudo-label, that is, the label predicted by the model; S indicates the entire sample set in the data set; indicates the i-th sample in the data set; is a data item, indicating the data cost of a single sample , which is related to the difference between the true label and the predicted label of the sample; N indicates a neighborhood set, indicating the adjacency relationship between samples; indicates that there is an adjacency relationship between samples and ; indicates a smoothing term, indicating the smoothing cost between adjacent samples and , which is used to ensure that the predicted labels of adjacent samples are as consistent as possible. In addition, the pixels in the unlabeled region U are diffused through graph cut, and inherit the categories of the labeled region (such as the pixels inside the cultural relic contour are diffused as “cultural relics”, and the pixels near the building structure are diffused as “building”), to generate a full-image pseudo-label as a weak supervision signal.
[0065] We constructed mixed training data by manually annotating key regions (high confidence, weighted to 1.0) and pseudo-labeled regions generated by graph cuts (low confidence regions were filtered out, for example, pixels with a model-predicted probability > 0.7 were retained, with a weight of 0.5). In the loss function, the total loss = supervision loss + 0.3 consistency regularization loss. We updated the model by generating 50 pseudo-labeled images and merging them with the original manually annotated data. We then fine-tuned the model (freezing the first two encoder layers and updating only the decoder and CRF parameters). We used the AdamW optimizer (learning rate set to 1e-5) and trained for 5 epochs to avoid overfitting to pseudo-label noise.
[0066] Active learning is guided by entropy. For each pixel in the unlabeled area U, the entropy value H(p) of the category probability distribution is calculated. High entropy areas (e.g., H(p)>0.8) indicate high uncertainty in the model prediction (e.g., blurred areas of cultural relic inscriptions, structures in the shadows of buildings). Low entropy areas (e.g., H(p)<0.3) indicate accurate model predictions and can be directly used as pseudo-labels. When selecting regions, the regions are sorted in descending order of entropy value, and the top 20% of high-entropy regions are extracted. A rectangular box or polygonal mask is generated, and the user is prompted to label it ("Please confirm whether this blurred area is a cultural relic inscription"). For continuous high-entropy regions (e.g., large blurred areas on the surface of a cultural relic), the user only needs to mark the boundaries or key points, and the interior is automatically filled in using graph cuts.
[0067] For connected region feature extraction, OpenCV is used to extract the following features from the segmented binary mask: Area: number of pixels A (to distinguish small areas such as inscription spots (A<100) from large areas such as building walls (A>1000). Perimeter: number of contour pixels P (to evaluate edge complexity). Eccentricity: eccentricity of ellipse fitting (to determine whether the shape of the region is regular. Select the appropriate kernel size. For small areas (A<200): use 3×3 kernel erosion or dilation (to protect details and avoid over-smoothing). For large areas (A≥200): use 5×5 kernel or Gaussian kernel (to smooth large area edges, such as noise spots on building walls). Erosion can remove small noise blocks (such as isolated pixels outside the outline of cultural relics), and dilation can connect adjacent areas (such as broken lines of building structures).
[0068] For effect evaluation and iteration, the entropy reduction rate of the processed area is calculated (target: entropy reduction of high entropy area ≥ 30%) to ensure the integrity of the contour after edge detection (Canny edge detection, calculation of the overlap rate between the predicted contour and the manually labeled contour). For areas that still have high entropy or discontinuous contours after processing, they are re-added to the active learning queue and the labeling-training-processing process is repeated.
[0069] Step S212 : performing a segmentation operation on the marked pre-processed image based on the image segmentation model to segment the marked pre-processed image into a plurality of connected regions.
[0070] Specifically, based on the image segmentation model, a segmentation operation is performed on the labeled pre-processed image to divide the labeled pre-processed image into multiple connected regions. Specifically, the U-Net++ neural network with the SENet channel attention mechanism is first used to perform full-image prediction on the pre-processed image with key area annotations to generate a semantic probability map. Subsequently, a fully connected conditional random field is introduced based on the probability map to optimize edge continuity and improve the structural consistency of the detailed contours. Based on the segmentation results, a binary mask map is constructed, and the regions of different semantic categories in the mask map are segmented using the connected domain labeling algorithm to extract all connected regions. Based on the spatial adjacency of the pixels, each connected region is numbered and contour traced to form an independent structural unit for subsequent structural feature analysis, kernel scale regulation, and uncertainty screening. By combining deep semantic modeling with spatial topology recognition, this process not only ensures the complete expression of semantic boundaries, but also lays the foundation for accurate partitioning for subsequent adaptive processing based on regional granularity.
[0071] In step S213, based on the characteristic parameters of each connected area, an erosion or expansion operation of the corresponding core is performed on the connected area, and a high uncertainty area in multiple connected areas is obtained through a preset evaluation operation. The characteristic parameters include area parameters, perimeter parameters, aspect ratio parameters, and edge complexity parameters.
[0072] Specifically, geometric features are first extracted from each connected region in the image segmentation result, including counting the number of pixels in the region as an area parameter, counting the pixels along the contour path to form a perimeter parameter, calculating the aspect ratio through the minimum enclosing rectangle to obtain the length-to-height ratio parameter, and constructing the edge complexity parameter in combination with the amplitude of the contour gradient change. Based on these parameters, the morphological scale and structural characteristics of the region are determined, and the appropriate morphological processing kernel size and type are selected. For regions with an area smaller than a preset threshold or high edge complexity, a smaller erosion kernel (such as 3×3) is used to remove isolated noise points, while for regions with a larger area or insufficient contour continuity, a larger dilation kernel (such as 5×5) is used to connect broken boundary structures. Subsequently, based on the changes in the contour structure and information entropy value before and after processing, combined with the edge detection accuracy and regional consistency index, a multidimensional evaluation function is constructed to identify regions with unstable processing effects, blurred boundaries, damaged structures, or significant semantic splitting tendencies. These regions are identified as high-uncertainty regions and prioritized for subsequent optimization iterations, achieving a joint improvement in segmentation accuracy and structural expression integrity.
[0073] Step S214 , performing optimization iterations on the high uncertainty region to complete the image segmentation operation, and outputting a segmented image.
[0074] Specifically, for all high-uncertainty areas, we first construct a graph cut energy minimization model based on the initial segmentation probability map generated by the model and the user interactive annotation results. We use the spatial position and color similarity between pixels to define bilateral weights, perform pseudo-label diffusion on unlabeled areas, and generate an initial pseudo-label map covering the entire image. Subsequently, based on the confidence level of the pixel prediction probability in the pseudo-label, we filter out low-confidence areas and only retain those with a probability higher than a set threshold (such as 0.7) as training samples. We then construct a hybrid training dataset based on the original manually annotated areas, introduce supervision loss and consistency regularization loss into the loss function, and assign pseudo-labels a lower loss weight (such as 0.5) to suppress the propagation of pseudo-label noise. During the training process, a number of pseudo-label images (such as 50) are generated cumulatively each time. , that is, triggering a model fine-tuning, only updating the decoder and conditional random field parameters, freezing the front-layer structure of the encoder, controlling the model's generalization ability and preventing overfitting; at the same time, calculating the predicted entropy value of the pixels in the unlabeled area, extracting the high-uncertainty area according to the descending order of the entropy value, and preferentially guiding the user to locally annotate the boundaries or key points of the area, forming an active learning closed loop of "high entropy priority and minimum interaction"; for continuous high-entropy areas, combining contextual texture, color and boundary information, automatically generating complete labels through graph cut propagation and semantic guidance; through multiple rounds of interactive optimization and model adaptive update, the segmentation boundaries of all high-uncertainty areas tend to converge, and finally output a segmented image with clear semantics, complete boundaries and consistent regional structure, meeting the accuracy requirements of subsequent edge extraction and vector fitting.
[0075] Step S103: performing edge extraction and fitting operations on the segmented image to obtain image edge features.
[0076] Specifically, by constructing an edge extraction mechanism based on gradient analysis and structural domain tracking, we first identify the connected areas of each sub-element in the segmented image one by one, use the eight-neighbor difference calculation of pixels to extract the gradient amplitude map, and extract the fine edge contour through non-maximum suppression and maximum entropy adaptive threshold method; then apply the connected domain labeling algorithm to perform structural recognition and directionality segmentation on the edge contour, and extract the edge pixel coordinate sequence; on this basis, we introduce the vector data compression and curvature analysis mechanism, use the cubic Bezier curve to perform curve fitting processing on the edge coordinate sequence, and ensure the continuity of contour curvature and smooth edges through control point adjustment to construct a high-fidelity image edge vector expression; this processing not only improves the geometric expression accuracy and compression efficiency of the structural contour, but also establishes a highly consistent shape foundation for subsequent color assignment, semantic encoding and vector graphics output.
[0077] In a possible implementation, step S103 further includes the following sub-steps: S311 to S312.
[0078] Step S311 : performing an edge extraction operation on each sub-element in the segmented image to obtain sub-element information.
[0079] Specifically, the sub-element information includes the sub-element outline; the differential calculation is performed on the 8 adjacent sub-element pixels to obtain the gradient amplitude of the sub-element, the non-maximum suppression processing is performed on the gradient amplitude of the sub-element, and the maximum entropy adaptive algorithm is used to perform edge detection to obtain the sub-element outline; the sub-element information also includes the coordinates of the sub-element edge pixels; the connected domain labeling algorithm is used to label the connected areas of the sub-element outline to obtain 8 connected areas, and the vector data compression algorithm is used for each marked area to extract the coordinates of the sub-element edge pixels.
[0080] For contour extraction, eight-neighbor difference calculations can be performed on each subelement (such as cultural relic fragments or architectural components) to calculate the gradient amplitude. Non-maximum suppression can be used to refine the edges, and then a maximum entropy adaptive algorithm (which automatically determines the edge detection threshold) can be used to extract the precise contour. For coordinate extraction, a connected domain labeling algorithm can be used to divide the contour into eight connected regions (corresponding to eight adjacent directions). Vector data compression algorithms (such as the Douglas-Peucker algorithm) can then be used to extract the coordinates of key edge pixels to reduce data redundancy.
[0081] Step S312: Using a cubic Bezier curve to perform a fitting operation on the sub-element information to obtain image edge features.
[0082] Specifically, curve fitting is performed based on the sub-element information, and a cubic Bezier curve is used to fit the extracted sub-element edge pixel points.
[0083] In this embodiment, a cubic Bezier curve may be used to fit edge pixels, and the curvature of the curve may be adjusted by controlling vertices to ensure smooth representation of complex contours (such as patterns on cultural relics and arcs of ancient buildings) and avoid polyline distortion.
[0084] Step S104 : performing a color attribute assignment operation based on the image edge features to obtain a color calibrated image.
[0085] Specifically, a color analysis model based on the mapping of pixel coordinates and color values is constructed around each connected area defined by the edge of the image. First, the RGB value distribution characteristics are extracted through pixel statistics in the area, and the average color of the area is calculated to determine the fill color. A color clustering compression algorithm such as octree structure or K-means clustering is introduced to control the number of color types to improve the compactness of the vector graphics file; further, the brightness of the fill color is adjusted according to the boundary structure, and a dark version is generated for outline stroke to enhance the visual separation of the structure boundary; finally, a color fine-tuning mechanism is provided in combination with user interaction operations, allowing manual intervention and correction of the fill color and stroke color to adapt to personalized art style or the requirements of restoring the true color tone of cultural relics, thereby generating a color-calibrated image with realism, editability and semantic consistency, laying a visual foundation for vector graphics semantic encoding and standard output.
[0086] In a possible implementation, step S104 further includes the following sub-steps: S411 to S414.
[0087] Step S411 : Based on the coordinates of the image edge and the mapping relationship between the coordinates and the color values, pixel values are extracted from each connected area in the sub-element, and the average color is calculated as the filling color.
[0088] Specifically, color attributes include fill color and stroke color. Based on the coordinates of the image edge and the mapping relationship between coordinates and color values, the RGB values of all pixels in each connected area of the sub-element are extracted and the average color is calculated as the fill color. Specifically, the following is included: ; Where N is the number of pixels in the area; use octree or K-means clustering to compress the color table to ensure that the number of vector graphics colors is controllable.
[0089] Step S412: Fill the cultural tourism image with color based on the filling color, and determine whether the color of the image after color filling meets the preset requirements.
[0090] Specifically, a darker version of the area color (e.g., 30% lower brightness) is used for edge stroke.
[0091] In this embodiment, for fill colors, the RGB average of all pixels within a connected region is calculated, and the color table is compressed using octree or K-means clustering (for example, compressing 24-bit true color to 256 colors), ensuring lightweight vector image files. For stroke colors, the fill color's brightness is reduced by 30%, converting RGB (200, 150, 100) to RGB (140, 105, 70), enhancing edge contrast.
[0092] Step S413: If the image color does not meet the preset requirements, the color is increased or decreased through a preset interactive method.
[0093] Specifically, the user determines whether the image color meets the requirements. If not, they can interactively increase or decrease the color using the mouse. Users can click on areas to increase or decrease color (for example, manually adjusting the green hue of rust on cultural relics or the saturation of ancient building paintings), achieving personalized color calibration.
[0094] In step S414, if the image color meets the preset requirements, the image after color filling is used as the color calibrated image.
[0095] Specifically, the user determines whether the image color meets the requirements. If it meets the requirements, the image after color filling is used as the color calibrated image.
[0096] S105 , performing a semantic encoding operation based on the color calibration image to obtain a vector map corresponding to the cultural tourism image.
[0097] Specifically, the edge curves of each connected area in the color calibration image are first parameterized into the form of Bezier curves, and the corresponding color attributes, structural features and spatial position information are uniformly encoded into an XML data structure; then a semantic tag system is introduced to attach semantic tags with cultural attributes to each area based on the image content type and structural composition, such as cultural relic category, building component type or landform element name, to enhance the comprehensibility and retrieval value of the graphics; finally, the multidimensional structure containing edge geometry, color attributes and semantic information is output as a vector graphics image in the SVG standard format, achieving highly compressible, high-resolution, scalable and content-oriented structured representation, supporting subsequent efficient applications in scenarios such as cultural heritage digital modeling, graphic re-creation and semantic search.
[0098] In a possible implementation, step S105 further includes the following sub-steps: S511 to S512.
[0099] S511 , encoding the color attributes corresponding to the color calibration image into an XML format, and adding corresponding semantic tags in the XML format.
[0100] Specifically, the Bezier curve is parameterized and the color attributes are encoded in XML format, and semantic tags are added to the XML: first, each connected area after edge fitting in the color calibration image is parsed, and the corresponding cubic Bezier curve control point parameters are extracted, including the starting point, end point and coordinate information of the two control points. The curve parameters are expressed in a standardized manner and organized as path definition nodes in the XML format; then the color attributes of the area are encoded, including the RGB values of the fill color and the stroke color, and uniformly inserted into the style definition segment of the XML in hexadecimal format to ensure visual consistency when the graphics are rendered; after the structure encoding is completed, a semantic tag layer is constructed based on the classification results of the previous image semantic segmentation stage or user-specified tags, and a semantic tag field is attached to each path node in the XML document, such as <artifact type> bronze ware< / 文物类型> <Architectural component> bracket< / 建筑部件> , thereby realizing the structural-visual-semantic trinity expression of graphic data; the final generated XML not only has the renderability and editability of the SVG standard, but also has a semantic structure for content understanding and retrieval, supporting in-depth application and intelligent linkage in graphic editing software, cultural tourism databases or cultural asset management systems.
[0101] S512: Output the color-calibrated image with added semantic tags as an SVG vector image file to obtain a vector image corresponding to the cultural tourism image.
[0102] Specifically, after completing the embedding of Bezier curve control points, color attributes and semantic tags in the XML structure, all path data are converted to <path>The SVG document body is written in node form, and the coordinate system, view window and graphic unit are set uniformly to ensure that the image geometry and color remain consistent in different rendering environments; the fill color attribute is marked through the fill parameter, the stroke color is defined by the stroke parameter, and the stroke-width is set to control the edge thickness, forming a graphic style that meets the visual perception requirements; semantic tags are added by adding a custom namespace (such as xmlns:semantic) under each path element or embedding the tag field as a comment to achieve semantic parsability and cross-system recognizability; during the output process, UTF-8 encoding is used to save the SVG file to ensure compatibility and international support; the final generated SVG vector file has precise structural boundary expression, color restoration effect and complete semantic information, which can be used for post-editing in graphic design software, and can also be directly embedded in cultural and tourism application systems for scenarios such as cultural relics display, digital twin modeling or semantic retrieval interaction, to achieve a deep integration of graphic expression and cultural semantic communication.
[0103] In this embodiment, the Bezier curve parameters (control point coordinates) and color attributes (fill or stroke color) are encoded into XML format, and semantic tags (such as <artifact type> bronze artifact) are added.< / 文物类型> <Architectural component> bracket< / 建筑部件> ), thereby improving the semantic information and editability of vector graphics. The output image is converted to an SVG vector file, which supports unlimited scaling without distortion, making it easy to reuse in CAD and illustration software.
[0104] This application adopts the above method, through collection, preprocessing, interactive segmentation, edge fitting, color allocation and semantic encoding, taking into account both automation efficiency and manual controllability, and solves the difficult problems of complex details, variable lighting, and irregular geometric structures of cultural and tourism images. The generated vector graphics have high precision, high semantics and high editability, providing key technical support for the digital inheritance and innovative application of cultural heritage.
[0105] Please refer to Figure 2 , which shows a module schematic diagram of an interactive cultural tourism image vectorization device provided by an embodiment of the present application, the device includes an acquisition module 21 and a processing module 22, wherein, The acquisition module 21 is used to acquire cultural and tourism images and preprocess the cultural and tourism images to obtain preprocessed images.
[0106] The processing module 22 is used to segment the preprocessed image to obtain a segmented image; perform edge extraction and fitting operations on the segmented image to obtain image edge features; perform color attribute assignment operations based on the image edge features to obtain a color calibration image; perform semantic encoding operations based on the color calibration image to obtain a vector diagram corresponding to the cultural and tourism image.
[0107] In a possible embodiment, the acquisition module 21 is used to acquire cultural and tourism images, specifically including: acquiring cultural and tourism images through preset acquisition equipment, the preset acquisition equipment including a digital camera and a scanner; acquiring the image type of the cultural and tourism image, and setting differentiated acquisition parameters based on the image type, the image type including cultural relics image type, architectural image type and landscape image type; based on the differentiated acquisition parameters, constructing the coordinate mapping relationship corresponding to the cultural and tourism image through geometric correction operations and diffraction mapping operations.
[0108] In a possible embodiment, the acquisition module 21 is used to preprocess the cultural and tourism image to obtain a preprocessed image, specifically including: converting the cultural and tourism image into a grayscale image through a weighted averaging method; determining the texture entropy of the cultural and tourism image based on the grayscale image, and performing hierarchical noise reduction on the cultural and tourism image through texture entropy; performing image enhancement processing on the denoised cultural and tourism image, and the image enhancement processing includes Retinex processing and CLAHE enhancement processing; based on the peak signal-to-noise ratio and structural similarity of the cultural and tourism image after image enhancement, performing an evaluation operation on the enhanced cultural and tourism image; if the evaluation operation is confirmed as the first evaluation result, returning the original cultural and tourism image and marking the noise area in the cultural and tourism image; if the evaluation operation is confirmed as the second evaluation result, adjusting the filter kernel size in the cultural and tourism image and enhancing the number of iterations; if the evaluation operation is confirmed as the third evaluation result, outputting the image-enhanced cultural and tourism image as a preprocessed image.
[0109] In one possible embodiment, the processing module 22 is used to segment the preprocessed image to obtain a segmented image, specifically including: constructing an image segmentation model based on a neural network structure and conditional random field that introduces an attention mechanism, and interactively marking the preprocessed image through the image segmentation model; performing a segmentation operation on the marked preprocessed image based on the image segmentation model to divide the marked preprocessed image into multiple connected regions; based on the characteristic parameters of each connected region, performing an erosion or expansion operation of the corresponding kernel on the connected region, and obtaining high-uncertainty regions in the multiple connected regions through a preset evaluation operation, wherein the characteristic parameters include area parameters, perimeter parameters, aspect ratio parameters, and edge complexity parameters; optimizing and iterating the high-uncertainty regions to complete the image segmentation operation, and outputting the segmented image.
[0110] In one possible embodiment, the processing module 22 is used to perform edge extraction and fitting operations on the segmented image to obtain image edge features, specifically including: performing edge extraction operations on each sub-element in the segmented image to obtain sub-element information; and using a cubic Bezier curve to perform a fitting operation on the sub-element information to obtain image edge features.
[0111] In one possible embodiment, the processing module 22 is used to perform color attribute assignment operations based on image edge features to obtain a color-calibrated image, specifically including: extracting pixel values for each connected area in the sub-element based on the coordinates of the image edge and the mapping relationship between the coordinates and the color values, and calculating the average color as the filling color; filling the cultural and tourism image with color based on the filling color, and judging whether the color of the image after color filling meets the preset requirements; if the image color does not meet the preset requirements, adding or subtracting the color through a preset interactive method; if the image color meets the preset requirements, the image after color filling is used as the color-calibrated image.
[0112] In one possible implementation, the processing module 22 is used to perform semantic encoding operations based on the color calibration image to obtain a vector diagram corresponding to the cultural and tourism image, specifically including: encoding the color attributes corresponding to the color calibration image into XML format, and adding corresponding semantic tags in the XML format; outputting the color calibration image with added semantic tags as an SVG vector diagram file to obtain a vector diagram corresponding to the cultural and tourism image.
[0113] It should be noted that the above embodiments provide devices that implement their functions using only the division of the above functional modules as examples. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0114] This application also provides an electronic device. Figure 3 , Figure 3 3 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. The electronic device may include: at least one processor 301, at least one communication bus 302, a user interface 303, at least one network interface 304, and a memory 305.
[0115] The communication bus 302 is used to implement the connection and communication between these components.
[0116] The user interface 303 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 303 may also include a standard wired interface and a wireless interface.
[0117] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).
[0118] The processor 301 may include one or more processing cores. Using various interfaces and circuits, the processor 301 connects to various components within the server. It executes instructions, programs, code sets, or instruction sets stored in the memory 305, as well as accesses data stored in the memory 305, to perform various server functions and process data. Optionally, the processor 301 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 301 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing content displayed on the display screen; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 301 but implemented as a separate chip.
[0119] Among them, the memory 305 may include a random access memory (RAM) or a read-only memory (Read-Only Memory). Optionally, the memory 305 includes a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 305 may also optionally be at least one storage device located away from the aforementioned processor 301. Refer to Figure 3 , as a computer storage medium, the memory 305 may include an operating system, a network communication module, a user interface module, and an interactive cultural and tourism image vectorization application.
[0120] exist Figure 3 In the electronic device shown, the user interface 303 is mainly used to provide an input interface for the user and obtain the data input by the user; and the processor 301 can be used to call the interactive cultural and tourism image vectorization application stored in the memory 305. When executed by one or more processors 301, the electronic device executes one or more of the methods described in the above embodiments. It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should know that this application is not limited to the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required for this application.
[0121] The present application also provides a computer-readable storage medium storing instructions, which, when executed by one or more processors, enable an electronic device to execute one or more of the methods described in the above embodiments.
[0122] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0123] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely schematic, such as the division of units, which is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interface, and the indirect coupling or communication connection of devices or units can be electrical or other forms.
[0124] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0125] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0126] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of this application. The aforementioned memory includes various media that can store program code, such as USB flash drives, mobile hard drives, magnetic disks, or optical disks.
[0127] The above descriptions are merely exemplary embodiments disclosed in this application and are not intended to limit the scope of this application. That is, any equivalent changes and modifications made based on the teachings disclosed in this application are still within the scope of this application.
[0128] This application is intended to cover any modifications, uses or adaptations disclosed in this application, which follow the general principles disclosed in this application and include common knowledge or customary technical means in the technical field not disclosed in this application.< / path>
Claims
1. An interactive cultural tourism image vectorization method, characterized in that: The method comprises: Acquire a cultural tourism image, and preprocess the cultural tourism image to obtain a preprocessed image; Segmenting the preprocessed image to obtain a segmented image; Performing edge extraction and fitting operations on the segmented image to obtain image edge features; Performing a color attribute assignment operation based on the image edge features to obtain a color calibrated image; A semantic encoding operation is performed based on the color calibration image to obtain a vector map corresponding to the cultural and tourism image.
2. The method according to claim 1, characterized in that The obtaining of cultural tourism images specifically includes: Acquire the cultural tourism image through a preset acquisition device, wherein the preset acquisition device includes a digital camera and a scanner; Obtaining the image type of the cultural tourism image, and setting differentiated acquisition parameters based on the image type, wherein the image type includes a cultural relic image type, a building image type, and a landscape image type; Based on the differentiated acquisition parameters, a coordinate mapping relationship corresponding to the cultural and tourism image is constructed through geometric correction operations and diffraction mapping operations.
3. The method according to claim 1, characterized in that Preprocessing the cultural tourism image to obtain a preprocessed image specifically includes: Converting the cultural tourism image into a grayscale image by a weighted average method; Determining the texture entropy of the cultural tourism image based on the grayscale image, and performing hierarchical noise reduction on the cultural tourism image according to the texture entropy; Performing image enhancement processing on the cultural tourism image after noise reduction, wherein the image enhancement processing includes Retinex processing and CLAHE enhancement processing; performing an evaluation operation on the enhanced cultural and tourism image based on the peak signal-to-noise ratio and structural similarity of the enhanced cultural and tourism image; If the evaluation operation is confirmed as the first evaluation result, returning the original cultural tourism image and marking the noise area in the cultural tourism image; If the evaluation operation is confirmed as the second evaluation result, adjusting the filter kernel size in the cultural tourism image and increasing the number of iterations; If the evaluation operation is confirmed as the third evaluation result, the cultural and tourism image after image enhancement is output as the pre-processed image.
4. The method according to claim 1, wherein Segmenting the preprocessed image to obtain a segmented image specifically includes: Building an image segmentation model based on a neural network structure and conditional random field with an attention mechanism, and interactively labeling the preprocessed image through the image segmentation model; Based on the image segmentation model, performing a segmentation operation on the marked pre-processed image to segment the marked pre-processed image into a plurality of connected regions; Based on characteristic parameters of each of the connected regions, performing an erosion or dilation operation corresponding to a core on the connected regions, and obtaining high uncertainty regions in the multiple connected regions through a preset evaluation operation, wherein the characteristic parameters include an area parameter, a perimeter parameter, an aspect ratio parameter, and an edge complexity parameter; An optimization iteration is performed on the high uncertainty region to complete an image segmentation operation, and the segmented image is output.
5. The method according to claim 1, wherein The performing edge extraction and fitting operations on the segmented image to obtain image edge features specifically includes: Performing the edge extraction operation on each sub-element in the segmented image to obtain sub-element information; The fitting operation is performed on the sub-element information using a cubic Bezier curve to obtain the image edge feature.
6. The method according to claim 5, characterized in that The performing of a color attribute assignment operation based on the image edge features to obtain a color calibrated image specifically includes: Based on the coordinates of the image edge and the mapping relationship between the coordinates and the color values, pixel values are extracted from each connected area in the sub-element, and the average color is calculated as the filling color; Filling the cultural tourism image with color based on the filling color, and determining whether the color of the image after color filling meets the preset requirements; If the image color does not meet the preset requirements, then increase or decrease the color through a preset interactive method; If the color of the image meets the preset requirement, the image after color filling is used as the color calibration image.
7. The method according to claim 1, characterized in that The performing of a semantic encoding operation based on the color calibration image to obtain a vector graph corresponding to the cultural tourism image specifically includes: Encoding the color attributes corresponding to the color calibration image into XML format, and adding corresponding semantic tags in the XML format; The color calibration image after adding semantic tags is output as an SVG vector image file to obtain the vector image corresponding to the cultural tourism image.
8. An interactive cultural tourism image vectorization device, characterized in that: The device includes an acquisition module and a processing module, wherein: The acquisition module is used to acquire cultural tourism images and preprocess the cultural tourism images to obtain preprocessed images; The processing module is used to segment the preprocessed image to obtain a segmented image; perform edge extraction and fitting operations on the segmented image to obtain image edge features; perform color attribute assignment operations based on the image edge features to obtain a color calibration image; and perform semantic encoding operations based on the color calibration image to obtain a vector diagram corresponding to the cultural and tourism image.
9. An electronic device, characterized in that: The electronic device comprises a processor, a communication bus, a user interface, a network interface and a memory, wherein the memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1 to 7 is performed.
Citation Information
Patent Citations
Three-dimensional image segmentation method based on double-path attention coding and decoding network
CN113643303A
Fabric image recoloring method and system
CN114067014A
Pixel picture vectorization method for textiles
CN116912338A
Computer vision-based travel lighting display effect evaluation method and system
CN118968234A
Image processing and vectorisation
WO2008003944A2