Marine surveying and mapping picture data vectorization method
Through the methods of chart image registration and sensitive band selection, image processing, point element and line element extraction, the problem of low recognition accuracy of point element and line element in vectorization of marine surveying and mapping pictures is solved, efficient chart vectorization processing is achieved, and the application value of chart data is enhanced.
Patent Information
- Application Number
- CN202510561989.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-12
AI Technical Summary
The existing marine mapping image vectorization technology has low accuracy and low automation when identifying point and line elements on the chart, making it difficult to achieve efficient vectorization processing.
The methods of chart image registration, sensitive band selection, image processing, point feature extraction and line feature extraction are adopted, including coarse registration and precision registration, optimal index index algorithm, brightness contrast enhancement, Gaussian blur, sharpening, color clustering, Tesseract engine training model and disconnection processing, which improves point feature positioning accuracy and line feature continuity.
It improves the accuracy and automation of vectorization of marine surveying and mapping pictures, ensures high readability and availability of chart information, broadens the application areas of chart data, and supports in-depth research such as coastal zone change research, waterway analysis, and marine trade analysis.
Smart Images

Figure CN120471778A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of ocean information, and in particular to a method for vectorizing ocean surveying and mapping picture data. Background Art
[0002] With the widespread adoption of computer technology and the increasing application of GIS, the demand for nautical chart data is increasing. Nautical charts support a wide range of applications, including maritime navigation, marine resource development, and marine environmental protection. For example, they provide precise navigation information to help avoid hazardous areas; they provide detailed seafloor topography and geographic information to aid resource exploration, wind farm construction, and the demarcation of marine protected areas. Vectorization of marine surveying and mapping imagery is the process of converting raster images (such as nautical charts and seafloor topographic maps) into vector data for further analysis and application in GIS. This process primarily involves accurately converting chart information, including soundings, depth contours, and coastlines, into digital format. Vectorization of nautical chart images is of great significance. Firstly, it improves the display quality of nautical chart information and enhances data accuracy. It maintains high clarity even when zooming, eliminates pixelation, and provides high readability and usability. Chart data can be dynamically rendered and interactive, allowing for customized display based on user needs, enhancing the user experience. Secondly, it broadens the application of nautical charts. They can be imported into geographic information systems, facilitating in-depth research and analysis of nautical chart data, expanding its application areas, such as coastal zone change research, waterway analysis, maritime trade analysis, geopolitical analysis, and precise geographic information support for military navigation. Therefore, high-quality, GIS-compliant nautical chart vectorization products are an important foundational task in establishing GIS.
[0003] Currently, one of the primary methods for producing nautical chart vectorization products is to convert existing paper charts into images and then perform vectorization. The most common point and line features on chart images are bathymetric points and depth contours. Due to the complexity and diversity of features in nautical charts, existing vectorization technologies suffer from low recognition accuracy for these features. Point features are inaccurately positioned, line features are difficult to fully extract, and the level of automation and real-time performance are limited. Summary of the Invention
[0004] To achieve the above-mentioned and other related purposes, the present invention discloses a method for vectorizing marine surveying and mapping image data, comprising the following steps: S1: Chart Image Registration: Perform coarse and fine registration of historical chart images to achieve spatial alignment between the image and the geographic coordinate system; S2: Sensitive Band Selection: Based on the Optimal Index (OIF) algorithm, the Sensitive Band (SWB) is selected from the red, green, and blue bands and characteristic indices, including the Enhanced Water Index (EWI), the Extra Green Index (EXG), the Normalized Green-Blue Difference Index (NGBDI), and the Visible Vegetation Index (GLI); S3: Image processing: Perform brightness contrast enhancement, Gaussian blur, sharpening, color clustering and binarization on SWB to generate a binary image with enhanced features; S4: Point feature extraction: Use the Tesseract engine to train a custom model to identify the characters and positions of point features in binary images and generate vector point data and attribute tables; S5: Line feature extraction: Delete the point features extracted from the binary image. After deleting the point features, perform line connection processing on the binary image and generate continuous vector line features based on the endpoint distance and tangent direction.
[0005] Furthermore, the implementation of the sensitive band selection in step S2 includes: Calculate the standard deviation and correlation coefficient of each candidate band, determine the OIF value based on the formula and sort in descending order: ; in is the standard deviation of the i-th band, is the correlation coefficient between bands i and j; The band with the highest OIF value is selected as SWB.
[0006] Furthermore, the image processing in step S3 includes: Through the linear transformation formula: ; Adjust image contrast ( ) and brightness ( ); Blurring and sharpening images; The image is converted to HSV color space for K-Means clustering, and the color is compressed and restored to RGB space to generate a binary image.
[0007] Furthermore, the midpoint feature extraction in step S4 includes: Model training is performed using Tesseract's LSTM training tool; Use the trained Tesseract LSTM model to recognize characters in the slice, use the center of the minimum bounding rectangle as the coordinates of the point feature, export it as vector data, and automatically fill in the recognized numbers and letters in the corresponding attribute table columns of the vector data.
[0008] Furthermore, the training of the Tesseract model includes: Mark water depth point numbers and letter samples, and perform rotation, scaling, and denoising enhancement; Set the maximum number of iterations and error rate to perform LSTM training; After the model training is completed, it is tested and verified. If the model accuracy does not meet the preset requirements, the number of samples and training parameters are adjusted and training is performed again until the model accuracy meets the preset requirements.
[0009] Furthermore, the specific method of disconnecting the line in step S5 is: Extract line feature endpoints and calculate the distance between endpoints and tangent direction angle ; satisfy and When , connected pixels are generated, where is the preset dashed line interval threshold.
[0010] Furthermore, the implementation of color clustering includes: Color space conversion, normalize the RGB grayscale values, calculate the maximum and minimum values of the normalized grayscale values, and obtain the hue H, saturation S, and lightness V based on the maximum and minimum values. The image color is converted from the RGB color space to the HSV color space; Initialize the cluster center, generate the initial cluster center by randomly scattering points, and obtain the pixel color of the initial cluster center: ; in, It is Initial cluster centers; Aggregate pixels to the nearest cluster center and classify them by calculating the Euclidean distance of each pixel from all cluster centers: ; in, Indicates the cluster to which the target pixel x belongs, It is The first iteration cluster centers, is the pixel x and the cluster center The Euclidean distance between Update the cluster center, for each cluster , calculate the average value of all its pixels and update the cluster center: ; The classified pixel colors are uniformly replaced with the cluster center pixel colors to perform color compression and eliminate noise; Color space restoration, calculate auxiliary variables, determine RGB components according to hue H, and finally calculate the final RGB value to restore the HSV color space to RGB color space.
[0011] On the other hand, the present invention discloses a computer-readable storage medium having a computer program stored thereon, which implements the above method when executed by a processor.
[0012] By adopting the above technical solutions, the spatial position accuracy is improved through chart image registration, the optimal index indexing algorithm is used to select sensitive bands to compress the data volume and enhance the separability of features, and brightness contrast adjustment, Gaussian blur, sharpening and K-Means color clustering are combined to eliminate noise and color interference. Tesseract LSTM model training is used to achieve high-precision recognition of characters in multiple fonts, directions and sizes. The broken lines are automatically connected based on the endpoint distance and tangent direction to solve the problem of broken dotted line intervals, improve the positioning accuracy of point features and the continuity of line features, reduce manual intervention, and adapt to the needs of automated vectorization of complex chart features. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. The accompanying drawings are provided for a better understanding of the present disclosure and do not constitute a limitation of the present disclosure. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, among which: Figure 1 is a flow chart of the present invention; Figure 2 Schematic diagram of the disconnection connection method of the present invention; Figure 3 Extract results for line and point features. DETAILED DESCRIPTION
[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0015] Reference Figure 1 The embodiment of the present invention provides a method for vectorizing marine surveying and mapping image data, comprising the following steps: S1: Chart Image Registration: Perform coarse and fine registration of historical chart images to achieve spatial alignment between the image and the geographic coordinate system; S2: Sensitive Band Selection: Based on the Optimal Index (OIF) algorithm, the Sensitive Band (SWB) is selected from the red, green, and blue bands and characteristic indices, including the Enhanced Water Index (EWI), the Extra Green Index (EXG), the Normalized Green-Blue Difference Index (NGBDI), and the Visible Vegetation Index (GLI); S3: Image processing: Perform brightness contrast enhancement, Gaussian blur, sharpening, color clustering and binarization on SWB to generate a binary image with enhanced features; S4: Point feature extraction: Use the Tesseract engine to train a custom model to identify the characters and positions of point features in binary images and generate vector point data and attribute tables; S5: Line feature extraction: Delete the point features extracted from the binary image. After deleting the point features, perform line connection processing on the binary image and generate continuous vector line features based on the endpoint distance and tangent direction.
[0016] The above steps in this embodiment specifically include: Chart image registration. Historical chart images are mostly sourced from historical electronic data and paper charts from national hydrographic surveying departments. Coarse and fine registration is performed on photographed or scanned images to improve vectorization accuracy.
[0017] Sensitive band selection. The nautical chart image is a three-band visible light image with a wide coverage area. It contains three bands: red, green, and blue. The surface elements are mainly sea areas and land areas. Among them, the sea areas are blue with different brightness after symbolization, and the land areas are yellow with different brightness after symbolization. The linear elements are mainly isobaths and contour lines, which are dark colors such as black, purple, and blue after symbolization. The optimal index (OIF) algorithm is comprehensively used to extract the optimal band. The OIF calculation formula is: 2-1: Based on the spectral characteristics of the symbolized chart image, we select indices that are sensitive to dark colors and blue elements in the image, including the Enhanced Water Index (EWI), the Extra Green Index (EXG), the Normalized Green-Blue Difference Index (NGBDI), and the Visible Light Vegetation Index (GLI). These indices and the red, green, and blue bands are considered as candidate bands. The calculation formulas for these indices are as follows:
[0018]
[0019]
[0020]
[0021] Among them: Red, Green, and Blue represent the red, green, and blue bands respectively.
[0022] 2-2: Calculate the standard deviation and correlation coefficient of each candidate band. Based on the standard deviation and correlation coefficient, calculate the OIF value of each candidate band and sort them in descending order by OIF value. The OIF calculation formula is as follows:
[0023] in: Representative The standard deviation of the candidate bands, Indicates the Candidate bands and The correlation coefficient between candidate bands, Indicates the number of candidate bands.
[0024] 2-3: Based on the OIF sequence, the band with the highest OIF value is selected as the sensitive band (SWB) for subsequent vectoring.
[0025] Image processing: Perform image processing operations such as contrast enhancement, sharpening and smoothing on SWB to improve the clarity of linear features and enhance their separability.
[0026] 3-1: Brightness and contrast enhancement: The brightness and contrast of SWB are enhanced by linear changes.
[0027]
[0028] in: For the enhanced image, is the contrast adjustment factor, which increases the contrast when it is greater than 1 and decreases when it is less than 1; It is the brightness adjustment factor. If it is greater than 0, the brightness increases; if it is less than 0, the brightness decreases.
[0029] 3-2: Gaussian Blur and Sharpening. As these are historical materials from a long time ago, the images inevitably contain some pixel contamination, such as mosaics and brightness / luminance noise. To address these issues, we first perform a Gaussian blur to weaken the contaminated pixels, then filter to remove color noise, and finally perform sharpening to improve the pixel purity and line clarity of the target elements in the image.
[0030] Gaussian blur, set the standard deviation, generate a Gaussian kernel, set the standard deviation to 1 and the Gaussian kernel size to 5×5 to retain the details of the linear features, and perform Gaussian blur processing on the enhanced image.
[0031] Image filtering, bilateral filtering combines the spatial proximity and pixel value similarity of the image. Using bilateral filtering can remove noise while retaining edge details. The calculation formula of bilateral filtering is as follows:
[0032]
[0033] Among them, p and q are the coordinates of the center pixel and the neighborhood pixel respectively, is the standard deviation in the spatial domain, is the spatial domain weight, is the range weight, x and y are the target pixel locations, is the neighborhood window of the filter, i and j are row and column numbers respectively, is a normalization factor that ensures the weights sum to 1, is the filtered image.
[0034] Image sharpening judges the pixel grayscale and sharpens the pixels that meet the conditions, while leaving the rest unprocessed. The calculation formula is as follows:
[0035]
[0036] Where D is the sharpened image, h is a constant, and T is the threshold.
[0037] 3-2: Color clustering and rendering: Cluster similar colors in the image to eliminate the influence of noise; then render the image to highlight the linear and point features in the image.
[0038] Color clustering first converts the image from RGB color space to HSV color space, then divides the image into different regions, uses the K-Means algorithm to group the colors in the image, and finally maps the colors in the original image to the selected color set to achieve color simplification.
[0039] Step 1: Color space conversion. First, normalize the RGB grayscale values, then calculate the maximum and minimum values of the normalized grayscale values. According to the maximum and minimum values, the hue H, saturation S, and lightness V are obtained. The image color is converted from the RGB color space to the HSV color space.
[0040] Step 2: Initialize the cluster center, generate the initial cluster center by randomly scattering points, and obtain the pixel color of the initial cluster center:
[0041] in, It is Initial cluster centers.
[0042] Step 3: Gather pixels to the nearest cluster center and classify the pixels by calculating the Euclidean distance of each pixel from all cluster centers:
[0043] in, Indicates the cluster to which the target pixel x belongs, It is The first iteration cluster centers, is the pixel x and the cluster center The Euclidean distance between .
[0044] Step 4: Update the cluster center for each cluster , calculate the average value of all its pixels and update the cluster center:
[0045] Step 5: Replace the classified pixel colors with the cluster center pixel colors to perform color compression and eliminate noise.
[0046] Step 6: Color space restoration. First, calculate the auxiliary variables, then determine the RGB components according to the hue H, and finally calculate the final RGB value to restore the HSV color space to the RGB color space.
[0047] Feature rendering generates edge lines based on the edge and texture direction of the features on the image, retains the key details of the features on the image, and finally generates a binary image.
[0048] Step 1: Use the Sobel operator to detect the edge and texture direction in the image and obtain the gradient image and gradient direction;
[0049]
[0050] in, and are the horizontal and vertical filtering operators respectively, G is the gradient image, is the gradient direction.
[0051] Step 2: Generate edge lines based on the gradient image and gradient direction, divide the image into multiple regions based on the watershed algorithm and the initial lines, and use pixel mean replacement in each region.
[0052] Step 3: Low-frequency information suppression. This method retains high-frequency information such as edges and textures, while suppressing low-frequency information such as smooth areas. The image is Gaussian blurred to obtain low-frequency information. The difference between the images before and after Gaussian blurring is then used to obtain high-frequency information. This emphasizes linear, point-like, and other symbolic elements, while minimizing the impact of surface areas of varying colors. The processed image is then binarized using the Otsu algorithm.
[0053] Step 4: Perform morphological closing operation on the binarized image to fill the broken parts of numbers and symbols on the image and smooth the boundaries.
[0054] Point feature extraction. The most common point features on nautical charts are bathymetric points. Other features include those related to navigation, marine resources, and the marine environment, such as lighthouses, buoys, shipwrecks, and reefs. After symbolization, bathymetric points appear as numbers of varying fonts, orientations, and sizes on the binary image. Tesseract is used to identify bathymetric values and locations, as well as other alphabetic features. Tesseract is an open-source text recognition (ORC) engine that supports text recognition in over 100 languages. It can perform custom training and LSTM training, achieving high recognition accuracy.
[0055] 4-1: Tesseract model selection and environment deployment.
[0056] Download the latest version installation package, add the Tesseract installation path to the system environment variable PATH, and verify the installation; select the TesseractOCR Docker image, pull the required image and run the container, install the Python library, and configure the Tesseract path.
[0057] 4-2: Training sample data generation.
[0058] Mark the symbolized water depth points and use a rectangular box to select the characters. When selecting, ensure that multiple numbers and consecutive letters representing a point feature are within a rectangular box. The rectangular box is the minimum circumscribed rectangle. The number label is 1 and the letter label is 2. Clean and enhance the labeled sample data to remove unclear and incomplete sample data, and perform enhancement operations on the cleaned sample data, including rotation, scaling, and denoising operations of varying magnitudes, to improve recognition accuracy; Convert the enhanced samples into .box files and verify the .box files to ensure the accuracy of the annotated bounding boxes.
[0059] 4-3: Model training and optimal model generation.
[0060] The model is trained using Tesseract's LSTM training tool. The corresponding character set and dictionary are selected based on the data type on the chart. During training, the maximum number of iterations is not less than 5000, the expected error rate is not higher than 0.05, and the debug print level is set to -5.
[0061] After training is completed, testing and verification are carried out. If the model accuracy does not meet the requirements, the number of samples is increased, and the training parameters are adjusted and trained again until the model accuracy meets the requirements.
[0062] 4-4: Point feature and letter feature extraction.
[0063] Coarsely locate and slice the target area. Because nautical charts are large, overall image reasoning is inefficient and subject to numerous interference factors. First, locate the target area where numbers and letters are concentrated. Then, remove the invalid areas and slice the valid areas.
[0064] Step 1: Select clear, distinctive, and multi-angled sample images of numbers and letters as template images; Step 2: Calculate the correlation coefficient between each connected domain and the template image , if the correlation coefficient is greater than 0.5, it is marked as the target connected domain, and the geometric center of the target connected domain is marked.
[0065]
[0066] in, For template pictures, is the picture to be identified, are the means of the template image and the image to be identified, respectively, and i, j are the horizontal and vertical coordinates of the pixel.
[0067] Step 3: Delete non-target connected domains, perform cluster analysis on the geometric centers of the marked target connected domains using the K-means algorithm, divide the geometric centers into multiple clusters, generate a minimum bounding rectangle for each cluster, and crop the image according to the minimum bounding rectangle.
[0068] Step 4: Slice the cropped image, setting the horizontal and vertical overlap between each slice to 10%. When slicing, ensure that multiple numbers and consecutive letters representing a point feature are within one image to improve image inference speed.
[0069] Number and letter recognition. Use the trained optimal model to identify each slice one by one, read the minimum bounding rectangle and value of the recognized numbers and letters, take the center point of the minimum bounding rectangle as the point feature, export it as vector data, and automatically fill in the corresponding attribute table column of the vector data with the recognized number and letter value.
[0070] Line feature extraction. The most common line features on a nautical chart are coastlines and depth contours, while others include roads, depth limits, and waterways. After symbolization, line features appear as straight lines of different colors or dashed lines with different intervals on the image. First, clean and repair the binary image to eliminate interference from other elements, connect the dashed lines with intervals or the lines that are disconnected after binarization, and finally vectorize the processed image. See the attached diagram for a schematic diagram of the broken line connection method. Figure 2 .
[0071] 5-1: Image cleaning. Delete the point features extracted from the binary image and clean the binary raster data; 5-2: Connecting Broken Lines. Use the ArcScan tool to vectorize the binary raster data. After vectorization, convert it to a binary image to eliminate the influence of line thickness in the binary image. Then, perform a neighborhood check on the pixels on the line segment. If there is only one pixel in the neighborhood, it is the line segment endpoint. Based on the endpoint, locate the only pixel in the neighborhood. This is the pixel connecting the endpoint. Based on these two pixels, determine the tangent line at the endpoint and determine whether to generate a connecting line.
[0072]
[0073]
[0074]
[0075] in,( , )、( , ) is the pixel coordinates of the endpoints of a line feature and the connecting endpoints, ( , )、( , ) is the pixel coordinates of the endpoint of another line feature and the connecting endpoint. If ≤45°, ≤ ( is a dotted line interval), a connected pixel is generated between the two pixels.
[0076] 5-3: Line feature generation: The binary image after connecting the broken lines is vectorized again to generate line features.
[0077] Examples of point and line feature extraction are attached. Figure 3 .
[0078] Those skilled in the art will understand that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those skilled in the art in the art to which the present invention pertains. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with those in the context of the prior art and, unless specifically defined, will not be interpreted in an idealized or overly formal sense.
[0079] For simplicity of description, the method embodiments are described as a series of actions. However, those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because certain steps can be performed in other orders or simultaneously according to the embodiments of the present invention. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0080] Through the description of the above embodiments, it can be seen that those skilled in the art can clearly understand that the present application can be implemented by means of software plus the necessary general hardware platform. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application or certain parts of the embodiments.
[0081] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for vectorizing marine surveying and mapping image data, characterized in that: The following steps are involved: S1: Chart Image Registration: Perform coarse and fine registration of historical chart images to achieve spatial alignment between the image and the geographic coordinate system; S2: Sensitive Band Selection: Based on the Optimal Index (OIF) algorithm, the Sensitive Band (SWB) is selected from the red, green, and blue bands and characteristic indices, including the Enhanced Water Index (EWI), the Extra Green Index (EXG), the Normalized Green-Blue Difference Index (NGBDI), and the Visible Vegetation Index (GLI); S3: Image processing: Perform brightness contrast enhancement, Gaussian blur, sharpening, color clustering and binarization on SWB to generate a binary image with enhanced features; S4: Point feature extraction: Use the Tesseract engine to train a custom model to identify the characters and positions of point features in binary images and generate vector point data and attribute tables; S5: Line feature extraction: Delete the point features extracted from the binary image. After deleting the point features, perform line connection processing on the binary image and generate continuous vector line features based on the endpoint distance and tangent direction.
2. The method according to claim 1, characterized in that The implementation of the sensitive band selection in step S2 includes: Calculate the standard deviation and correlation coefficient of each candidate band, determine the OIF value based on the formula and sort in descending order: ; in is the standard deviation of the i-th band, is the correlation coefficient between bands i and j, and M is the number of bands; The band with the highest OIF value is selected as SWB.
3. The method according to claim 1, characterized in that The image processing in step S3 includes: Through the linear transformation formula: ; Adjust image contrast ( ) and brightness ( ); Blurring and sharpening images; The image is converted to HSV color space for K-Means clustering, and the color is compressed and restored to RGB space to generate a binary image.
4. The method according to claim 1, wherein The midpoint feature extraction in step S4 includes: Model training is performed using Tesseract's LSTM training tool; Use the trained Tesseract LSTM model to recognize characters in the slice, use the center of the minimum bounding rectangle as the coordinates of the point feature, export it as vector data, and automatically fill in the recognized numbers and letters in the corresponding attribute table columns of the vector data.
5. The method according to claim 4, characterized in that Training the Tesseract model involves: Mark water depth point numbers and letter samples, and perform rotation, scaling, and denoising enhancement; Set the maximum number of iterations and error rate to perform LSTM training; After the model training is completed, it is tested and verified. If the model accuracy does not meet the preset requirements, the number of samples and training parameters are adjusted and training is performed again until the model accuracy meets the preset requirements.
6. The method according to claim 1, wherein The specific method of disconnecting the line in step S5 is: Extract line feature endpoints and calculate the distance between endpoints and tangent direction angle ; satisfy and When , connected pixels are generated, where is the preset dashed line interval threshold.
7. The method according to claim 3, characterized in that The implementation of color clustering includes: Color space conversion, normalize the RGB grayscale values, calculate the maximum and minimum values of the normalized grayscale values, and obtain the hue H, saturation S, and lightness V based on the maximum and minimum values. The image color is converted from the RGB color space to the HSV color space; Initialize the cluster center, generate the initial cluster center by randomly scattering points, and obtain the pixel color of the initial cluster center: ; in, It is Initial cluster centers; Aggregate pixels to the nearest cluster center and classify them by calculating the Euclidean distance of each pixel from all cluster centers: ; in, Indicates the cluster to which the target pixel x belongs, It is The first iteration cluster centers, is the pixel x and the cluster center The Euclidean distance between Update the cluster center, for each cluster , calculate the average value of all its pixels and update the cluster center: ; The classified pixel colors are uniformly replaced with the cluster center pixel colors to perform color compression and eliminate noise; Color space restoration, calculate auxiliary variables, determine RGB components according to hue H, and finally calculate the final RGB value to restore the HSV color space to RGB color space.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.