Image character edge optimization method, device and system and storage medium

By combining multimodal noise reduction and edge-aware detection with a deep learning model, subpixel-level repair and refinement are achieved, solving the problems of brokenness, blurring and noise in image text edge processing in existing technologies. This enables efficient and accurate text edge optimization, suitable for various application scenarios.

CN121032844APending Publication Date: 2025-11-28BEIJING TIANJIU RENHE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511140422.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing image text edge processing methods are prone to producing broken or false edges in low-quality images. Traditional methods are sensitive to noise, while deep learning methods rely on a large amount of labeled data and have high model complexity, making it difficult to adapt to multilingual and multi-font scenarios, resulting in obvious jagged edges.

Method used

The process employs multimodal denoising, edge-aware detection, subpixel-level repair and thinning, multi-directional smoothing, and quality assessment. It combines the Canny algorithm and the SegFormer neural network, integrates semantic information and geometric features, achieves subpixel-level localization through Zernike moments and least squares, enhances stroke continuity with Gabor multi-directional filtering and maximum value fusion, and balances denoising and detail preservation with bilateral filtering and three-level smoothing. Quantitative evaluation is performed based on edge integrity, peak signal-to-noise ratio, and structural similarity.

Benefits of technology

It achieves more accurate text edge extraction in complex scenes, eliminates jagged edges, adapts to multilingual and multi-font scenarios, improves the accuracy and efficiency of edge processing, and is suitable for ancient book restoration, industrial inspection, and natural scene text recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121032844A_ABST
    Figure CN121032844A_ABST
Patent Text Reader

Abstract

The invention discloses an image character edge optimization method, device and system and a storage medium, and the method comprises the steps: carrying out the preprocessing of an input image, and the preprocessing comprises the multi-mode noise reduction, graying and contrast enhancement; performing edge perception detection on the preprocessed image, wherein the edge perception detection comprises edge extraction and particle swarm optimization edge linking by using an edge perception neural network model; edge repairing and refining are carried out on an edge image obtained through edge perception detection; edge smoothing and enhancement are carried out on the repaired and refined edge image; post-processing is carried out on the smoothed and enhanced edge image, the post-processing comprises binarization, sharpening and quality evaluation, and if the quality evaluation does not reach the standard, edge perception detection is carried out again. According to the technical scheme, the edge processing effect is optimized, edge extraction is more accurate, and the sawtooth phenomenon is eliminated.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to an image text edge optimization method, device, system and storage medium. BACKGROUND

[0002] The existing image text edge processing method is mainly based on traditional algorithms such as Canny, Sobel, etc. edge detection operator, combined with morphological processing or deep learning method such as HED, RCF, etc. edge detection network. Although these methods can extract the text edge, they still have the following shortcomings: the traditional method is sensitive to noise, and is easy to produce fracture or false edge in low-quality image; the method based on deep learning relies on a large amount of labeled data, and the model complexity is high, which is difficult to adapt to multi-language and multi-font scene; the existing method is mostly limited to pixel-level edge extraction, resulting in obvious sawtooth phenomenon. These problems limit the application effect of text edge processing in complex scenes. SUMMARY

[0003] The main purpose of the present application is to provide an image text edge optimization method, device, system and storage medium, which aims to optimize the edge processing effect, and the edge extraction is more accurate and the sawtooth phenomenon is eliminated.

[0004] To achieve the above purpose, the present application provides an image text edge optimization method, which comprises:

[0005] S100, pre-processing the input image, the pre-processing including multi-modal noise reduction, graying and contrast enhancement;

[0006] S200, edge perception detection on the pre-processed image, the edge perception detection including edge extraction by edge perception neural network model and edge link optimization by particle swarm optimization;

[0007] S300, edge repair and thinning on the edge image obtained by edge perception detection;

[0008] S400, edge smoothing and enhancement on the edge image after repair and thinning;

[0009] S500, post-processing on the edge image after smoothing and enhancement, the post-processing including binarization, sharpening and quality assessment, if the quality assessment is not up to standard, returning to step S200 to re-perform edge perception detection.

[0010] In a possible implementation, the step S100 comprises:

[0011] S110, removing impulse noise by adaptive median filtering, the window size is dynamically adjusted according to noise density, and when noise pixels are detected, the window size is increased until the proportion of non-noise pixels in the window is more than 60%;

[0012] S120, smoothing Gaussian noise by combining Gaussian filtering;

[0013] S130, converting the RGB image into a grayscale image;

[0014] S140, enhancing local contrast by using an adaptive histogram equalization algorithm, dividing the image into multiple sub-blocks, and performing histogram equalization on each sub-block.

[0015] In a possible implementation, the step S200 of using an edge-aware neural network model to extract edges includes:

[0016] S210, first extracting global edges using the Canny algorithm, then filtering non-text edges through a lightweight text detection model, and finally multiplying the text mask and the Canny edges pixel by pixel to obtain the edges of the text region;

[0017] S220, based on the SegFormer architecture, fusing the text region edges and the original image features, and extracting local features by dividing the image into overlapping regions for feature extraction;

[0018] S230, fusing multi-level features and outputting a text mask.

[0019] In a possible implementation, the step S300 includes:

[0020] S310, first dilating to expand the target boundary in the image to connect broken edges, and then eroding to shrink the target boundary in the image, and iterating at least once to remove burrs;

[0021] S320, calculating the sub-pixel coordinates of the edge points based on the Zernike moment function template, and fitting the character contour using the least squares method;

[0022] S330, performing three-level bilateral filtering on the edge image.

[0023] In a possible implementation, the step S400 includes:

[0024] S410, using a Gabor filter to perform multi-directional filtering on the high-pass image after bilateral filtering;

[0025] S420, calculating the response values of each directional filter, fusing each directional response through a maximum value fusion strategy, and enhancing the continuity of the character strokes.

[0026] S430, calculating the sub-pixel offset based on the Zernike moment function template;

[0027] S440, point-by-point correction is performed on the fitted character contour to generate sub-pixel level edge coordinates, and the distance variation rate of adjacent points is ensured to be less than 10% during the correction process.

[0028] In a possible implementation, the quality evaluation in the step S500 includes:

[0029] S510, edge integrity, peak signal-to-noise ratio, and structural similarity are calculated.

[0030] S520, it is judged whether the calculated value reaches a preset threshold value, if not, the iteration optimization is automatically returned to the step S200, and the iteration number is at most 3 times, if the threshold value is not reached after 3 iterations, the current optimal result is output and it is prompted that the image quality is too low.

[0031] The application further provides an image character edge optimization system, comprising:

[0032] A preprocessing unit is configured to preprocess the input image, and the preprocessing includes multi-modal noise reduction, grayscale, and contrast enhancement.

[0033] An edge detection unit is configured to perform edge perception detection on the preprocessed image, and the edge perception detection includes edge extraction using an edge perception neural network model and particle swarm optimization edge linking.

[0034] An edge repair unit is configured to perform edge repair and thinning on the edge image obtained by the edge perception detection.

[0035] An edge enhancement unit is configured to perform edge smoothing and enhancement on the edge image after repair and thinning.

[0036] A post-processing unit is configured to perform post-processing on the edge image after smoothing and enhancement, and the post-processing includes binarization, sharpening, and quality evaluation, and if the quality evaluation does not reach the threshold value, the edge detection unit is returned to perform edge perception detection again.

[0037] The application further provides an electronic device, comprising:

[0038] A memory is configured to save a computer program.

[0039] A processor is configured to execute the computer program to implement the image character edge optimization method according to any one of the above.

[0040] The application further provides a computer readable storage medium configured to store a computer program, and the computer program is executed by a processor to implement the image character edge optimization method according to any one of the above.

[0041] The technical scheme of the present application realizes the following core advantages by adopting a complete process of multi-modal noise reduction, edge perception detection, sub-pixel level repair and refinement, multi-direction smoothing and quality evaluation: combining Canny algorithm and SegFormer neural network, fusing semantic information and geometric features, accurately separating text edges; realizing sub-pixel level positioning through Zernike moment and least square method, eliminating jaggies, and adapting to inclined and curved characters; Gabor multi-direction filtering and maximum fusion enhance stroke continuity, three-level smoothing of bilateral filtering balances noise reduction and detail preservation; based on edge integrity, peak signal-to-noise ratio and structural similarity quantitative evaluation, automatic iterative optimization ensures that the output quality meets the standard. The problems of breakage, blur, noise and geometric distortion in text edge processing are systematically solved, and the system has efficiency and precision, and is suitable for ancient book restoration, industrial detection, natural scene text recognition and other applications. BRIEF DESCRIPTION OF DRAWINGS

[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from the structures shown in the drawings without creative labor.

[0043] Figure 1 The step flowchart of the image text edge optimization method of the present application is shown in the figure.

[0044] Figure 2 The step flowchart of the pre-processing of the input image of the present application is shown in the figure.

[0045] Figure 3 The step flowchart of the edge perception detection of the pre-processed image of the present application is shown in the figure.

[0046] Figure 4 The step flowchart of the edge repair and refinement of the edge image obtained by edge perception detection of the present application is shown in the figure.

[0047] Figure 5 The step flowchart of the edge smoothing and enhancement of the edge image after repair and refinement of the present application is shown in the figure.

[0048] Figure 6 The step flowchart of the post-processing of the edge image after smoothing and enhancement of the present application is shown in the figure.

[0049] The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0050] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application.

[0051] With reference to Figure 1 The present application provides an image text edge optimization method, comprising:

[0052] S100, pre-processing the input image, the pre-processing including multi-modal noise reduction, grayscale and contrast enhancement;

[0053] S200, edge perception detection on the pre-processed image, the edge perception detection including edge extraction using an edge perception neural network model and particle swarm optimization edge linking;

[0054] S300, edge repair and thinning on the edge image obtained by edge perception detection;

[0055] S400, edge smoothing and enhancement on the repaired and thinned edge image;

[0056] S500, post-processing on the smoothed and enhanced edge image, the post-processing including binarization, sharpening and quality assessment, and if the quality assessment is not up to standard, returning to step S200 to re-perform edge perception detection.

[0057] It can be understood that in S100, multi-modal noise reduction, i.e. using a combination of multiple denoising techniques such as Gaussian filtering, non-local mean denoising, etc., eliminates noise such as scanning noise points, shooting interference, etc. in the image while trying to preserve the text edge details. Grayscale, i.e. converting a color image to a grayscale image, reduces color channel interference and simplifies subsequent processing. Contrast enhancement can be through histogram equalization, gamma correction, etc. to enhance the contrast between text and background, making the edge more easily identifiable. This provides a high-quality input image for subsequent edge detection.

[0058] In step S200, a deep learning model such as U-Net, HED, etc. is used to extract the text edge. Such models can learn edge features in complex scenes and are more suitable for fuzzy or low-contrast text than traditional operators such as Canny. Then, through particle swarm optimization (PSO) edge linking, the broken edges are intelligently connected. PSO is a bionic optimization algorithm that adjusts the edge path by simulating group behavior to find the optimal connection method such as minimizing the gap between broken edges or curvature changes. This generates a coherent edge profile, avoiding broken edges or burrs.

[0059] In step S300, edge repair fills in small gaps or holes in the edge, and thinning converts coarse edges to single-pixel width, ensuring that the edge lines accurately fit the text shape. The effect is to improve the geometric accuracy and readability of the edge.

[0060] In step S400, first, the sawtooth or uneven part is removed to make the edge natural and smooth; then the edge strength is highlighted to strengthen the boundary between the text and the background. The visual effect of the optimized edge is adapted to the needs of the human eye or OCR recognition.

[0061] In the final step S500, the edge image is first converted into a black and white binary image for storage or further analysis; the edge sharpness is enhanced to compensate for the possible blurring caused by smoothing; and the result is judged by quantitative indicators such as edge continuity score, signal-to-noise ratio, or manual rules. If the result does not meet the standard, the edge detection parameters or the model output are adjusted again in S200. The final result is ensured to meet the application standard (such as OCR recognition rate, human eye perception).

[0062] Reference Figure 2 In an embodiment of the present application, step S100 includes:

[0063] S110, adaptive median filtering is used to remove impulse noise, and the window size is dynamically adjusted according to the noise density. When noise pixels are detected, the window size is increased until the proportion of non-noise pixels in the window exceeds 60%.

[0064] S120, Gaussian filtering is used to smooth Gaussian noise;

[0065] S130, the RGB image is converted into a grayscale image;

[0066] S140, an adaptive histogram equalization algorithm is used to enhance the local contrast, and the image is divided into multiple sub-blocks, and histogram equalization is performed on each sub-block.

[0067] It can be understood that the target of S110 is to eliminate impulse noise such as salt and pepper noise in the image, which appears as black and white point-like interference. A small window is initially used to detect noise pixels, and if noise pixels such as extremely bright or extremely dark outliers are detected in the current window, the window size is dynamically increased. The termination condition is that the proportion of non-noise pixels in the window exceeds 60%, ensuring that enough normal pixels participate in filtering. Then the median value of the non-noise pixels in the window is used to replace the noise pixels, avoiding the edge blurring caused by direct mean filtering. Compared with fixed window median filtering, it can better balance the de-noising effect and detail preservation, especially suitable for high noise density areas.

[0068] The target of S120 is to eliminate Gaussian noise, specifically random distributed fine grain noise, which is often seen in low light shooting. A Gaussian filter is used to perform weighted smoothing on the image. The standard deviation of the Gaussian kernel can control the smoothing strength. The larger the standard deviation, the stronger the de-noising effect, but the edge may be more blurred. In this example, a small standard deviation is selected to preserve the text edge. The impulse noise is removed first in the previous step, and then the Gaussian noise is processed to avoid interference of the impulse noise with the Gaussian filtering.

[0069] The goal of S130 is to reduce the data dimension and simplify the subsequent processing. The RGB three-channel image is converted into a single-channel grayscale image. It should be noted that if the original image is a grayscale image, this step can be skipped.

[0070] The goal of S140 is to solve the problem of local overexposure or under-enhancement that may be caused by global histogram equalization. First, the image is divided into multiple sub-blocks, and the histogram of each sub-block is calculated and equalized independently to enhance the contrast of the region. To avoid discontinuity between sub-blocks, the edge pixels of the sub-blocks are smoothed using bilinear interpolation. This can significantly improve the visibility of dark or bright areas, and compared with global equalization, it can better preserve local details.

[0071] Reference Figure 3 In an embodiment of the present application, the edge extraction using the edge-aware neural network model in step S200 includes:

[0072] S210, first extract the global edge using the Canny algorithm, then filter the non-text edge through the lightweight text detection model, and finally multiply the text mask and the Canny edge pixel by pixel to obtain the text region edge;

[0073] S220, based on the SegFormer architecture, the text region edge and the original image feature are fused, and the local feature is extracted by dividing the image into overlapping regions for feature extraction;

[0074] S230, fuse multi-level features and output a text mask.

[0075] It can be understood that S210 can quickly locate the text region edge in the image and filter non-text interference such as background texture and natural object edge. First, use the Canny operator to detect all possible edges in the image, including text and non-text, and output a binary edge map; use a lightweight model to output a text mask, with white areas having high text probability and black areas being non-text areas; multiply the Canny edge map and the text mask by pixel to retain the edges of the text region and eliminate non-text edges, and output a preliminary text edge map that may have breaks or redundancies.

[0076] S220 can fuse the text rough edge and the original image features through the SegFormer model of the Transformer architecture, enhance local details such as blurred text and small font edges, etc. First, the text edge map obtained by S210 is input, and then the original preprocessed image is input to provide context information such as color and texture. The image is divided into multiple overlapping local regions to avoid loss of boundary information; each region extracts multi-scale features through the Mix Transformer Encoder (MiT) of the SegFormer.

[0077] S230 can integrate features at different levels to generate a high-precision text edge mask. The multi-level features of the SegFormer encoder are gradually upsampled and fused through the MLP decoder, and finally a binary mask with the same resolution as the input image is output, with white pixels representing text edges and black pixels representing non-edges.

[0078] Reference Figure 4 In an embodiment of the present application, step S300 includes:

[0079] S310, first inflation, expanding the target boundary in the image to connect the broken edges; then erosion, shrinking the target boundary in the image, and iterating at least once to remove burrs;

[0080] S320, calculating the sub-pixel coordinates of the edge points based on the Zernike moment function template, and fitting the text contour using the least squares method;

[0081] S330, performing three-level bilateral filtering on the edge image.

[0082] It can be understood that S310 is used to repair edge breaks and remove burrs while trying to maintain the original edge shape. First, inflation expands the edge region to connect broken edges caused by noise or low contrast, and a structure element such as a 3x3 rectangular kernel is used to traverse the image. If there is a white pixel in the kernel coverage area, the center pixel will be set to white, which can make the edge thicker and the broken edges bridged. Then, erosion shrinks the edge region to eliminate redundant pixels introduced by inflation, and a structure element is also used. Only when the kernel completely covers the white area, the center pixel is retained to make the edge return to near the original width, but the burrs are removed.

[0083] S320 is used to promote the pixel-level edge to sub-pixel accuracy to make the contour more consistent with the true text shape. Zernike moments are a set of orthogonal moment functions that are not sensitive to image rotation and noise. The Zernike moments are calculated in the neighborhood of the edge points, and the sub-pixel offset is determined through the phase information of the moments to fit the sub-pixel edge points into a geometric curve.

[0084] S330 is used for smoothing edges while retaining sharp turns, combining a Gaussian kernel with a gray difference for weighted smoothing. The first stage removes small burrs; the second stage is medium smoothing, balancing smoothing and details; and the third stage is weak smoothing, retaining sharp corners.

[0085] Referring to Figure 5 In an embodiment of the present application, step S400 comprises:

[0086] S410, using a Gabor filter to perform multi-direction filtering on the high-pass image after bilateral filtering;

[0087] S420, calculating the response value of each direction filter, fusing each direction response through a maximum value fusion strategy to enhance the continuity of the text strokes.

[0088] S430, calculating the sub-pixel offset based on the Zernike moment function template;

[0089] S440, performing point-by-point correction on the fitted text contour to generate sub-pixel level edge coordinates, and ensuring that the distance change rate of adjacent points does not exceed 10% during the correction process.

[0090] It can be understood that S410 can enhance the direction consistency of the text strokes while suppressing noise and non-text edge interference. The Gabor filter is a filter with spatial and frequency domain characteristics, which can extract texture features of a specific direction and frequency. Usually, 4-8 directions are selected to cover the possible direction of the text strokes, and the high-pass image after bilateral filtering is filtered to highlight the stroke structure. The response map of each direction is output, and the high response value area represents the edge consistent with the direction.

[0091] S420 can fuse multi-direction responses to enhance the continuity of the text strokes and avoid the breakage caused by direction selection deviation. First, for each pixel position, the Gabor response values of all directions are compared, and the maximum value is retained as the final response. Only the maximum response point is retained in the local window, and the edge width is refined to a single-pixel level. It can automatically match the stroke direction, avoid missing detection caused by fixed direction filtering, and the fused edge is more coherent.

[0092] S430 further improves the edge positioning accuracy to the sub-pixel level and reduces the sawtooth effect. First, calculate the Zernike moment in the neighborhood of the edge point, and use the phase information of the moment to determine the sub-pixel offset. This method is suitable for text edges of any direction and is not sensitive to local noise.

[0093] S440 optimizes the spatial distribution of sub-pixel edge points to ensure geometric rationality. The sub-pixel edge points are fitted into a parameterized curve, and then resampled according to the curvature to eliminate abnormal points. The Euclidean distance change rate of adjacent edge points is calculated, and the change rate is forcibly limited to ≤10%. If the limit is exceeded, the coordinates are adjusted through linear interpolation or local re-fitting. This avoids the sudden gathering or scattering of edge points and also maintains the natural geometric shape of the text strokes.

[0094] Referring to Figure 6 In an embodiment of the present application, the quality evaluation in step S500 includes:

[0095] S510, calculating edge completeness, peak signal-to-noise ratio, and structural similarity;

[0096] S520, judging whether the calculated values reach the preset threshold value. If not, automatically returning to step S200 for iterative optimization, with a maximum of 3 iterations. If the threshold value is not reached after 3 iterations, outputting the current optimal result and prompting that the image quality is too low.

[0097] Understandably, S510 evaluates the quality of the edge image through three core indicators. Edge completeness evaluates the continuity and degree of breakage of the edge, reflecting whether the text strokes are connected completely. The specific method is to perform connected component analysis on the edge image, label all independent edge segments, calculate the total length of the effective edge, estimate based on the area of the text region, or predict the theoretical ideal edge length through a pre-trained model. The threshold value is usually set to more than 90% of the edge continuity.

[0098] Peak signal-to-noise ratio is used to measure the noise level and clarity of the edge image, evaluating the effect of noise reduction and edge enhancement. The current edge image is compared with the ideal reference edge to calculate the mean square error (MSE) and peak signal-to-noise ratio (PSNR).

[0099] Structural similarity is used to evaluate the structural fidelity of the edge image, checking whether the geometric shape of the text strokes is distorted. The luminance (L), contrast (C), and structure (S) before and after are compared, and the closer the value is to 1, the more similar the structure is.

[0100] The judgment logic of S520 can be single-indicator judgment, triggering iterative optimization if any indicator does not meet the threshold. It can also be comprehensive judgment, with a weighted calculation of the comprehensive score if multiple indicators are close to the threshold but not all meet the threshold. If the threshold is not met, re-perform edge perception detection, focusing on optimizing the failed indicators of the previous iteration, such as enhancing edge connectivity if EC is low. The maximum number of iterations is 3 to avoid infinite loops. The termination condition is to judge whether the threshold is met, and the final edge image can be output. If the threshold is not met after 3 iterations, the current optimal result is output, and a prompt is given that the image quality is too low. Possible reasons include low original image resolution, excessive noise, and severely damaged text, etc.

[0101] The technical scheme of the present application realizes the following core advantages through the complete process of multi-modal noise reduction, edge perception detection, sub-pixel level repair and refinement, multi-direction smoothing and quality evaluation: combining Canny algorithm and SegFormer neural network, fusing semantic information and geometric features, accurately separating text edges; realizing sub-pixel level positioning through Zernike moment and least square method, eliminating jaggies, adapting to inclined and curved characters; Gabor multi-direction filtering and maximum fusion enhance stroke continuity, three-level smoothing of bilateral filtering balances noise reduction and detail preservation; based on edge integrity, peak signal-to-noise ratio and structural similarity quantitative evaluation, automatic iterative optimization, ensuring that the output quality meets the standards. The problems of breakage, blur, noise and geometric distortion in character edge processing are systematically solved, and the system has efficiency and precision, and is suitable for ancient book restoration, industrial detection, natural scene character recognition and other applications.

[0102] The present application also proposes an image character edge optimization system, comprising:

[0103] The preprocessing unit is used for preprocessing the input image, and the preprocessing includes multi-modal noise reduction, graying and contrast enhancement.

[0104] The edge detection unit is used for edge perception detection on the preprocessed image, and the edge perception detection includes edge extraction by an edge perception neural network model and edge link optimization by a particle swarm.

[0105] The edge repair unit is used for edge repair and refinement of the edge image obtained by edge perception detection.

[0106] The edge enhancement unit is used for edge smoothing and enhancement of the edge image after repair and refinement.

[0107] The post-processing unit is used for post-processing of the edge image after smoothing and enhancement, and the post-processing includes binarization, sharpening and quality evaluation, and if the quality evaluation is not up to standard, the edge detection unit is returned to perform edge perception detection again.

[0108] In the present application, various messages / information / devices / network elements / systems / devices / actions / operations / processes / concepts and other types of objects may be named, and it can be understood that these specific names do not constitute a limitation on the related objects, and the names assigned can be changed according to factors such as scene, context or usage habits. The technical meaning of the technical terms in the present application should be mainly determined from the function and technical effect embodied / executed in the technical scheme.

[0109] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system, device and unit can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.

[0110] In the embodiments of the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There can be another division manner for the actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different units, or the among different units, can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.

[0111] The units described as separated components can or can not be physically separated, and the components displayed as units can or can not be physical units. That is, they can be located in one place, or can also be distributed on a plurality of network units. In actual implementation, some or all of the units can be selected according to the actual needs to achieve the purposes of the embodiments of the present application.

[0112] Those skilled in the art can clearly understand the unit and algorithm steps of each example described in combination with the embodiments disclosed in the present application can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software mode depends on the specific application and design constraints of the technical solutions. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0113] It should also be understood that in each embodiment of the present application, the first, second, etc. are only to represent that a plurality of objects are different. For example, the first time window and the second time window are only to represent different time windows. There should be no impact on the time window itself, and the above first, second, etc. should not limit the embodiments of the present application.

[0114] It should also be understood that in each embodiment of the present application, the terms and / or descriptions of different embodiments have consistency and can be mutually referred to if there is no special description and logical conflict. The technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationship.

[0115] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the parts that make contributions to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a computer readable storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned computer readable storage medium includes: a U disk, a mobile hard disk, a read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0116] The present application also provides a computer program product, which includes instructions that, when executed, cause the network switch and the network system to perform the operations of the network switch and the network system corresponding to the above method.

[0117] The present application also provides a network system, which includes:

[0118] One or more memories for storing instructions; and

[0119] One or more processors for invoking and running the instructions from the memory to perform the method as described above.

[0120] The present application also provides a chip system, which includes a processor for implementing the functions involved in the above description, such as generating, receiving, sending, or processing the data and / or information involved in the above method.

[0121] The chip system can be composed of a chip, or can include a chip and other discrete devices.

[0122] The processor mentioned in any of the above can be a CPU, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the above-mentioned feedback information transmission method.

[0123] In a possible design, the chip system further includes a memory, which is configured to store necessary program instructions and data. The processor and the memory can be decoupled and arranged on different devices, and connected through wired or wireless means to support the chip system to realize various functions in the above embodiments. Alternatively, the processor and the memory can also be coupled on the same device.

[0124] Optionally, the computer instructions are stored in a memory.

[0125] Optionally, the memory is a storage unit within the chip, such as a register, a cache, etc. The memory can also be a storage unit outside the chip within the terminal, such as a ROM or other type of static storage device that can store static information and instructions, a RAM, etc.

[0126] It can be understood that the memory in the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories.

[0127] The non-volatile memory can be a ROM, a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory.

[0128] The volatile memory can be a RAM used as an external cache. There are many different types of RAM, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous dynamic RAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and direct Rambus RAM.

[0129] The embodiments of the specific implementation are the preferred embodiments of the present application, and are not intended to limit the protection scope of the present application. Therefore, any equivalent changes made in the structure, shape, and principle of the present application should be covered within the protection scope of the present application.

Claims

1. A method for optimizing image text edges, characterized in that, include: S100, preprocess the input image, the preprocessing including multimodal noise reduction, grayscale conversion and contrast enhancement; S200, perform edge-aware detection on the preprocessed image, the edge-aware detection including edge extraction using an edge-aware neural network model and particle swarm optimization of edge links; S300 performs edge restoration and thinning on the edge image obtained by edge sensing detection; S400 performs edge smoothing and enhancement on the repaired and refined edge image; S500: Post-process the smoothed and enhanced edge image. The post-processing includes binarization, sharpening, and quality assessment. If the quality assessment is not up to standard, return to step S200 to re-perform edge detection.

2. The image text edge optimization method according to claim 1, characterized in that, Step S100 includes: S110 uses adaptive median filtering to remove impulse noise. The window size is dynamically adjusted according to the noise density. When a noise pixel is detected, the window size is increased until the proportion of non-noise pixels in the window exceeds 60%. S120, combined with Gaussian filtering to smooth Gaussian noise; S130, converts RGB images to grayscale images; S140 employs an adaptive histogram equalization algorithm to enhance local contrast by dividing the image into multiple sub-blocks and performing histogram equalization on each sub-block.

3. The image text edge optimization method according to claim 2, characterized in that, The edge extraction using the edge-aware neural network model in step S200 includes: S210 first uses the Canny algorithm to extract global edges, then uses a lightweight text detection model to filter non-text edges, and finally multiplies the text mask and Canny edge pixel by pixel to obtain the text region edge. S220, based on the SegFormer architecture, fuses text region edges with original image features and extracts local features by segmenting the image into overlapping regions. S230 integrates multi-level features and outputs a text mask.

4. The image text edge optimization method according to claim 3, characterized in that, Step S300 includes: S310 first dilates the image to expand the target boundary and connect the broken edges; then iterates to shrink the target boundary and repeats at least once to remove burrs. S320 calculates the sub-pixel coordinates of edge points based on the Zernike moment function template and fits the text outline using the least squares method. S330 performs three-stage bilateral filtering on the edge image.

5. The image text edge optimization method according to claim 4, characterized in that, Step S400 includes: S410 uses a Gabor filter to perform multi-directional filtering on the high-pass image after bilateral filtering; S420 calculates the response values ​​of the filters in each direction and fuses the responses in each direction using a maximum value fusion strategy to enhance the continuity of the character strokes. S430 calculates sub-pixel offsets based on the Zernike moment function template; S440 performs point-by-point correction on the fitted text outline, generating sub-pixel level edge coordinates, ensuring that the distance change rate between adjacent points does not exceed 10% during the correction process.

6. The image text edge optimization method according to claim 5, characterized in that, The quality assessment in step S500 includes: S510 calculates edge integrity, peak signal-to-noise ratio, and structural similarity. S520: Determine whether the calculated value reaches the preset threshold. If not, automatically return to step S200 for iterative optimization. The maximum number of iterations is 3. If the value still does not meet the standard after 3 iterations, output the current optimal result and prompt that the image quality is too low.

7. An image text edge optimization system, characterized in that, include: Preprocessing unit: used to preprocess the input image, the preprocessing including multimodal noise reduction, grayscale conversion and contrast enhancement; Edge detection unit: used to perform edge-aware detection on the preprocessed image, the edge-aware detection including edge extraction using an edge-aware neural network model and particle swarm optimization of edge links; Edge restoration unit: used to restore and refine the edge image obtained by edge sensing detection; Edge enhancement unit: used to smooth and enhance the edges of the repaired and refined edge image; Post-processing unit: Used to perform post-processing on the smoothed and enhanced edge image. The post-processing includes binarization, sharpening and quality assessment. If the quality assessment is not up to standard, the image is returned to the edge detection unit for edge perception detection to be performed again.

8. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the image text edge optimization method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, Used to store a computer program, which, when executed by a processor, performs the image text edge optimization method as described in any one of claims 1 to 6.