Image classification method and device based on connected region analysis, equipment and medium
By converting the image into a grayscale image, calculating the grayscale difference value and performing connected area analysis, the problems of low image classification efficiency and large resource utilization in the prior art are solved, and an efficient and highly adaptable image classification method is realized.
Patent Information
- Application Number
- CN202510651311.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-26
AI Technical Summary
When facing images from different sources, it is difficult to achieve efficient image classification based on the connectivity characteristics of structural areas in the image without model training, resulting in unstable classification accuracy and excessive resource utilization.
By converting the image to be classified into a grayscale image, traversing adjacent pixel pairs and calculating the grayscale difference value, generating a structural binary image, performing a connection area analysis, determining the pixel proportion of the maximum connected foreground area, and comparing it with the classification threshold to output the image classification type.
It realizes efficient image classification without training samples or model reasoning, has the advantages of simple algorithms, high execution efficiency and strong adaptability, and is suitable for business scenarios such as financial technology and medical health.
Smart Images

Figure CN120543933A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to an image classification method, device, equipment and storage medium based on connected region analysis. Background Art
[0002] In the field of visual document processing, especially document classification tasks, how to efficiently and accurately distinguish between electronically generated documents and photographed image documents has always been an important issue in practical applications. Traditional methods rely on deep learning models to train a large number of labeled samples to extract high-level semantic features for classification. However, such methods not only have extremely high requirements for data quality and quantity, but also require a large amount of computing resources during model deployment and operation, making it difficult to meet the processing needs of lightweight document systems or terminal-side devices. In addition, since photographed documents often have problems such as image distortion, background interference, and shadow overlap, the generalization ability of existing models is insufficient, and they often show the defect of unstable classification accuracy in actual scenarios.
[0003] In the fintech business sector, document recognition and classification are widely used in processes such as automated report archiving, customer data entry, and invoice review. The sources of documents uploaded by financial institutions are complex, including standardized electronic documents exported by the system and images of various paper materials taken offline by customers. Because electronic documents have a regular structure and clear background, while photographed documents often have significant background noise, existing deep learning-based classification models are prone to mistakenly identifying photographic images with simple structures but strong background noise as electronic documents when processing batches of documents from mixed sources. This can lead to classification errors and OCR extraction failures in subsequent processes.
[0004] In the healthcare sector, medical institutions need to structuredly manage and archive massive amounts of documents, including medical records, examination reports, and medical insurance materials. Standard documents generated by electronic medical record systems are mixed with scanned or photographed paper documents uploaded offline by patients. Existing technologies struggle to accurately identify document collection methods while ensuring real-time performance and resource control. This results in some photographed documents being misidentified as electronic documents, leading to image processing failures, information extraction bias, and document indexing confusion, severely impacting the overall operational efficiency and data accuracy of medical information systems. Summary of the Invention
[0005] The main purpose of the present invention is to provide an image classification method, device, equipment and storage medium based on connected region analysis, aiming to solve the technical problem that the existing technology is difficult to achieve efficient image classification based on the connectivity characteristics of structural regions in the image without the need for model training when facing images from different sources.
[0006] To achieve the above object, the present invention provides an image classification method based on connected component analysis, comprising:
[0007] Convert the image to be classified into a grayscale image;
[0008] Traversing adjacent pixel pairs in the grayscale image and determining the grayscale difference between each pair of adjacent pixels;
[0009] Comparing the grayscale difference with a difference threshold, and generating a structural binary image according to the comparison result;
[0010] Performing connected region analysis based on the structural binary image to determine all connected foreground regions;
[0011] Determine the pixel ratio of the maximum connected foreground area in the structured binary image;
[0012] The pixel ratio is compared with the classification threshold, and the corresponding image classification type is output according to the comparison result.
[0013] Furthermore, to achieve the above-mentioned object, the present invention provides an image classification device based on connected component analysis, comprising:
[0014] Grayscale conversion module, used to convert the image to be classified into a grayscale image;
[0015] a grayscale difference analysis module, configured to traverse adjacent pixel pairs in the grayscale image and determine the grayscale difference between each pair of adjacent pixels;
[0016] A structural binary image generating module is used to compare the grayscale difference with the difference threshold and generate a structural binary image according to the comparison result;
[0017] A connected region analysis module, configured to perform connected region analysis based on the structural binary image to determine all connected foreground regions;
[0018] A pixel ratio analysis module, used to determine the pixel ratio of the maximum connected foreground area in the structural binary image;
[0019] The image classification determination module is used to compare the pixel ratio with the classification threshold and output the corresponding image classification type according to the comparison result.
[0020] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and an image classification program based on connected component analysis stored in the memory and run on the processor. When the image classification program based on connected component analysis is executed by the processor, the steps of the image classification method based on connected component analysis as described above are implemented.
[0021] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which is stored an image classification program based on connected component analysis. When the image classification program based on connected component analysis is executed by a processor, the steps of the image classification method based on connected component analysis as described above are implemented.
[0022] Beneficial effects: The present invention relates to the field of image processing technology and can be applied to business scenarios such as financial technology and medical health. A method for image classification based on connected region analysis is disclosed, comprising: converting an image to be classified into a grayscale image, traversing adjacent pixel pairs in the grayscale image and calculating grayscale differences, comparing the grayscale differences with a difference threshold to generate a structural binary image, and forming a foreground area at positions where the grayscale differences in the structural binary image exceed the difference threshold; performing connected region analysis on the structural binary image, determining all connected areas in the foreground area, calculating the pixel ratio of the largest connected area in the structural binary image, and comparing the ratio with a classification threshold to output the classification type of the image. The present invention generates a structural binary image through structural differences and analyzes image region distribution characteristics based on connectivity. It can realize automatic classification of images with different structural features without the need for training samples or model inference processes, and has the advantages of simple algorithm, high execution efficiency and strong adaptability. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:
[0024] Figure 1 A schematic diagram of an application environment of an image classification method based on connected component analysis in one embodiment of the present invention;
[0025] Figure 2 This is a flow chart of an embodiment of an image classification method based on connected component analysis according to the present invention;
[0026] Figure 3 Schematic diagram of functional modules of a preferred embodiment of an image classification device based on connected component analysis of the present invention;
[0027] Figure 4 A schematic diagram of the structure of a computer device according to an embodiment of the present invention;
[0028] Figure 5 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0029] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0030] The image classification method based on connected component analysis provided by the embodiment of the present invention can be applied in the following fields: Figure 1 In an application environment, the user terminal communicates with the server terminal via a network. The server terminal can convert the image to be classified into a grayscale image through the user terminal, traverse adjacent pixel pairs in the grayscale image and calculate the grayscale difference, compare the grayscale difference with the difference threshold to generate a structural binary image, and the position where the grayscale difference in the structural binary image exceeds the difference threshold forms a foreground area; perform connected region analysis on the structural binary image, determine all connected regions in the foreground area, calculate the pixel ratio of the largest connected region in the structural binary image, and compare the ratio with the classification threshold to output the classification type of the image. The present invention generates a structural binary image through structural difference and analyzes the distribution characteristics of image regions based on connectivity. It can realize automatic classification of images with different structural features without the need for training samples or model inference processes, and has the advantages of simple algorithm, high execution efficiency and strong adaptability. The user terminal can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server terminal can be implemented as an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.
[0031] See also Figure 2 , Figure 2 This is a flow chart of an embodiment of an image classification method based on connected component analysis provided by the present invention. It should be noted that although a logical order is shown in the flow chart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0032] like Figure 2 As shown, the image classification method based on connected component analysis proposed in the present invention includes the following steps:
[0033] S10, converting the image to be classified into a grayscale image;
[0034] In this embodiment, the process of converting the image to be classified into a grayscale image is the basic stage of the entire image structure analysis and classification process. This process aims to extract the intensity structure characteristics of the image by removing the color dimension information, so that the subsequent pixel difference analysis and connected area identification can be carried out based on a unified data channel. In this link, the "image to be classified" refers to the input image source, which can include electronic images output by image scanning equipment, photo files generated by mobile devices, web page screenshots, file image records in the database, etc. The input has certain format specifications, such as JPG, PNG, TIFF, BMP, etc. The action of obtaining the image to be classified can be implemented based on file path reading, camera device data stream access, system interface upload, or database image field decoding.
[0035] Converting to grayscale involves mapping the input image from a multi-channel format (typically a three-channel RGB image or YUV image) to a single-channel grayscale structure. "Color channel information" here refers to the multiple dimensions carried by each pixel in the image used to represent color, such as the intensity values of the red, green, and blue channels or the brightness and chrominance components in YUV. This conversion is performed by weighted fusion of the values of multiple channels. For example, RGB to grayscale can be converted using the weighted formula Y = 0.299R + 0.587G + 0.114B, forming a grayscale intensity matrix in which only the brightness information is retained for each pixel. This conversion can be implemented using image processing libraries, such as the cvtColor function in OpenCV or the convert("L") method in the Pillow library.
[0036] In order to improve the clarity of the expression of structural boundaries in grayscale space, structural contour enhancement processing can be performed based on its pixel gradient or edge response after the grayscale image is formed. The main goal of this processing is to make the structural boundaries in the image more prominent, so that they can be effectively identified in subsequent difference extraction and regional analysis. This enhancement can use gradient operators (such as Sobel, Scharr, etc.) to detect grayscale changes in the horizontal and vertical directions, and then generate a clearer structural image by fusing gradient responses; Laplace filtering, nonlinear sharpening, image enhancement convolution, etc. can also be used to improve the responsiveness of texture and structural contours. In addition, morphological processing (such as closing operations, boundary expansion) can also be used to structurally enhance the edge lines formed by text boundaries or graphic frames.
[0037] After the grayscale image conversion is completed, the resulting grayscale image should have the same size as the original Figure 1 The grayscale image is a consistent two-dimensional matrix structure, with pixel values typically ranging from 0 to 255. This grayscale image serves as the input for subsequent analysis of adjacent pixel pairs. In the specific logic, pixel positions in the grayscale image correspond one-to-one with spatial coordinates, providing a coordinate reference for subsequent pixel pair traversal and grayscale difference calculations. This process unifies diverse and complex color image data into a structure-oriented single-channel input format.
[0038] In the adaptation of different image source types, this processing can be optimized and implemented in a variety of ways. For electronic document type images, the image frame can be directly extracted from the PDF parsing module and converted into a grayscale image without image compression or sharpening adjustment. For images in mobile phone camera scenes, the edge structure can be optimized after conversion to grayscale in combination with exposure compensation and local contrast enhancement technology to offset the boundary blurring caused by changes in ambient light in the image. For image data containing noise interference or background patterns, such as medical bills, seals in financial contracts, or background texture interference layers, a pre-filtering step (such as median filtering, bilateral filtering) can be added before grayscale conversion to smooth the interference area, so that the grayscale image is more focused on the content structure.
[0039] To improve computational efficiency in high-resolution document images, downsampling can be performed before grayscale conversion and structural contour enhancement. The results can then be remapped based on the magnified coordinates during the subsequent connected region annotation phase. When deployed on embedded devices or mobile processing platforms, grayscale conversion and enhancement can be accelerated through hardware instruction level acceleration, combined with OpenCL or GPU CUDA modules, to reduce processing latency while preserving image features.
[0040] Example: In the healthcare field, if automatic recognition of physical examination report images is required, the system can convert the PDF electronic report and photographs into grayscale images, effectively counteracting interference from color legends, background watermarks, etc., and extract medical record boundaries and structural diagrams, thereby supporting subsequent report structure analysis and archiving classification.
[0041] In the financial business field, when the system processes image materials such as paper bills, electronic invoices, and financial contracts, it first eliminates interference caused by color variations through grayscale conversion, highlights the text border and seal structure, and focuses more on the text and graphic distribution logic in subsequent structural analysis and risk identification, which helps to complete the type determination and archiving classification of electronic bills and photographed contracts.
[0042] By converting the image to be classified into a grayscale image and combining it with a structural contour enhancement strategy, the color interference dimension is effectively removed, retaining only the brightness intensity signal that affects structural analysis. This provides a unified, low-redundancy input data foundation for subsequent pixel-pair interpolation operations, differential intensity extraction, and foreground structure identification. This approach improves the contrast clarity of structural information, making the connected foreground regions in the structural binary image more concentrated and complete, thereby achieving more stable and accurate results in the connected region identification and proportion calculation stages.
[0043] S20, traversing adjacent pixel pairs in the grayscale image, and determining the grayscale difference between each pair of adjacent pixels;
[0044] In this embodiment, traversing adjacent pixel pairs in a grayscale image and determining the grayscale difference between each pair of adjacent pixels is a key operation link in image structural feature extraction. The core purpose is to reveal the spatial structural characteristics of the image by detecting the degree of grayscale change between local pixels and to construct a primitive measurement basis for structural judgment. A grayscale image is essentially a two-dimensional matrix, in which each position corresponds to a pixel value representing brightness intensity. In the spatial structure, the difference between the grayscale value of any pixel position and its adjacent position can reflect the change trend or edge attribute of the local image. Therefore, by comparing the grayscale values in pairs and recording their differences, a high-dimensional gradient distribution view of the overall structure of the image can be formed.
[0045] During the traversal process, the selection of adjacent pixel pairs must be organized according to well-defined directions. To facilitate structure recognition, adjacent pixel pairs are usually extracted from both horizontal and vertical perspectives. Horizontally, pixel pairs consisting of left and right pixels in each row of the image are selected, while vertically, pixel pairs consisting of upper and lower pixels in each column are selected. This approach ensures that the vast majority of pixel adjacencies in the image are covered while preserving the image's horizontal and vertical structural features. The grayscale difference of each pixel pair is calculated by calculating the difference in absolute values of the grayscale values at the two locations, generating a single value representing the intensity of the brightness change between the two points. This set of values forms a difference matrix in two directions, representing the local grayscale changes in the horizontal and vertical directions, respectively.
[0046] Difference calculations should be based on the integer space of pixel grayscale values. For example, in an 8-bit image, grayscale values range from 0 to 255, with a maximum difference of 255 and a minimum of 0. This operation can be performed by traversing the image pixels row by row and column by column and calculating the absolute difference between adjacent index positions. This operation can be completed using efficient vectorized processing methods such as image processing libraries such as NumPy and OpenCV. The above-mentioned difference extraction does not perform a structural transformation on the original image itself, but rather an intermediate calculation result based on the grayscale image. This result will serve as the basic input for subsequent difference intensity determination, structural foreground extraction, and connectivity analysis.
[0047] The specific implementation of the pixel pair traversal and grayscale difference extraction can be adjusted in multiple dimensions according to image characteristics and processing requirements. When processing images with dense horizontal text layout, the horizontal pixel pair extraction logic can be used first, and with a higher sampling density, difference calculations can be performed between 1 to 3 adjacent pixels in each row to obtain a finer structural response; in vertically laid out structural diagrams or invoice form images, the vertical pixel pair collection can be encrypted. For high-resolution image data, the size of the pixel pair selection area can be controlled by block sampling or using a sliding window mechanism to improve processing efficiency. In scenarios where edge transitions are blurred or image noise is strong, pre-processing filtering steps such as mean filtering, median filtering, or Gaussian blurring can be added to reduce the interference of pseudo-edge responses and enhance the accuracy of grayscale difference expression.
[0048] Different platforms can use different computational methods to optimize this operation. For example, on mobile devices or embedded systems, grayscale difference extraction can be converted into assembly instructions for parallel execution. In GPU-accelerated environments, grayscale difference calculation can be integrated into the image convolution kernel operation process to further improve batch processing capabilities. In text image processing scenarios, a blocking strategy can be introduced to divide the image into multiple small grids and perform structural difference statistics on the local pixels within each grid to improve the stability of connected structure recognition.
[0049] Example description: In healthcare scenarios, this processing can be used to process scanned medical record images or hospitalization record forms. The traversal and calculation of grayscale differences can effectively capture text boundaries, grid outlines, and title structures, allowing the structural binary image to retain key content even in the presence of watermarks and background interference.
[0050] In financial applications, such as processing credit reports, bill images, or printed forms, this method leverages local contrast enhancement, formed by the difference between adjacent pixels, to effectively extract stable structures such as edges, seals, and frames, providing clear regional contours for tasks like bill recognition and invoice structure analysis. This processing, which does not rely on external models or training data, is suitable for financial and medical image processing scenarios with high requirements for privacy, security, and computational efficiency.
[0051] By systematically traversing and extracting differences between pairs of adjacent pixels in a grayscale image, it is possible to capture structural change boundaries within the image based solely on local brightness gradients, without introducing semantic models or prior knowledge. This processing approach avoids interference from overall image intensity differences in the analysis results, focusing on the image's spatial variation trends, thereby establishing a highly reliable input foundation for subsequent structural extraction and connectivity determination. Directional extraction of grayscale differences and their absolute value quantization not only enhances boundary response but also provides a clearly distinguishable numerical input channel for subsequent difference screening and binary mapping.
[0052] S30, comparing the grayscale difference with a difference threshold, and generating a structural binary image according to the comparison result;
[0053] In this embodiment, the grayscale difference is compared with the difference threshold, and a structural binary image is generated based on the comparison result. This is a key processing process for performing structural boundary visualization based on brightness gradient intensity. The grayscale difference is a numerical set of local grayscale change intensities extracted in the previous step. This set is presented in matrix form, and the value at each position represents the intensity change level of the pixel in its adjacent direction. The difference threshold serves as the classification boundary for dividing the foreground and background. The numerical value should be set in combination with the grayscale dynamic range of the image and the desired structural complexity to be retained. A suitable range is usually selected between 0 and 255.
[0054] When comparing the difference value to the threshold, a point-by-point determination mechanism is used. Each pixel coordinate in the grayscale difference set is traversed and the grayscale difference at that location is determined to be greater than a preset difference threshold. If so, it is considered that there is a sufficiently strong brightness jump between the pixel and its neighbors, indicating that it may be located at a structural edge or content separation zone, and the location is marked as the foreground region in the structural binary image. If it is not greater than the threshold, the brightness change is considered insignificant, and the location is considered non-structural, corresponding to the background region. This comparison behavior achieves the goal of converting a continuous intensity image into a discrete structure image.
[0055] The initialization of the structural binary image should maintain the same size as the original grayscale image, and a binary Boolean matrix should be used to record the foreground and background division status. Typically, each pixel has only two possible values, such as 0 for background and 1 for foreground. In practical applications, to enhance the structural representation of the processed image, directional information is often used to set the marker position. For example, based on the direction of adjacent pixels, the marker operation is placed at the back of the current adjacent pixel pair, such as marking the right pixel for horizontal centering and marking the bottom pixel for vertical centering, to ensure structural integrity and image coherence.
[0056] This processing process does not rely on a learning model, but instead achieves structural foreground extraction based on the saliency contrast of the local structure of the image. It can maintain good boundary separation effects in a variety of image scenes and is suitable for visual content types with obvious structures, such as tables, text blocks, and contour maps.
[0057] The difference threshold can be set dynamically based on fixed empirical values, image adaptive algorithms, or statistical features. A fixed constant value, such as 60, can be used, which is suitable for scanned images with significant brightness variations. Alternatively, a threshold can be adaptively generated based on the image gradient histogram distribution. For example, the grayscale difference density curve can be used to determine the cumulative coverage of 85% of the structural boundaries as the threshold. In scenes with non-uniform image brightness, a local dynamic threshold strategy can be introduced to set a specific threshold for local sub-blocks based on the grayscale mean or standard deviation of different image regions, thereby improving the ability to identify boundaries in low-contrast areas.
[0058] The generation process of the structural binary image can also control the update region through the mask matrix, avoiding repeated judgment and labeling. Combined with the direction judgment strategy, each grayscale difference can be pre-labeled with its source direction, and the direction label can be used to guide the assignment of the label position. The label image can be initialized with an integer matrix, allowing future steps to perform further operations such as connected region labeling and feature extraction based on the structural image.
[0059] In terms of computational efficiency, difference comparison can achieve batch parallel judgment through bit operations and vectorized logic, making it particularly suitable for structural extraction tasks in high-resolution images. Furthermore, in embedded or edge devices, the grayscale difference judgment logic can be compiled into lightweight hardware instructions to improve processing response speed.
[0060] Example: In healthcare scenarios, such as processing scanned diagnostic report images from hospital systems, images often contain key structural areas such as patient information, medical order forms, and examination data, along with background watermarks, anti-counterfeiting layers, or shadows caused by scanning. To distinguish these structural areas from the background, an appropriate grayscale difference threshold must be set during processing so that only the structural boundaries can be effectively identified in subsequent analysis. The setting of the difference threshold directly determines the accuracy of structure extraction and subsequent region classification.
[0061] The process of setting the grayscale difference threshold begins with acquiring a grayscale image. By calculating the grayscale difference between each pair of adjacent pixels, we obtain horizontal and vertical difference matrices, which are then fused together using the maximum value of each coordinate to create a unified grayscale difference matrix. All differences in this matrix are then expanded into a linear sequence, and their frequency distribution within the range of 0 to 255 is calculated. A histogram is then constructed to observe the global distribution pattern of the differences.
[0062] In structured document images, this histogram often exhibits two distinct regions: one concentrated in low-value areas with small differences, representing background or areas of subtle change; the other distributed in high-value areas with large differences, representing significant structural areas such as text, table lines, and graphic boundaries. To determine the appropriate threshold, it is necessary to calculate the cumulative distribution function (CDF). This involves accumulating the number of pixels in each difference interval and finding a turning point that serves as the boundary between background and foreground.
[0063] For example, in a medical data table image, analysis found that when the difference value was 38, 88% of the total pixels were covered, which can be used as a preliminary global threshold candidate. Subsequently, multiple image regions were selected for visual verification, comparing the structural image effects extracted at thresholds of 30, 38, and 45. If the threshold is too low (such as 30), weak background textures in the image may be mistakenly identified as structures; if the threshold is too high (such as 45), some font boundaries will be lost, affecting the integrity of table recognition. Therefore, the threshold was set around 38 and fine-tuned, and the optimal value was selected based on the structural integrity and noise suppression effect.
[0064] To improve the adaptability of the threshold to different image types, a regional adaptive mechanism can be introduced. The image is divided into multiple fixed-size windows (e.g., 64×64 or 128×128 pixels). The local mean and variance of the grayscale difference within each window are calculated, and the threshold for that region is dynamically adjusted based on local characteristics. For example, a window with a standard deviation greater than twice the global mean can use a higher threshold to avoid strong light interference or shadowed areas. Regions with lower standard deviations can maintain the original threshold or adjust it appropriately to preserve subtle structures.
[0065] By comparing grayscale differences with a difference threshold and mapping them into a structural binary image, we can quickly extract structural edges and content separation zones within an image without relying on a learning model. This allows the logical structure of images such as documents, tables, and text graphs to be accurately abstracted and expressed in the image dimension. This processing approach significantly reduces the reliance on training data and image quality, ensuring that subsequent connected region extraction and proportional analysis have clear input conditions with well-defined structures and boundaries, thereby improving the accuracy and efficiency of overall structural analysis.
[0066] S40, performing connected region analysis based on the structural binary image to determine all connected foreground regions;
[0067] In this embodiment, a structural binary image is generated by comparing grayscale differences with a set threshold. It is used to emphasize pixel regions with significant structural features in the image. In this image, the foreground region is typically identified by a single value (such as 1 or 255), while the background region is identified by another value (such as 0). The purpose of connected component analysis is to identify sets of pixels in these foreground regions that have continuous and spatially adjacent values. By marking their connectivity, the structural regions can be independently segmented.
[0068] A connected foreground region is a group of foreground pixels that are coherently connected according to the pixel adjacency rule. Depending on the application, either the four-neighborhood rule or the eight-neighborhood rule can be used. The four-neighborhood rule only considers pixel connections in the up, down, left, and right directions, making it suitable for extracting strictly linear structures. The eight-neighborhood rule also considers diagonal connections, making it easier to identify blocky regions with more complex contours. The choice of connectivity rule depends on the trade-off between structural complexity and target accuracy in the image being processed.
[0069] The core process of region labeling begins with an unlabeled foreground pixel, expands the search to all connected foreground pixels, and assigns a unique identifier to each set. To ensure processing efficiency and labeling consistency, the labeling matrix uses the same row and column dimensions as the structural binary image to record the connectivity status of each pixel. Initially, all label values are set to 0, indicating that no region has been assigned. As the image is scanned, any unlabeled foreground pixel is encountered, triggering the region growing process.
[0070] Region growing can be performed using either a breadth-first search (BFS) or a depth-first search (DFS) to ensure complete spatial coverage of all connected pixels. Each time a connected region is discovered, it is assigned the next available region identifier and written into the label matrix, completing the labeling of the region. After the entire image is scanned, each non-zero value in the label matrix represents the connected region number to which it belongs. The process also accumulates all pixel positions associated with that number, facilitating subsequent structural analysis.
[0071] A two-pass scanning algorithm can be used in this implementation. The first pass traverses the image row by row, determining whether the current pixel is connected to its left and upper neighbors based on adjacency rules, and propagating labels based on existing identifiers. If the current pixel has no adjacent foreground pixels, a new number is assigned. If there are adjacent foreground pixels with different identifiers, an equivalence table is created and the standard numbers are uniformly replaced in the second pass.
[0072] Alternatively, a seed-filling approach can be used: starting with each unlabeled foreground pixel, its coordinates are pushed into a queue. When popping a pixel, all adjacent locations are checked. If the adjacent pixel is foreground and unlabeled, it is added to the queue and its identifier is updated. This ultimately forms a set of pixels representing all connected regions, and its geometric properties, such as area and boundary, can be calculated in real time.
[0073] The region growing process can also be parallelized based on the graphics processing acceleration framework, which is especially suitable for large-format scanned images or high-definition graphic images. The processing efficiency can be improved through parallel partitioning of image blocks and cross-block merging mechanism.
[0074] Example: In healthcare data processing scenarios, received images may include drug inserts, test reports, or registration receipts. Connected component analysis can quickly extract the outline blocks of text, tables, or diagnostic charts, allowing for structural hierarchical segmentation without requiring OCR recognition. This provides a basis for subsequent system determination of whether the image is a scanned or photographed image or an electronically exported image.
[0075] In the financial industry, documents such as bank statements, insurance policies, and receipts often have dense structures, with foreground text blending into the background. Connected Region Analysis can effectively extract large document borders, barcodes, or table areas, and use this to determine whether the image source is a machine-scanned electronic document or a photographed image, improving document classification automation and data extraction efficiency.
[0076] Connected region analysis in structural binary images enables rapid and accurate spatial classification of foreground structures within an image, separating multiple image regions with independent structural boundaries. This approach eliminates the need for traditional semantic feature-based deep models and relies on complex modules such as text detection and OCR. While maintaining lightweight processing, it provides a high-fidelity restoration of the image's structural distribution. By precisely labeling the connected blocks of each structure, a structured foundation is laid for the subsequent calculation of features such as size, density, and position, enhancing the interpretability and scalability of the entire classification task.
[0077] S50, determining the pixel ratio of the maximum connected foreground area in the structured binary image;
[0078] In this embodiment, after connected component labeling, the structural binary image is typically divided into multiple foreground regions with independent connectivity. Each foreground region corresponds to a set of spatially connected foreground pixels with the same value. To assess the spatial proportion occupied by the region with the highest concentration of structure in the entire image, the region containing the largest number of pixels is identified from the multiple foreground regions obtained as the most connected foreground region.
[0079] During connectivity analysis, each foreground region has its corresponding attribute record established, including its area and boundary range. The area is used to measure the total number of foreground pixels contained in the connected region. The area value is derived from the number of pixels corresponding to the same identifier in the foreground pixel label matrix. This method uses discrete counting logic and does not involve image pixel value calculations, resulting in high robustness.
[0080] The total number of pixels in a structured binary image is determined by the image dimensions—the product of the number of rows and columns—which is typically determined during image input or initialization. Pixel fraction calculations are based on a mathematical ratio model: the number of pixels in the largest connected region is used as the numerator, while the total number of pixels in the entire image is used as the denominator. The ratio between the two is calculated to determine the spatial fraction of the largest connected foreground region in the image.
[0081] In the context of image processing, this percentage reflects the spatial weight of the most concentrated structural regions within an image. It offers stable discrimination capabilities and is particularly suitable for determining image types based on contrasting structural density. The pixel percentage is expressed as a floating-point value between 0 and 1, serving as a comparable numerical indicator in the threshold determination logic during subsequent classification decisions.
[0082] The calculation of pixel ratio can take the connected region feature set as input, traverse the area attribute value of each region in the set, and record the connected region number and area corresponding to the maximum value. This process can be completed with a single linear scan, with low time complexity and linearly related to the number of regions. While performing connected region marking, a region area mapping table can be established, using each identifier as a key and the corresponding number of foreground pixels as a value. Finally, by traversing the mapping table, the maximum value item is selected, which is the largest connected region. The total number of pixels in a structural binary image can be calculated directly using its size parameter: if the image width is W and the height is H, then the total number of pixels is W×H. This parameter does not need to be parsed from the image content and can be obtained directly from the image carrier properties, making it more universal and stable.
[0083] When calculating the percentage, use double-precision floating-point division to divide the maximum area by the total number of pixels, retaining 3 to 5 decimal places of precision to ensure sufficient sensitivity to numerical boundaries in subsequent classification decisions. Alternatively, convert the foreground area percentage to a percentage and use a percentage comparison mode when setting the classification logic to match common data formats in business systems.
[0084] Example: In healthcare scenarios, electronic documents such as outpatient registration forms, laboratory test forms, and medication lists often have large structural areas, such as black-and-white tables or printed text blocks. These documents present a highly continuous and compact foreground area in the structured binary image, and their maximum connected area pixel ratio is generally high. In contrast, photographed medical records often have fragmented foreground areas due to lighting, background interference, or blurred edges, resulting in poor connectivity and a low pixel ratio. The aforementioned ratio judgment logic can quickly distinguish between clearly structured and unstructured inputs, thereby assisting in the classification and determination of the document's source.
[0085] In financial transactions, the main structure of documents such as electronic statements and electronic invoices is often concentrated in a large, centrally located area. Photographed documents such as bank card receipts and handwritten forms tend to be more fragmented in the foreground. Evaluating the proportion of the largest connected area can serve as an important criterion for determining the structural integrity of documents. This can be used in automated classification processes to identify objects requiring further OCR recognition or enhanced processing, optimizing computing resource allocation within the overall processing chain.
[0086] By quantifying the percentage of pixels in the maximum connected foreground area, we can extract discriminative statistical features from the image structure without introducing high-complexity models or relying on semantic recognition. This approach is suitable for images with high noise, noisy images, or complex structures, where traditional text detection methods fail, enhancing the system's discriminative power under weak supervision. The extraction logic for the percentage parameter is repeatable and interpretable, providing a clear quantitative basis for image structural density classification and improving the reliability of subsequent classification decisions.
[0087] S60: Compare the pixel ratio with a classification threshold, and output a corresponding image classification type according to the comparison result.
[0088] In this embodiment, after obtaining the pixel ratio of the maximum connected foreground area, this ratio needs to be classified to further determine the image type. The core of the classification process is to use a predefined classification threshold as the judgment standard and compare the ratio value with the threshold to determine the feature attribution of the input image at the structural density level.
[0089] The classification threshold is an adjustable parameter, typically expressed as a percentage or floating-point number, ranging from 0 to 1. Essentially, it serves as a demarcation metric for structural density. In technical applications, this threshold is used to distinguish images with a high structural density from those with a low structural density, corresponding to the typical data types of structured and unstructured documents. The decision logic generally considers images with a value greater than or equal to the threshold as structurally significant, while images with a value less than the threshold as structurally sparse, thus enabling image classification.
[0090] The judgment rules are based on numerical comparisons, requiring no model training or manual annotation. The judgment process is independent of image resolution or compression ratio, resulting in good generalization stability. In implementation, comparisons are typically performed using floating-point comparison instructions, and the classification output is a predefined classification identifier, indicating the category label to which the image belongs.
[0091] The classification identifier can be an enumeration structure, such as TEXT_DOMINANT, GRAPHIC_DOMINANT, PHOTO_SCANNED, or SCREEN_CAPTURE, or a numeric or label string. Throughout the entire process, the pixel percentage serves as the intermediate decision basis, the classification threshold serves as a static parameter, and the classification identifier serves as the final decision output, forming a complete decision chain.
[0092] In practice, the system can load a preset classification threshold from a configuration item, database, or model parameter set. This threshold can be set by the system management module or operations personnel based on experience or business characteristics. The previously calculated pixel ratio of the maximum connected foreground area is then used as input and directly compared with the classification threshold.
[0093] The comparison logic can be implemented using language-independent operators, such as the >= operator in Python or the std::greater_equal function object in C++, to determine whether the pixel ratio reaches the classification boundary. When the ratio is greater than or equal to a threshold, the system determines that the image is a structured document, such as a standardized electronically generated form or a clearly photographed printed document. In this case, a first-class identifier such as "structured" or "TEXT_DOMINANT" can be output. When the ratio is lower than the threshold, the system outputs a second-class identifier such as "photo_like" or "GRAPHIC_DOMINANT", indicating that the image is closer to unstructured content or low-density typesetting.
[0094] Multiple classification thresholds can also be introduced to form multi-level classification logic. For example, three structural proportion intervals of low, medium, and high can be set, corresponding to three different complexities or image source types, to adapt to more complex data distribution characteristics.
[0095] Example: In a healthcare scenario, after receiving image data from a hospital information system, the system calculates the pixel ratio of the largest connected area to be 0.72, and the classification threshold is set to 0.65. Because 0.72>0.65, the system identifies the image as an electronically generated structured prescription and automatically places it into the structured document processing queue. However, the medical record image taken from a mobile terminal has a ratio of only 0.41, which is lower than the classification threshold. The system then identifies it as an unstructured image and enters the image enhancement and text extraction process.
[0096] In the financial business scenario, if the system receives a scanned credit report, the calculated pixel ratio is 0.78, which is much higher than the 0.60 threshold set in the system. The system determines that it is a well-formatted electronic document and is suitable for directly entering the OCR analysis process. If the system receives a bill photographed by the user, its ratio is only 0.35, which is classified as an image-type document. The system guides it to the image cleaning and preprocessing module before performing the next step of recognition.
[0097] In this way, the structural features of the image are accurately converted into input parameters in the numerical comparison process, and a high-efficiency, low-misjudgment image type automatic recognition mechanism based on pixel density is constructed.
[0098] By quantifying structural density information as pixel percentages and comparing the values against a preset threshold, this method allows for rapid determination of document image type with minimal computational complexity, significantly improving image classification efficiency and accuracy. This approach eliminates the need for model training, sample feature extraction, and normalization, effectively avoiding the generalization issues of traditional image classification methods when faced with inconsistent data dimensions or diverse input formats. The judgment process relies on connected region structural information, resulting in high robustness and interpretability, making it suitable for deployment in high-concurrency, low-latency document processing systems.
[0099] The present invention relates to the field of image processing technology and can be applied to business scenarios such as financial technology and medical health. A method for image classification based on connected region analysis is disclosed, comprising: converting an image to be classified into a grayscale image, traversing adjacent pixel pairs in the grayscale image and calculating grayscale differences, comparing the grayscale differences with a difference threshold to generate a structural binary image, and forming a foreground region at locations in the structural binary image where the grayscale differences exceed the difference threshold; performing connected region analysis on the structural binary image to determine all connected regions in the foreground region, calculating the pixel ratio of the largest connected region in the structural binary image, and comparing the ratio with a classification threshold to output the classification type of the image. The present invention generates a structural binary image through structural differences and analyzes image region distribution characteristics based on connectivity. Automatic classification of images with different structural features can be achieved without the need for training samples or model inference processes, and has the advantages of simple algorithm, high execution efficiency, and strong adaptability.
[0100] In one embodiment, the above step S10 includes:
[0101] S101, obtaining a target image to be classified from a data source of images to be classified;
[0102] S102, eliminating color channel information of the target image to be classified;
[0103] S103, converting the target image to be classified after eliminating the color channel information into a single-channel grayscale image;
[0104] S104: Perform structure contour enhancement processing on the single-channel grayscale image to generate the grayscale image.
[0105] In this embodiment, converting an image to a grayscale image is a prerequisite for structural region analysis. Its main purpose is to remove the interference of color information on the calculation of structural differences and retain the core features of the image in terms of spatial structure and brightness distribution. In this operation, the target image must first be extracted from the data source. The data source may include an image database, a real-time acquisition channel, a document scanning device, an image transmission interface, etc. The target image refers to the image to be classified for subsequent connected region identification. Its source may be diverse, such as desktop screenshots, mobile device images, scanner images, or image compression and decoding results.
[0106] After acquiring an image, color channel information must be removed. Color images typically consist of three color channels, such as in RGB or BGR formats, each recording a specific color component. Preserving this channel information interferes with the balanced distribution of grayscale intensity, so it must be uniformly converted to a brightness-dominated representation. This can be achieved through weighted averaging, maximum component extraction, or projection onto a specific color model. Commonly used coefficient combinations in weighted averaging are 0.299, 0.587, and 0.114, corresponding to the red, green, and blue channels, respectively, which better align with the human eye's brightness perception.
[0107] After color channel removal, the image must be converted to a single-channel grayscale image, where only the grayscale intensity value is retained for each pixel. This can be accomplished by reconstructing the image matrix or using conversion functions in an image processing library. The grayscale image provides a uniform brightness basis, facilitating subsequent calculations of differences between adjacent pixels.
[0108] To further enhance structural edges and make it easier for potential foreground areas in the image to form connected block structures during binarization, structural contour enhancement is performed on the grayscale image after conversion. Structural contour enhancement refers to the use of specific image enhancement algorithms to improve the contrast and clarity of structural boundaries in an image. Common enhancement methods include gradient operators (such as Sobel and Scharr), edge detection algorithms (such as Canny), Laplacian operators, and unsharp masks. These methods can significantly enhance the strength of structural boundaries while maintaining regional grayscale continuity, reducing misclassification issues caused by unclear differences in foreground and background.
[0109] The output image after structural contour enhancement is a standard grayscale image used for difference calculation and structural analysis, and its pixel value space has been significantly weighted for potential structures through enhancement processing.
[0110] Grayscale conversion and structural contour enhancement of the target image can be performed based on the following methods:
[0111] In one implementation, a JPEG image is loaded from an image database. The target image data is read using the cv2.imread function in the Python OpenCV library. The color channel conversion operation is then performed using the cv2.cvtColor function with the cv2.COLOR_BGR2GRAY parameter to generate a single-channel grayscale image. Gradient enhancement is then performed using the cv2.Sobel function in both the X and Y directions. Finally, a weighted overlay method is used to generate an edge-enhanced image, which serves as the input image for structural analysis.
[0112] In another implementation, if the image comes from a real-time camera or is uploaded by a mobile device, the original pixel matrix can be constructed through the image decoding interface after receiving the image byte stream. In the GPU or NPU environment, built-in instructions (such as TensorRT or OpenCL kernel) are used to perform Y channel extraction on the image matrix, and the built-in unsharp mask kernel function in the low-level image processing module is called to perform local contrast enhancement. The output image is transmitted through a pipeline to the image difference module in the next stage.
[0113] It is also possible to directly use the Y channel as the grayscale channel based on the YCbCr image format output by the image compression decoder without the need for additional color conversion overhead, and then combine it with morphological gradient enhancement (such as the difference between corrosion and dilation) to achieve structural edge extraction. This method is suitable for financial invoice images with high image clarity and mild edge weakening.
[0114] By performing grayscale conversion and structural contour enhancement, this embodiment significantly improves image clarity and segmentation capabilities at the structural boundary level, ensuring that subsequent pixel difference calculations and structural region identification processes are based on a consistent and high-quality pixel intensity foundation. Compared to direct color image processing, the dual grayscale and structural enhancement steps improve the accuracy and stability of structurally connected region identification while maintaining the pixel computational complexity. This is particularly true for images with complex backgrounds, unstable image quality, or blurred boundaries, making it easier to form clearly connected structural blocks and significantly reducing the missegmentation rate of binary images.
[0115] In one embodiment, the above step S20 includes:
[0116] S201, traversing all left and right adjacent pixel pairs in the grayscale image along the horizontal direction, and determining the grayscale absolute value difference between each pair of left and right adjacent pixels and storing it as a horizontal difference matrix;
[0117] S202, traversing all pairs of upper and lower adjacent pixels in the grayscale image along the vertical direction, and determining the grayscale absolute value difference between each pair of upper and lower adjacent pixels and storing the difference as a vertical difference matrix;
[0118] S203, traversing each identical row and column coordinate position in the horizontal difference matrix and the vertical difference matrix, and comparing the element value of the current row and column coordinate position in the horizontal difference matrix with the element value of the current coordinate position in the vertical difference matrix;
[0119] S204, selecting a larger value between the element value at the current row and column coordinate position in the horizontal difference matrix and the element value at the current row and column coordinate position in the vertical difference matrix;
[0120] S205, taking the larger value as the fused grayscale difference value of the current row and column coordinate position;
[0121] S206, writing the fused grayscale difference into the corresponding row and column coordinate positions of the fused difference matrix;
[0122] S207, summing up the fused grayscale differences of all row and column coordinate positions to generate a grayscale difference set of all adjacent pixel pairs.
[0123] In this embodiment, the core purpose of traversing adjacent pixel pairs in a grayscale image is to capture the local grayscale variation characteristics of the image in the spatial distribution dimension, providing a differential basis for subsequent binary structure determination. In specific operations, the grayscale absolute value difference between adjacent pixels is extracted in two orthogonal directions: horizontally and vertically.
[0124] Horizontal traversal involves pairing two consecutive pixels within each row. From left to right, each pair of adjacent pixels is selected and the absolute difference between their grayscale values is calculated. This value represents the degree of local horizontal brightness jump. The differences between all pairs of adjacent pixels in all rows are sequentially recorded to form a two-dimensional matrix, which is the horizontal difference matrix. The matrix size is usually the number of rows in the image multiplied by the number of columns minus one.
[0125] Similar logic is used for traversal in the vertical direction, but the pixels in each column are processed from top to bottom. The pixel points between two consecutive rows in each column are processed from top to bottom. The resulting difference matrix is called the vertical difference matrix, and its size is the number of rows minus one multiplied by the number of columns.
[0126] After constructing the two directional difference matrices, to unify the boundary salience of each pixel in the image structure, a comparison operation is performed on the horizontal and vertical differences at the same coordinate position. This operation is performed using a pixel index traversal method, extracting the current element value at each matrix coordinate position, that is, the difference response value of the two directions at that position. By comparing the two values, the larger one is selected as the fused grayscale difference. The logic behind this is that stronger directional differences between pixels tend to better represent structural boundaries, so retaining the stronger directional response in adjacent pixel pairs can enhance structural perception.
[0127] The fused grayscale differences are written into a newly created fused difference matrix. The matrix size is consistent with the input grayscale image to ensure that the subsequent structural region judgment has complete spatial coordinate coverage. Each position in the fused difference matrix records the local maximum grayscale difference response at that position.
[0128] Finally, by performing a summary operation on all elements in the fused difference matrix, a fused grayscale difference set can be constructed. This set not only contains the specific difference values, but also comes with its corresponding pixel coordinate index, which is used for positioning and classification processing in subsequent difference threshold comparison and structure labeling operations.
[0129] Pixel difference calculations can be performed based on two-dimensional array indexing. In one implementation, Python and NumPy are used. After reading the grayscale image matrix, a slicing operation is performed on each row and column. The image is horizontally shifted right by one column and vertically shifted down by one row. Horizontal and vertical difference matrices are generated by pairwise subtraction and taking the absolute value.
[0130] GPUs can also be used in parallel to perform batch grayscale difference calculations on image pixels. In high-resolution image processing scenarios, CUDA kernels can be used to parallelize the calculations for adjacent horizontal and vertical pixels, with the results written to a dedicated storage cache. This shared memory mechanism significantly reduces memory access latency.
[0131] In another implementation adapted for embedded image processing devices, the two-dimensional pixel array sliding window module arranged in the FPGA can be used to perform row cache and pixel window comparison on the input stream image in real time. Structurally, point-by-point difference operations and size comparisons can be implemented through LUT and addition and subtraction circuits, and the fusion results are output in real time and written into RAM to form a difference set cache.
[0132] This embodiment uses bidirectional pixel difference extraction and maximum fusion mechanisms to sensitively capture regions of structural change within an image. It can also accurately identify significant boundary points even in images with uneven distribution of structural information in different directions, avoiding issues such as broken or weak boundary recognition caused by a single direction. The fused grayscale difference set provides stable, highly responsive region identification for constructing a structural binary image, effectively improving boundary integrity and noise robustness in subsequent region growing and connectivity determination.
[0133] In one embodiment, the above step S30 includes:
[0134] S301, initializing and generating a full background binary image of the same size as the grayscale image, and initializing all pixel positions in the full background binary image as background areas;
[0135] S302, obtaining a preset difference threshold, traversing the grayscale difference between each pair of adjacent pixels, and determining whether the grayscale difference between the current pair of adjacent pixels is greater than the difference threshold;
[0136] S303, when the grayscale difference between the current pair of adjacent pixels is greater than the difference threshold, determining a corresponding mark position according to the direction type of the current pair of adjacent pixels, and updating the mark position in the full background binary image to a foreground area;
[0137] S304: when the grayscale difference between the current pair of adjacent pixels is not greater than the difference threshold, determining a corresponding mark position according to the direction type of the current pair of adjacent pixels, and maintaining the mark position in the full background binary image as a background area;
[0138] S305 , integrating the foreground areas and background areas of all marked positions in the full-background binary image to generate a complete structural binary image.
[0139] In this embodiment, to determine image structural regions based on grayscale differences, an output image is constructed to record the distribution of structural regions. This image is initialized to a matrix with the same dimensions as the input grayscale image, with all pixels initially set to background values. In practice, this output image can be initialized to a two-dimensional array with a value of 0, indicating that all locations are initially unstructured.
[0140] The structure determination process relies on a pre-defined difference threshold, which determines whether the grayscale change between adjacent pixels is sufficient to constitute a potential structural boundary. This threshold can be determined through historical experience, scene statistics analysis, or data-driven methods. The typical value range may vary depending on the image source type, such as scanned, photographic, or composite images.
[0141] When traversing the fused grayscale difference set, for each pair of adjacent pixel differences, determine whether the difference is greater than the set difference threshold. If so, it indicates that there may be a structural boundary at that location, and it should be marked as a foreground region in the structural binary image; otherwise, it remains as a background region.
[0142] Because grayscale differences are calculated by comparing adjacent pixels, it's impossible to directly mark the position between two pixels in a structural binary image. Instead, a target marking position must be determined based on the position and orientation of the pixel pair. For horizontal pixel pairs, the position of the right pixel is used as the structural marking position. This strategy ensures that the marking position falls in the extension direction of the structural boundary. For vertical pixel pairs, the position of the lower pixel is used as the marking position, also based on the geometric rationality of the structural boundary growth direction.
[0143] When the difference is not greater than the threshold, the marker position analysis step is still performed, but the initial background area remains unchanged at this location. This direction-sensitive pixel mapping method can achieve the distinction and expression of the saliency direction of the structural area, and enhance the geometric expression accuracy of the structural image.
[0144] After the traversal is complete, each pixel in the structural binary image has been updated to either foreground or background based on its performance in the grayscale difference analysis of adjacent pixels. The image as a whole exhibits a sparse structure distribution, with the foreground region roughly covering the boundary regions of grayscale abrupt changes in the original image, forming the structural foundation for subsequent connected region analysis.
[0145] During the processing, the initialization phase generates a structure map array with the same size as the grayscale image by constructing an all-zero matrix. Each position in this array stores a binary state value, which is used to distinguish the edge area of the structure from the background area.
[0146] Obtaining the difference threshold can be achieved through configuration file reading, interface input, or dynamic adaptive algorithm derivation. In static configuration, the threshold is set to a fixed constant, such as 30 in document scanning. In dynamic adaptive solutions, the grayscale gradient variance of the image grayscale histogram is calculated based on its distribution, and the optimal threshold is determined by combining the Otsu algorithm or other minimum inter-class variance methods.
[0147] When traversing the fused grayscale difference set, take out a fused difference and its index position in the image each time, combine the direction information of the difference source, determine the coordinates of the position to be marked, and use the conditional statement to determine whether the difference exceeds the threshold. According to the judgment result, update the value of the corresponding position in the structure diagram array.
[0148] We can further implement rightward and downward structural mapping operations based on the convolution kernel. For example, we can construct template matrices in two directions, convolve the fused difference matrix with the right-biased and downward-biased templates respectively to obtain a preliminary labeled image, and then binarize it according to the threshold, which can speed up the processing of large images.
[0149] This embodiment combines grayscale difference with a difference threshold, combined with a structural direction mapping strategy, to generate a structural binary image that accurately locates image regions containing structural elements such as document boundaries, field outlines, and graphic edges, avoiding the issues of blurred boundaries and content loss associated with traditional full-image grayscale thresholding. It also avoids complex filter design and feature extraction modeling, relying solely on local difference intensity and structural position mapping, resulting in a simple, adaptable, and robust implementation.
[0150] In one embodiment, the above step S40 includes:
[0151] S401, initializing and generating a region labeling matrix with the same size as the structured binary image, wherein all elements in the region labeling matrix are initialized to an unlabeled state;
[0152] S402, scanning each foreground pixel in the structured binary image row by row;
[0153] S403: When an unmarked foreground pixel is scanned, region growing is performed starting from the current unmarked foreground pixel according to the four-connectivity or eight-connectivity judgment rule, all foreground pixels covered by the current connected region are determined, a unique region identifier is assigned to the current connected region, all foreground pixels covered by the current connected region are updated as marked in the region marking matrix, and the corresponding region identifiers are annotated.
[0154] S404, recording attribute parameters of the connected area;
[0155] S405, repeatedly performing the row-by-row scanning and connectivity marking operations until all foreground area pixels in the area marking matrix are in a marked state;
[0156] S406: Integrate the attribute parameters of all marked connected regions to generate a foreground region feature set.
[0157] In this embodiment, a structured binary image typically contains multiple foreground regions of varying shapes and sizes, distributed at different locations. To identify the boundaries and composition of these regions at the pixel level, a connectivity analysis process is performed to segment, number, and attribute all foreground regions in the image.
[0158] First, a region labeling matrix is constructed that is identical in size to the structural binary image to manage whether each pixel has been processed. This matrix is initialized to unlabeled and is often implemented as an all-zero matrix. Each element in this matrix corresponds to a pixel position in the structural image, and the value stores the region identification or unlabeled status. The labeling matrix ensures that each foreground pixel is accessed only once, avoiding duplicate labeling and conflicts.
[0159] When traversing a structured binary image, a row-by-row scan is used to access each pixel position in sequence. In practical implementation, a two-layer nested loop structure can be used, with the first layer traversing the row index and the second layer traversing the column index. During the scanning process, if the current pixel is detected as the foreground region and is unmarked in the region marking matrix, the connected region growing operation is triggered.
[0160] The key to connected region growing lies in the choice of connectivity judgment rules. Common connectivity rules include four-connectivity and eight-connectivity. Four-connectivity establishes connectivity between the current pixel and its upper, lower, left, and right neighbors, while eight-connectivity adds four diagonal pixels to the four-connectivity. The choice of rule affects the shape representation of region boundaries and the sensitivity of region separation. In practice, switching between the two methods can be supported through parameter settings.
[0161] Connected region growing is typically implemented recursively or using a stack / queue approach. Starting from a seed pixel, the search expands layer by layer to adjacent pixels that meet the connectivity criteria, marking these pixels as belonging to the same connected region. Region identifiers are typically auto-incrementing integer sequences to ensure the uniqueness of each connected region. Each pixel marking process updates the value at that position in the labeling matrix, simultaneously recording the region to which it belongs.
[0162] During the labeling process, attribute parameters for each region must be simultaneously counted. The region area refers to the number of pixels marked by the region identifier and can be accumulated during the labeling process. The bounding coordinate set represents the spatial distribution of the region within the image, typically represented by the coordinates of the top-left and bottom-right corners of the minimum bounding rectangle. This attribute can be dynamically recorded during the growth process to avoid post-processing overhead.
[0163] Once a region is marked and its area and boundary information are recorded, the system returns to the outer scan to continue detecting the next unmarked foreground pixel and repeating the same connectivity labeling operation until there are no unmarked foreground pixel positions in the entire structural image.
[0164] The attribute parameters of all marked regions are integrated into a structured foreground region feature set, typically a dictionary or table structure that records region identifiers, areas, and boundary coordinates. This set will serve as the basic data source for subsequent maximum connected region determination and type determination.
[0165] In actual implementation, the region labeling matrix can be initialized to -1 using an integer matrix of the same size as the structural image, representing an unlabeled state. When a new foreground pixel is found and the growth operation is initiated, a region identification variable (for example, numbered 5) can be created, and a recursive depth-first search (DFS) or queued breadth-first search (BFS) method can be used to explore all adjacent foreground pixels that meet the connectivity criteria.
[0166] For example, using a queue to implement BFS, the queue is initialized to the current pixel's coordinates. After dequeuing a pixel, its four / eight adjacent pixels are checked to see if they meet the criteria and are not marked. If so, the pixel is added to the queue and the marking matrix is updated simultaneously. During this process, each time a pixel is marked, its position is added to the coordinate set of the current region, and its horizontal and vertical coordinates in the image are recorded for subsequent bounding box calculation.
[0167] During the region growing process, four variables—minimum row, maximum row, minimum column, and maximum column—are dynamically maintained to update the boundary coordinates of the current connected region in real time. The region's area can be simply calculated by counting the number of marked pixels using a count variable. Connected region attributes are ultimately output as a dictionary structure, including fields such as region ID, area value, and boundary coordinates.
[0168] The above process can be completed in a single-threaded serial manner, or in a multi-threaded manner to process different image blocks in parallel in large-scale image processing tasks, and a lock mechanism is used to prevent data competition when writing the label matrix.
[0169] This embodiment converts the structural binary image into a region labeling matrix and performs pixel-by-pixel connectivity analysis, which can effectively identify all structural foreground areas in the image and assign a unique identifier and complete attribute parameters to each area. Compared with the traditional full-image convolution boundary recognition method, this method has higher flexibility and robustness in terms of computing resource consumption, region isolation accuracy, and subsequent statistical indicator extraction. The region area and boundary information provide a direct basis for the subsequent determination of the largest connected area, and can also be used in other scenarios for functions such as structural block segmentation, field extraction, or image layout reconstruction.
[0170] In one embodiment, the above step S50 includes:
[0171] S501, traversing the area of all connected foreground regions;
[0172] S502, selecting a connected region with the largest area from all connected foreground regions as the maximum connected foreground region;
[0173] S503, obtaining the area value of the maximum connected foreground region;
[0174] S504, determining the total number of pixels of the structured binary image;
[0175] S505 : Divide the area value of the maximum connected foreground region by the total number of pixels to generate a pixel ratio of the maximum connected foreground region in the structural binary image.
[0176] In this embodiment, after connected component labeling, a structural binary image forms multiple foreground regions. Each region has its own independent region identifier and corresponding region attributes, including the number of pixels and boundary coordinates. To determine the degree of image structure concentration, it is necessary to calculate the pixel ratio of the largest connected foreground region to the entire image, i.e., the pixel fraction.
[0177] Before calculation, it is necessary to traverse all marked connected foreground regions and extract their area from the attribute parameters of each region. The area is the number of pixels covered by the same connectivity marker. This data can be dynamically maintained during the connectivity marking stage or obtained by traversing the cumulative number of each marker number in the region marker matrix.
[0178] After obtaining the areas of all connected foreground regions, a comparison operation is performed to determine the region with the largest area value. This selection process is typically completed in a single iteration, during which a variable is used to record the current maximum value and its corresponding region identification number. The maximum region area is the number of pixels in the largest connected foreground region.
[0179] Next, we need to calculate the total number of pixels in the structural binary image. Since this image is a fixed-size image generated using structural rules, the total number of pixels is equal to the product of its height and width. These two parameters can be directly read from the image's dimensional information or recorded when initializing the structural image.
[0180] Finally, the number of pixels in the largest connected area is divided by the total number of pixels in the entire image to get the pixel ratio of the largest connected foreground area. The result is usually in decimal form, and the result value ranges from 0 to 1, representing the proportion of the area in the entire image.
[0181] This value serves as an important indicator of the overall compactness of an image. A larger pixel ratio generally indicates a relatively concentrated content area within the document, such as large blocks of text, graphic areas, or formula blocks. A smaller pixel ratio, on the other hand, may indicate a loose document structure, incomplete edges, high page noise, or the appearance of a photographic document.
[0182] This embodiment can accurately extract the key structural areas that occupy the main body of the image by traversing and maximizing the areas of connected regions in the structural binary image one by one, and quantify their occupancy ratio as a pixel ratio index through standardization. This index serves as an input parameter for subsequent classification and does not rely on image content semantics, color information, or complex models. It can complete the structural judgment of the document type only through pixel-level structural statistics. It has the advantages of strong stability, low computational overhead, and a wide range of applicable scenarios. In scenarios where structural features are significant or document layout is standard, this method can significantly improve the accuracy and processing efficiency of image classification tasks.
[0183] In one embodiment, the above step S60 includes:
[0184] S601, obtaining document type classification threshold parameters from a preset classification threshold database;
[0185] S602, comparing the pixel ratio with the document type classification threshold parameter;
[0186] S603: When the pixel ratio is greater than or equal to the document type classification threshold parameter, determine that the document corresponding to the image to be classified is a text-dominated document and output a first classification identifier;
[0187] S604: When the pixel ratio is less than the document type classification threshold parameter, determine that the document corresponding to the image to be classified is a graphics-intensive document and output a second classification identifier;
[0188] S605: Use the first classification identifier or the second classification identifier as a final image classification result.
[0189] In this embodiment, after obtaining the pixel ratio of the maximum connected foreground area in the structural binary image, it is necessary to further compare the ratio value with the pre-defined classification threshold to achieve automatic classification of the image. First, the document type classification threshold parameters are extracted from the classification threshold database configured in the system. The database can be a local configuration file, a parameter mapping table, an external service interface or a persistent storage system. The classification threshold parameters provided are used to distinguish document images with different structural feature types. They are usually decimals between 0 and 1, such as 0.45 or 0.65, which are used for quantitative comparison with the pixel ratio.
[0190] The system determines the numerical relationship between the proportion of pixels in the maximum connected foreground area and the threshold parameter. The judgment process does not involve fuzzy matching, model reasoning, or image semantic understanding, but only relies on a single indicator for logical judgment, ensuring the clarity of the calculation path and the interpretability of the results.
[0191] When the pixel ratio is greater than or equal to the classification threshold, it indicates a high concentration of primary structures in the structural image, a relatively clustered foreground area, and typical image characteristics such as orderly layout, regular structure, and dense text blocks. Such images are typically electronic documents or scanned documents exported from native systems. The output is then labeled as a text-dominant document and a first classification identifier is generated. This identifier can be a specific label (such as "TEXT-DOMINANT") or a coded value (such as 1), indicating that text content is the primary information carrier.
[0192] When the pixel ratio is less than the classification threshold, the image's largest structural region accounts for a low proportion, indicating a lack of concentrated structural information, a dispersed content distribution, or the presence of significant background information. Such images are more consistent with the typical characteristics of photographed documents, such as printed materials photographed with a mobile phone, with unclear edges and significant background noise. The system labels the image as a graphics-intensive document and outputs a secondary classification identifier, such as "GRAPHIC-INTENSIVE" or a code value of 2, for subsequent classification, circulation, or OCR configuration adjustments.
[0193] Regardless of whether the first or second classification identifier is output, the system will use it as the final image classification result of the current image to be classified, and write it into the image metadata or pass it to subsequent modules as a process control condition.
[0194] Example: In a financial scenario, the system receives a batch of image files uploaded by corporate clients. These images include PDF screenshots of VAT invoices, original contracts taken with mobile phones, and scanned statements. First, the images are decoded and loaded as color images. The image content to be classified is extracted from them, and the image processing module converts the color images into single-channel grayscale images. During this process, the RGB channels are removed, retaining only the brightness information for subsequent difference calculations. The system then iterates over all adjacent pixel pairs in the grayscale image, calculating the absolute values of the pixel differences in the left-right and top-bottom directions. The results generate horizontal and vertical difference matrices, respectively. For each pixel coordinate, the system compares the differences in the two directions, selects the larger value as the fused grayscale difference, and writes it into the fused difference matrix. Each value in the fused difference matrix is then compared with a set difference threshold (e.g., 35). If the value exceeds the difference threshold, the system marks the pixel position to the right or below as foreground, based on the direction of the adjacent pixel pair. The corresponding position in the full background binary image is updated as foreground, while the remaining positions remain background. After the structural binary image is generated, the system performs a connected region analysis operation on it, initializes the labeling matrix, and scans the image line by line. Once unlabeled foreground area pixels are found, region growth is initiated based on the four-connectivity rule, all connected foreground pixels are marked, region identifiers are assigned, and the area of the region and the coordinates of the circumscribed rectangle are recorded. This operation is repeated until all foreground areas in the image are identified and marked. In the constructed foreground area feature set, the system traverses all area areas, identifies the largest connected area, and divides its area by the total number of pixels in the image (height multiplied by width) to calculate the pixel ratio of the area. For example, a 640×960 statement screenshot has a maximum area of 228,300 pixels and a total number of pixels of 614,400, which is approximately 37.15%. The classification threshold for invoice documents (such as set to 30%) is retrieved from the preset classification threshold database, and the actual pixel ratio is compared with the threshold. Since the maximum connected area ratio of the current image is higher than the threshold, the system determines that it is a text-dominated document, outputs the first classification identifier, enters the OCR structure recognition process, and skips the image denoising and border correction steps, thereby improving the overall recognition efficiency.
[0195] A physician assistant uploaded a set of patient medical records to the medical information system, including electronic discharge summaries (system-generated PDFs), mobile phone-photographed laboratory reports, and scanned images of handwritten outpatient records. The system sequentially loaded the image data, removing color information from each input image to generate a grayscale image. Due to uneven illumination and complex backgrounds in photographic images, the system performed edge enhancement operations, such as using the Sobel operator to enhance text block outlines, to enhance structural boundary features in the image. This enhanced the accuracy of subsequent structure extraction. During the structural difference analysis phase, each pair of adjacent pixels in the image was processed to generate horizontal and vertical difference matrices. The system then iterated through the coordinate locations, selecting the direction with the largest difference as the fused difference and writing it into the fused matrix. All differences were then aggregated to form a grayscale difference set. The system set a medical image difference threshold of 40. During the binarization process, for coordinate locations where the fused difference exceeded this threshold, the corresponding location in the structural binary image was updated to the foreground region based on the direction of the adjacent pixel pairs, while the remaining locations were retained as the background region. This resulted in a binary image that reflects structural edge features. Subsequently, the connected region marking operation is performed, the marking matrix is initialized, and region growth is carried out from the unmarked foreground pixel positions using the eight-connectivity rule. The identification and marking status update of all connected foreground regions are completed, and the pixel area and boundary coordinate data of each region are recorded. The system locates the largest region in the feature set, records its area (e.g., 142,800 pixels), calculates the total number of pixels in the image (e.g., 720×1280=921,600), and obtains a pixel ratio of 15.49% for the largest connected region. The judgment threshold setting for inspection documents (e.g., 18%) is loaded from the classification threshold database, and the system compares the ratio value with the threshold. Because the pixel ratio does not reach the classification threshold, the system classifies the image as a graphic-intensive document and outputs a second classification identifier, indicating that its source may be a photograph or a low-structured document. The image enters the manual review auxiliary recognition process, and simultaneously calls background enhancement, image dedistortion and other modules for OCR preparation to ensure recognition accuracy.
[0196] This embodiment achieves automatic classification of structured documents without the support of deep learning models or large-scale training data by directly comparing the pixel ratio of the maximum connected foreground area in the image with the preset classification threshold. It has the advantages of strong configurability, low computational overhead, and high tolerance to structural noise. The classification decision results are clear and concise, which facilitates the system to perform process branching, document routing, or OCR model parameter switching. The attribution of structural features also makes the system more interpretable of the judgment results, meeting the needs of regulatory compliance or business traceability.
[0197] In one embodiment, an image classification device based on connected component analysis is provided, and the image classification device based on connected component analysis corresponds one-to-one to the image classification method based on connected component analysis in the above embodiment. Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the image classification device based on connected component analysis of the present invention. These include a grayscale conversion module 10, a grayscale difference analysis module 20, a structure binary map generation module 30, a connected component analysis module 40, a pixel ratio analysis module 50, and an image classification determination module 60. Each functional module is described in detail below:
[0198] A grayscale conversion module 10 is used to convert the image to be classified into a grayscale image;
[0199] a grayscale difference analysis module 20 for traversing adjacent pixel pairs in the grayscale image and determining the grayscale difference between each pair of adjacent pixels;
[0200] a structure binary image generating module 30, configured to compare the grayscale difference with a difference threshold and generate a structure binary image according to the comparison result;
[0201] A connected region analysis module 40 is configured to perform connected region analysis based on the structural binary image to determine all connected foreground regions;
[0202] a pixel ratio analysis module 50 for determining the pixel ratio of the maximum connected foreground region in the structured binary image;
[0203] The image classification determination module 60 is used to compare the pixel ratio with the classification threshold and output the corresponding image classification type according to the comparison result.
[0204] In one embodiment, the grayscale conversion module 10 is specifically configured to:
[0205] Obtaining a target image to be classified from a data source of images to be classified;
[0206] Eliminating color channel information of the target image to be classified;
[0207] Convert the target image to be classified after eliminating the color channel information into a single-channel grayscale image;
[0208] Performing structure contour enhancement processing on the single-channel grayscale image to generate the grayscale image.
[0209] In one embodiment, the grayscale difference analysis module 20 is specifically configured to:
[0210] Traversing all left and right adjacent pixel pairs in the grayscale image along the horizontal direction, and determining the grayscale absolute value difference between each pair of left and right adjacent pixels and storing it as a horizontal difference matrix;
[0211] Traversing all pairs of upper and lower adjacent pixels in the grayscale image along the vertical direction, and determining the grayscale absolute value difference between each pair of upper and lower adjacent pixels and storing the difference as a vertical difference matrix;
[0212] Traversing each identical row and column coordinate position in the horizontal difference matrix and the vertical difference matrix, and comparing the element value of the current row and column coordinate position in the horizontal difference matrix with the element value of the current coordinate position in the vertical difference matrix;
[0213] Selecting a larger value between the element value at the current row and column coordinate position in the horizontal difference matrix and the element value at the current row and column coordinate position in the vertical difference matrix;
[0214] The larger value is used as the fused grayscale difference value of the current row and column coordinate position;
[0215] Writing the fused grayscale difference into the corresponding row and column coordinate positions of the fused difference matrix;
[0216] The fused grayscale differences of all row and column coordinate positions are summarized to generate a set of grayscale differences of all adjacent pixel pairs.
[0217] In one embodiment, the structure binary image generating module 30 is specifically configured to:
[0218] Initializing and generating a full background binary image of the same size as the grayscale image, and initializing all pixel positions in the full background binary image as background areas;
[0219] Obtain a preset difference threshold, traverse the grayscale difference between each pair of adjacent pixels, and determine whether the grayscale difference between the current pair of adjacent pixels is greater than the difference threshold;
[0220] When the grayscale difference between the current pair of adjacent pixels is greater than the difference threshold, determining a corresponding mark position according to the direction type of the current pair of adjacent pixels, and updating the mark position in the full background binary image to a foreground area;
[0221] When the grayscale difference between the current pair of adjacent pixels is not greater than the difference threshold, determining a corresponding mark position according to the direction type of the current pair of adjacent pixels, and keeping the mark position in the full background binary image as the background area;
[0222] The foreground area and the background area of all marked positions in the full-background binary image are integrated to generate a complete structural binary image.
[0223] In one embodiment, the connected region analysis module 40 is specifically configured to:
[0224] Initializing and generating a region labeling matrix having the same size as the structure binary image, wherein all elements in the region labeling matrix are initialized to an unlabeled state;
[0225] Scanning each foreground pixel in the structured binary image line by line;
[0226] When scanning to an unmarked foreground area pixel, region growing is performed starting from the current unmarked foreground pixel according to the four-connectivity or eight-connectivity judgment rule, all foreground pixels covered by the current connected region are determined, a unique region identifier is assigned to the current connected region, all foreground pixels covered by the current connected region are updated to a marked state in the region marking matrix, and the corresponding region identifiers are marked;
[0227] Recording attribute parameters of the connected area;
[0228] Repeating the row-by-row scanning and connectivity marking operations until all foreground area pixels in the area marking matrix are in a marked state;
[0229] The attribute parameters of all marked connected regions are integrated to generate a foreground region feature set.
[0230] In one embodiment, the pixel ratio analysis module 50 is specifically configured to:
[0231] Traverse the area of all connected foreground regions;
[0232] Select the connected region with the largest area from all connected foreground regions as the maximum connected foreground region;
[0233] Obtaining the area value of the maximum connected foreground region;
[0234] Determining the total number of pixels of the structured binary image;
[0235] The area value of the largest connected foreground region is divided by the total number of pixels to generate a pixel ratio of the largest connected foreground region in the structural binary image.
[0236] In one embodiment, the image classification and determination module 60 is specifically configured to:
[0237] Obtain document type classification threshold parameters from a preset classification threshold database;
[0238] Comparing the pixel ratio with the document type classification threshold parameter;
[0239] When the pixel ratio is greater than or equal to the document type classification threshold parameter, determining that the document corresponding to the image to be classified is a text-dominated document and outputting a first classification identifier;
[0240] When the pixel ratio is less than the document type classification threshold parameter, determining that the document corresponding to the image to be classified is a graphics-intensive document and outputting a second classification identifier;
[0241] The first classification identifier or the second classification identifier is used as a final image classification result.
[0242] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, memory, network interface and database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the server side of an image classification method based on connected region analysis.
[0243] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps of the user side of an image classification method based on connected region analysis.
[0244] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0245] Convert the image to be classified into a grayscale image;
[0246] Traversing adjacent pixel pairs in the grayscale image and determining the grayscale difference between each pair of adjacent pixels;
[0247] Comparing the grayscale difference with a difference threshold, and generating a structural binary image according to the comparison result;
[0248] Performing connected region analysis based on the structural binary image to determine all connected foreground regions;
[0249] Determine the pixel ratio of the maximum connected foreground area in the structured binary image;
[0250] The pixel ratio is compared with the classification threshold, and the corresponding image classification type is output according to the comparison result.
[0251] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0252] Convert the image to be classified into a grayscale image;
[0253] Traversing adjacent pixel pairs in the grayscale image and determining the grayscale difference between each pair of adjacent pixels;
[0254] Comparing the grayscale difference with a difference threshold, and generating a structural binary image according to the comparison result;
[0255] Performing connected region analysis based on the structural binary image to determine all connected foreground regions;
[0256] Determine the pixel ratio of the maximum connected foreground area in the structured binary image;
[0257] The pixel ratio is compared with the classification threshold, and the corresponding image classification type is output according to the comparison result.
[0258] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0259] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0260] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0261] It should be noted that if any software tools or components other than those of the Company appear in the embodiments of this application, they are merely for illustration and do not represent actual use. The above embodiments are intended only to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. An image classification method based on connected component analysis, characterized in that: The following steps are involved: Convert the image to be classified into a grayscale image; Traversing adjacent pixel pairs in the grayscale image and determining the grayscale difference between each pair of adjacent pixels; Comparing the grayscale difference with a difference threshold, and generating a structural binary image according to the comparison result; Performing connected region analysis based on the structural binary image to determine all connected foreground regions; Determine the pixel ratio of the maximum connected foreground area in the structured binary image; The pixel ratio is compared with the classification threshold, and the corresponding image classification type is output according to the comparison result.
2. The image classification method based on connected component analysis according to claim 1, wherein: Convert the image to be classified into a grayscale image, including: Obtaining a target image to be classified from a data source of images to be classified; Eliminating color channel information of the target image to be classified; Convert the target image to be classified after eliminating the color channel information into a single-channel grayscale image; Performing structure contour enhancement processing on the single-channel grayscale image to generate the grayscale image.
3. The image classification method based on connected component analysis according to claim 1, wherein: Traversing adjacent pixel pairs in the grayscale image and determining the grayscale difference between each pair of adjacent pixels, comprising: Traversing all left and right adjacent pixel pairs in the grayscale image along the horizontal direction, and determining the grayscale absolute value difference between each pair of left and right adjacent pixels and storing it as a horizontal difference matrix; Traversing all pairs of upper and lower adjacent pixels in the grayscale image along the vertical direction, and determining the grayscale absolute value difference between each pair of upper and lower adjacent pixels and storing the difference as a vertical difference matrix; Traversing each identical row and column coordinate position in the horizontal difference matrix and the vertical difference matrix, and comparing the element value of the current row and column coordinate position in the horizontal difference matrix with the element value of the current coordinate position in the vertical difference matrix; Selecting a larger value between the element value at the current row and column coordinate position in the horizontal difference matrix and the element value at the current row and column coordinate position in the vertical difference matrix; The larger value is used as the fused grayscale difference value of the current row and column coordinate position; Writing the fused grayscale difference into the corresponding row and column coordinate positions of the fused difference matrix; The fused grayscale differences of all row and column coordinate positions are summarized to generate a set of grayscale differences of all adjacent pixel pairs.
4. The image classification method based on connected component analysis according to claim 1, wherein: Comparing the grayscale difference with a difference threshold, and generating a structural binary image according to the comparison result, including: Initializing and generating a full background binary image of the same size as the grayscale image, and initializing all pixel positions in the full background binary image as background areas; Obtain a preset difference threshold, traverse the grayscale difference between each pair of adjacent pixels, and determine whether the grayscale difference between the current pair of adjacent pixels is greater than the difference threshold; When the grayscale difference between the current pair of adjacent pixels is greater than the difference threshold, determining a corresponding mark position according to the direction type of the current pair of adjacent pixels, and updating the mark position in the full background binary image to a foreground area; When the grayscale difference between the current pair of adjacent pixels is not greater than the difference threshold, determining a corresponding mark position according to the direction type of the current pair of adjacent pixels, and keeping the mark position in the full background binary image as the background area; The foreground area and the background area of all marked positions in the full-background binary image are integrated to generate a complete structural binary image.
5. The image classification method based on connected component analysis according to claim 1, wherein: Performing connected region analysis based on the structural binary image to determine all connected foreground regions, including: Initializing and generating a region labeling matrix having the same size as the structure binary image, wherein all elements in the region labeling matrix are initialized to an unlabeled state; Scanning each foreground pixel in the structured binary image line by line; When scanning to an unmarked foreground area pixel, region growing is performed starting from the current unmarked foreground pixel according to the four-connectivity or eight-connectivity judgment rule, all foreground pixels covered by the current connected region are determined, a unique region identifier is assigned to the current connected region, all foreground pixels covered by the current connected region are updated to a marked state in the region marking matrix, and the corresponding region identifiers are marked; Recording attribute parameters of the connected area; Repeating the row-by-row scanning and connectivity marking operations until all foreground area pixels in the area marking matrix are in a marked state; The attribute parameters of all marked connected regions are integrated to generate a foreground region feature set.
6. The image classification method based on connected component analysis according to claim 1, wherein: Determining the pixel ratio of the maximum connected foreground area in the structured binary image includes: Traverse the area of all connected foreground regions; Select the connected region with the largest area from all connected foreground regions as the maximum connected foreground region; Obtaining the area value of the maximum connected foreground region; Determining the total number of pixels of the structured binary image; The area value of the largest connected foreground region is divided by the total number of pixels to generate a pixel ratio of the largest connected foreground region in the structural binary image.
7. The image classification method based on connected component analysis according to claim 1, wherein: Compare the pixel ratio with the classification threshold, and output the corresponding image classification type based on the comparison result, including: Obtain document type classification threshold parameters from a preset classification threshold database; Comparing the pixel ratio with the document type classification threshold parameter; When the pixel ratio is greater than or equal to the document type classification threshold parameter, determining that the document corresponding to the image to be classified is a text-dominated document and outputting a first classification identifier; When the pixel ratio is less than the document type classification threshold parameter, determining that the document corresponding to the image to be classified is a graphics-intensive document and outputting a second classification identifier; The first classification identifier or the second classification identifier is used as a final image classification result.
8. An image classification device based on connected component analysis, characterized in that: The image classification device based on connected component analysis includes: Grayscale conversion module, used to convert the image to be classified into a grayscale image; a grayscale difference analysis module, configured to traverse adjacent pixel pairs in the grayscale image and determine the grayscale difference between each pair of adjacent pixels; A structural binary image generating module is used to compare the grayscale difference with the difference threshold and generate a structural binary image according to the comparison result; A connected region analysis module, configured to perform connected region analysis based on the structural binary image to determine all connected foreground regions; A pixel ratio analysis module, used to determine the pixel ratio of the maximum connected foreground area in the structural binary image; The image classification determination module is used to compare the pixel ratio with the classification threshold and output the corresponding image classification type according to the comparison result.
9. A computer device, characterized in that: The computer device includes a memory, a processor, and an image classification program based on connected component analysis stored in the memory and capable of running on the processor. When the image classification program based on connected component analysis is executed by the processor, the steps of the image classification method based on connected component analysis as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that The storage medium stores an image classification program based on connected component analysis, which, when executed by a processor, implements the steps of the image classification method based on connected component analysis according to any one of claims 1 to 7.