Scanner image paper fold detection methods, systems, media and products
Patent Information
- Application Number
- CN202511945722.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-22
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2045-12-22
AI Technical Summary
[0004]但是在实际应用中,由于文档扫描过程中存在多种干扰因素,单一检测方式难以有效处理文档边角区域的复杂特征,检测结果容易受到影响,造成误差,从而降低了纸张折角检测的准确性
通过采用上述技术方案,首先对未裁切原始图像进行边缘检测并通过双向扫描提取边界点集合,再对边界点集合进行区域划分以获取角区域边界点子集并构建角度直方图,从而可以基于主导角度序列进行第一次折角检测;若第一次检测未发现折角,则进一步通过直线检测和旋转校正得到摆正图像,并基于纸张边界坐标进行预设裁切处理,再从裁切后图像中提取角区域子图像进行凸多边形拟合实现第二次折角检测,最后综合两次检测结果得出最终的折角检测结果。这种双重检测机制既可以通过角度直方图快速识别明显的折角特征,又能够通过图像校正和局部区域分析发现不易察觉的细微折角,有效克服了单一检测方式难以处理文档边角区域复杂特征的问题,提高了纸张折角检测的准确性。
Smart Images

Figure CN121907963B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, specifically to a method, system, medium, and product for detecting paper folding angles in images from a scanner device. Background Technology
[0002] With the rapid development of office automation and document digitization, document scanning equipment has become an essential piece of office equipment. In the daily document scanning process, the quality of the paper document directly affects the scanning image effect and the accuracy of subsequent document processing.
[0003] Currently, document scanning devices typically employ a single image processing technique to perform quality checks on scanned documents. These methods, after acquiring the scanned image, analyze and process it using only one detection algorithm to determine whether the document has quality issues.
[0004] However, in practical applications, due to various interference factors during document scanning, a single detection method is difficult to effectively handle the complex features of document corner areas, and the detection results are easily affected, resulting in errors and reducing the accuracy of paper corner detection. Summary of the Invention
[0005] This application provides a method, system, medium, and product for detecting paper folds in images from a scanner device, which can improve the accuracy of paper fold detection.
[0006] The first aspect of this application provides a method for detecting paper fold corners in an image from a scanner device, comprising: The uncropped original image captured by the scanner is acquired, and the uncropped original image is converted into a grayscale image and edge detection is performed to obtain an edge-detected image; The edge detection image is scanned bidirectionally to extract boundary points and generate a boundary point set, which includes a left boundary point set and a right boundary point set. The boundary point set is divided into regions, the angular region boundary point subsets of the boundary point set are extracted, and the angular region boundary point subsets are constructed. Filter the dominant angle sequence in the angle histogram, and output the first angle detection signal based on the number of angles and the angle relationship of the dominant angle sequence; If the first fold detection signal does not detect a fold, then straight line detection is performed on the edge detection image to determine grouped straight lines. A reference baseline is selected to perform rotation correction on the uncropped original image to obtain a straightened image. The paper boundary coordinates are determined according to the position of the grouped straight lines in the straightened image. Based on the paper boundary coordinates, the straightened image is subjected to preset cropping processing to obtain a cropped image. Extract corner region sub-images from the cropped image, perform convex polygon fitting on the corner region sub-images, and output a second corner detection signal; The angle detection result is determined based on the first angle detection signal and the second angle detection signal.
[0007] By employing the above technical solution, edge detection is first performed on the uncropped original image, and a set of boundary points is extracted through bidirectional scanning. Then, the set of boundary points is divided into regions to obtain a subset of corner boundary points and an angle histogram is constructed. This allows for the first corner detection based on the dominant angle sequence. If no corner is detected in the first detection, a straightened image is obtained through line detection and rotation correction. A pre-defined cropping process is then performed based on the paper boundary coordinates. A second corner detection is achieved by extracting corner region sub-images from the cropped image and fitting convex polygons. Finally, the results of both detections are combined to obtain the final corner detection result. This dual detection mechanism can quickly identify obvious corner features through angle histograms and also discover subtle, hard-to-detect corners through image correction and local region analysis. This effectively overcomes the problem of single detection methods being unable to handle complex features in document corner regions, thus improving the accuracy of paper corner detection.
[0008] Optionally, the edge detection image is scanned line by line. For each line, pixels with gradient values exceeding a preset gradient threshold from the left boundary to the center are searched as left boundary candidate points, and pixels with gradient values exceeding the preset gradient threshold from the right boundary to the center are searched as right boundary candidate points. The horizontal distance between the left boundary candidate points and the right boundary candidate points in each line is calculated. If the horizontal distance is within a preset distance range, the left boundary candidate point corresponding to the currently traversed line is stored in the left boundary point set, and the right boundary candidate point is stored in the right boundary point set. If the horizontal distance is outside the preset distance range, the currently traversed line is marked as an invalid line, and the next line is traversed until all lines in the edge detection image have been scanned. The left boundary point set and the right boundary point set are combined into the boundary point set.
[0009] Optionally, multiple boundary points are selected sequentially from the subset of boundary points in the corner region according to a preset step size, and boundary vectors are constructed based on adjacent boundary points; the included angle between adjacent boundary vectors is calculated, and boundary vectors with included angles less than a preset included angle threshold are taken as valid boundary vectors; the valid boundary vectors are subjected to angle normalization processing to obtain the target angle mean; and an angle histogram is generated based on the target angle mean and the corresponding angle interval.
[0010] Optionally, the frequency of each angle interval in the angle histogram is obtained; angle intervals with frequencies lower than a preset frequency threshold are filtered out, and the angles corresponding to the preset number of angle intervals with the highest retained frequencies are taken as the dominant angle sequence; if the dominant angle sequence contains only one angle, it is determined that no angle is detected; if there is an angle in the dominant angle sequence whose deviation from the right angle is less than a preset deviation threshold, and adjacent angles in the dominant angle sequence are perpendicular, it is determined that no angle is detected; if the dominant angle sequence does not contain only one angle, or if there is an angle in the dominant angle sequence whose deviation from the right angle is less than a preset deviation threshold, and adjacent angles in the dominant angle sequence are perpendicular, it is determined that an angle is detected; a first angle detection signal is generated based on the angle determination result, and a stop scanning command is returned to the scanner control terminal.
[0011] Optionally, from the grouped straight lines, straight lines located in the left region of the aligned image are selected as the left boundary line candidate set, and straight lines located in the right region of the aligned image are selected as the right boundary line candidate set; the straight line with the smallest x-coordinate in the left boundary line candidate set is selected as the left boundary line, and the straight line with the largest x-coordinate in the right boundary line candidate set is selected as the right boundary line; the coordinates of the intersection point of the left boundary line and the upper boundary of the aligned image are calculated as the upper left corner coordinates, and the coordinates of the intersection point of the left boundary line and the lower boundary of the aligned image are calculated as the lower left corner coordinates; the coordinates of the intersection point of the right boundary line and the upper boundary of the aligned image are calculated as the upper right corner coordinates, and the coordinates of the intersection point of the right boundary line and the lower boundary of the aligned image are calculated as the lower right corner coordinates; the upper left corner coordinates, the lower left corner coordinates, the upper right corner coordinates, and the lower right corner coordinates are combined to form the paper boundary coordinates.
[0012] Optionally, the corner region sub-image is binarized and morphologically filtered to obtain a filtered binary image; multiple connected contours are extracted from the filtered binary image, and the area of each connected contour is calculated; connected contours with areas within a preset area range are selected as valid contours, and convex polygon fitting is performed on each valid contour; if the fitting result is a preset shape, it is determined that the corner region sub-image has a folded corner; if the fitting result is a target shape other than the preset shape, and the ratio of the area of the target shape to the area of the smallest circumscribed triangle is greater than a preset threshold, it is determined that the corner region sub-image has a folded corner; a second folded corner detection signal is generated based on the folded corner determination result of the corner region sub-image.
[0013] Optionally, the background type parameter of the scanner is obtained; if the background type parameter is a black background, pixels with brightness values less than a preset brightness threshold in the corner region sub-image are set as white missing corner regions, and pixels with brightness values greater than or equal to the preset brightness threshold are set as black paper, thus obtaining a binarized image; if the background type parameter indicates a gray background, the RGB three-channel components of each pixel in the corner region sub-image are extracted, and the minimum and maximum values of the RGB three-channel components are calculated. When the minimum value is greater than a preset background brightness threshold or the maximum value is less than a preset dark spot threshold, the corresponding pixel is set as black paper; otherwise, the corresponding pixel is set as a white missing corner region, thus obtaining a binarized image; a morphological opening operation is performed on the binarized image to remove noise points and isolated pixels, thus obtaining the filtered binary image.
[0014] In a second aspect, embodiments of this application provide a scanner device image paper folding corner detection system, the scanner device image paper folding corner detection system comprising: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code including computer instructions, the one or more processors calling the computer instructions to cause the scanner device image paper folding corner detection system to perform the method as described in the first aspect and any possible implementation thereof.
[0015] Thirdly, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on a scanner device image paper folding detection system, cause the scanner device image paper folding detection system to perform the method described in the first aspect and any possible implementation thereof.
[0016] Fourthly, embodiments of this application provide a computer program product containing instructions that, when the computer program product is run on a scanner device image paper folding detection system, cause the scanner device image paper folding detection system to perform the method described in the first aspect and any possible implementation thereof.
[0017] In summary, one or more technical solutions provided in this application have at least the following technical effects or advantages: By employing the above technical solution, edge detection is first performed on the uncropped original image, and a set of boundary points is extracted through bidirectional scanning. Then, the set of boundary points is divided into regions to obtain a subset of corner boundary points and an angle histogram is constructed. This allows for the first corner detection based on the dominant angle sequence. If no corner is detected in the first detection, a straightened image is obtained through line detection and rotation correction. A pre-defined cropping process is then performed based on the paper boundary coordinates. A second corner detection is achieved by extracting corner region sub-images from the cropped image and fitting convex polygons. Finally, the results of both detections are combined to obtain the final corner detection result. This dual detection mechanism can quickly identify obvious corner features through angle histograms and also discover subtle, hard-to-detect corners through image correction and local region analysis. This effectively overcomes the problem of single detection methods being unable to handle complex features in document corner regions, thus improving the accuracy of paper corner detection. Attached Figure Description
[0018] Figure 1 This is a schematic flowchart of a paper folding angle detection method for an image scanner disclosed in an embodiment of this application; Figure 2 This is another schematic flowchart of a method for detecting paper folds in an image using a scanner, as disclosed in an embodiment of this application. Figure 3 This is a schematic diagram of the structure of a system provided in an embodiment of this application.
[0019] Explanation of reference numerals in the attached drawings: 301, Central Processing Unit; 302, Read-Only Memory; 303, Random Access Memory; 304, Bus; 305, Input / Output Interface; 306, Input Section; 307, Output Section; 308, Storage Section; 309, Communication Section; 310, Driver; 311, Removable Media. Detailed Implementation
[0020] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0021] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.
[0022] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0023] This application provides a method for detecting paper folding angles in images from a scanner device, referring to... Figure 1 , Figure 1 This is a schematic flowchart illustrating a method for detecting paper fold corners in an image from a scanner, as provided in an embodiment of this application. The method is applied to a system, which refers to a hardware and software integrated platform capable of executing a scanner image paper fold corner detection program. The system can execute a scanner image paper fold corner detection program, and the method includes steps 101 to 107, as follows: Step 101: Acquire the uncropped original image captured by the scanner, convert the uncropped original image into a grayscale image and perform edge detection to obtain the edge-detected image.
[0024] An uncropped original image refers to a complete scanned image directly output by the scanner, containing all image information of the scanned document and its surrounding background area. A grayscale image is an image obtained by converting the three RGB color channels of a color image into a single brightness channel, where each pixel contains only a grayscale value between 0 and 255. An edge-detected image is a gradient image obtained after processing with an edge detection operator; the gradient value represents the edge strength at that point. Pixel values are larger in edge regions and smaller in non-edge regions.
[0025] Specifically, the complete scanned image data is first acquired through the scanner's image acquisition interface. For RGB color images, a weighted average method is used to convert them into grayscale images. The conversion formula is: Gray = 0.299R + 0.587G + 0.114B, where R, G, and B represent the pixel values of the red, green, and blue channels, respectively. Then, edge detection is performed on the grayscale image, using the Sobel, Canny, or Roberts operators. Taking the Sobel operator as an example, the gradients in the horizontal and vertical directions are calculated separately: using [-1, 0, 1; -2, 0, 2; -1, 0, 1] as the horizontal convolution kernel and [-1, -2, -1; 0, 0, 0; 1, 2, 1] as the vertical convolution kernel, the image is convolved to obtain Gx and Gy. Then, the total gradient value G = sqrt(Gx^2 + Gy^2) is calculated. When the gradient value G is greater than a set threshold, the corresponding pixel is marked as an edge point. The final output is an edge detection image, which highlights the document's contour features and provides basic data for subsequent corner detection.
[0026] Step 102: Perform bidirectional scanning on the edge detection image to extract boundary points and generate a boundary point set, which includes a left boundary point set and a right boundary point set.
[0027] Bidirectional scanning refers to the process of simultaneously searching line by line towards the center from both the left and right sides of the image. Boundary points are the first pixels encountered during the scan that satisfy the gradient threshold condition, representing the positional features of document edges. The boundary point set is a set of all detected boundary points, where the left boundary point set contains boundary feature points from the left side of the image, and the right boundary point set contains boundary feature points from the right side of the image. The gradient threshold is a critical value used to determine whether a pixel is a boundary point.
[0028] Specifically, first, a gradient threshold is set, typically 10% to 30% of the maximum gradient value of the image. The edge detection image is scanned line by line. For each line, a search is performed simultaneously from the left boundary (x=0) and the right boundary (x=image width - 1) towards the center. During the left-side search, the coordinates (x1, y) of the first pixel with a gradient value greater than the gradient threshold are recorded as a candidate left boundary point; during the right-side search, the coordinates (x2, y) of the first pixel with a gradient value greater than the gradient threshold are recorded as a candidate right boundary point. The horizontal distance d = x2 - x1 between the candidate left and right boundary points is calculated and compared to a preset reasonable range. If the distance d falls within the preset range, the current candidate left boundary point is added to the left boundary point set, and the candidate right boundary point is added to the right boundary point set; if the distance d exceeds the preset range, the current line is marked as invalid, and the process continues to the next line. After scanning all lines, the left and right boundary point sets are merged to form a complete boundary point set, providing a data foundation for subsequent corner region analysis.
[0029] In one possible implementation, the edge detection image is scanned bidirectionally to extract boundary points and generate a boundary point set, specifically including steps 1021-1023, as follows: Step 1021: Scan the edge detection image line by line. For each line, search for pixels whose gradient value exceeds a preset gradient threshold from the left boundary to the center as candidate points for the left boundary, and search for pixels whose gradient value exceeds a preset gradient threshold from the right boundary to the center as candidate points for the right boundary.
[0030] The gradient value represents the degree of change in an image at a given point, calculated as the grayscale difference between that point and its neighboring pixels. The preset gradient threshold is a standard value used to determine whether a pixel is a boundary point; it is typically set as a statistical value based on the overall gradient value distribution of the image. Boundary candidate points are pixels whose gradient values exceed the preset threshold for the first time during scanning, representing potential edge locations in the document.
[0031] Specifically, for each row y (from 0 to image height - 1) in the edge detection image, searches are performed in both the left and right directions. During the left-side search, the x-coordinate increments from 0, and the Canny gradient value at each pixel (x, y) is calculated. When a gradient value g(x, y) greater than a preset gradient threshold is detected, the point is recorded as a candidate point for the left boundary, and its coordinates (x, y) are stored. The right-side search follows the same principle, with the x-coordinate decreasing from the image width - 1. Gradient values are calculated and compared with the preset threshold, and the first pixel meeting the condition is selected as a candidate point for the right boundary. The preset gradient threshold is determined as follows: first, the gradient values of all pixels in the image are calculated, and the histogram distribution of the gradient values is obtained. Then, 50% of the peak value of the histogram is taken as the threshold. For example, for an 8-bit grayscale image, if the maximum gradient value is 255 and the peak value of the resulting gradient histogram is 100, then the preset gradient threshold is set to 50. This bidirectional search method effectively captures the left and right boundary features of the document, providing basic data for subsequent boundary analysis.
[0032] Step 1022: Calculate the horizontal distance between the left and right boundary candidate points of each row; if the horizontal distance is within the preset distance range, store the left boundary candidate point corresponding to the currently traversed row into the left boundary point set, and store the right boundary candidate point into the right boundary point set.
[0033] Horizontal distance refers to the coordinate difference between the left and right boundary candidate points along the x-axis. The preset distance range is a reasonable boundary spacing interval set according to the standard document size, used to filter out abnormal boundary point pairs. The boundary point set is a data structure that stores boundary points that meet the conditions, divided into a left boundary point set and a right boundary point set, used to store the coordinate information of valid boundary feature points in the document.
[0034] Specifically, for the left boundary candidate point (x1, y) and the right boundary candidate point (x2, y) detected in the current scanning line y, the horizontal distance d=x2-x1 is calculated. The preset distance range is determined based on the actual size of a standard A4 paper under the scanning resolution. Assuming the scanning resolution is 300 dpi and the width of A4 paper is 210 mm, the standard width of an A4 paper in the image is 2480 pixels. In consideration of the deviation of paper placement during actual scanning, the preset distance range is set to [0.85×2480, 1.15×2480] pixels, that is, [2108, 2852] pixels. When the calculated horizontal distance d falls within this range, it indicates that the pair of boundary candidate points of the current line represents the real boundary of the document. At this time, the left boundary candidate point (x1, y) is added to the left boundary point set, and the right boundary candidate point (x2, y) is added to the right boundary point set. Each boundary point is stored in the form of a coordinate pair (x, y), and the gradient value and line number information of the point are recorded at the same time to form a complete boundary point data structure. This distance-based screening method can effectively remove false boundary points caused by noise, shadow and other factors, and improve the accuracy of boundary detection.
[0035] Step 1023: If the horizontal distance is outside the preset distance range, mark the currently traversed line as an invalid line, and traverse the next line until the scanning of all lines in the edge detection image is completed; combine the left boundary point set and the right boundary point set into a boundary point set.
[0036] The boundary point set is a complete boundary feature point dataset formed after merging the left boundary point set and the right boundary point set, and contains all valid feature point information of the document edge.
[0037] Specifically, when scanning each line, first check whether the horizontal distance d between the left and right boundary candidate points is within the preset range [dmin, dmax]. When d<dmin or d>dmax, it indicates that the boundary detection result of this line is unreliable, and it will not be added to the boundary point set. The program continues to process the scanning of the next line until the scanning of the entire image is completed (y reaches image height - 1). After the scanning of all lines is completed, a left boundary point set L and a right boundary point set R are obtained, which are used for subsequent corner folding detection analysis.
[0038] Step 103: Perform region division on the boundary point set, extract the corner region boundary point subset of the boundary point set, and construct an angle histogram of the corner region boundary point subset.
[0039] The corner region boundary point subset is a boundary point set located in the four corner regions of the image, and usually contains the feature information of document folded corners. The angle histogram is a statistical graph that counts the angle distribution between adjacent boundary points in the corner region, and is used to characterize the direction characteristics of boundary line segments.
[0040] Specifically, first, the range of the angular region is defined. The left boundary point set L stores the coordinates of the left boundary points from top to bottom. The first 1 / 4 is taken as the top-left boundary point set, and the last 1 / 4 is taken as the bottom-left boundary point set. The right boundary point set R stores the coordinates of the right boundary points from top to bottom. The first 1 / 4 is taken as the top-right boundary point set, and the last 1 / 4 is taken as the bottom-right boundary point set. Then, an angle histogram is constructed: the angle interval step size is set to 10 degrees, and the 180 degrees are divided into 18 equal intervals. Based on the step size, three adjacent points P1, P2, and P3 are taken respectively.<P1,P2> and<P2,P3> If two line segments have similar angles, they are considered valid boundary points. If a point is a valid boundary point, then calculation is performed.<P1,P2> and<P2,P3> The average angle of the two line segments is used as the current angle (normalized to between 0 and 180 degrees). This angle value is added to the corresponding angle interval count. For example, when the calculated included angle α = 75 degrees, the count of those falling into the 70-80 degree interval is incremented by 1. After completing the angle calculation for all adjacent point pairs, an angle histogram H
[18] is obtained, which records the distribution of the direction of the boundary line segments, where H[i] represents the number of boundary line segments in the i-th angle interval. This angle histogram reflects the directional characteristics of the angular region boundary and provides a basis for subsequent angle judgment.
[0041] In one possible implementation, an angular histogram of the subset of boundary points of the angular region is constructed, specifically including steps 1031-1033, as follows: Step 1031: Select multiple boundary points sequentially from the set of boundary points in the corner region according to a preset step size, and construct a boundary vector based on the adjacent boundary points.
[0042] The preset step size is the interval between selections of boundary points in the subset of boundary points in the corner region, used to control the density of sampling points. A boundary vector is a directed line segment formed by two adjacent boundary points, containing direction and magnitude information, used to describe the local features of the boundary. Boundary points are feature points located at the edge of the document, containing their spatial coordinate information.
[0043] Specifically, first, determine the value of the preset step size s, which is usually set to 2~5. Starting from the starting position of the subset C of boundary points in the corner region, select a boundary point every s points, denoted as P[i] (i=0, 1, 2, ..., n). For adjacent boundary points P[i] and P[i+1], construct the boundary vector V[i]=P[i+1]-P[i]. The mathematical representation of the vector V[i] is: V[i].x=P[i+1].xP[i].x, V[i].y=P[i+1].yP[i].y. At the same time, calculate the length L[i]=sqrt(V[i].x^2+V[i].y^2) and the direction angle θ[i]=arctan(V[i].y / V[i].x). To ensure the consistency of vector direction, the direction angle needs to be corrected according to the vector's position in different quadrants: when V[i].x < 0, θ[i] = θ[i] + π; when V[i].x > 0 and V[i].y < 0, θ[i] = θ[i] + 2π. All boundary vector information (starting point coordinates, ending point coordinates, length, and direction angle) is stored in an array, forming a boundary vector set {V[0], V[1], ..., V[n-1]}. This sampling method based on a fixed step size reduces data redundancy while retaining the main feature information of the boundary, providing basic data for subsequent angle analysis.
[0044] Step 1032: Calculate the angle between adjacent boundary vectors, and take the boundary vectors with an angle less than a preset angle threshold as valid boundary vectors.
[0045] Adjacent boundary vectors refer to two consecutive vectors in a boundary vector sequence. The included angle is the angular difference between two adjacent boundary vectors. A preset included angle threshold is a standard value used to determine whether a boundary vector is valid. A valid boundary vector is a boundary vector whose included angle with its adjacent vectors satisfies the preset threshold condition, representing the continuity of the boundary.
[0046] Specifically, for each pair of adjacent vectors V[i] and V[i+1] in the boundary vector set {V[0], V[1], ..., V[n-1]}, the angle α[i] between them is calculated. The angle is calculated using the vector dot product formula: α[i] = arccos((V[i].x × V[i+1].x + V[i].y × V[i+1].y) / (|V[i]| × |V[i+1]|)), where |V[i]| represents the magnitude of vector V[i]. To ensure the accuracy of the calculation results, a numerical correction is required: when the dot product result is greater than 1, it is set to 1; when it is less than -1, it is set to -1. The preset angle threshold is usually set to 10 degrees (π / 18 radians). For each boundary vector V[i], check its angle with its preceding and following vectors: if α[i] < π / 18, mark V[i] as a valid boundary vector and record it in the valid vector set E; if the angle is greater than or equal to π / 18, mark V[i] as an invalid vector. Each vector in the valid vector set E stores the following information: starting coordinates (x1, y1), ending coordinates (x2, y2), vector length L, direction angle θ, and angles α1 and α2 with the preceding and following vectors. This angle-based filtering method can remove abrupt changes on the boundary, retain boundary features with good continuity, and provide a reliable data foundation for subsequent angle statistics.
[0047] Step 1033: Perform angle normalization on the effective boundary vector to obtain the target angle mean; generate an angle histogram based on the target angle mean and the corresponding angle interval.
[0048] Angle normalization is a standardization process that maps the direction angles of different effective boundary vectors to the interval [0, 180). The target angle mean is the average angle obtained by statistically analyzing the normalized angle values over a local area. An angle interval is a range of angles obtained by dividing the 180-degree range into equal parts. An angle histogram is a frequency distribution chart that statistically analyzes the number of boundary vectors within each angle interval.
[0049] Specifically, the orientation angle θ of each vector in the effective boundary vector set E is first normalized. The normalization formula is: when θ < 0, θ_norm = θ + 180; when θ ≥ 180, θ_norm = θ - 180; when 0 ≤ θ < 180, θ_norm = θ. Local smoothing is then performed on the normalized angle values: the angle values of each vector and its k neighboring vectors (k = 1) are taken, and their arithmetic mean is calculated as the target angle mean θ_mean[i] = (θ_norm[ik] + ... + θ_norm[i] + ... + θ_norm[i + k]) / (2k + 1). Then, the 180-degree range is divided into 18 angle intervals, each spanning 10 degrees. An array H of length 18 is created to store the angle histogram, where H[j] represents the number of boundary vectors falling within the j-th angle interval [10j, 10(j+1)). For each valid boundary vector's target angle mean θ_mean[i], calculate its corresponding angle interval index j=floor(θ_mean[i] / 10), and increment the count of H[j]. The resulting angle histogram H reflects the distribution characteristics of the boundary vector directions, with peak positions corresponding to the main directions of the boundary, and the shape of the histogram reflecting the continuity and variation characteristics of the boundary. This statistical method can effectively describe the overall directional characteristics of the boundary, providing an important basis for subsequent angle determination.
[0050] Step 104: Filter the dominant angle sequence in the angle histogram, and output the first angle detection signal based on the number of angles and the angle relationship of the dominant angle sequence.
[0051] The dominant angle sequence is an ordered sequence of the most frequent angle values in the angle histogram, representing the main directional features of the boundary. The number of angles refers to the number of angle values in the dominant angle sequence. The angle relationship refers to the geometric relationship between adjacent angles in the dominant angle sequence. The first bend detection signal is a binary output signal that determines whether a bend exists based on the characteristics of the dominant angle sequence.
[0052] Specifically, firstly, the frequency values in the angle histogram H are sorted in descending order, and the top 3 most frequent angle intervals are selected as candidate dominant angles. For the retained angle intervals, their center angle value is calculated: α[i] = 10 × i + 5, where i is the angle interval index. The angle values that meet the frequency requirements are arranged in descending order of frequency to form the dominant angle sequence A. The length n of the dominant angle sequence is counted, representing the number of dominant directions. When n = 1, it means that there is only one dominant direction at the boundary, and the no-angle signal (0) is directly output; when n ≥ 2, the included angle between adjacent dominant angles is calculated: β[i] = min(|A[i+1] - A[i]|, 360 - |A[i+1] - A[i]|). Check whether these included angles are close to 90 degrees: for each included angle β[i], its deviation from 90 degrees is calculated: δ[i] = |β[i] - 90|. The deviation threshold is set to 10 degrees. If all δ[i] are less than 10 degrees, it indicates that the adjacent dominant directions are perpendicular, and the output signal is no fold (0); otherwise, it indicates that there is a non-orthogonal turning point at the boundary, and the output signal is fold (1). The detection signal is sent to the scanner through the control interface. When a fold is detected (signal is 1), a scanning stop command is triggered. This detection method based on angle statistics can effectively identify abnormal folding features of document corners.
[0053] In one possible implementation, the dominant angle sequence in the angle histogram is filtered, and a first angle detection signal is output based on the number of angles and the angle relationship of the dominant angle sequence. Specifically, this includes steps 1041-1043, as follows: Step 1041: Obtain the frequency of each angle interval in the angle histogram; filter out angle intervals with frequencies lower than a preset frequency threshold, and use the angles corresponding to the preset number of angle intervals with the highest frequencies as the dominant angle sequence.
[0054] Angle interval frequency refers to the number of boundary vectors contained in each interval of the angle histogram. The preset frequency threshold is the minimum frequency value used to filter valid angle intervals. The dominant angle sequence is an ordered array consisting of the center angle values of the most frequent angle intervals. The preset quantity refers to the number of dominant angles to be retained, used to limit the length of the dominant angle sequence.
[0055] Specifically, first, the angle histogram array H[0:18] is traversed, and the frequency value freq[i] = H[i] of each angle interval is extracted. A frequency-angle pair array `pairs` is created, where each element contains the angle interval index i and the corresponding frequency value freq[i]. The `pairs` array is sorted in descending order of frequency value. A preset quantity N = 3 is set, and the angle intervals corresponding to the top N highest frequencies are selected from the remaining frequency-angle pairs. For each retained angle interval i, its center angle value is calculated: angle[i] = 10 × i + 5. For example, for the angle interval [30, 40), its center angle is 35 degrees. These center angle values are arranged in descending order of their corresponding frequencies to form the dominant angle sequence A. Each dominant angle stores the following information: angle value, corresponding frequency value, and interval index in the original histogram. The final dominant angle sequence A contains the main features of the boundary direction, providing basic data for subsequent angle judgment.
[0056] Step 1042: If the dominant angle sequence contains only one angle, it is determined that no angle has been detected; if the dominant angle sequence contains an angle whose deviation from the right angle is less than a preset deviation threshold, and the adjacent angles in the dominant angle sequence are perpendicular, it is determined that no angle has been detected; if the dominant angle sequence does not contain only one angle, or if the dominant angle sequence contains an angle whose deviation from the right angle is less than a preset deviation threshold, and the adjacent angles in the dominant angle sequence are perpendicular, it is determined that an angle has been detected.
[0057] The dominant angle sequence length refers to the number of angle values in the sequence. Right angles refer to standard right angle values such as 90 degrees, 180 degrees, 270 degrees, and 360 degrees. The preset deviation threshold is the tolerance range for judging whether an angle is close to a right angle. Perpendicular relationship refers to the difference between two angles being close to 90 degrees. The angle determination result is a binary judgment result based on angle feature analysis.
[0058] Specifically, first, obtain the length n of the dominant angle sequence A. When n=1, directly return the no-angle judgment result (0), which indicates that the boundary has only a single direction. When n≥2, perform the following judgment process: First, construct a standard right angle array R=[90, 180, 270, 360], and set the preset deviation threshold δ=10 degrees. For each angle A[i] in the dominant angle sequence, calculate its minimum deviation from all standard right angles: min_diff[i]=min(|A[i]-R[j]|), j=0, 1, 2, 3. If there exists min_diff[i]<δ, then mark the angle as an approximate right angle. For example, when A[0]=92 degrees, its deviation from R[0]=90 degrees is 2 degrees, which is less than the threshold of 10 degrees, so A[0] is marked as an approximate right angle. Next, check the perpendicular relationship between adjacent angles: calculate the absolute value of the difference between adjacent angles, diff[i] = |A[i+1] - A[i]|, and standardize diff[i] to the interval [0, 90]. If diff[i] > 180 degrees, then diff[i] = 360 degrees - diff[i]. Calculate the deviation of the standardized angle difference from 90 degrees: angle_diff[i] = |diff[i] - 90|. If all angle_diff[i] are less than the preset deviation threshold δ, then the adjacent angles are determined to be perpendicular. The final judgment logic is: if there is a right angle approximation in the sequence and the adjacent angles are perpendicular, then return the no-angle judgment result (0); otherwise, return the angle-bending judgment result (1). This judgment method based on angle relationship can accurately distinguish between normal right angle boundaries and abnormal angle boundaries.
[0059] Step 1043: Generate the first angle detection signal based on the angle determination result, and return a stop scanning command to the scanner control terminal.
[0060] The angle detection result is a binary logic judgment obtained from the previous steps; 1 indicates an angle was detected, and 0 indicates no angle was detected. The first angle detection signal is the output that converts the angle detection result into a standard level signal. The scanner control terminal is the hardware interface responsible for receiving and executing control commands. The stop scanning command is a control command that instructs the scanner to terminate the current scanning task.
[0061] Specifically, a standard-format first angle detection signal is constructed based on the angle determination result. The signal structure includes the following fields: signal type (type=1, indicating first-level detection), timestamp (recording the detection time), detection result (result, 0 or 1), and confidence (a reliability index calculated based on the frequency of the dominant angle). The detection signal is encapsulated in binary format: the first byte is the signal type, bytes 2-5 are the timestamp, the sixth byte is the detection result, and the seventh byte is the confidence. When the detection result is 1, a stop scanning command is sent through the scanner's control interface. The stop command format is: start byte (0xAA), command type (0x01 indicating stop scanning), parameter length (0x00), and checksum byte. Before sending the command, the checksum is calculated: all byte values except the checksum byte are added together, and the lowest byte is taken as the checksum. For example, a complete stop command sequence is [0xAA, 0x01, 0x00, 0xAB]. After sending the command, the system waits for the scanner to return an acknowledgment signal (ACK). If no acknowledgment signal is received within the preset timeout period (usually 100ms), the command is resent, with a maximum of 3 retries. This standard signal and command processing mechanism ensures that the detection results are reliably transmitted to the scanner control system and that the scanning process of documents with folds is interrupted in a timely manner.
[0062] Step 105: If the first fold detection signal does not detect a fold, then perform straight line detection on the edge detection image, determine the grouped straight lines, select a reference baseline to perform rotation correction on the uncropped original image to obtain a straightened image, determine the paper boundary coordinates based on the position of the grouped straight lines in the straightened image, and perform preset cropping processing on the straightened image based on the paper boundary coordinates to obtain the cropped image.
[0063] Grouped lines are sets of lines categorized based on their position and orientation features. The reference baseline is a baseline line used for image rotation correction. The straightened image is the image after rotation correction. Paper boundary coordinates are the positional information of the four corner points of the document. The preset cropping process removes non-document areas from the image based on the boundary coordinates.
[0064] Specifically, firstly, Hough transform is applied to the edge detection image for line detection. The parameters of Hough transform are set as follows: the angle resolution is 1 degree, the distance resolution is 1 pixel, and the accumulator threshold is 20% of the diagonal length of the image. The detected lines are represented by polar coordinates (ρ, θ), where ρ is the vertical distance from the line to the origin, and θ is the included angle between the perpendicular line and the x-axis. The detected lines are grouped according to the angle θ, and lines with an angle difference of less than 5 degrees are classified into the same group. The average angle θ_mean and the average distance ρ_mean are calculated for each group of lines, so as to obtain the representative line of the group. The longest representative line in the horizontal direction (|θ| < 5 degrees) is selected as the reference baseline. The inclination angle α of the reference baseline is calculated as α=arctan(θ), and the rotation transformation matrix M is constructed as M=[[cos(α), -sin(α)], [sin(α), cos(α)]]. Affine transformation is applied to the uncropped original image: for each pixel point (x, y), the rotated coordinate (x', y')=M×(x, y) is calculated, and the pixel value at the new coordinate is calculated through bilinear interpolation to obtain a straightened image. In the straightened image, the grouped lines are divided into a left boundary line group (x<width / 3) and a right boundary line group (x>2×width / 3) according to their positions. The outermost line in each group is selected as the document boundary line. The intersection points between the boundary lines and the upper and lower boundaries of the image are calculated, and the coordinates of the four corner points of the document (left_top, left_bottom, right_top, right_bottom) are obtained. According to these coordinates, the cropping area is defined on the straightened image: left boundary=min(left_top.x, left_bottom.x)+margin, right boundary=max(right_top.x, right_bottom.x)-margin, upper boundary=min(left_top.y, right_top.y)+margin, lower boundary=max(left_bottom.y, right_bottom.y)-margin, wherein margin is a preset margin (usually 10 pixels). The image in this area is extracted to obtain the final cropped image. This image correction and cropping method based on line detection can effectively remove the tilted and redundant areas in the image.
[0065] In a possible implementation, determining the paper boundary coordinates according to the positions of the grouped lines in the straightened image specifically includes step 1051 to step 1054, the above steps are as follows: Step 1051: screening out lines located in the left area of the straightened image from the grouped lines as a left boundary line candidate set, and screening out lines located in the right area of the straightened image as a right boundary line candidate set.
[0066] The grouped straight lines are a set of straight lines obtained after angle clustering. The left image region refers to the range from the left boundary to 1 / 2 of the image width in the horizontal direction of the image. The right image region refers to the range from 1 / 2 of the image width to the right boundary in the horizontal direction of the image. The left boundary straight line candidate set is a set of straight lines located in the left region whose directions are close to vertical. The right boundary straight line candidate set is a set of straight lines located in the right region whose directions are close to vertical. The upper image region refers to the range from the upper boundary to 1 / 2 of the image height in the vertical direction of the image. The lower image region refers to the range from 1 / 2 of the image height to the lower boundary of the image in the vertical direction of the image. The upper boundary straight line candidate set is a set of straight lines located in the upper region whose directions are close to horizontal. The lower boundary straight line candidate set is a set of straight lines located in the lower region whose directions are close to horizontal.
[0067] Specifically, first obtain the size information (width, height) of the image, and define the range of the left region [0, width / 2] and the range of the right region [width / 2, width]. For each straight line L in the vertical straight line grouping set, calculate the midpoint coordinate of the straight line xc=(x1+x2) / 2. Classify the straight line according to the position of the midpoint coordinate xc: when xc<width / 2, check whether the angle of the straight line is close to vertical (85°<θ<95° or 265°<θ<275°), if the condition is satisfied, add the straight line to the left boundary straight line candidate set L_left; when xc>width / 2, also check the angle condition, if the condition is satisfied, add the straight line to the right boundary straight line candidate set L_right. For each straight line in the candidate sets, record the following information: starting point coordinates (x1, y1), end point coordinates (x2, y2), midpoint coordinates (xc, yc), angle θ, and length length. This straight line screening method based on position and direction can effectively identify the left and right boundary lines of the document, and provide reliable candidate straight lines for subsequent boundary positioning.
[0068] Step 1052: selecting the straight line with the smallest abscissa from the left boundary straight line candidate set as the left boundary straight line, and selecting the straight line with the largest abscissa from the right boundary straight line candidate set as the right boundary straight line.
[0069] The abscissa refers to the x-axis position value of a vertical straight line in the image coordinate system. The left boundary straight line is the straight line closest to the left edge of the document selected from the left boundary candidate set. The right boundary straight line is the straight line closest to the right edge of the document selected from the right boundary candidate set.
[0070] Specifically, the actual boundaries of the paper are determined by analyzing the positional characteristics of straight lines. In practice, first, each line in the candidate set of left boundary lines is traversed, and the x-coordinate value of the intersection point of each line with the upper boundary of the image is calculated. These x-coordinate values are compared, and the line with the smallest x-coordinate value is selected as the final left boundary line. This is because in the image coordinate system, a smaller x-coordinate indicates a position further to the left, and the leftmost line is most likely the true left boundary of the paper. Similarly, each line in the candidate set of right boundary lines is traversed, and the x-coordinate value of their intersection points with the upper boundary of the image is calculated. The line with the largest x-coordinate value is selected as the final right boundary line, because a larger x-coordinate indicates a position further to the right. In this way, the left and right boundary positions of the paper can be accurately located.
[0071] Step 1053: Calculate the coordinates of the intersection point of the left boundary line and the upper boundary of the straightened image as the coordinates of the upper left corner; calculate the coordinates of the intersection point of the left boundary line and the lower boundary of the straightened image as the coordinates of the lower left corner; calculate the coordinates of the intersection point of the right boundary line and the upper boundary of the straightened image as the coordinates of the upper right corner; calculate the coordinates of the intersection point of the right boundary line and the lower boundary of the straightened image as the coordinates of the lower right corner.
[0072] The left and right boundary lines are two straight lines selected in the previous steps to represent the document edges. The intersection point coordinates are the (x, y) values of the point where the line intersects the boundary line. The coordinates of the top-left, bottom-left, top-right, and bottom-right corners together constitute the position information of the four vertices of the document.
[0073] Specifically, first, determine the intersection of the left boundary line and the top of the aligned image, and record the coordinates of this point as the top-left corner coordinates. Then, determine the intersection of the left boundary line and the bottom of the image, and record the coordinates of this point as the bottom-left corner coordinates. Next, using the same method, find the intersection of the right boundary line and the top of the image as the top-right corner coordinates, and the intersection of the right boundary line and the bottom of the image as the bottom-right corner coordinates. This gives us the accurate positions of the four corner points of the document. If any intersection point is outside the image boundaries, adjust it to the image boundaries. In this way, we obtain precise boundary coordinates for subsequent image cropping.
[0074] Step 1054: Combine the coordinates of the top left corner, bottom left corner, top right corner, and bottom right corner to form the paper boundary coordinates.
[0075] The top-left, bottom-left, top-right, and bottom-right corner coordinates are the two-dimensional coordinate values of the four vertices of the document, calculated in the previous steps. The paper boundary coordinates are a data structure composed of these four vertex coordinates in a specific order, used to describe the complete boundary information of the document in the image. The combination process integrates discrete coordinate points into an ordered set of coordinates in a unified format.
[0076] Specifically, firstly, define the data structure for the boundary coordinates, which includes the following fields: vertex array vertices[4], each vertex containing x and y coordinates; boundary type type (fixed value is "rectangle"); vertex order order (fixed value is "clockwise"). Create a vertex array and store the coordinates of the four vertices in clockwise order: vertices[0]={x: x_left_top, y: y_left_top}, vertices[1]={x: x_right_top, y: y_right_top}, vertices[2]={x: x_right_bottom, y: y_right_bottom}, vertices[3]={x: x_left_bottom, y: y_left_bottom}. Calculate additional boundary features: diagonal length = sqrt((x_right_bottom - x_left_top)^2 + (y_right_bottom - y_left_top)^2), bounding box width = max(x_right_top, x_right_bottom) - min(x_left_top, x_left_bottom), bounding box height = max(y_left_bottom, y_right_bottom) - min(y_left_top, y_right_top). Construct a complete boundary coordinate data structure: boundary = {vertices: vertices, type: "rectangle", order: "clockwise", diagonal: diagonal, width: width, height: height}. This standardized boundary coordinate representation not only includes basic positional information but also provides the geometric features of the boundary, facilitating subsequent image cropping and processing operations.
[0077] In a preferred embodiment, the process further includes dividing the grouped straight lines and determining the paper boundary coordinates, specifically including: The grouped straight lines are divided into horizontal straight line groups (which may be inclined) and vertical straight line groups (which may be inclined) according to their slopes. The horizontal straight line groups are sorted from top to bottom by the vertical coordinate of the midpoint of the line; the vertical straight line groups are sorted from left to right by the horizontal coordinate of the midpoint of the line.
[0078] From the vertical line group, lines located in the left region of the original image are selected as the left boundary line candidate set, with the image center point as the boundary. Lines located in the right region of the straightened image are selected as the right boundary line candidate set. The line with the smallest x-coordinate in the left boundary line candidate set is selected as the left boundary line, and the line with the largest x-coordinate in the right boundary line candidate set is selected as the right boundary line.
[0079] From the grouped horizontal lines, lines located in the upper region of the image are selected as the upper boundary line candidate set, with the image center point as the boundary. Lines located in the lower region of the image are selected as the lower boundary line candidate set. The line with the smallest y-coordinate in the upper boundary line candidate set is selected as the upper boundary line, and the line with the largest y-coordinate in the lower boundary line candidate set is selected as the lower boundary line. Here, the y-coordinate refers to the y-axis position value of the horizontal line in the image coordinate system. The upper boundary line is the line closest to the upper edge of the document selected from the upper boundary candidate set. The lower boundary line is the line closest to the lower edge of the document selected from the lower boundary candidate set. The y-coordinate of the line is represented by the average of the starting and ending y-coordinate points of the line. The upper and lower boundary lines are the two lines selected to represent the upper and lower edges of the document.
[0080] A reference baseline is found among the four boundary lines: upper, lower, left, and right. The reference baseline is a boundary line determined by the algorithm; it can be either the upper or lower boundary line. Similarly, it can be either the left or right boundary line. The boundary line with perpendicular lines at both its starting and ending positions is selected as the baseline.
[0081] Estimate the image tilt angle based on the aforementioned reference baseline, and rotate the image around the baseline center point to straighten it. Ensure that any region of the image remains in the new image after rotation; that is, the image size may increase after rotation, which will be cropped in later steps.
[0082] The coordinate position of the boundary line in the straightened image is recalculated based on the rotation angle determined by the aforementioned reference baseline.
[0083] The coordinates of the intersection point of the left boundary line and the upper boundary line of the straightened image are calculated as the coordinates of the upper left corner, and the coordinates of the intersection point of the left boundary line and the lower boundary line of the straightened image are calculated as the coordinates of the lower left corner.
[0084] The coordinates of the intersection point of the right boundary line and the upper boundary line of the straightened image are calculated as the upper right corner coordinates, and the coordinates of the intersection point of the right boundary line and the lower boundary line of the straightened image are calculated as the lower right corner coordinates.
[0085] When a boundary line is not detected, the boundary coordinate position is obtained by grayscale projection.
[0086] The coordinates of the top left corner, the bottom left corner, the top right corner, and the bottom right corner are combined to form the paper boundary coordinates.
[0087] Step 106: Extract the corner region sub-image from the cropped image, perform convex polygon fitting on the corner region sub-image, and output the second corner detection signal.
[0088] The cropped image is an image containing only the document region obtained by boundary cropping. Corner region sub-images are local image patches extracted from the four corners of the cropped image. Convex polygon fitting is the process of fitting the corner region boundary points into convex polygons. The second corner detection signal is a corner judgment signal output based on the convex polygon fitting result.
[0089] Specifically, the extraction range of the corner regions is first defined. Using the four corner points of the cropped image as centers, square regions with a side length of min(width, height) / 5 are extracted. For an 800×600 pixel cropped image, the corner region size is 120×120 pixels. The extraction process uses coordinate mapping: top-left corner region ROI_lt=image[0:120, 0:120], top-right corner region ROI_rt=image[0:120, width-120:width], bottom-left corner region ROI_lb=image[height-120:height, 0:120], bottom-right corner region ROI_rb=image[height-120:height, width-120:width]. The following processing is performed on each corner region sub-image: First, the region sub-image is binarized. If the background type parameter is black, pixels with brightness values less than a preset brightness threshold in the corner region sub-image are set to white, and pixels with brightness values greater than or equal to the preset brightness threshold are set to black, resulting in a binarized image. If the background type parameter indicates a gray background, the RGB three-channel components of each pixel in the corner region sub-image are extracted, and the minimum and maximum values of the RGB three-channel components are calculated. When the minimum value is greater than a preset background brightness threshold or the maximum value is less than a preset dark spot threshold, the corresponding pixel is set to black; otherwise, the corresponding pixel is set to white, resulting in a binarized image. A morphological opening operation is performed on the binarized image to remove noise points and isolated pixels, resulting in a filtered binary image. Contour point detection is performed on the binary image, and the contour point set contour is extracted. The Douglas-Peucker algorithm is used to approximate the contour into a polygon, with an approximation accuracy ε set to 2% of the contour perimeter. The convex hull of the approximated polygon is calculated to obtain a convex polygon hull. Analyze the characteristics of convex polygons: calculate the number of vertices n_vertices, the array of interior angles angles[], and the size of each interior angle α[i] = arccos((v[i-1]·v[i]) / (|v[i-1]|·|v[i]|)), where v[i] is the vector of the adjacent side. Determine angles based on these characteristics: when the number of vertices n_vertices = 3 or n_vertices = 4 and the ratio of the area of the vertex to the area of the smallest circumscribed triangle is > 0.85, it indicates the existence of abnormal angles, and a detection signal signal = {type: 2, result: 1, corner_id: i, angle: α[i]} is generated; otherwise, signal = {type: 2, result: 0} is generated. This detection method based on geometric features can accurately identify abnormal shapes in corner regions.
[0090] Step 107: Determine the angle detection result based on the first angle detection signal and the second angle detection signal.
[0091] The first angle detection signal is a preliminary detection result obtained based on angle histogram analysis. The second angle detection signal is a precise detection result obtained based on convex polygon fitting. The final angle detection result is a comprehensive judgment obtained by combining the two detection signals, including the determination of the existence of the angle and specific angle feature information.
[0092] Specifically, the information fields of the first corner detection signal S1 and the second corner detection signal S2 are extracted first. S1 includes: detection type (type=1), detection result (result) (0 or 1), and confidence. S2 includes: detection type (type=2), detection result (result) (0 or 1), corner number (corner_id), and angle value (angle). A weighted fusion strategy is used to determine the final result, with the weights allocated as follows: w1=0.4 (weight of the first detection signal), w2=0.6 (weight of the second detection signal). The weighted score is calculated as score=w1×S1.result+w2×S2.result. A decision threshold of threshold=0.5 is set; when the score>threshold, a corner is determined to exist. A complete detection result data structure is constructed: result = {detection_result: (score>threshold?1: 0), confidence_score: score, first_detection: {result: S1.result, confidence: S1.confidence}, second_detection: {result: S2.result, corner_id: S2.corner_id, angle: S2.angle}, timestamp: current_time}. When a corner is detected (detection_result=1), additional detailed information about the corner is recorded: the location of the corner point, the corner angle, and the detection confidence level. This multi-level detection-based fusion judgment method fully utilizes the complementarity of different features, improving the accuracy and reliability of detection.
[0093] In the above embodiments, a basic corner detection framework was implemented by fitting boundary point angle histograms and convex polygons. To further improve the accuracy of corner recognition and reduce the influence of background type on image segmentation, this application also provides an enhanced image processing method. This method analyzes the binarized features of the corner region, studies the shape rules of connected contours, and performs background adaptive correction, enabling the system to more accurately handle corner detection requirements in complex scanning environments. The following section combines... Figure 2Another method for detecting paper fold angles in an image using a scanner, as described in this application, is as follows: Please see Figure 2 This is a flowchart illustrating a method for detecting paper folds in an image using a scanner, as described in this application.
[0094] Step 201: Perform binarization and morphological filtering on the diagonal sub-image to obtain a filtered binary image.
[0095] Corner sub-images are local image regions extracted from the four corners of a cropped image. Binarization is the process of converting a grayscale image into one containing only black and white pixel values. Morphological filtering is the operation of enhancing or suppressing local features of an image using mathematical morphological operators. A filtered binary image is a black and white image with sharp boundaries obtained after binarization and morphological processing.
[0096] Specifically, firstly, binarization is performed on the diagonal sub-image. The Otsu adaptive thresholding method is used to determine the binarization threshold T: calculate the gray-level histogram H
[256] of the image, traverse all possible thresholds (0-255), calculate the inter-class variance σ_b^2(t) for each threshold t, and select the value that maximizes the inter-class variance as the final threshold T. Binarize each pixel p(x,y) in the image: when p(x,y)>T, the output value is 255 (white); when p(x,y)≤T, the output value is 0 (black). The preliminary binary image binary_image is obtained. Then, morphological processing is performed: first, a 3×3 structuring element kernel= [[1,1,1], [1,1,1], [1,1,1]]. The opening operation is performed: erosion followed by dilation. Erosion operation: For each pixel in the binary image, align the center of the structuring element with that pixel. If all pixels within the area covered by the structuring element have a value of 255, output 255; otherwise, output 0. Dilation operation: Similar to erosion, but output 255 as long as there is a pixel with a value of 255 within the area covered by the structuring element. Opening followed by closing operation: Dilation is performed first, then erosion, using the same structuring element. The final filtered binary image (filtered_binary_image) has the following characteristics: small-area noise is removed, boundary contours are smoothed, but the main shape features are preserved. This combination of binarization and morphological processing can effectively extract key shape features in corner regions, providing clear image data for subsequent corner detection.
[0097] In one possible implementation, the diagonal region sub-image is binarized and morphologically filtered to obtain a filtered binary image, specifically including steps 2011-2013, as follows: Step 2011: Acquire a background type parameter of a scanner; if the background type parameter indicates a black background, set pixels with a brightness value less than a preset brightness threshold in the corner region sub-image as a white missing corner region, and set pixels with a brightness value greater than or equal to the preset brightness threshold as black paper, thereby obtaining a binarized image.
[0098] The background type parameter is a setting value of the currently used background plate type of the scanner. The corner region sub-image is a partial image region extracted from four corners of a cropped image. The brightness value refers to the gray intensity value of an image pixel, which ranges from 0 to 255. The preset brightness threshold is a brightness critical value for distinguishing a foreground from a background. A binarized image is an image that only contains two pixel values: black (0) and white (255).
[0099] Specifically, the background type parameter background_type is first read through a device interface of the scanner. A standard communication protocol is adopted for parameter reading: a query command [0xAA, 0x02, 0x01, 0xAD] is sent, and return data is received. When background_type = 0, it indicates a black background. For the black background condition, reverse binarization processing is performed. The preset brightness threshold T is calculated first: a grayscale histogram H
[256] is calculated for the corner region sub-image, and a double-peak method is used to determine the threshold. The specific steps are: find the positions of two main peaks p1 and p2 in the histogram, and take the midpoint thereof as an initial threshold T0 = (p1 + p2) / 2. Iterative optimization: calculate the average pixel value μ1 of pixels with a value less than T0 and the average pixel value μ2 of pixels with a value greater than T0, update the threshold T = (μ1 + μ2) / 2, and repeat this process until the change of T is less than 1. For example, for a typical scanned image with a black background, the first peak p1 is usually in the range of 30-50, and the second peak p2 is usually in the range of 180-200, and the finally obtained threshold T is about 110-130. Perform binarization conversion: traverse each pixel point p(x, y) in the corner region sub-image, when p(x, y)<T, set the output value to 255 (white missing corner region); when p(x, y) ≥ T, set the output value to 0 (black paper). In the converted binarized image binary_image, the document area is displayed as black paper and the background area is displayed as white missing corner area. This adaptive binarization method based on background type can effectively solve the image segmentation problem under black background, and provide clear boundary information for subsequent corner folding detection.
[0100] Step 2012: if the background type parameter indicates a gray background, extract RGB three-channel components of each pixel in the corner region sub-image, calculate the minimum value and the maximum value among the RGB three-channel components, when the minimum value is greater than a preset background brightness threshold or the maximum value is less than a preset dark point threshold, set the corresponding pixel as black paper, otherwise set the corresponding pixel as a white missing corner region, thereby obtaining a binarized image.
[0101] The three RGB channel components refer to the red, green and blue color component values of an image pixel, and each component ranges from 0 to 255. The minimum value and the maximum value refer to the minimum and maximum values among the three RGB components of a single pixel. The preset background brightness threshold is a standard value for judging whether a pixel belongs to a light-colored background. The preset dark spot threshold is a standard value for judging whether a pixel belongs to a dark foreground. A binarized image is a black-and-white binary image obtained after converting a gray background image.
[0102] Specifically, first confirm that the background type is gray background (background_type = 1). For each pixel p(x, y) in the diagonal region sub-image, extract its RGB components: R = p(x, y, 0), G = p(x, y, 1), B = p(x, y, 2). Calculate the minimum value min_rgb = min(R, G, B) and the maximum value max_rgb = max(R, G, B) of the RGB components. Set the preset background brightness threshold background_threshold = 140, and the preset dark spot threshold dark_threshold = 50. Perform binarization judgment: when min_rgb > background_threshold, it indicates that all components of the pixel are relatively bright, the pixel is determined as a paper area, and the output value is set to 0 (black paper); when max_rgb < dark_threshold, it indicates that all components of the pixel are relatively dark, the pixel is also determined as a shadow area at the edge of the paper, and the output value is set to 0 (black paper); in other cases, the output value is set to 255 (white corner missing area). For example, for a background pixel with an RGB value of (190, 185, 188), its min_rgb = 185 > 140, so it is set as black paper; for a document edge pixel with an RGB value of (45, 48, 42), its max_rgb = 48 < 50, so it is also set as black paper. Through this binarization method based on RGB component analysis, the image segmentation problem under a gray background can be effectively handled, and it is particularly suitable for handling the gray gradient phenomenon at document edges and corner folding areas. In the finally obtained binarized image binary_image, the document and the background area are clearly separated, which provides a good foundation for subsequent morphological processing.
[0103] Step 2013: perform a morphological opening operation on the binarized image to remove noise points and isolated pixels in the binarized image, so as to obtain a filtered binarized image.
[0104] A binarized image is an image containing only black and white pixel values, obtained from previous steps. Morphological opening is a composite operation that first performs erosion and then dilation. Noise points are isolated pixels in the image that are not part of the main object but appear as foreground. Isolated pixels are single or a small number of connected pixels whose values differ from their surrounding pixels. A filtered binary image is a binary image with sharp boundaries obtained after morphological processing.
[0105] Specifically, first, define the structuring element `kernel` for morphological operations. Use a 3×3 square structuring element: `kernel = [[1, 1, 1], [1, 1, 1], [1, 1, 1]]`. Perform the erosion operation `erode`: slide the structuring element `kernel` on the binary image, and for each pixel position (x, y), check the 3×3 neighborhood centered on that pixel. If all pixel values in the neighborhood are 255 (white missing corner area), the output pixel value is set to 255; otherwise, it is set to 0 (black paper). The mathematical expression for the erosion operation is: `eroded(x, y) = min{binary(x+i, y+j): (i, j)∈kernel}`. For example, if the center pixel value is 255, but there is a pixel with a value of 0 in the 3×3 neighborhood, the center pixel value becomes 0 after erosion. After the erosion operation, perform the dilation operation: again, slide the 3×3 structuring element on the eroded image. If there is a pixel with a value of 255 in the neighborhood, the output pixel value is set to 255; otherwise, it is set to 0. The mathematical expression for the dilation operation is: dilated(x, y) = max{eroded(x+i, y+j): (i, j)∈kernel}. The entire opening operation can be represented as: filtered = dilate(erode(binary_image, kernel) , kernel). This morphological processing can effectively remove the following types of noise: isolated white points with an area smaller than the structuring element, small connecting bridges, and edge burrs. For example, when there are 2×2 isolated white points in the binary image, the erosion operation will completely eliminate these points; while the boundary of the main target region basically retains its original shape after erosion and dilation. The final filtered binary image filtered_binary_image has clear boundaries and less noise interference, providing a high-quality data foundation for subsequent contour extraction.
[0106] Step 202: Extract multiple connected contours from the filtered binary image and calculate the area of each connected contour; select connected contours with areas within a preset area range as valid contours, and perform convex polygon fitting on each valid contour.
[0107] A connected contour is the boundary of a set of adjacent pixels with the same pixel value in a binary image. The area of a connected contour is the number of pixels within the region enclosed by the contour. A preset area range is an area threshold interval used to filter valid contours. A valid contour is a connected contour whose area meets the preset range requirement. Convex polygon fitting is the process of approximating a set of contour points as a convex polygon with the fewest vertices.
[0108] Specifically, a boundary tracking algorithm is first used to extract connected contours from the filtered binary image. An 8-connectedness tracking method is employed: scanning begins from the top left corner of the image. When a pixel with a value of 255 is encountered, the search continues clockwise, starting from that point, to find adjacent boundary pixels. For the current boundary point (x, y), its eight neighboring pixels are checked, and the coordinates of the next boundary point are recorded, until the process returns to the starting point, completing the extraction of a closed contour. The extracted contour point sequence is stored in the array `contours`, where each contour is represented as an ordered set of points {(x1, y1), (x2, y2), ..., (xn, yn)}. The area of each contour is calculated using Green's formula: area = 0.5 × |Σ(xi × yi + 1 - xi + 1 × yi)|, where i ranges from 1 to n-1. A preset area range is set: the lower limit is 5% of the corner area (min_area = 0.05 × ROI_area), and the upper limit is 30% of the corner area (max_area = 0.3 × ROI_area). For a 120×120 pixel corner region, the area ranges from [720, 4320] pixels. Contours with areas within this range are selected as valid contours. For each valid contour, convex polygon fitting is performed: first, the convex hull of the contour is calculated using the Graham scan algorithm. The fitting precision is set to ε = 0.02 × contour perimeter. The Douglas-Peucker algorithm is then applied for polygon approximation: straight lines are constructed from the first and last points, and the maximum distance d_max from other points to the lines is calculated. If d_max < ε, the contour is directly represented by a straight line segment; otherwise, the contour is divided at the point corresponding to d_max, and the two sub-intervals are processed recursively. The resulting convex polygon represents the main shape features of the contour with the fewest vertices. This method, based on area selection and shape fitting, effectively extracts key contour features from corner regions.
[0109] Step 203: If the fitting result is a preset shape, then the corner region sub-image is determined to have a folded corner; if the fitting result is a target shape other than the preset shape, and the ratio of the area of the target shape to the area of the smallest circumscribed triangle is greater than a preset threshold, then the corner region sub-image is determined to have a folded corner.
[0110] The preset shape refers to the standard geometric shape under typical angle conditions, usually a triangle, quadrilateral, or pentagon. The target shape is the actual contour shape obtained by fitting a convex polygon. The minimum circumscribed triangle is the triangle with the smallest area that completely contains the target shape. The area ratio is the ratio of the area of the target shape to the area of its minimum circumscribed triangle. The preset threshold is the standard value used to determine the area ratio.
[0111] Specifically, the geometric characteristics of the fitted convex polygon are analyzed first. The vertex sequence of the polygon, vertices = {v1, v2, ..., vn}, and the number of sides, n, are obtained. When n = 5 or n = 6, it is identified as a preset shape. For the preset shape, its interior angle sequence is calculated: α[i] = arccos((vi-1 - vi)·(vi+1 - vi) / (|vi-1 - vi|·|vi+1 - vi|)). The interior angle sequence is checked to see if it satisfies the angle-folding characteristic: there exists an interior angle α[k] < 75°, and its adjacent interior angles 75° < α[k±1] < 135°. If the condition is met, an angle-folding feature is directly determined. When n ≠ 5 and n ≠ 6, area ratio analysis is required. First, the area A1 of the target shape is calculated: using the polygon area formula A1 = 0.5 × |Σ(xi × yi + 1 - xi + 1 × yi)|. Then, the minimum circumscribed triangle is calculated: using the rotating caliper algorithm, the farthest point pair (p1, p2) of the polygon is found. Other vertices are traversed to find the point p3 farthest from line segment p1p2. These three points form the minimum circumscribed triangle. The area A2 of the triangle is calculated: A2 = 0.5 × |x1(y2 - y3) + x2(y3 - y1) + x3(y1 - y2)|. The area ratio r = A1 / A2 is calculated. A preset threshold T = 0.85 is set; when r > T, a folded angle is determined to exist. This dual-judgment method based on shape features and area ratio can adapt to different types of folded angles, improving the accuracy and robustness of detection. The final judgment result includes the following information: whether a folded angle exists (is_folded), the folded angle type (fold_type: "preset" or "other"), and key parameters (vertex_count, angles, area_ratio).
[0112] Step 204: Generate a second corner detection signal based on the corner determination result of the corner region sub-image.
[0113] The corner determination result of the corner region sub-image is the corner existence judgment obtained from the shape analysis in the previous step. The second corner detection signal is the output signal that converts the corner determination result into a standard format. The generation process is the operation of encapsulating the determination result and its related feature parameters into structured data.
[0114] First, a standard signal data structure is constructed, containing the following fields: signal type (type=2, indicating second-level detection), detection result (result, 0 or 1), timestamp, corner region identifier (corner_id, value range 0-3, representing the top left, top right, bottom left, and bottom right corners respectively), detection confidence, and shape feature parameters (shape_features). For the detection result field, the corner determination result is directly used: result=1 when a corner is determined to exist, otherwise result=0. The timestamp uses the current system time, accurate to milliseconds: timestamp = current_time_ms. The detection confidence is calculated based on the shape analysis result: when the corner type is a preset shape, confidence = min(1.0, |90° - min_angle| / 15°), where min_angle is the minimum interior angle value; when the corner type is other shapes, confidence = min(1.0, (area_ratio - 0.75) / 0.25), where area_ratio is the area ratio. The shape feature parameters record detailed geometric information: vertex_count (number of vertices), angles (sequence of interior angles), area_ratio (area ratio), min_angle (minimum interior angle), and fold_type (fold type). This information is organized into a JSON-formatted data packet: signal = {type: 2, result: result, timestamp: timestamp, corner_id: corner_id, confidence: confidence, shape_features: {vertex_count: n, angles: [α1, α2, ..., αn], area_ratio: r, min_angle: α_min, fold_type: type}}. This standardized signal format contains both basic detection results and detailed feature information, facilitating subsequent result analysis and processing. This signal is passed to the upper-level processing module through the system interface for the final fold determination decision.
[0115] The following describes a scanner device image paper folding corner detection system according to an embodiment of the present invention from the perspective of hardware processing. Please refer to [link / reference]. Figure 3 This is a schematic diagram of the structure of a scanner device image paper folding corner detection system in an embodiment of this application.
[0116] It should be noted that, Figure 3The structure of the image paper fold detection system shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.
[0117] like Figure 3 As shown, a scanner device image paper fold detection system includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage section 308 into a random access memory (RAM) 303, such as executing the method described in the above embodiments. The RAM 303 also stores various programs and data required for system operation. The CPU 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0118] The following components are connected to I / O interface 305: input section 306 including audio input devices, push-button switches, etc.; output section 307 including a liquid crystal display (LCD) and audio output devices, indicator lights, etc.; storage section 308 including a hard disk, etc.; and communication section 309 including a network interface card such as a LAN (Local Area Network) card, modem, etc. Communication section 309 performs communication processing via a network such as the Internet. Drive 310 is also connected to I / O interface 305 as needed. Removable media 311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 310 as needed so that computer programs read from them can be installed into storage section 308 as needed.
[0119] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing computer programs for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 309, and / or installed from removable medium 311. When the computer program is executed by central processing unit (CPU) 301, it performs the various functions defined in the present invention.
[0120] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0121] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those shown in the drawings.
[0122] Specifically, the scanner device image paper fold corner detection system of this embodiment includes a processor and a memory. The memory stores a computer program. When the computer program is executed by the processor, it implements the scanner device image paper fold corner detection method provided in the above embodiment.
[0123] In another aspect, the present invention also provides a computer-readable storage medium, which may be included in the scanner device image paper fold corner detection system described in the above embodiments; or it may exist independently and not assembled into the scanner device image paper fold corner detection system. The storage medium carries one or more computer programs, which, when executed by a processor of the scanner device image paper fold corner detection system, cause the scanner device image paper fold corner detection system to implement the scanner device image paper fold corner detection method based on IoT data encryption transmission provided in the above embodiments.
Claims
1. A method of detecting a corner of a paper sheet in a scanner apparatus, characterized by, The method includes: The uncropped original image captured by the scanner is acquired, and the uncropped original image is converted into a grayscale image and edge detection is performed to obtain an edge-detected image; The edge detection image is scanned bidirectionally to extract boundary points and generate a boundary point set, which includes a left boundary point set and a right boundary point set. The boundary point set is divided into regions, the angular region boundary point subsets of the boundary point set are extracted, and the angular region boundary point subsets are constructed. Filter the dominant angle sequence in the angle histogram, and output the first angle detection signal based on the number of angles and the angle relationship of the dominant angle sequence; If the first fold detection signal does not detect a fold, then straight line detection is performed on the edge detection image to determine grouped straight lines. A reference baseline is selected to perform rotation correction on the uncropped original image to obtain a straightened image. The paper boundary coordinates are determined according to the position of the grouped straight lines in the straightened image. Based on the paper boundary coordinates, the straightened image is subjected to preset cropping processing to obtain a cropped image. Extract corner region sub-images from the cropped image, perform convex polygon fitting on the corner region sub-images, and output a second corner detection signal; The angle detection result is determined based on the first angle detection signal and the second angle detection signal.
2. The method of claim 1, wherein, The step of bidirectional scanning of the edge detection image to extract boundary points and generate a boundary point set includes: The edge detection image is scanned line by line. For each line, pixels whose gradient values exceed a preset gradient threshold are searched from the left boundary to the center as candidate points for the left boundary, and pixels whose gradient values exceed the preset gradient threshold are searched from the right boundary to the center as candidate points for the right boundary. Calculate the horizontal distance between the left boundary candidate point and the right boundary candidate point in each row; If the horizontal distance is within the preset distance range, the left boundary candidate point corresponding to the currently traversed row is stored in the left boundary point set, and the right boundary candidate point is stored in the right boundary point set. If the horizontal distance is outside the preset distance range, the currently traversed row is marked as an invalid row, and the next row is traversed until all rows in the edge detection image have been scanned. The set of left boundary points and the set of right boundary points are combined to form the set of boundary points.
3. The method of claim 1, wherein, The construction of the angle histogram of the subset of boundary points of the angular region includes: Multiple boundary points are selected sequentially from the set of boundary points in the corner region according to a preset step size, and a boundary vector is constructed based on the adjacent boundary points. Calculate the angle between adjacent boundary vectors, and take the boundary vectors whose angle is less than a preset angle threshold as valid boundary vectors; The effective boundary vector is normalized to obtain the target angle mean. An angle histogram is generated based on the mean of the target angle and the corresponding angle interval.
4. The method of claim 1, wherein, The step of filtering the dominant angle sequence in the angle histogram and outputting a first fold angle detection signal based on the number of angles and the angle relationship of the dominant angle sequence includes: Obtain the frequency of each angle interval in the angle histogram; The angle intervals with frequencies lower than a preset frequency threshold are filtered out, and the angles corresponding to the preset number of angle intervals with the highest frequencies are taken as the dominant angle sequence. If the dominant angle sequence contains only one angle, it is determined that no angle was detected; If there is an angle in the dominant angle sequence whose deviation from the right angle is less than a preset deviation threshold, and adjacent angles in the dominant angle sequence are perpendicular to each other, then it is determined that no angle has been detected. If the dominant angle sequence does not contain only one angle, or if the dominant angle sequence contains an angle whose deviation from the right angle is less than a preset deviation threshold, and if adjacent angles in the dominant angle sequence are perpendicular, then a fold angle is detected. Based on the angle determination result, a first angle detection signal is generated, and a stop scanning command is returned to the scanner control terminal.
5. The method of claim 1, wherein, Determining the paper boundary coordinates based on the position of the grouped straight lines in the alignment image includes: From the grouped straight lines, straight lines located in the left region of the straightened image are selected as the left boundary line candidate set, and straight lines located in the right region of the straightened image are selected as the right boundary line candidate set; The line with the smallest x-coordinate in the candidate set of left boundary lines is selected as the left boundary line, and the line with the largest x-coordinate in the candidate set of right boundary lines is selected as the right boundary line. The coordinates of the intersection point of the left boundary line and the upper boundary of the straightened image are calculated as the coordinates of the upper left corner, and the coordinates of the intersection point of the left boundary line and the lower boundary of the straightened image are calculated as the coordinates of the lower left corner. The coordinates of the intersection point of the right boundary line and the upper boundary of the straightened image are calculated as the upper right corner coordinates, and the coordinates of the intersection point of the right boundary line and the lower boundary of the straightened image are calculated as the lower right corner coordinates. The coordinates of the top left corner, the bottom left corner, the top right corner, and the bottom right corner are combined to form the paper boundary coordinates.
6. The method of claim 1, wherein, The step of performing convex polygon fitting on the corner region sub-image and outputting a second corner detection signal includes: The corner region sub-image is binarized and morphologically filtered to obtain a filtered binary image; Multiple connected contours are extracted from the filtered binary image, and the area of each connected contour is calculated. Filter connected contours whose area is within a preset area range as valid contours, and perform convex polygon fitting on each valid contour; If the fitting result is a preset shape, then it is determined that the corner region sub-image has a folded corner; If the fitting result is a target shape other than the preset shape, and the ratio of the area of the target shape to the area of the smallest circumscribed triangle is greater than the preset threshold, then the corner region sub-image is determined to have a folded corner. A second corner detection signal is generated based on the corner determination result of the corner region sub-image.
7. The method of claim 6, wherein, The step of performing binarization and morphological filtering on the corner region sub-image to obtain a filtered binary image includes: Obtain the background type parameter of the scanner; If the background type parameter is black background, then the pixels in the corner region sub-image with brightness values less than the preset brightness threshold are set as white missing corner regions, and the pixels with brightness values greater than or equal to the preset brightness threshold are set as black paper, thus obtaining a binarized image. If the background type parameter indicates a gray background, the RGB three-channel components of each pixel in the corner region sub-image are extracted, and the minimum and maximum values of the RGB three-channel components are calculated. When the minimum value is greater than a preset background brightness threshold or the maximum value is less than a preset dark spot threshold, the corresponding pixel is set to black paper; otherwise, the corresponding pixel is set to white missing corner region to obtain a binarized image. A morphological opening operation is performed on the binarized image to remove noise points and isolated pixels, resulting in the filtered binary image.
8. A scanner apparatus image paper corner angle detection system characterized by comprising: The scanner device image paper fold detection system includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the scanner device image paper fold detection system to perform the method as described in any one of claims 1-7.
9. A computer-readable storage medium comprising instructions, characterized in that, When the instruction is executed on the scanner device image paper fold detection system, the scanner device image paper fold detection system performs the method as described in any one of claims 1-7.
10. A computer program product, characterised in that, When the computer program product is run on the scanner device image paper folding detection system, the scanner device image paper folding detection system performs the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Folded banknote identification method
CN108320372A
Method and device for automatic tilt correction of document image of scanning equipment
CN118823789A