Automatic scanning device and method for digitizing documents based on ai vision and pneumatic assistance
By employing an AI-assisted vision and pneumatic-assisted automatic document digitization scanning method, which combines image processing and deep learning models to identify page gaps and generate page-turning trajectory instructions, and utilizing pneumatic technology for non-contact page turning and static electricity neutralization, this method solves the problems of low efficiency and high risk of damage in existing document digitization technologies, achieving safe and efficient document digitization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 广东蓝莺高科有限公司
- Filing Date
- 2025-09-16
- Publication Date
- 2026-07-21
AI Technical Summary
Existing methods for digitizing documents are inefficient, labor-intensive, and prone to causing physical damage to documents. They are also poorly suited for documents that are thin, brittle, adhered, or have sensitive surfaces.
An automatic document digitization scanning method based on AI vision and pneumatic assistance is adopted. By acquiring dynamic image data of the book's flip edges, combining image processing algorithms and deep learning models to identify the gaps between target pages, generating page-turning trajectory instructions, and using pneumatic technology for non-contact page turning and static charge neutralization, the book pages are lifted and scanned.
It enables safe, efficient, and high-quality digitization of documents, reduces the risk of physical damage to fragile pages, and improves the efficiency and image quality of automated scanning.
Smart Images

Figure CN121442047B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of document digitization and automated processing technology, and in particular to an automatic document digitization scanning device and method based on AI vision and pneumatic assistance. Background Technology
[0002] Traditional document digitization methods mainly rely on manual operation or semi-automatic scanning equipment, which suffers from problems such as low efficiency, high labor intensity, easy physical damage to documents (such as paper tearing and ink abrasion), and poor adaptability to old documents. Existing automated scanning equipment mostly uses mechanical suction cups or grippers to directly contact the page for page turning, which can easily cause document damage and has poor adaptability to thin, brittle, stuck, or surface-sensitive documents, resulting in unsatisfactory processing effects. Summary of the Invention
[0003] In view of this, the main objective of the embodiments of the present invention is to provide an automatic document digitization scanning device and method based on AI vision and pneumatic assistance, in order to solve at least one of the problems of the prior art and achieve safe, efficient and high-quality digitization of documents.
[0004] To achieve the above objectives, one aspect of the present invention provides an automatic document digitization scanning method based on AI vision and pneumatic assistance, the method comprising:
[0005] Acquire dynamic image data of the book's flip-page area;
[0006] Based on image processing algorithms and deep learning models, the dynamic image data is identified to obtain the target page gap;
[0007] Based on the target page gaps, generate page-turning motion trajectory instructions;
[0008] The static charge between the pages of the book is neutralized, and the current target page and the next target page are pre-separated.
[0009] According to the page-turning trajectory instruction, the guide line is inserted into the gap between the current target page and the next target page, and the current target page is lifted.
[0010] Scan the previously mentioned target page and the current target page to obtain a first page image of the previously mentioned target page and a second page image of the current target page;
[0011] Turn the current target page, set the current target page as the previous target page, and return to the step of neutralizing the static charge between the pages of the book until the book scanning is complete.
[0012] In some embodiments, the step of identifying the dynamic image data based on image processing algorithms and deep learning models to obtain the target page gap includes the following steps:
[0013] Based on image processing algorithms, the first page gap and the first confidence level are obtained;
[0014] Based on a deep learning model, the second page gap and the second confidence level are obtained;
[0015] The first confidence level and the second confidence level are compared to obtain the comparison results;
[0016] Obtain the geometric consistency verification result between the first page gap and the second page gap;
[0017] Based on the comparison results and the geometric consistency verification results, the target page gap is obtained.
[0018] In some embodiments, obtaining the first page gap and the first confidence level based on the image processing algorithm includes the following steps:
[0019] The dynamic image data is subjected to noise reduction and edge contrast enhancement operations to obtain a first intermediate image;
[0020] Edge detection is performed on the first intermediate image to obtain the page edge features;
[0021] Based on prior knowledge of the book's flip-edge, a flip-edge curve is fitted according to the page edge features;
[0022] A strip-shaped search area is established inside the curve of the flip-end;
[0023] The vertical projection calculation is performed on the strip search area, and the local minimum point of the projection curve is obtained as the first page gap;
[0024] Obtain the first confidence level of the first page gap.
[0025] In some embodiments, obtaining the second page gap and the second confidence level based on a deep learning model includes the following steps:
[0026] The dynamic image data is subjected to noise reduction and edge contrast enhancement operations to obtain a second intermediate image;
[0027] The second intermediate image is input into a pre-trained semantic segmentation model, which outputs a segmentation mask; each pixel in the segmentation mask is classified into a background class, a page class, or a gap class.
[0028] A morphological closing operation is performed on the target region corresponding to the pixel of the gap class to obtain the gap region;
[0029] The second page gap is extracted from the gap region using a linear fitting algorithm;
[0030] Obtain the second confidence level of the second page gap.
[0031] In some embodiments, obtaining the target page gap based on the comparison result and the geometric consistency verification result includes the following steps:
[0032] When the second confidence level is greater than the first preset threshold, the second page gap is taken as the target page gap;
[0033] When the second confidence level is less than or equal to the first preset threshold, the first confidence level is greater than the second preset threshold, and the geometric difference between the first page gap and the second page gap is less than the first error range, the first page gap is taken as the target page gap.
[0034] When the second confidence level is less than or equal to the first preset threshold, and the first confidence level is less than or equal to the second preset threshold, return to the step of obtaining dynamic image data of the book's flip-page area.
[0035] In some embodiments, generating the page-turning motion trajectory instruction based on the target page gap includes the following steps:
[0036] Obtain the first point cloud data of the three-dimensional spatial position of the gap between the target pages;
[0037] The outliers in the first point cloud data are removed to obtain the second point cloud data;
[0038] The B-spline curve fitting algorithm is used to fit the second point cloud data to generate a trajectory curve in three-dimensional space.
[0039] Based on preset kinematic parameters, S-shaped velocity planning is performed on the trajectory curve to generate a position-time series;
[0040] The position-time sequence is converted into the page-turning motion trajectory instruction through inverse kinematics operations.
[0041] In some embodiments, neutralizing the static charge between pages of the book and pre-separating the current target page from the next target page includes the following steps:
[0042] The pneumatic actuator is controlled to perform ion air jetting to neutralize the static charge between the pages of the book;
[0043] The pneumatic actuator is controlled to perform micro-airflow injection to pre-separate the current target page from the next target page.
[0044] To achieve the above objectives, another aspect of the present invention proposes an automatic document digitization scanning device based on AI vision and pneumatic assistance, the device comprising:
[0045] The first AI visual servo module is used to acquire dynamic image data of the book's flip-page area;
[0046] The second AI visual servo module is used to identify the dynamic image data based on image processing algorithms and deep learning models to obtain the target page gap;
[0047] The third AI visual servo module is used to generate page-turning motion trajectory instructions based on the gaps in the target page.
[0048] The pneumatic actuator module is used to neutralize the static charge between the pages of the book and to pre-separate the current target page from the next target page.
[0049] The first mechanical execution module is used to insert the guide line into the gap between the current target page and the next target page according to the page turning motion trajectory instruction, and to lift the current target page;
[0050] The image acquisition module is used to scan the previously mentioned target page and the current target page to obtain a first page image of the previously mentioned target page and a second page image of the current target page;
[0051] The second mechanical execution module is used to turn the current target page, use the current target page as the previous target page, and return to the step of neutralizing the static charge between the pages of the book until the book scanning is completed.
[0052] To achieve the above objectives, another aspect of the present invention provides an electronic device, the electronic device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method described above.
[0053] To achieve the above objectives, another aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods described above.
[0054] To achieve the above objectives, another aspect of the present invention provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions to cause the computer device to perform the aforementioned method.
[0055] The embodiments of the present invention include at least the following beneficial effects: The present invention provides an automatic document digitization scanning device and method based on AI vision and pneumatic assistance. This scheme provides real-time image data by acquiring dynamic image data of the book's page-turning area; based on image processing algorithms and deep learning models, the dynamic image data is identified to obtain the target page gap, laying a foundation for subsequent visual recognition and automated page turning; according to the target page gap, page-turning motion trajectory instructions are generated, and the visual recognition results are converted into an executable path for precision machinery, ensuring smooth, accurate, and efficient page-turning actions; the static charge between the book's pages is neutralized, and the current target page is pre-separated from the next target page. Through non-contact pneumatic technology, static charge is first eliminated to prevent page adhesion, and then initial separation is gently created using micro-airflow, greatly reducing the risk of physical damage to fragile pages. Based on the page-turning trajectory instructions, the guide line is inserted into the gap between the current and next target pages, and the current target page is lifted. This lifting action is achieved with minimal contact area and mechanical stress, combined with pneumatic assistance, enabling gentle handling of valuable documents. The system scans both the previous and current target pages, simultaneously acquiring high-resolution images of both pages to obtain the first page image of the previous target page and the second page image of the current target page, ensuring the quality and efficiency of the digitized images. The current target page is then turned, becoming the previous target page, and the system returns to the step of neutralizing the static charge between the pages until the book scanning is complete. This forms a fully automated closed-loop control cycle, continuously and reliably completing the digitization process of the entire book without manual intervention, achieving safe, efficient, and high-quality digitization of documents. Attached Figure Description
[0056] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0057] Figure 1 This is a flowchart of an automatic document digitization scanning method based on AI vision and pneumatic assistance provided in an embodiment of the present invention;
[0058] Figure 2 This is a schematic diagram of the hybrid decision-making process provided in an embodiment of the present invention;
[0059] Figure 3 This is a schematic diagram of the operation process of the gas-electric integrated nozzle provided in an embodiment of the present invention;
[0060] Figure 4 This is a schematic diagram of the image acquisition system provided in an embodiment of the present invention;
[0061] Figure 5 This is a schematic diagram of the mobile terminal interactive interface of the document digitization automatic scanning system provided in the embodiment of the present invention;
[0062] Figure 6 This is a schematic diagram of one view of the monitoring area in the mobile terminal interactive interface of the document digitization automatic scanning system provided in this embodiment of the invention;
[0063] Figure 7 This is a schematic diagram of another view of the monitoring area in the mobile terminal interactive interface of the document digitization automatic scanning system provided in this embodiment of the invention;
[0064] Figure 8 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present invention.
[0065] Figure label:
[0066] Horizontal stage 1; horizontal page camera 2; vertical page camera 3. Detailed Implementation
[0067] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this invention; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this invention as detailed in the appended claims.
[0068] It should be noted that although functional modules are divided in the system diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the system or the order in the flowchart. The terms "first / S100" and "second / S200" in the specification, claims, and the foregoing drawings may be used herein to describe various concepts, but unless specifically stated otherwise, these concepts are not limited by these terms. These terms are used only to distinguish one concept from another. For example, first information may also be referred to as second information without departing from the scope of the embodiments of the invention, and similarly, second information may also be referred to as first information. Depending on the context, the words "if" or "when" as used herein may be interpreted as "when," "in response to a determination," or "in the event of a determination."
[0069] The terms “at least one,” “multiple,” “each,” “any,” etc., used in this invention, “at least one” includes one, two, or more than two; “multiple” includes two or more than two; “each” refers to each of the corresponding multiple; and “any” refers to any one of the multiple.
[0070] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein is for the purpose of describing embodiments of the invention only and is not intended to limit the invention.
[0071] Before providing a detailed description of the embodiments of the present invention, some of the nouns and terms involved in the embodiments of the present invention will be explained first. The nouns and terms involved in the embodiments of the present invention are subject to the following interpretations.
[0072] The folding edge, also called the outer edge, refers to the outer end opposite the spine of the book. It is the side that can be opened and closed freely and is also the outermost edge of the folding area, the outermost edge of the book.
[0073] The turning area refers to the entire movable page area from the spine to the outer edge of the book. This is the core moving area of the page; when the book is turned, the pages in this area will bend, arch, and move accordingly.
[0074] Page gaps refer to the groove formed along the spine between the left and right pages when the book is closed, or the "V"-shaped space formed at the base of the pages when the book is open.
[0075] Existing book digitization equipment mostly relies on manual operation of various types of scanners or robotic page-turning devices with limited automation. The former is inefficient and has high labor costs; the latter often uses contact-based page-turning methods such as suction cups, friction wheels, or mechanical fingers, where mechanical parts directly contact and apply force to the paper, easily causing scratches, tears, or permanent creases to fragile or aged book pages. Furthermore, contact-based page-turning methods cannot effectively handle books of different thicknesses, materials, sizes, and storage conditions (such as sticking or curling), resulting in weak generalization ability.
[0076] In view of this, embodiments of the present invention provide an automatic document digitization scanning method based on AI vision and pneumatic assistance. By combining image processing algorithms and deep learning models, the method accurately identifies the target page gaps in the book's flip-page area and generates page-turning trajectory instructions accordingly. Then, after neutralizing the static charge between the pages and performing pre-separation using an integrated pneumatic-electric nozzle, the mechanical execution system guides the guide line into the gap and lifts the pages according to the instructions. After simultaneously scanning the double-page image, the page turning is completed. This process is repeated until the entire book is scanned, thereby achieving safe, efficient, and high-quality digitization of documents.
[0077] Figure 1 This is an optional flowchart of an automatic document digitization scanning method based on AI vision and pneumatic assistance provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S100 to S700:
[0078] Step S100: Obtain dynamic image data of the book's flip-page area;
[0079] Step S200: Based on image processing algorithms and deep learning models, identify dynamic image data and obtain the target page gap;
[0080] Step S300: Generate page-turning motion trajectory instructions based on the target page gaps;
[0081] Step S400: Neutralize the static charge between the pages of the book and pre-separate the current target page from the next target page;
[0082] Step S500: According to the page turning trajectory instruction, insert the guide line into the gap between the current target page and the next target page, and lift the current target page;
[0083] Step S600: Scan the previous target page and the current target page to obtain the first page image of the previous target page and the second page image of the current target page;
[0084] Step S700: Turn the current target page, set the current target page as the previous target page, and return to the step of neutralizing the static charge between the pages of the book until the book scanning is completed.
[0085] In steps S100 to S700 of some embodiments, a hybrid decision-making process combining image processing algorithms and deep learning models is used to identify page gaps, improving accuracy and robustness under different conditions. The use of pneumatic technology and non-contact or minimal-contact mechanical operations significantly reduces the risk of physical damage to fragile documents. From image recognition, trajectory planning, electrostatic treatment, page turning to image acquisition, a complete automated closed-loop process is formed, requiring no manual intervention and improving the efficiency of automatic document digitization scanning.
[0086] In step S100 of some embodiments, the folding edge region refers to the side of the book opposite the spine, i.e., the area where the unbound free edge of the page is located. When the book is laid flat naturally, this area will naturally form a concave curve due to the flexibility of the paper and gravity. Dynamic image data refers to video streams or image sequences captured continuously or at specific time intervals by an image acquisition device, which can reflect the changes in the morphology of the folding edge region over time. It not only contains spatial information but also contains temporal changes in the page during the folding process, used to track the page movement trajectory and detect instantaneous separation gaps. By capturing the morphological changes of the folding edge region in real time and in high definition, stable and reliable raw data is provided for subsequent AI recognition. For example, a high-resolution industrial camera equipped with a macro lens can be used to ensure that the microscopic edges and tiny gaps of the page at the folding edge can be clearly captured. The camera has high frame rate characteristics to meet the capture requirements of dynamic processes. Optionally, a ring-shaped LED shadowless lamp source and a side auxiliary light source at a specific angle can also be integrated. A ring light source provides uniform base lighting, reducing shadows; a side light source projects light at a low angle, enhancing the outline of the page edges and creating a "blade line" effect that highlights subtle height differences, greatly optimizing the contrast of edge features.
[0087] In step S200 of some embodiments, a hybrid decision-making system is employed. This system combines the speed and accuracy of image processing algorithms with the strong adaptability and robustness of deep learning models by running two recognition engines in parallel. The results are then fused and arbitrated through rigorous decision-making logic to ultimately output the optimal target page gap recognition result. This effectively combines the interpretability and computational efficiency of image processing algorithms with the robustness and high accuracy of deep learning models, overcoming the limitations of single methods in complex scenarios. This significantly improves the success rate and accuracy of page gap recognition under different paper materials, lighting conditions, and page states.
[0088] In some embodiments, step S200 may include, but is not limited to, steps S210 to S250:
[0089] Step S210: Based on the image processing algorithm, obtain the first page gap and the first confidence level;
[0090] Step S220: Based on the deep learning model, obtain the second page gap and the second confidence level;
[0091] Step S230: Compare the first confidence level and the second confidence level to obtain the comparison result;
[0092] Step S240: Obtain the geometric consistency verification result between the first page gap and the second page gap;
[0093] Step S250: Based on the comparison results and the geometric consistency verification results, obtain the target page gap.
[0094] In step S210 of some embodiments, the input dynamic image data is preprocessed, followed by edge detection and feature extraction. Then, the search range is narrowed to a narrow strip region by region focusing. The focused region is vertically projected to determine the candidate gap, i.e., the first page gap. A comprehensive evaluation is performed based on the low-level visual features of the candidate gap to obtain a first confidence level. Image processing technology is used to quickly obtain a preliminary gap location and its reliability assessment.
[0095] In some embodiments, step S210 may include, but is not limited to, steps S211 to S216:
[0096] Step S211: Perform noise interference removal and edge contrast enhancement operations on the dynamic image data to obtain the first intermediate image;
[0097] Step S212: Perform edge detection on the first intermediate image to obtain the page edge features;
[0098] Step S213: Based on prior knowledge of the book's flip edge, fit the flip edge curve according to the page edge features;
[0099] Step S214: Establish a strip-shaped search area inside the curve at the flip end;
[0100] Step S215: Perform vertical projection calculation on the strip search area and obtain the local minimum point of the projection curve as the first page gap.
[0101] Step S216: Obtain the first confidence level of the first page gap.
[0102] In step S211 of some embodiments, the acquired dynamic image data of the book's flip edge area is subjected to grayscale conversion, filtering, and illumination equalization to eliminate noise interference and enhance edge contrast, laying the foundation for subsequent feature extraction. Optionally, the color dynamic image data is converted into a grayscale image to reduce the amount of data required for subsequent calculations. Then, Gaussian filtering or median filtering algorithms are used to eliminate interference from image sensor noise and the texture of the paper itself. Next, algorithms such as contrast-limited adaptive histogram equalization (CLAHE) are used to eliminate local overexposure or underexposure caused by uneven illumination and enhance the contrast of the page edges. Exemplarily, the dynamic image data is sent to a preprocessing thread, which sequentially performs grayscale conversion (cv2.cvtColor(Color2Gray), Gaussian blur (cv2.GaussianBlur(ksize=5)), and contrast-limited adaptive histogram equalization (cv2.createCLAHE(clipLimit=2.0)) operations, outputting an optimized grayscale image as a first intermediate image for use in subsequent steps. The `cv2.cvtColor(Color2Gray)` function converts an RGB color image to a grayscale image using OpenCV functions. Essentially, it calculates the brightness value of each pixel. Since the human eye has varying sensitivities to different colors, a weighted average method is used.
[0103] ;
[0104] In the formula, R, G, and B represent the intensity values of a pixel in the red, green, and blue channels, respectively (ranging from 0 to 255); Gray is the grayscale value, a single value calculated using a specific weighted combination (0.299·R + 0.587·G + 0.114·B), used to represent the brightness or darkness of this colored pixel. The weighting coefficients (0.299, 0.587, 0.114) are standard values set based on the differences in human eye sensitivity to different wavelengths of light.
[0105] For Gaussian blur `cv2.GaussianBlur(ksize=5)`, this operation blurs (smooths) the image using Gaussian filtering to suppress noise. Its core parameter `ksize=5` specifies the size of the Gaussian kernel used for convolution as 5×5 pixels. This size determines the neighborhood range considered for each pixel during filtering; a larger kernel produces a stronger blurring effect, while 5×5 is a commonly used value that strikes a balance between effectively suppressing noise and preserving important edge features. OpenCV automatically calculates the standard deviation of the Gaussian kernel based on this size.
[0106] The contrast-limited adaptive histogram equalization `cv2.createCLAHE(clipLimit=2.0)` enhances local image contrast through contrast-limited adaptive histogram equalization. The core parameter `clipLimit=2.0` sets the threshold for histogram clipping. This value limits the number of pixels at any given gray level in a local region, preventing over-enhancement and noise amplification caused by a small percentage of gray levels being too high.
[0107] In step S212 of some embodiments, edge detection operators such as Canny and Sobel are used to highlight edge features in all vertical directions in the image, replacing simple grayscale scanning, which more robustly highlights structural edges in the image and provides higher quality data points for subsequent curve fitting.
[0108] In step S213 of some embodiments, based on prior knowledge of the book's flip edge, the search range for page gaps is narrowed to a narrow strip-shaped area to improve processing efficiency and anti-interference capability. Optionally, the prior knowledge includes that page gaps typically appear in a specific area of the flip edge curve, and are dark seams between two bright edges (the edges of the page paper). Exemplarily, a set of pixels belonging to the outermost edge of the book's flip edge is extracted from the page edge features. A curve model (such as a quadratic curve, polynomial curve, or spline curve) is used to fit these discrete points, thereby finding a smooth curve that best represents the overall trend of these points, thus obtaining the flip edge curve.
[0109] In step S214 of some embodiments, the fitted folded edge curve is offset by a fixed distance along its normal direction towards the inside of the book (i.e., towards the spine), thereby generating a new, "parallel" inner curve. The area between the original folded edge curve and the offset curve forms a narrow, curved band. By transforming the problem of "finding a curve in the entire image" into the problem of "finding a straight line in a narrow band," the search difficulty and computational load are greatly reduced.
[0110] In step S215 of some embodiments, a vertical projection calculation is performed on the strip-shaped search region, and the center position of the candidate gap is determined by finding the local minimum point (valley) on the projection curve. For example, for the image matrix of the strip-shaped search region, the sum of gray values in each column (x-coordinate), Projection(x), is calculated using the following formula:
[0111] ;
[0112] In the formula, I(x,y) is the gray value of the pixel located at (x,y); and The boundary of the strip region in the y-direction is defined. Inside the page, the paper is mainly the background color (e.g., white), with high pixel grayscale values, resulting in a large summation projection(x) value. At the page gaps, there are dark shadows with low pixel grayscale values, leading to a small summation projection(x) value. Therefore, the page gaps appear as a distinct "valley" or "local minimum" on the projection curve. All local minima are found on the obtained vertical projection curve (Projection(x)) using a sliding window or derivative detection algorithm. True page gaps can be filtered based on the depth (difference between the minimum and the average values on either side) and width of the valley, eliminating false valleys caused by paper texture, binding holes, or noise. The x-coordinate of each valid local minimum point corresponds to the center position of a page gap within the strip region. Combining the y-coordinate of the strip region, a precise, curved page gap can be drawn in the original image.
[0113] In step S216 of some embodiments, a first confidence level is calculated based on the sharpness of the troughs of the projection curve at the center of the first page gap, the consistency of the spacing between adjacent troughs, and the clarity of edge features. The value of this first confidence level is between 0 and 1. For example, a comprehensive evaluation is made based on the low-level visual features of the first page gap: the sharper and deeper the trough, the more drastic the grayscale change at that point, and the higher the probability that it is a real gap. This can be quantified by calculating the second derivative or gradient rate of change of the trough point. Since book page numbers or margins typically have roughly equal physical spacing, the standard deviation of the current page gap and the spacing of adjacent historical gaps is calculated; the smaller the difference, the higher the confidence level. The average gradient magnitude of the edges near the page gap is obtained; the larger the value, the clearer the edge, and the higher the confidence level. A first confidence level between 0 and 1 is calculated by weighted averaging the sharpness of the troughs, the consistency of the spacing between adjacent troughs, and the clarity of edge features.
[0114] In step S220 of some embodiments, the input dynamic image data is preprocessed, and then the preprocessed image is input into a pre-trained semantic segmentation model through model inference. Morphological closing operations are then used for post-processing to form a coherent gap region. A second page gap is then extracted from this coherent gap region. Based on the powerful feature learning capabilities of deep learning models, another independent gap recognition result and probabilistic confidence level are obtained.
[0115] In some embodiments, step S220 may include, but is not limited to, steps S221 to S225:
[0116] Step S221: Perform noise interference removal and edge contrast enhancement operations on the dynamic image data to obtain the second intermediate image;
[0117] Step S222: Input the second intermediate image into the pre-trained semantic segmentation model and output a segmentation mask; each pixel in the segmentation mask is divided into background, page, or gap classes.
[0118] Step S223: Perform morphological closing operation on the target region corresponding to the pixel of the gap class to obtain the gap region;
[0119] Step S224: Extract the second page gap from the gap region using a straight line fitting algorithm;
[0120] Step S225: Obtain the second confidence level of the second page gap.
[0121] In step S221 of some embodiments, the acquired dynamic image data of the book's flip edge area is subjected to grayscale conversion, filtering, and illumination equalization to eliminate noise interference and enhance edge contrast, laying the foundation for subsequent model inference. Optionally, the color dynamic image data is converted into a grayscale image to reduce the amount of data required for subsequent calculations. Gaussian filtering or median filtering algorithms are used to eliminate interference from image sensor noise and the texture of the paper itself. Then, algorithms such as contrast-limited adaptive histogram equalization are used to eliminate local overexposure or underexposure caused by uneven illumination, enhancing the contrast of the page edges. Exemplarily, the dynamic image data is sent to a preprocessing thread, which sequentially performs grayscale conversion (cv2.cvtColor(Color2Gray), Gaussian blur (cv2.GaussianBlur(ksize=5)), and contrast-limited adaptive histogram equalization (cv2.createCLAHE(clipLimit=2.0)) operations, outputting an optimized grayscale image as a second intermediate image for use in subsequent steps.
[0122] In step S222 of some embodiments, the preprocessed second intermediate image is input into a pre-trained semantic segmentation model (such as U-Net). This model is able to classify each pixel in the image and output a segmentation mask, where each pixel is labeled as a category such as "background", "page", or "gap".
[0123] In step S223 of some embodiments, a morphological closing operation (dilation followed by erosion) is performed on the “gap” category region output by the model to connect any possible breakpoints, fill the voids, and form a coherent gap region.
[0124] In step S224 of some embodiments, a skeletonization algorithm or least squares linear fitting is used to extract a smooth, single-pixel-wide center line from the continuous gap region as the recognition result of the second page gap.
[0125] In step S225 of some embodiments, the deep learning model typically outputs a probability value for each pixel's category prediction. The average probability value of all pixels in the second page gap belonging to the "gap" category is taken as the second confidence level.
[0126] In step S230 of some embodiments, comparing the first confidence level and the second confidence level is the first step of the decision-making logic, and preliminary screening is performed based on the inherent reliability of the two results. Two preset thresholds are established: the first preset threshold is a high threshold (…). ) and the second preset threshold is a low threshold ( ),and .
[0127] In step S240 of some embodiments, obtaining the geometric consistency verification result between the first page gap and the second page gap is the second step of the decision logic, verifying whether the two results match in geometric space. The average Euclidean distance between the first page gap (usually represented as a straight line) and the second page gap (another straight line) is calculated. (or overlap). Set a maximum permissible error. (e.g., 2 pixels). If If the result is correct, the verification passes and the results are considered consistent; otherwise, they are inconsistent.
[0128] In some embodiments, step S250 may include, but is not limited to, steps S251 to S253:
[0129] Step S251: When the second confidence level is greater than the first preset threshold, the second page gap is taken as the target page gap.
[0130] Step S252: When the second confidence level is less than or equal to the first preset threshold, the first confidence level is greater than the second preset threshold, and the geometric difference between the first page gap and the second page gap is less than the first error range, the first page gap is taken as the target page gap.
[0131] Step S253: When the second confidence level is less than or equal to the first preset threshold and the first confidence level is less than or equal to the second preset threshold, return to the step of obtaining dynamic image data of the book's flip-page area.
[0132] In steps S251 to S253 of some embodiments, when the second confidence level is greater than the first preset threshold, the result of the deep learning engine (second page gap) is adopted as the target page gap. When the second confidence level is less than or equal to the first preset threshold, the first confidence level is greater than the second preset threshold, and the geometric difference between the first page gap and the second page gap is less than the first error range, it indicates that the results of the two engines are consistent and mutually corroborate each other. In this case, the result of the image processing algorithm (first page gap) is adopted as the final output because its edge localization is usually more pixel-level accurate. When the second confidence level is less than or equal to the first preset threshold, the first confidence level is greater than the second preset threshold, and the geometric difference between the first page gap and the second page gap is greater than or equal to the first error range, the page gap corresponding to the higher confidence level (first page gap or second page gap) is adopted as the target page gap. In case of conflicting results, the more reliable page gap is selected. When the second confidence level is less than or equal to the first preset threshold, and the first confidence level is less than or equal to the second preset threshold, a redundancy processing mechanism is activated. This could involve returning to the step of acquiring dynamic image data of the book's flip area, adjusting the lighting, fine-tuning the book's posture, and then re-shooting, or using the high-confidence result of the previous frame to predict the current frame gap position.
[0133] For example, such as Figure 2 As shown, the pre-processed image of the book's flip edge region is simultaneously fed into two independent recognition engines for processing. The image has undergone grayscale conversion, Gaussian filtering, and contrast-limited adaptive histogram equalization to eliminate noise interference and enhance edge contrast. In the image processing algorithm engine, edge detection (e.g., using the Canny operator) is first performed to obtain page edge features. Then, based on prior knowledge of the book's concave flip edge, a curve is fitted to the flip edge, and a narrow strip-shaped search region is established inside it. Finally, by performing vertical projection calculations on this strip-shaped region, local minima (valleys) are found on the projection curve, thus determining the gap line L1 (i.e., the first page gap). The engine calculates a first confidence level Conf1 based on low-level visual features such as the depth and sharpness of the valleys. In the deep learning engine, the model performs pixel-level classification on the input image and outputs a segmentation mask, where each pixel is classified as "background," "page," or "gap." Post-processing, such as morphological closing operations, is performed on the pixel regions of the "gap" category to form coherent gap regions. Finally, the gap line L2 (i.e., the second page gap) is extracted through skeletonization or straight line fitting. The engine uses the average probability value of all pixels on the gap line L2 belonging to the "gap" category as the second confidence level Conf2.
[0134] refer to Figure 2After the dual engines complete the recognition, the process enters the arbitration stage, where the logic makes the final decision based on confidence comparison and geometric consistency verification. Upon receiving the outputs of the two engines (L1, Conf1, L2, Conf2), it first determines whether Conf2 > the high threshold. If so, the result L2 from the deep learning engine is directly adopted as the final output. If Conf2 ≤ the high threshold, but Conf1 > the low threshold, the average Euclidean distance between the two gap lines L1 and L2 is calculated. . judge <Allowable error? If yes, adopt the result L1 of the traditional algorithm as the final output; if not, directly adopt the result corresponding to the higher of Conf1 and Conf2. If the confidence of both engines is very low (e.g., Conf2 ≤ high threshold and Conf1 ≤ low threshold), then activate the redundant recognition mechanism. Optionally, this mechanism may include: the instruction execution mechanism fine-tunes the book's posture or lighting conditions and then re-acquires the image for recognition; or using temporal information, predicting the gap position of the current frame based on the high confidence result of the previous frame, and performing high-intensity calculations near the predicted position.
[0135] After the arbitration logic completes its decision, the process outputs the final, verified target page gap for use in subsequent motion trajectory generation steps. This process, through dual-engine parallelism and intelligent arbitration, maximizes the respective advantages of the image processing algorithm's precise positioning and the deep learning model's robustness, forming a complementary strength that ensures the system can stably and reliably identify page gaps in various complex scenarios.
[0136] In step S300 of some embodiments, the page-turning motion trajectory command refers to a series of low-level drive commands that control the mechanical execution system (such as a page-turning mechanical clamp, page-turning guide line, etc.) to complete the page-turning action. These commands typically contain information such as position, speed, acceleration, or torque, and are sent to the servo driver or motor through a specific communication protocol (such as EtherCAT, CANopen) to drive the mechanical unit to complete a precise trajectory movement. By converting the geometric information (target page gap) representing the ideal intervention position identified by the upstream vision system into low-level drive commands that the mechanical execution system can directly understand and execute, this process fully considers the geometric characteristics of three-dimensional space, the smoothness and stability of the motion process, and the kinematic constraints of specific mechanical structures, ensuring the accuracy, smoothness, and reliability of the page-turning action.
[0137] In some embodiments, step S300 may include, but is not limited to, steps S310 to S350:
[0138] Step S310: Obtain the first point cloud data of the three-dimensional spatial position of the gap between the target pages;
[0139] Step S320: Remove outliers from the first point cloud data to obtain the second point cloud data;
[0140] Step S330: Use the B-spline curve fitting algorithm to fit the second point cloud data to generate a trajectory curve in three-dimensional space.
[0141] Step S340: Based on preset kinematic parameters, perform S-shaped velocity planning on the trajectory curve to generate a position-time series;
[0142] Step S350: Convert the position-time series into page-turning motion trajectory instructions through inverse kinematics operations.
[0143] In step S310 of some embodiments, the target page gap identification result in the two-dimensional image is mapped to the three-dimensional physical world. For each feature pixel point p(u,v) on the target page gap identified in the image, its corresponding three-dimensional world coordinates P(X,0,Z) are solved by back-projection calculation using camera calibration parameters (intrinsic parameter matrix K and extrinsic parameter matrix [R|t]), utilizing the constraint that it is located on the book plane (Y=0). All calculated three-dimensional points {P} i The set of} constitutes the first point cloud data representing the gaps in that page.
[0144] In step S320 of some embodiments, due to visual recognition errors or back-projection calculation errors, the first point cloud may contain a few outliers that significantly deviate from the actual gap positions. A filtering algorithm is used to remove these outliers from the first point cloud data, thereby obtaining the second point cloud data. For example, a statistical filtering algorithm is used. First, the average distance between each point in the point cloud and its k nearest neighbors is calculated; assuming these distances follow a Gaussian distribution, all distances exceeding the mean are... 10 ... Points with a value between 1.0 and 2.0 are typically considered outliers and removed. This process effectively filters outlier noise points, retaining the point set representing the true page gap structure, i.e., the second point cloud data. Optionally, radius filtering or pass-through filtering can be used as auxiliary methods to further remove invalid points that clearly exceed the physical space range.
[0145] In step S330 of some embodiments, a B-spline curve fitting algorithm is used to fit the second point cloud data to generate a smooth, continuous three-dimensional spatial B-spline curve. For example, the points in the second point cloud data are parameterized, typically by calculating the cumulative chord length based on the spatial distribution of the points. Then, node vectors are defined based on the parameterization results to determine the structure of the B-spline curve. Next, the least squares method is used to solve for the control points of the B-spline curve, minimizing the sum of squared errors between the curve and the second point cloud data, ultimately generating a three-dimensional spatial trajectory curve that accurately describes the centerline of the ideal motion path of the mechanical end effector (such as a guide wire).
[0146] In step S340 of some embodiments, the generated three-dimensional spatial trajectory curve is parameterized by arc length to obtain the relationship s(u) between the curve length s and the parameter u. Then, S-shaped velocity planning is performed in the arc length-time domain (st domain). The planning process ensures that the generated velocity curve v(t) and acceleration curve a(t) are continuous and smooth, and the jerk j(t) is bounded. Finally, through integration and sampling, a series of discrete time points and their corresponding desired position points (arc lengths on the curve) or three-dimensional coordinate points are obtained, forming a position-time series.
[0147] In step S350 of some embodiments, the path in Cartesian space (operation space) can be converted into instructions in joint space (drive space) through inverse kinematics operations. For a specific mechanical actuator, an inverse kinematics mathematical model is established. This model describes the mapping relationship between the pose of the end effector (such as a page-turning guide line, guide line fixing rod, etc.) and the angles (or displacements) of each joint. For each three-dimensional coordinate point in the generated position-time series (and the required pose at that point, usually determined by the tangent direction of the curve), the inverse kinematics solver is invoked to calculate a set of joint angle values. The joint angle values corresponding to all time points are arranged in a time series, and interpolation and smoothing are performed considering the servo control cycle, ultimately generating a page-turning motion trajectory instruction that can be directly sent to the servo driver. This instruction contains information such as target position, velocity, or torque.
[0148] In step S400 of some embodiments, inter-page static charge refers to the accumulation of static electricity in book pages (especially paper pages) during printing, storage, and flipping due to friction, contact, and separation. Like charges repel each other, while unlike charges attract each other, causing pages to "stick" or "adhere," making reliable separation difficult and contributing significantly to automated page turning. A pneumatic actuator neutralizes the inter-page static charge, reducing the amount of static charge on the object's surface to a level insufficient to cause interference. Furthermore, the pneumatic actuator pre-separates the current target page from the next target page; that is, before the main page-turning action (such as the intervention of a guide line), a small, localized separation or loosening is created between the current target page (the page to be turned) and the next target page (the page immediately following it). This creates more favorable initial conditions for subsequent page-turning actions, significantly reducing the risk of page tearing or multiple pages being turned simultaneously. By using a non-contact physical method, a low-static-state, easily separable page is created before the page-turning mechanical action is executed. This combines static electricity elimination technology with fluid dynamic control technology, improving the robustness and success rate of the entire system.
[0149] In some embodiments, step S400 may include, but is not limited to, steps S410 to S420:
[0150] Step S410: Control the pneumatic actuator to perform ion air jetting to neutralize the static charge between the pages of the book;
[0151] Step S420: Control the pneumatic actuator to perform micro-airflow injection to pre-separate the current target page from the next target page.
[0152] In step S410 of some embodiments, such as Figure 3 As shown, the pneumatic-electric integrated nozzle in the control pneumatic actuator system simultaneously activates the high-pressure module and the air valve to generate and spray an ion wind to neutralize the static charge between the pages of the book. Exemplarily, the ion generator is connected to high-voltage AC power, generating corona discharge at the tip of its needle-shaped or wire-shaped electrode, ionizing the surrounding air and generating a large number of positive and negative ions. An air pump generates compressed air, the pressure and flow rate of which are controlled by a flow regulating valve to ensure a stable and gentle airflow, avoiding disturbing the pages. The compressed air flows through the ion generator, "enveloping" the positive and negative ions to form an ion wind for neutralization. The timing of the ion wind's spray is controlled by a solenoid valve, and the nozzle is aimed at the target page gaps at the flip opening, thereby quickly eliminating the static charge between the pages and fundamentally weakening the electrostatic force that causes the pages to adhere. Optionally, the system can integrate an electrostatic sensor to monitor the electrostatic potential of the page surface in real time. Based on sensor feedback, the controller dynamically adjusts the spray time of the ion wind or the operating power of the ion generator to achieve closed-loop control and ensure the neutralization effect.
[0153] In step S420 of some embodiments, such as Figure 3 As shown, the integrated pneumatic-electric nozzle is controlled to shut off the high-pressure module while keeping the air valve open, generating and spraying a dry micro-airflow to pre-separate the current target page from the next target page. Exemplarily, under the control of the generated trajectory command, the micro-airflow nozzle is precisely moved above the edge of the current target page. The airflow pressure (e.g., 0.1 to 0.3 MPa) and spray duration (e.g., 50 to 200 ms) are preset according to the book page material (e.g., paper weight, smoothness). The controller issues a command to quickly open the precision solenoid valve. A brief, precisely controlled micro-airflow is ejected from the nozzle and blown into the gap between the current and next target pages. The pressure difference and shear force generated by the airflow act on the page edge, sufficient to overcome the weak van der Waals forces, adsorption forces, and residual electrostatic forces between the pages, causing a minute separation (possibly only 0.5 to 2 mm) between the top two or three pages that is imperceptible to the naked eye, creating an entry point for the subsequent page-turning guide line. This physical pre-separation of the pages significantly reduces the force and difficulty of subsequent mechanical page turning, increasing the success rate of single-page turning.
[0154] In step S500 of some embodiments, a specially designed ultra-fine, ultra-smooth guide line is used as the execution end. Under precise motion trajectory command control, the separation and lifting of the page is gently completed through "line contact" rather than "surface contact," thereby minimizing the contact area and interaction force with the pages of the valuable document and fundamentally avoiding physical damage such as scratches and creases that may be caused by rigid actuators. For example, the intervention of the guide line to lift the target page may include the following steps:
[0155] Step 1, Precise Positioning: The motion controller receives and parses the page-turning trajectory instructions, controls the guide line fixing rod and the guide line it carries to move to the starting point of the trajectory (usually located above the outer edge of the upper right corner of the page).
[0156] Step 2, Oblique Intervention: The controller guides the guide wire along a pre-calculated oblique trajectory (the angle of which is optimized to facilitate sliding into the gap), causing the tip of the guide wire to gradually penetrate into the pre-separated V-shaped opening or natural gap. During this process, the pneumatic actuator continuously injects micro-airflow to assist the guide wire in sliding smoothly, forming an "air cushion" effect to further reduce friction.
[0157] Step 3, Horizontal Deflection: After the guide line is inserted to the predetermined depth (such as near the spine), the controller instructs the guide line fixing rod to perform a small horizontal deflection movement (such as moving along the X-axis) to ensure that the guide line is in an ideal position below the page that is conducive to lifting.
[0158] Step 4: Vertical Lifting: The controller ultimately instructs the guide wire fixing rod to move upwards at a constant speed along the vertical direction (Z-axis). Since the guide wire is located below the target page, its upward movement will arch the single target page upwards, forming a distinct and stable raised fold. This lifting height is typically 3 to 5 centimeters to ensure reliable gripping by the subsequent mechanical clamp.
[0159] In step S600 of some embodiments, an orthogonal dual-camera system is used to scan the previous target page and the current target page. For example... Figure 4 As shown, a book to be scanned is placed on a horizontal stage 1. The orthogonal dual cameras include a horizontal page camera 2 and a vertical page camera 3. The horizontal page camera 2 has a vertical optical axis and captures the previous target page laid flat on the horizontal stage to obtain the first page image. The vertical page camera 3 has a horizontal optical axis and captures the current target page perpendicular to the horizontal plane to obtain the second page image, thus achieving synchronous, distortion-free, high-definition acquisition of both pages.
[0160] In step S700 of some embodiments, after a complete "discharge-separation-pickup-scanning" process is completed, the page-turning mechanical gripper picks up the lifted page edge, completing the "page-turning" action to update the book's state. The updated state variable ("current target page" becomes "previous target page") is then passed to the next loop, while all actuators are reset. Thus, the system can process the entire book page by page without manual intervention until the entire scanning task is completed. Exemplarily, this loop and control process may include the following steps:
[0161] Step 1: Perform page turning: After scanning is complete, the page holder is released. The page turning mechanism holds the edge of the current target page and moves along an optimized trajectory, turning it over the spine from the current side (e.g., the right side) and smoothly placing it on the platform on the other side (the left side).
[0162] Step 2, Status Update: In the system's internal counter, increment the page number that was just processed (e.g., page N) by one (N+1). The new page becomes the new "current target page", while the original "current target page" (page N) becomes the "previous target page".
[0163] Step 3, Actuator Reset: The page-turning mechanical clamp releases the page and returns to the initial position; the guide wire fixing rod descends and returns to the initial position.
[0164] Step 4, Loop Judgment and Jump: The system determines whether the current page number has exceeded the total number of pages in the book. If not, the program jumps (returns) to the step of neutralizing the static charge between pages and pre-separating the current target page from the next target page, and begins the static neutralization and pre-separation of the new page; if the total number of pages has exceeded the limit, the entire scanning process ends.
[0165] In some embodiments, an AI-based vision and pneumatic-assisted automatic document digitization scanning system is provided for implementing an AI vision and pneumatic-assisted automatic document digitization scanning method. The system includes a book fixing system, an AI vision system, a pneumatic execution system, a mechanical execution system, and an image acquisition system. Specifically, it includes the following:
[0166] 1. Book fixing system
[0167] Horizontal stage 1 (e.g.) Figure 4 (As shown): Used to support books, providing uniform planar support.
[0168] Spine clamp: Located on one side of the platform, it is used to vertically clamp the spine of the book, gently but firmly securing the book. Its clamping surface is covered with a flexible protective material (silicone, soft velvet) to gently and firmly clamp the spine of the book, fixing it to the horizontal platform.
[0169] Flexible book-pressing strips: These strips position and press down on non-working areas from the book's edge to prevent the book from shifting during page turning.
[0170] 2. AI Vision Servo System
[0171] High-definition macro camera: Captures the pages of books, obtaining real-time images of their working status.
[0172] Built-in machine vision large model (based on Gemma 3 27B model): for page gap recognition: identifying the edge boundary between adjacent pages through a hybrid decision system (fusion of image processing and deep learning); pose calculation: calculating the spatial position and pose of books, guide lines, and pages in real time; motion planning: generating optimal motion trajectory instructions for page-turning guide lines and page-turning mechanical clamps; closed-loop control: dynamically adjusting the actions of the actuators based on real-time image feedback to form a visual servo closed loop.
[0173] 3. Pneumatic actuator system
[0174] Integrated pneumatic-electric nozzle: Integrates ion wind generation and micro-airflow output functions. First, it releases ion wind to neutralize the static charge on the page surface; then, it outputs a dry micro-airflow with adjustable pressure for pre-separating pages, assisting in lifting guide lines, and assisting in page turning, creating an "air bearing" effect to reduce friction.
[0175] Drying micro-airflow generator: includes a micro air compressor, a drying filter, an air tank, an electro-proportional valve (for stepless adjustment of output air pressure), a high-speed switching solenoid valve, and a pressure sensor, forming a closed-loop control to ensure that the airflow is dry, clean, and its intensity is precisely controllable.
[0176] 4. Mechanical Actuation System
[0177] Page turning guide: Uses ultra-smooth, wear-resistant fine thread (such as 0.1mm fluorocarbon coated nylon thread) as the execution end for intervening and lifting the page.
[0178] Guide line fixing rod: Installed on the high-precision linear module, it can control the guide line to perform oblique insertion, horizontal deflection and vertical lifting movements.
[0179] Page-turning mechanical gripper: It uses flexible grippers to grab the edge of the lifted page and complete the page-turning action.
[0180] Page retainer: Used to hold a page in place after it has been turned on, preventing it from springing back.
[0181] 5. Image Acquisition System
[0182] Orthogonal dual-camera setup: Vertical page camera with horizontal optical axis: photographs the opened page that is perpendicular to the horizontal plane; Horizontal page camera with vertical optical axis: photographs the page laid flat on the horizontal stage.
[0183] This invention also provides a visual interaction method between the document digitization automatic scanning system and a smart terminal, which may include, but is not limited to, tablet computers, laptops, and mobile phones. Figure 5 As shown, taking a mobile smart terminal as an example, this embodiment of the invention provides a mobile smart terminal interactive interface for an automatic document digitization scanning system.
[0184] refer to Figure 5 The mobile smart terminal interface of the document digitization automatic scanning system includes a status area 4, a control area 5, and a monitoring area 6. Status area 4, located at the top of the interface, centrally displays the system's core status and metadata information, including but not limited to the title of the book currently being scanned (e.g., "XXXX"). When the title is too long, the interface will automatically partially hide it, allowing users to expand the full text by clicking the title. The system's current working status is displayed in real-time using text labels (e.g., "Page Turning," "Ready," "Paused," "Scanning"), allowing users to quickly perceive the device's operational stage. The scanning progress indicator accurately displays the currently processed page number (e.g., "Page 208"), providing users with intuitive progress feedback. In addition, status area 4 includes function entry buttons, providing two core function buttons: "Settings" and "View Page Turning List." The "Settings" button takes users to the system parameter configuration page, where they can adjust advanced parameters such as scan resolution, page turning speed, and airflow pressure. The "View Page Turning List" button takes users to the task management page, displaying the queue of books to be scanned, the history of completed scans, and corresponding document information.
[0185] like Figure 5The control area 5 shown is located at the bottom of the interface and integrates a series of touch buttons for directly controlling the scanning process. These may include, but are not limited to, a start / pause button for starting or pausing the automatic scanning process; a page forward / page backward button for performing a single page turn in manual or semi-automatic mode; an emergency stop button for immediately stopping all mechanical movements and entering a safe state in case of an abnormality; an image acquisition button for manually triggering an image capture command; and a view settings button for setting the view format of the information displayed in the monitoring area.
[0186] like Figure 5 The monitoring area 6 shown is used for real-time visualization of key information during the scanning process. Based on the view settings buttons in control area 5, the view format in monitoring area 6 can be set. Optionally, monitoring area 6 simulates a book's double-page spread format, such as... Figure 6 As shown, the area is divided into two sub-areas by a vertical dotted line (corresponding to the gap between the target pages). The left sub-area 61 is used to display and locate the first page image (e.g., page 208) of the book currently laid flat on the horizontal platform (i.e., the previous target page), which is the page image captured by the horizontal page camera with the optical axis perpendicular to the horizontal plane. The right sub-area 62 is used to display and locate the second page image (e.g., page 209) of the book currently opened and perpendicular to the horizontal plane (i.e., the current target page), which is the page image captured by the vertical page camera with the optical axis horizontal to the horizontal plane. Optionally, the monitoring area 6 only displays the first page image of the book currently laid flat on the horizontal platform (e.g., page 209). Figure 7 (as shown in page 208), or simply display an image of the second page (as shown in page 209) that is perpendicular to the horizontal plane and positioned relative to the currently opened page.
[0187] refer to Figure 5 , Figure 6 and Figure 7 A visual control method for document digitization scanning applied to smart terminals may include the following steps:
[0188] The display scan control interface includes a status area, a control area, and a monitoring area.
[0189] The status area includes various status indicator components, which display the title of the book currently being scanned, the current working status of the automatic scanning system, and the scanning progress indicator, but are not limited to these. The status area also includes function entry buttons "Settings" and "View Page List Books".
[0190] Clicking the "Settings" button will take you to the system parameter configuration page, where you can adjust advanced parameters such as scan resolution, page turning speed, and airflow pressure. Clicking the "View Page Turning List Books" button will take you to the task management page, which displays the queue of books to be scanned, the history of completed scans, and the corresponding literature information.
[0191] The control area includes, but is not limited to, start / pause buttons, page forward / page backward buttons, emergency stop buttons, image acquisition buttons, and view settings buttons.
[0192] In response to clicking the "Start / Pause" button, the automatic scanning process is started or paused; in response to clicking the "Page Forward / Page Back" button, a single page turn is performed; in response to clicking the "Emergency Stop" button, all mechanical movements are stopped; in response to clicking the "Image Acquisition" button, an image capture is performed; in response to selecting the "View Settings" button, the view format of the information displayed in the monitored area is set, including a double-page view of a book and a single-page view of a book, wherein the single-page view of a book includes a view of the page currently laid flat on the horizontal platform and a view of the page currently perpendicular to the horizontal plane.
[0193] In response to a selection operation in the first view mode, the scanning book's double-page view is displayed in real time in the monitoring area; in response to a selection operation in the second view mode, the page image captured by the horizontal page camera with the optical axis perpendicular to it is displayed in real time in the monitoring area; in response to a selection operation in the third view mode, the page image captured by the vertical page camera with the optical axis perpendicular to it is displayed in real time in the monitoring area.
[0194] By combining the interactive interface of status area, monitoring area and control area, a mobile control terminal with comprehensive information display, intuitive visual feedback and clear operation logic is constructed, which greatly improves the visualization of the document digitization scanning process and the user's operating experience.
[0195] This invention also provides an automatic document digitization scanning device based on AI vision and pneumatic assistance, which can realize the above-mentioned automatic document digitization scanning method based on AI vision and pneumatic assistance. The device includes:
[0196] The first AI visual servo module is used to acquire dynamic image data of the book's flip-page area;
[0197] The second AI visual servo module is used to identify dynamic image data and obtain the gaps in the target page based on image processing algorithms and deep learning models.
[0198] The third AI visual servo module is used to generate page-turning motion trajectory instructions based on the gaps in the target page.
[0199] The pneumatic actuator module is used to neutralize the static charge between the pages of a book and to pre-separate the current target page from the next target page.
[0200] The first mechanical execution module is used to insert the guide line into the gap between the current target page and the next target page according to the page turning trajectory instruction, and to lift the current target page;
[0201] The image acquisition module is used to scan the previous target page and the current target page to obtain the first page image of the previous target page and the second page image of the current target page;
[0202] The second mechanical execution module is used to turn the current target page, use the current target page as the previous target page, and return to the step of neutralizing the static charge between the pages of the book until the book scanning is completed.
[0203] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0204] This invention also provides an electronic device, which includes a processor and a memory. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including a tablet computer, an in-vehicle computer, or similar device.
[0205] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0206] refer to Figure 8 , Figure 8 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:
[0207] The processor 801 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention.
[0208] The memory 802 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 802 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 802 and is called and executed by the processor 801.
[0209] The 803 input / output interface is used to implement information input and output.
[0210] The communication interface 804 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0211] Bus 805 transmits information between various components of the device (e.g., processor 801, memory 802, input / output interface 803, and communication interface 804);
[0212] The processor 801, memory 802, input / output interface 803, and communication interface 804 are connected to each other within the device via bus 805.
[0213] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0214] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0215] This invention also provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions to cause the computer device to perform the aforementioned method.
[0216] In summary, the document digitization automatic scanning device and method based on AI vision and pneumatic assistance of the present invention has the following advantages:
[0217] 1. By integrating the electrostatic neutralization of ion wind with the physical separation of drying micro-airflow into a single nozzle, and through precise timing control to operate in stages, it not only eliminates electrostatic interference but also achieves non-contact physical separation, greatly reducing the risk of damage to fragile pages.
[0218] 2. By combining the advantages of image processing algorithms (high interpretability and pixel-level accuracy) and deep learning models (high adaptability and strong robustness), intelligent arbitration is performed through confidence assessment and geometric consistency verification, which significantly improves the accuracy and reliability of page gap recognition.
[0219] 3. By utilizing controllable micro-airflow to form an air film between the page and the guide line / mechanical clamp, friction and mechanical stress are effectively reduced, making the page turning process smoother.
[0220] 4. Based on monocular vision and prior constraints (the book is fixed on the stage and the page is in a known plane), the precise three-dimensional pose of the page gap is calculated in real time, and a smooth mechanical motion trajectory is generated, realizing high-precision servo control.
[0221] 5. The image acquisition system uses orthogonal dual cameras to achieve synchronous, distortion-free high-definition acquisition of both pages.
[0222] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order shown in the operation diagrams. For example, depending on the functions / operations involved, two consecutively shown blocks may actually be executed substantially simultaneously, or the blocks may sometimes be executed in reverse order. Furthermore, the embodiments presented and described in the flowcharts of this invention are provided by way of example to provide a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logic flows presented herein. Alternative embodiments are contemplated in which the order of various operations is altered and sub-operations described as part of a larger operation are executed independently.
[0223] Furthermore, although the invention has been described in the context of functional modules, it should be understood that, unless otherwise stated, one or more of the described functions and / or features may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in a separate physical device or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the invention. Rather, given the properties, functions, and internal relationships of the various functional modules in the apparatus disclosed herein, the actual implementation of the module will be understood within the scope of conventional skill of an engineer. Therefore, those skilled in the art can implement the invention as set forth in the claims using ordinary techniques without excessive experimentation. It is also understood that the specific concepts disclosed are merely illustrative and not intended to limit the scope of the invention, which is determined by the full scope of the appended claims and their equivalents.
[0224] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0225] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-including system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0226] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0227] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0228] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0229] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.
[0230] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of the present invention.
Claims
1. An automatic document digitization scanning method based on AI vision and pneumatic assistance, characterized in that, Includes the following steps: Acquire dynamic image data of the book's flip-page area; Based on image processing algorithms and deep learning models, the dynamic image data is identified to obtain the target page gap; Based on the target page gap, a page-turning motion trajectory instruction is generated, including: acquiring first point cloud data of the three-dimensional spatial position of the target page gap; removing outliers from the first point cloud data to obtain second point cloud data; using a B-spline curve fitting algorithm to fit the second point cloud data to generate a trajectory curve in three-dimensional space; performing S-shaped velocity planning on the trajectory curve based on preset kinematic parameters to generate a position-time series; and converting the position-time series into the page-turning motion trajectory instruction through inverse kinematics operations. The static charge between the pages of the book is neutralized, and the current target page and the next target page are pre-separated. According to the page-turning trajectory instruction, the guide line is inserted into the gap between the current target page and the next target page, and the current target page is lifted. Scan the previously mentioned target page and the current target page to obtain a first page image of the previously mentioned target page and a second page image of the current target page; Turn the current target page, set the current target page as the previous target page, and return to the step of neutralizing the static charge between the pages of the book until the book scanning is complete.
2. The method according to claim 1, characterized in that, The method of identifying the dynamic image data and obtaining the target page gap based on image processing algorithms and deep learning models includes the following steps: Based on image processing algorithms, the first page gap and the first confidence level are obtained; Based on a deep learning model, the second page gap and the second confidence level are obtained; The first confidence level and the second confidence level are compared to obtain the comparison results; Obtain the geometric consistency verification result between the first page gap and the second page gap; Based on the comparison results and the geometric consistency verification results, the target page gap is obtained.
3. The method according to claim 2, characterized in that, The method for obtaining the first page gap and the first confidence level based on the image processing algorithm includes the following steps: The dynamic image data is subjected to noise reduction and edge contrast enhancement operations to obtain a first intermediate image; Edge detection is performed on the first intermediate image to obtain the page edge features; Based on prior knowledge of the book's flip-edge, a flip-edge curve is fitted according to the page edge features; A strip-shaped search area is established inside the curve of the flip-end; The vertical projection calculation is performed on the strip search area, and the local minimum point of the projection curve is obtained as the first page gap; Obtain the first confidence level of the first page gap.
4. The method according to claim 2, characterized in that, The method for obtaining the second page gap and the second confidence level based on a deep learning model includes the following steps: The dynamic image data is subjected to noise reduction and edge contrast enhancement operations to obtain a second intermediate image; The second intermediate image is input into a pre-trained semantic segmentation model, which outputs a segmentation mask; each pixel in the segmentation mask is classified into a background class, a page class, or a gap class. A morphological closing operation is performed on the target region corresponding to the pixel of the gap class to obtain the gap region; The second page gap is extracted from the gap region using a linear fitting algorithm; Obtain the second confidence level of the second page gap.
5. The method according to claim 2, characterized in that, The step of obtaining the target page gap based on the comparison result and the geometric consistency verification result includes the following steps: When the second confidence level is greater than the first preset threshold, the second page gap is taken as the target page gap; When the second confidence level is less than or equal to the first preset threshold, the first confidence level is greater than the second preset threshold, and the geometric difference between the first page gap and the second page gap is less than the first error range, the first page gap is taken as the target page gap. When the second confidence level is less than or equal to the first preset threshold, and the first confidence level is less than or equal to the second preset threshold, return to the step of obtaining dynamic image data of the book's flip-page area.
6. The method according to claim 1, characterized in that, The neutralization of static charge between pages of the book and the pre-separation of the current target page from the next target page include the following steps: The pneumatic actuator is controlled to perform ion air jetting to neutralize the static charge between the pages of the book; The pneumatic actuator is controlled to perform micro-airflow injection to pre-separate the current target page from the next target page.
7. An automatic document digitization scanning device based on AI vision and pneumatic assistance, characterized in that, include: The first AI visual servo module is used to acquire dynamic image data of the book's flip-page area; The second AI visual servo module is used to identify the dynamic image data based on image processing algorithms and deep learning models to obtain the target page gap; The third AI visual servo module is used to generate page-turning motion trajectory instructions based on the target page gap; the third AI visual servo module is specifically used to: acquire the first point cloud data of the three-dimensional spatial position of the target page gap; remove abnormal points in the first point cloud data to obtain the second point cloud data; and use a B-spline curve fitting algorithm to fit the second point cloud data to generate a trajectory curve in three-dimensional space. Based on preset kinematic parameters, S-shaped velocity planning is performed on the trajectory curve to generate a position-time series; through inverse kinematics operation, the position-time series is converted into the page-turning motion trajectory command; The pneumatic actuator module is used to neutralize the static charge between the pages of the book and to pre-separate the current target page from the next target page. The first mechanical execution module is used to insert the guide line into the gap between the current target page and the next target page according to the page turning motion trajectory instruction, and to lift the current target page; The image acquisition module is used to scan the previously mentioned target page and the current target page to obtain a first page image of the previously mentioned target page and a second page image of the current target page; The second mechanical execution module is used to turn the current target page, use the current target page as the previous target page, and return to the step of neutralizing the static charge between the pages of the book until the book scanning is completed.
8. An electronic device, characterized in that, Including the processor and memory; The memory is used to store programs; The processor executes the program to implement the method as described in any one of claims 1 to 6.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 6.