Method and device for automatically shooting multi-page document with repeated content and stability detection

By employing deep learning-based object detection and content duplication detection algorithms, the system automatically identifies and captures stable and clear multi-page documents, solving the problems of low efficiency in manual user operation and stability detection in existing technologies, and achieving efficient automatic capture and processing of multi-page documents.

CN120877170APending Publication Date: 2025-10-31SHANGHAI HEHE INFORMATION TECH DEV +3
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510793339.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing technologies require users to manually click the camera button after each document page switch, and it is difficult to automatically detect the stability of the document and duplicate content during the shooting process, resulting in low operating efficiency.

Method used

A deep learning-based object detection algorithm is used to identify document and hand positions in video frames. By combining the vertex coordinate variance of the document position, the intersection ratio of the hand region and the document region, and the image blur value, the stability of the video frame is determined. A document content duplication detection algorithm is used to filter duplicate content and automatically trigger the capture of stable and clear document pages.

Benefits of technology

It enables automatic capture of multi-page documents without manual user operation, improving operational efficiency and ensuring the quality and uniqueness of captured images, resulting in clear electronic documents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877170A_ABST
    Figure CN120877170A_ABST
Patent Text Reader

Abstract

The invention discloses a method for automatically shooting a multi-page document with repeated content and stability detection. And outputting the extracted document position and hand position in each video frame by the trained document target detection network. Judging whether each video frame is in a document page switching state or a stable state; judging whether the document in the video frame at least in the stable state has a large-area hand shielding condition, and whether the document is clear or fuzzy; and meanwhile, the video frame meeting three conditions of a stable state, no large-area hand shielding and clearness is in a stable and photographable state. After the video frames in the document page state are switched, once the video frames in the stable and photographable state appear, shooting of the current document page is triggered immediately; and once the current document page is shot, waiting for the video frame for switching the document page state next time. And removing images with duplicate document contents by the trained document content duplicate detection network. The target detection algorithm is utilized to judge whether the current video frame is in a stable and photographable state or not.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to a method for automatically capturing images of multi-page documents. Background Technology

[0002] Existing non-automatic methods for photographing multi-page documents involve manually or mechanically switching between different document pages, with the operator manually determining whether each page is in a stable state, and photographing only when each page is in a stable state. The multi-page document includes unbound, discrete pages, as well as bound multi-page documents (e.g., books).

[0003] Chinese invention patent application CN119781801A, published on April 8, 2025, entitled "A Method and Apparatus for Automatically Capturing and Digitizing Books," discloses a method for automatically capturing multi-page documents. This method captures video of the book-turning process and extracts consecutive video frames, feeding them into two models. An optical flow estimation model estimates the optical flow between two consecutive frames, while a corner detection model obtains the corner points of each page. Based on the direction and amplitude of the optical flow and changes in the corner points, it determines whether the book is in a stable state. Only when the book is in a stable state is a stable, clear photograph of each page automatically captured for digitization. In other words, the time point at which a clear and stable photograph of each page can be captured is automatically determined based on the book-turning video. This method employs optical flow detection technology, primarily used in book-turning scenarios, and executes the automatic capturing logic based on the book-turning action and book stability. Summary of the Invention

[0004] The technical problem to be solved by this application is: how to automatically take pictures of multi-page documents without requiring the user to click the shooting button after each page switch, and to automatically perform stability detection and duplicate content detection during the shooting process, thereby improving operational efficiency.

[0005] To address the aforementioned technical problems, this application proposes a method for automatically capturing multi-page documents with duplicate content and stability detection, comprising the following steps: Step S1: Capture video of switching between different document pages in a multi-page document, and extract video frames sequentially from the video. Step S2: Feed each extracted video frame into a trained document object detection network, which outputs the document position and hand position in each video frame; the document position in each video frame is outlined by a rectangle, and the trained document object detection network outputs the coordinates of the four vertices of this rectangle; the hand position in each video frame is outlined by another rectangle, and the trained document object detection network outputs the coordinates of the four vertices of this rectangle. Step S3: Determine whether each video frame is in a document page switching state or a stable state; further determine whether the document in at least a stable state video frame has a large area of ​​hand occlusion; further determine whether the video frame in at least a stable state is clear or blurry; a video frame that simultaneously satisfies the three conditions of stable state, no large area of ​​hand occlusion, and clear is a stable and captureable state. Step S4: After the video frame switching document page states, once a stable, captureable video frame appears, immediately trigger the capture of the current document page; after capturing the current document page, wait for the next video frame switching document page states; repeat this step until every page of the multi-page document has been captured. Step S5: Feed all captured images into a trained document content duplication detection network to remove duplicate images. Step S6: Perform edge correction and sharpness enhancement on the deduplicated images, then assemble them sequentially to form an electronic document.

[0006] Furthermore, in step S1, one video frame is extracted every n video frames, where n is an integer between 0 and 5.

[0007] Furthermore, in step S2, the document position in each video frame is outlined by a rectangle, and the trained document object detection network outputs the coordinates of the four vertices of the rectangle; the hand position in each video frame is outlined by another rectangle, and the trained document object detection network outputs the coordinates of the four vertices of the rectangle; if different document pages are switched mechanically, the hand position is replaced by the position of the mechanical device that switches the document pages.

[0008] Further, in step S3, determining whether each video frame is in a document page switching state or a stable state is calculated based on the extracted n consecutive video frames; for the first n-1 video frames, they are all determined to be in a document page switching state; for the nth and subsequent video frames, the variance of the coordinates of the four vertices of the rectangle representing the document position in the n consecutive video frames ending with the current video frame is calculated; if any one or more variance values ​​are greater than a predetermined threshold, the current video frame is determined to be in a document page switching state; if all variance values ​​are less than the predetermined threshold, the current video frame is determined to be in a stable state.

[0009] Further, in step S3, the area of ​​the document region and the area of ​​the hand region in the video frame that is at least in a stable state are calculated, and then the area of ​​the overlapping region between the document region and the hand region in the video frame is calculated, and then the overlapping region area / document region area is calculated; if the overlapping region area / document region area ≤ a predetermined threshold, it means that there is no large area of ​​hand occlusion in the document in the video frame; if the overlapping region area / document region area > the predetermined threshold, it means that there is a large area of ​​hand occlusion in the document in the video frame.

[0010] Furthermore, in step S3, the Laplacian function of the OpenCV library is used to calculate the blur value of the video frame that is at least in a stable state; if the blur value is less than or equal to a predetermined threshold, it means that the video frame is clear; if the blur value is greater than or equal to the predetermined threshold, it means that the video frame is blurry.

[0011] Further, the training method for the document object detection network includes the following steps: Step S21: Collect videos of the process of switching between different document pages in a multi-page document, extract video frames sequentially from the videos, and manually annotate the document position and hand position in each extracted video frame; the extracted consecutive video frames with annotated document and hand positions are used as the training dataset for the document object detection network. Step S22: The document object detection network adopts an object detection model and is trained using the training dataset; during training, the document position and hand position in each video frame output by the document object detection network are required to be as close as possible to the annotated document and hand positions in that video frame; during training, classification loss, localization loss, and confidence loss are used as a hybrid loss function.

[0012] Further, the training method for the document content duplication detection network includes the following steps: Step S31: Collect multiple images with similar backgrounds but identical or different document content, and manually label the document category of each image; if the document content in the images is different, the document category labels are different; if the document content in the images is the same, the document category labels are the same; the images labeled with document category labels are used as the training dataset for the document content duplication detection network. Step S32: The document content duplication detection network adopts a Siamese network, and the training dataset is used to train the document content duplication detection network. During training, the determination result of whether the document content in any two images is the same is required to be as close as possible to the result of whether the document category labels of the two images are the same; during training, classification loss and prototype loss are used as a hybrid loss function.

[0013] Preferably, the camera used to capture the document page in step S4 is different from the camera used in step S1.

[0014] This application also proposes an automatic multi-page document capture device with duplicate content and stability detection, including a video capture unit, a target detection unit, a video frame state judgment unit, a document capture unit, a duplicate detection unit, and a post-processing unit. The video capture unit captures video of the multi-page document switching between different pages and extracts video frames sequentially from the video. The target detection unit feeds each extracted video frame into a trained document target detection network, which outputs the document position and hand position in each video frame. The document position in each video frame is outlined by a rectangle, and the trained document target detection network outputs the coordinates of the four vertices of this rectangle. The hand position in each video frame is outlined by another rectangle, and the trained document target detection network outputs the coordinates of the four vertices of this rectangle. The video frame state judgment unit determines whether each video frame is in a document page switching state or a stable state; it also determines whether the document in a video frame that is at least in a stable state has a large area of ​​hand obstruction; and it determines whether the video frame that is at least in a stable state is clear or blurry. A video frame that simultaneously meets the three conditions of being in a stable state, having no large area of ​​hand obstruction, and being clear is considered a stable and captureable state. The document capture unit is used to immediately trigger the capture of the current document page once a stable, captureable video frame appears after the document page state is switched. After capturing the current document page, it waits for the next video frame to switch document page states; this process is repeated until every page of a multi-page document is captured. The duplication detection unit is used to feed all captured images into a trained document content duplication detection network, which removes images with duplicate document content. The post-processing unit is used to perform edge correction and sharpness enhancement on the deduplicated images, and then assembles them sequentially to form an electronic document.

[0015] The technical effect achieved by this application is as follows: A deep learning-based object detection algorithm is used to identify and detect document and hand positions in video frames. The variance of vertex coordinates of the document position, the intersection ratio between the hand region and the document region, and the blur value of the document image are calculated to determine whether the current video frame is in a stable and captureable state. Then, a document duplication detection algorithm is applied to the captured images to filter out duplicate document content, resulting in non-duplicate multi-page document images. Finally, the multi-page document images undergo edge correction, sharpness enhancement, and other processing before being saved as electronic documents. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating the method for automatically capturing multi-page documents with duplicate content and stability detection as described in this application.

[0017] Figure 2 This is a flowchart illustrating the training method for a document object detection network.

[0018] Figure 3 This is a flowchart illustrating the training method for a document content duplication detection network.

[0019] Figure 4 This is a schematic diagram of the device for automatically capturing multi-page documents with duplicate content and stability detection according to this application.

[0020] The attached diagrams are labeled as follows: 1. Video capture unit; 2. Target detection unit; 3. Video frame status judgment unit; 4. Document capture unit; 5. Duplicate detection unit; 6. Post-processing unit. Detailed Implementation

[0021] Please see Figure 1 The method for automatically capturing multi-page documents with duplicate content and stability detection proposed in this application includes the following steps.

[0022] Step S1: Point the camera at the multi-page document and record a video of the process of switching between different pages. Extract video frames sequentially from the video. The multi-page document includes multiple unbound, discrete document pages, as well as bound multi-page documents (e.g., books). Switching between different document pages refers to the process of displaying different document pages one by one without obstruction. This step, for example, extracts one video frame every n video frames, where n is preferably an integer between 0 and 5. If n = 0, it indicates frame-by-frame extraction.

[0023] Step S2: Each extracted video frame is fed into the trained document object detection network, which outputs the document location and hand location in each video frame. The document location in each video frame is outlined by a rectangle, and the trained document object detection network outputs the coordinates of the four vertices of this rectangle. The hand location in each video frame is outlined by another rectangle, and the trained document object detection network outputs the coordinates of the four vertices of this rectangle. If different document pages are switched mechanically, the hand location is replaced by the location of the robotic arm (or other document page switching device).

[0024] Step S3: Calculate the variance of the coordinates of each vertex of the rectangle representing the document position in the extracted consecutive video frames, and determine whether each video frame is a switching document page state or a stable state.

[0025] This step requires the extraction of n consecutive video frames for calculation. The first n-1 video frames are all assumed to be in a document page switching state. For the nth video frame, the variance of the coordinates of the four vertices of the rectangle representing the document position from frame 1 to frame n is calculated, resulting in four variance values. If any one or more variance values ​​are greater than a predetermined threshold, it indicates a significant shift in the document position within the extracted n consecutive video frames, and the nth video frame is determined to be in a document page switching state. If all variance values ​​are less than the predetermined threshold, it indicates no significant shift in the document position within the extracted n consecutive video frames, and the nth video frame is determined to be in a stable state. For example, n is set to 10. The method for determining the nth video frame and subsequent video frames is the same as that for the nth video frame.

[0026] This step also calculates the document region area and hand region area in at least a stable video frame (or each video frame), then calculates the overlapping area of ​​the document region and hand region in the video frame, and finally calculates the overlapping area / document region area. If the overlapping area / document region area is less than or equal to a predetermined threshold, it indicates that there is no large-area hand occlusion in the document of the video frame. If the overlapping area / document region area is greater than the predetermined threshold, it indicates that there is a large-area hand occlusion in the document of the video frame.

[0027] This step also calculates the blur value for at least the stable video frames (or each video frame), for example, using a stable detection method employing the Laplacian operator. Preferably, the blur value of the video frame is calculated using the Laplacian function from the OpenCV library. This is a publicly available function that uses second-order derivatives to calculate edge information to reflect the degree of blur in the image. If the blur value is ≤ a predetermined threshold, the video frame is considered sharp. If the blur value is > a predetermined threshold, the video frame is considered blurry.

[0028] The predetermined thresholds in the above three judgment processes are independent of each other and are different. A video frame that simultaneously meets the three conditions of being in a stable state, having no large area of ​​hand obstruction, and being clear is considered to be in a stable and shootable state.

[0029] Step S4: After the video frame indicating "switching document page state," once a video frame indicating "stable and ready to be captured" appears, immediately trigger the capture of the current document page. Preferably, the device used to capture the document page in Step S4 is different from the camera device in Step S1. The camera device in Step S1 focuses more on video shooting performance, while the device used to capture the document page in Step S4 focuses more on high-definition photo shooting performance; the clarity of the video frame differs significantly from the resolution and clarity of the photo. Once the current document page is captured, wait for the next video frame indicating "switching document page state." Repeat this step until every page of a multi-page document has been captured. This mechanism ensures that the same document page will not be captured repeatedly in most application scenarios. However, if the camera device in Step S1 shakes excessively, or if the camera device in Step S1 moves significantly after capturing the document page, the same document page may be captured repeatedly. When the camera device in Step S1 is a mobile phone, and the user is holding the phone to record video, significant shaking or movement is likely to occur.

[0030] Step S5: Feed all captured images into the trained document content duplication detection network, which will remove images with duplicate document content.

[0031] Step S6: Perform edge correction and sharpness enhancement on the deduplicated images, and assemble them in sequence to form an electronic document.

[0032] Please see Figure 2 The training method for the document object detection network includes the following steps.

[0033] Step S21: Collect video footage of the process of switching between different document pages in a multi-page document. Extract video frames sequentially from the video, for example, extract one video frame every n video frames, where n is preferably an integer between 0 and 5. Manually label the document position and hand position in each extracted video frame. If a mechanical method is used to switch different document pages, the hand position is replaced by the position of a robotic arm (or other document page switching device). The document position in each video frame is enclosed by a rectangle, and the coordinates of the four vertices of the rectangle are labeled. The hand position in each video frame is enclosed by another rectangle, and the coordinates of the four vertices of the rectangle are labeled. The extracted consecutive video frames with labeled document and hand positions are used as the training dataset for the document object detection network.

[0034] Step S22: The document object detection network employs an object detection model, such as a YOLO series object detection model, preferably the small-sized YOLOv5n model because it is suitable for deployment on mobile devices. The document object detection network is trained using the aforementioned training dataset. During training, the document object detection network is required to output document and hand positions in each video frame as closely as possible to the labeled document and hand positions in that video frame.

[0035] During training, a hybrid loss function is used, comprising classification loss, localization loss, and confidence loss. The classification loss, for example, employs binary cross-entropy loss, which can be understood as a binary classification loss between two categories—document and hand. The localization loss, for example, employs complete intersection-union (IoU) loss. A core task of object detection is predicting the location of the bounding boxes for the document and hand. IoU loss considers the overlap of the bounding boxes, the distance between their centers, and their aspect ratio, and is currently the most widely used method for calculating IoU loss. The confidence loss measures the likelihood that an object exists within the predicted bounding box. The document object detection network derives its confidence loss by calculating the object confidence score within the predicted bounding box and the IoU value between the predicted bounding box and the corresponding ground truth box; for example, it also uses binary cross-entropy loss for calculation.

[0036] Please see Figure 3 The training method for the document content duplication detection network includes the following steps.

[0037] Step S31: Collect multiple images with similar backgrounds but identical or different document content. Manually label each image with a document category label. If the document content in the images is different, the document category labels are different. If the document content in the images is the same, the document category labels are the same. The images with labeled document category labels are used as the training dataset for the document content duplication detection network.

[0038] Step S32: The document content duplication detection network is implemented using a Siamese network, which is typically used to solve similarity comparison tasks. The document content duplication detection network is trained using the aforementioned training dataset. During training, the network is required to determine whether the document content of any two images is the same, and this determination should be as close as possible to the result of whether the document category labels of the two images are the same. During training, a hybrid loss function is used, consisting of classification loss and prototype loss. For example, binary cross-entropy loss is used for classification loss. For example, cosine similarity loss is used for prototype loss.

[0039] Please see Figure 4The device for automatically capturing multi-page documents with duplicate content and stability detection proposed in this application includes a video capturing unit 1, a target detection unit 2, a video frame state judgment unit 3, a document capturing unit 4, a duplicate detection unit 5, and a post-processing unit 6. Figure 4 The device shown corresponds to Figure 1 The method shown.

[0040] The video capturing unit 1 is used to capture video of switching between different document pages in a multi-page document, and to extract video frames sequentially from the video.

[0041] The target detection unit 2 is used to send each extracted video frame into the trained document target detection network, and the trained document target detection network outputs the document position and hand position in each video frame.

[0042] The video frame state judgment unit 3 is used to calculate the variance of the coordinates of each vertex of the rectangle representing the document position in the extracted consecutive video frames, and to determine whether each video frame is in a document page switching state or a stable state; it is also used to calculate the area of ​​the overlapping area between the document area and the hand area in the video frame that is at least in a stable state / the area of ​​the document area, and to determine whether there is a large area of ​​hand occlusion in the document in the video frame; it is also used to calculate the blur value of the video frame that is at least in a stable state; a video frame that simultaneously meets the three conditions of stable state, no large area of ​​hand occlusion, and clear is a stable and shootable state.

[0043] The document capturing unit 4 is used to immediately trigger the capture of the current document page once a video frame of "stable and ready to capture" appears after the video frame of "switching document page state". After capturing the current document page, it waits for the next video frame of "switching document page state" to appear. The above process is repeated until the capture of each document page of the multi-page document is completed.

[0044] The duplicate detection unit 5 is used to send all captured images into a trained document content duplicate detection network, which then removes images with duplicate document content.

[0045] The post-processing unit 6 is used to perform edge correction, sharpness enhancement, and other processing on the deduplicated images, and assemble them in sequence to form an electronic document.

[0046] Compared with the prior art, this application has the following technological innovations and beneficial effects.

[0047] First, this application proposes a logic for determining the document page switching state of video frames. During the initialization of continuous automatic shooting of multi-page documents, the first n-1 video frames are assumed to be in a document page switching state. The document location in each video frame is inferred using a document object detection model. The variance of the vertex coordinates of each document location is calculated over multiple consecutive video frames. Starting from the nth video frame onwards, if any variance value exceeds a preset threshold, the video frame is determined to be in a document page switching state.

[0048] Second, this application proposes a logic for determining a stable and captureable video frame. The logic is based on a comprehensive assessment of three factors: the variance of the vertex coordinates of the document position within multiple consecutive video frames, the ratio of the intersection area of ​​the hand region and the document region to the area of ​​the hand region, and the blur value of the document. If each factor is less than its respective threshold, the video frame is determined to be in a stable and captureable state.

[0049] Third, this application proposes a continuous automatic shooting logic for multi-page documents. Only when a stable, camera-ready video frame appears after a video frame switching document page states is the timing automatically determined as suitable for shooting the document page. After shooting the document page, the system waits for the next video frame switching document page states, avoiding repeatedly shooting the same document page when stable, camera-ready video frames appear consecutively.

[0050] Fourth, this application proposes a document content duplication detection logic. Document images captured by a document content duplication detection network are used to calculate the document content similarity differences between consecutively captured images. The deduplicated document images are then processed with edge correction and clarity enhancement to obtain the electronic document.

[0051] Fifth, compared with CN119781801A, this application adopts target detection technology, and the application scenarios cover book page turning and document (such as contract) page turning. It uses page stability, finger unobstructed, and page unblurred to make a comprehensive judgment on automatic shooting logic.

[0052] Sixth, existing devices for automatically capturing multi-page documents are usually bulky. This application supports deployment on mobile devices and covers various document scenarios, making it flexible and convenient to use.

[0053] The above are merely preferred embodiments of this application and are not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for automatically capturing multi-page documents with duplicate content and stability detection, characterized in that, Includes the following steps; Step S1: Capture a video of switching between different pages of a multi-page document, and extract video frames sequentially from the video; Step S2: Feed each extracted video frame into the trained document object detection network, and the trained document object detection network outputs the document position and hand position in each video frame; Step S3: Determine whether each video frame is in a document page switching state or a stable state; also determine whether there is a large area of ​​hand obstruction in the document of the video frame that is at least in a stable state; also determine whether the video frame that is at least in a stable state is clear or blurry; a video frame that meets all three conditions of being in a stable state, having no large area of ​​hand obstruction, and being clear is a stable state that can be filmed. Step S4: After the video frame that switches the document page state, once a stable video frame that can be captured appears, immediately trigger the capture of the current document page; once the current document page is captured, wait for the next video frame that switches the document page state; repeat this step until the capture of each document page of the multi-page document is completed; Step S5: Feed all captured images into the trained document content duplication detection network, which will remove images with duplicate document content. Step S6: After removing duplicates, the images are cropped and their clarity is enhanced. They are then assembled in sequence to form an electronic document.

2. The method for automatically capturing multi-page documents with duplicate content and stability detection according to claim 1, characterized in that, In step S1, one video frame is extracted every n video frames, where n is an integer between 0 and 5.

3. The method for automatically capturing multi-page documents with duplicate content and stability detection according to claim 1, characterized in that, In step S2, the document position in each video frame is outlined by a rectangle, and the trained document object detection network outputs the coordinates of the four vertices of the rectangle; the hand position in each video frame is outlined by another rectangle, and the trained document object detection network outputs the coordinates of the four vertices of the rectangle. If different document pages are switched mechanically, the hand position is changed to the position of the mechanical device used to switch the document pages.

4. The method for automatically capturing multi-page documents with duplicate content and stability detection according to claim 1, characterized in that, In step S3, determining whether each video frame is in a document page switching state or a stable state is calculated based on the extracted n consecutive video frames; for the first n-1 video frames, they are all determined to be in a document page switching state; for the nth and subsequent video frames, the variance of the coordinates of the four vertices of the rectangle representing the document position in the n consecutive video frames ending with the current video frame is calculated; if any one or more variance values ​​are greater than a predetermined threshold, the current video frame is determined to be in a document page switching state. If all variance values ​​are less than a predetermined threshold, the current video frame is determined to be in a stable state.

5. The method for automatically capturing multi-page documents with duplicate content and stability detection according to claim 1, characterized in that, In step S3, the area of ​​the document region and the area of ​​the hand region in the video frame that is at least in a stable state are calculated, and then the area of ​​the overlapping region between the document region and the hand region in the video frame is calculated, and then the overlapping region area / document region area is calculated; if the overlapping region area / document region area ≤ a predetermined threshold, it means that there is no large area of ​​hand occlusion in the document in the video frame; if the overlapping region area / document region area > the predetermined threshold, it means that there is a large area of ​​hand occlusion in the document in the video frame.

6. The method for automatically capturing multi-page documents with duplicate content and stability detection according to claim 1, characterized in that, In step S3, the Laplacian function of the OpenCV library is used to calculate the blur value of the video frame that is at least in a stable state; if the blur value is less than or equal to a predetermined threshold, the video frame is clear; if the blur value is greater than or equal to the predetermined threshold, the video frame is blurry.

7. The method for automatically capturing multi-page documents with duplicate content and stability detection according to claim 1, characterized in that, The training method for the document object detection network includes the following steps; Step S21: Collect video of the process of switching between different document pages in a multi-page document, extract video frames sequentially from the video, and manually annotate the document position and hand position in each extracted video frame; the extracted consecutive video frames with annotated document position and hand position are used as the training dataset for the document object detection network. Step S22: The document object detection network adopts an object detection model and is trained using the training dataset. During training, the document object detection network is required to make the document position and hand position in each video frame output by the document object detection network as close as possible to the document position and hand position marked in the video frame. During training, classification loss, localization loss and confidence loss are used as a hybrid loss function.

8. The method for automatically capturing multi-page documents with duplicate content and stability detection according to claim 1, characterized in that, The training method for the document content duplication detection network includes the following steps; Step S31: Collect multiple images with similar backgrounds but identical or different document content. Manually label the document category of each image. If the document content in the images is different, the document category labels are different. If the document content in the images is the same, the document category labels are the same. The images labeled with document category labels are used as the training dataset for the document content duplication detection network. Step S32: The document content duplication detection network adopts a Siamese network. The document content duplication detection network is trained using the training dataset. During training, it is required that the determination result of whether the document content in any two images is the same should be as close as possible to the result of whether the document category labels of the two images are the same. During training, classification loss and prototype loss are used as a hybrid loss function.

9. The method for automatically capturing multi-page documents with duplicate content and stability detection according to claim 1, characterized in that, The camera used to photograph the document page in step S4 is different from the camera used in step S1.

10. A device for automatically capturing multi-page documents with duplicate content and stability detection, characterized in that, It includes a video capture unit, a target detection unit, a video frame state judgment unit, a document capture unit, a duplicate detection unit, and a post-processing unit; The video capturing unit is used to capture video of switching between different pages of a multi-page document, and to extract video frames sequentially from the video. The target detection unit is used to send each extracted video frame into the trained document target detection network, and the trained document target detection network outputs the document position and hand position in each video frame; The video frame state determination unit is used to determine whether each video frame is in a document page switching state or a stable state; it is also used to determine whether there is a large area of ​​hand obstruction in the document of a video frame that is at least in a stable state; it is also used to determine whether a video frame that is at least in a stable state is clear or blurry; a video frame that simultaneously meets the three conditions of stable state, no large area of ​​hand obstruction, and clear is a stable state that can be filmed. The document capturing unit is used to immediately trigger the capture of the current document page once a stable video frame appears after the video frame of the document page state is switched; after capturing the current document page, it waits for the next video frame of the document page state to appear; repeat the above process until the capture of each document page of the multi-page document is completed. The duplicate detection unit is used to send all captured images into a trained document content duplicate detection network, which then removes images with duplicate document content. The post-processing unit is used to perform edge correction and clarity enhancement on the deduplicated images, and assemble them in sequence to form an electronic document.

Citation Information

Patent Citations

  • A method and device for searching useless files

    CN119781801A