Document - type Video Image Processing Method, Electronic Device and Computer Storage Medium

By extracting key video frames in document-like video clips and performing document positioning and correction, the problem of insufficient document content in document-like videos is solved, and high-quality document content extraction and editable document file generation are achieved.

CN117456406BActive Publication Date: 2025-06-20UCWEB
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311302617.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-09
Publication Date
2025-06-20
Estimated Expiration
2043-10-09

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively extract pure document content from document-like videos, because markings, paintings, rewriting, etc. during the explanation process will be recorded during the recording process, resulting in the intercepted document content not being clean enough and it is difficult to separate from the content in other situations.

Method used

By obtaining multiple document-type video clips, each clip is a single document content video clip, key video frames that meet the image quality standards are extracted, document positioning is performed to determine the image area where the document content is located, and correct the area to obtain a pure document content image, and finally generate an editable document file.

Benefits of technology

It effectively removes non-documentary content such as marks and paintings in the document content, ensures the purity and clarity of the document content, improves the quality of the document content images, and makes the generated editable document files clean and clear, with high file quality, which is convenient for subsequent use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117456406B_ABST
    Figure CN117456406B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a method for processing document - type video images, an electronic device, and a computer storage medium. The method for processing document - type video images includes: obtaining a plurality of document - type video segments, where each document - type video segment is a single - document content video segment; extracting key video frames that meet the image quality standard from each document - type video segment; performing document localization on the key video frames of each extracted document - type video segment to determine the image area where the document content is located; correcting the determined image area to obtain a pure document content image, and generating an editable document file based on the pure document content image. Through the embodiment of the present application, the quality of the document content image is improved, and the editable document file generated based on this pure document content image is clean and clear, has a high file quality, and is convenient for subsequent use.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of video technology, and in particular, to a method for processing document video images, an electronic device, and a computer storage medium. Background Art

[0002] Document videos refer to videos whose main content is document content (such as PPT, WORD, PDF, etc.). For example, in learning scenarios, training scenarios, meeting scenarios, etc., document videos that contain the entire process of an explainer explaining based on the document content are often formed.

[0003] For many users, there is often a need to obtain the document content from document videos for further use. In an existing method, by comparing the image content of video frames, video frames with different document contents are found, and then the image regions where the document content is located in these video frames are intercepted and spliced into a document file. However, for such videos, during their recording process, not only the document content is recorded, but also many other situations during the explanation process, such as the explainer's markings, scribbles, rewrites, occlusions, etc. on the document content, are recorded into the video together. Therefore, using the above method of intercepting document content is likely to result in the intercepted document content not being clean enough, that is, not only containing the document content, and it is not easy to separate from the content generated by the above other situations, thus making the content of the spliced document file not clean and clear, which is not convenient for subsequent use. Summary of the Invention

[0004] In view of this, the embodiments of the present application provide a solution for processing document video images to at least partially solve the above problems.

[0005] According to the first aspect of the embodiments of the present application, a method for processing document video images is provided, including: obtaining a plurality of document video segments, where each document video segment is a single-document content video segment; extracting key video frames that meet the image quality standard from each document video segment; performing document localization on the key video frames of each extracted document video segment to determine the image region where the document content is located; correcting the determined image region to obtain a pure document content image, and generating an editable document file based on the pure document content image.

[0006] According to the second aspect of the embodiments of the present application, an electronic device is provided, including: a processor, a memory, a communication interface, and a communication bus, where the processor, the memory, and the communication interface complete communication with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the method in the first aspect.

[0007] According to a third aspect of the embodiments of the present application, there is provided a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in the first aspect is implemented.

[0008] According to the solution provided by the embodiments of the present application, first, each of the multiple document-class video segments is a single-document content video segment, that is, each video segment corresponds to only one piece of document content. Thus, the subsequent extraction of video frames for the document content is more targeted and it is easier to extract the current document content. Second, when extracting key video frames for a piece of document content, the extracted key video frames need to meet certain image quality standards so that the subsequent extracted document content can be clearer and have a higher imaging quality. Furthermore, after the key video frames are extracted, a document positioning operation for the document content is performed. Thus, other parts that may affect the document content, such as portraits, or human hands, or chat boxes, etc., which are non-document content, will be effectively removed, that is, the interference data affecting the document content can be removed. In addition, for the image area where the document content part located in each key video frame is located, image correction for obtaining pure document content will also be performed, including but not limited to removing marks, scribbles, etc. on the document content. Thus, the non-document content in the image can be effectively removed, ensuring the purity, cleanliness, and clarity of the document in the obtained image, greatly improving the quality of the document content image, and also making the editable document file generated based on this pure document content image clean, clear, and have a high file quality, which is convenient for subsequent use. Description of the Drawings

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0010] Figure 1 Schematic diagram of an exemplary system for applying the document-class video image processing solution of the embodiments of the present application;

[0011] Figure 2 Flowchart of the steps of a document-class video image processing method according to an embodiment of the present application;

[0012] Figure 3 For Figure 2 Schematic diagram of a first quality type video frame in the shown embodiment;

[0013] Figure 4 For Figure 2 Schematic diagram of a second quality type video frame in the shown embodiment;

[0014] Figure 5 It is Figure 2 a schematic diagram of a third quality type video frame in the illustrated embodiment;

[0015] Figure 6 It is Figure 2 a schematic diagram of a fourth quality type video frame in the illustrated embodiment;

[0016] Figure 7 a schematic structural diagram of an electronic device according to an embodiment of the present application. Specific embodiments

[0017] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art shall fall within the protection scope of the embodiments of the present application.

[0018] The following further illustrates the specific implementation of the embodiments of the present application in conjunction with the accompanying drawings of the embodiments of the present application.

[0019] Figure 1 An exemplary system applicable to the solution of the embodiment of the present application is shown. As Figure 1 shown, the system 100 may include a cloud server 102, a communication network 104, and / or one or more user devices 106, Figure 1 and multiple user devices are exemplified herein. It should be noted that the solution of the embodiment of the present application can be completed by the cooperation of the cloud server 102 and the user device 106, or can be completed independently by the user device 106.

[0020] The cloud server 102 can be any suitable device for storing information, data, programs, and / or any other suitable type of content, including but not limited to distributed storage system devices, server clusters, computing cloud server clusters, etc. In some embodiments, the cloud server 102 can perform any suitable function. For example, when the cloud server 102 and the user device 106 cooperate to complete the solution of the embodiments of the present application, in some embodiments, the cloud server 102 can be used to generate a corresponding editable document file based on a document-like video. As an alternative example, in some embodiments, the cloud server 102 can first obtain multiple document-like video segments from the document-like video, where one document-like video segment corresponds to one document content, for example, one PPT page, etc.; then, key video frames with image quality meeting the image quality standard can be extracted from these document-like video segments; then, document localization is performed on the key video frames to determine the image area where the document content is located; then, the image area is corrected to obtain a pure document content image, that is, an image that does not contain situations such as human figures, human hands, chat boxes, scribbles, occlusions, etc., but only contains document content; then, an editable document file is generated based on the pure document content image. As another example, in some embodiments, the cloud server 102 can send the generated editable document file to the user device 106.

[0021] In some embodiments, the communication network 104 can be any suitable combination of one or more wired and / or wireless networks. For example, the communication network 104 can include any one or more of the following: the Internet, an intranet, a Wide Area Network (WAN), a Local Area Network (LAN), a wireless network, a Digital Subscriber Line (DSL) network, a Frame Relay network, an Asynchronous Transfer Mode (ATM) network, a Virtual Private Network (VPN), and / or any other suitable communication network. The user device 106 can be connected to the communication network 104 through one or more communication links (for example, the communication link 112), and the communication network 104 can be linked to the cloud server 102 through one or more communication links (for example, the communication link 114). The communication link can be any communication link suitable for transmitting data between the user device 106 and the cloud server 102, such as a network link, a dial-up link, a wireless link, a hardwired link, any other suitable communication link, or any suitable combination of such links.

[0022] The user device 106 may include any one or more user devices suitable for playing videos, presenting images, interacting with users, etc. When the solution of the embodiments of the present application is independently completed by the user device 106, the user device 106 can be used to generate a corresponding editable document file based on the document-based video. As an optional example, in some embodiments, the user device 106 may first obtain multiple document-based video segments from the document-based video, where one document-based video segment corresponds to one document content, for example, one PPT page, etc.; then, key video frames with image quality meeting the image quality standard can be extracted from these document-based video segments; then, document localization is performed on the key video frames to determine the image area where the document content is located; then, the image area is corrected to obtain a pure document content image, that is, an image that does not contain situations such as human figures, human hands, chat boxes, scribbles, occlusions, etc., but only contains the document content; then, an editable document file is generated based on the pure document content image. In some embodiments, the user device 106 may include any suitable type of device. For example, in some embodiments, the user device 106 may include a mobile device, a tablet computer, a laptop computer, a desktop computer, a game console, a media player, a vehicle system, and / or any other suitable type of user device.

[0023] Based on the above system, the document-based video image processing solution of the present application will be described below through multiple embodiments.

[0024] Refer to Figure 2 , which shows a flowchart of the steps of a document-based video image processing method according to an embodiment of the present application.

[0025] The document-based video image processing method of this embodiment includes the following steps:

[0026] Step S202: Obtain multiple document-based video segments.

[0027] In the embodiments of the present application, the document-based video segments are from the document-based video. However, different from traditional video segments, in the embodiments of the present application, each document-based video segment is a single-document content video segment, that is, multiple video frames in each document-based video segment correspond to the same document content, such as corresponding to the same page of PPT, or corresponding to the same WORD page or PDF page, or corresponding to the same part of the same WORD page or the same part of the same PDF page. Thus, the pertinence of subsequent document content processing can be effectively improved, and the efficiency of extracting document content can be improved.

[0028] The specific implementation of obtaining such video segments can be achieved by those skilled in the art in any appropriate manner according to actual needs. However, in order to improve the acquisition speed and efficiency, in a feasible manner, this step can be implemented as follows: dividing the document-type video into multiple segments in chronological order, performing parallel video frame extraction processing on each segment to obtain multiple candidate segments; determining at least one group of video frames with the same document content from each candidate segment; and obtaining multiple document-type video segments based on the multiple groups of video frames corresponding to the multiple candidate segments.

[0029] For a certain document-type video, it is first necessary to decompose it into a series of video frames in chronological order, that is, it is necessary to perform video frame extraction processing, also known as frame extraction. Exemplarily, FFmpeg can be used for video frame extraction. FFmpeg is an audio and video processing tool that includes numerous audio and video processing libraries and command-line tools. Among them is a frame extraction tool, through which a certain frame image can be extracted from the document-type video at a certain time interval for subsequent processing. However, when using the traditional frame extraction method, the speed and efficiency of frame extraction are not high. In order to improve the processing speed and efficiency of frame extraction, the embodiments of this application perform the following processing: (1) dividing the document-type video into multiple segments in chronological order, and using one CPU core for each segment to perform parallel frame extraction processing; (2) for larger document-type videos whose file size exceeds a preset size threshold (such as those with too long a duration, too high a resolution, etc.), reducing the frame extraction frequency. Among them, the preset size threshold can be flexibly set by those skilled in the art according to actual needs, and the embodiments of this application do not limit this, for example, 1G, etc. The specific implementation of reducing the frame extraction frequency can also be flexibly set by those skilled in the art, such as setting it to 2 times the original frame extraction frequency, etc. Thus, the document-type video can complete the frame extraction processing within a short time, such as 30 seconds.

[0030] For the sake of convenience of explanation, in a simple example, assume that a certain document-type video is divided into 4 segments A, B, C, and D in chronological order. For these 4 segments, parallel frame extraction processing is respectively performed to obtain the corresponding 4 candidate segments A', B', C', and D'. The difference between A', B', C', and D' and A, B, C, and D is that A' is composed of the video frames extracted from A, that is, part of the video frames in A. Similarly, B', C', and D' are also correspondingly composed of part of the video frames in B, C, and D.

[0031] Furthermore, for the candidate segments of A', B', C', and D', one or more groups of video frames with the same document content are respectively determined from each segment. For example, for A', a group of video frames is determined for it, and this group of video frames corresponds to the first page of the document content such as a PPT; for another example, for B', two groups of video frames are determined for it, one group of video frames corresponds to the second page of the document content such as a PPT, and the other group of video frames corresponds to the third page of the document content such as a PPT. Similarly, one or more groups of video frames are also determined for C' and D' respectively. Among them, each group of video frames can be regarded as a document-type video segment in step S202.

[0032] Among them, in a feasible manner, determining at least one group of video frames with the same document content from each candidate segment can be implemented as follows: for each candidate segment, taking two adjacent video frames in the candidate segment as a unit, fusing the two video frames in the channel dimension; performing semantic extraction processing based on the fused video frames, and classifying whether the document content of the two video frames is the same according to the semantic extraction processing result; according to the classification results corresponding to multiple pairs of adjacent video frames included in the candidate segment, the candidate segment is segmented into at least one group of video frames with the same document content. In this way, the video frames in each candidate segment can be decomposed into consecutive and non-overlapping video segments in chronological order, and each video segment only involves at most one document content, such as the content of one PPT page. At the same time, through the way of channel fusion, the local information of the video frame image can be not lost; and through semantic extraction, the interference of occlusion to the document content can be eliminated to a certain extent.

[0033] Since the traditional video segmentation algorithm based on global features will lose a large amount of local information and is not suitable for the document content extraction scenario, it is easy to cause a large number of documents with similar global layouts but completely different contents to be divided into the same video segment. And the traditional algorithm based on optical character recognition (OCR) only performs comparison based on text and cannot effectively handle problems such as human occlusion, which will cause a large number of document contents with the same but different occluded areas of people to be divided into different video segments. Based on this, in the embodiments of the present application, a video segmentation model is used, and in this model, a "shallow information fusion" method is adopted, and the adjacent front and rear two video frames are fused in the channel dimension at the video frame image input end to ensure that local information is not lost. At the same time, based on the semantic extraction ability of the deep neural network, the ability to ignore interference such as human occlusion is obtained in the high-dimensional semantic space. Then, the video segmentation model outputs the classification result of "whether the document content in the two video frame images is the same". According to this classification result, each candidate segment can be segmented into several consecutive and non-overlapping groups of video frames, that is, document-type video segments. Exemplarily, the video segmentation model can be implemented in the form of a convolutional neural network model.

[0034] In the training stage of the video segmentation model, a large number of paired training samples can be used, and these training samples come from a large number of different types of videos to enhance the generalization performance of the video segmentation model. However, a large amount of data in a video may be negative samples (for example, two video frame images belong to the same document content, such as belonging to the same PPT), while the proportion of positive samples is very low (for example, two video frame images belong to different document contents, such as belonging to different PPTs). For example, in a two-hour video, there may be only 40-50 images of different document contents, such as different PPT images. Therefore, the cost of manually annotating a large number of videos is too high, and it is difficult to obtain difficult positive and negative samples. For this reason, in the embodiments of the present application, a training sample acquisition scheme based on OCR and semi-supervised learning is adopted, including the following processes:

[0035] (1) Manually select videos of different document types (such as PPT, PDF, WORD, etc.) (a small amount, such as more than a dozen videos) for manual annotation, and call this batch of training samples cold start training samples;

[0036] (2) Train the video segmentation model based on the cold start training samples. And after the training is completed, perform inference on the videos of different document types manually selected (medium scale, such as more than a hundred videos), and manually select several videos with poor performance for manual annotation to form re-annotated training samples;

[0037] (3) Retrain the video segmentation model based on the two batches of training samples, namely the cold start training samples in (1) above and the re-annotated training samples in (2). After the training is completed, combined with the OCR scheme, perform inference on a large-scale document video (such as tens of thousands of videos). Extract the training sample pairs where the video segmentation model outputs "the same document content" but the OCR result shows "different document contents", and perform manual annotation;

[0038] (4) Retrain the video segmentation model based on the three batches of training samples, namely the cold start training samples in (1), the re-annotated training samples in (2), and the samples manually annotated in (3). After the training is completed, use the video segmentation model to perform inference on a large-scale video (about tens of thousands of videos), and extract the training samples with relatively high uncertainty output by the video segmentation model, and then perform manual annotation. Among them, if the video segmentation model outputs a score < 0.1, it means that two video frame images belong to the same document content; if the video segmentation model outputs a score > 0.9, it means that two video frame images belong to different document contents; and if the video segmentation model outputs data between 0.1 and 0.9, it means that the two video frame images are images with relatively high uncertainty, that is, the above-mentioned training samples with relatively high uncertainty.

[0039] Based on the above process, sufficient training samples required for training the video segmentation model can be obtained in a relatively short time with less annotation manpower, enabling the performance of the video segmentation model to reach the required level for effective classification of whether document contents are the same.

[0040] It can be seen that through this step, multiple document - type video segments can be effectively extracted from the document - type video, and each document - type video segment corresponds to a single document content.

[0041] It should be noted that in the embodiments of this application, unless otherwise specified, quantities related to "multiple", such as "multiple", "multiple types", "multiple sheets", etc., all mean two or more.

[0042] Step S204: Extract key video frames that meet the image quality standard from each document - type video segment.

[0043] Among them, the image quality standard can be flexibly set by those skilled in the art according to actual needs. Exemplarily, it can be set according to some or all of image brightness, blurriness, contrast, etc. Since each document - type video segment corresponds to a single document content, through this step, key video frames with good image quality and presentation effects that can effectively represent the document content can be extracted from each document - type video segment. It should be noted that generally, there may be one key video frame for a document - type video segment, but this is not limited thereto. There may also be multiple key video frames in a document - type video segment. For the case of multiple key video frames, the solutions of the embodiments of this application are also applicable. For example, processes such as merging and denoising of multiple key video frames can be performed to obtain the final key video frame; or, in the subsequent document localization and image region correction processes, each key video frame is processed in this way, and finally the corrected image regions are merged and denoised to obtain the final pure document content image, etc., all of which are within the protection scope of the embodiments of this application.

[0044] However, in order to make the obtained key video frames more accurate and effective, and at the same time make the image quality standard more objective, in a feasible way, multiple video frames included in each document-class video segment can be classified by quality to obtain the quality types corresponding to the multiple video frames respectively. Among them, the quality types include: a first quality type for indicating that the video frame is pure document content, a second quality type for indicating that the video frame does not contain document content, a third quality type for indicating that the document content in the video frame is dynamic document content, and a fourth quality type for indicating that the document content in the video frame has been painted; according to the quality types corresponding to the multiple video frames respectively, key video frames that meet the image quality standard are extracted therefrom. By classifying the image quality into these quality types, the document content situation in each video frame can be effectively distinguished, and the efficiency of extracting key video frames can be improved.

[0045] Exemplarily, a video frame of the first quality type is as Figure 3 shown. As can be seen from Figure 3 , the video frame contains the original document content and there are no situations such as occlusion or painting. A video frame of the second quality type is as Figure 4 shown. As can be seen from Figure 4 , the text in the video frame is handwritten and is not the original document content. A video frame of the third quality type is as Figure 5 shown. As can be seen from Figure 5 , there are foreground text and background text in the video frame. Among them, the foreground text gradually appears on top of the background text in an animated way. Figure 5 In , the background text is shown as gray font, and the foreground text is shown as white font. A video frame of the fourth quality type is as Figure 6 shown. As can be seen from Figure 6 , the document content in the video frame has been painted. It can be seen that through different quality types, the situations of video frames can be effectively distinguished, which is convenient for subsequent processing.

[0046] On this basis, extracting key video frames that meet the image quality standard according to the quality types corresponding to multiple video frames can be implemented as follows: from multiple video frames, exclude video frames of the second quality type, and exclude invalid frames that cannot be key video frames among video frames of the first quality type, the third quality type, or the fourth quality type; from the remaining video frames of the multiple video frames, use the video frame with the highest similarity to the video frame of pure document content as the key video frame that meets the image quality standard. Video frames of the second quality type are video frames that do not contain document content. Therefore, they can be excluded first. However, for video frames of the first, third, and fourth quality types, although they contain document content, there may be situations such as brightness and blurriness that cause the video frames to not effectively express the document content. Such video frames that cannot effectively express the document content are regarded as invalid frames and also need to be excluded. Furthermore, among the remaining video frames, select the video that is closest to the clean document content as the key video frame. Thus, the efficiency of key video frame extraction is further improved.

[0047] Exemplarily, excluding invalid frames that cannot be key video frames among video frames of the first quality type, the third quality type, or the fourth quality type includes: for each frame among video frames of the first quality type, the third quality type, or the fourth quality type, perform at least one of the following processes: if the brightness difference between the current video frame and other video frames in the video frame is greater than a preset brightness threshold, then confirm the current video frame as an invalid frame; if the blurriness between the current video frame and other video frames in the video frame is greater than a preset blurriness threshold, then confirm the current video frame as an invalid frame; if the number of video frames is greater than a preset number threshold, then confirm the first video frame and the last video frame in the video frame as invalid frames. Among them, the preset brightness threshold and the preset blurriness threshold can both be set by those skilled in the art according to actual needs. For example, the preset brightness threshold can be set to 50, the preset blurriness threshold can be set to 0.2, and so on. The embodiments of the present application do not limit this. In addition, the preset number threshold can also be set by those skilled in the art according to actual needs. For example, it can be 5 frames, etc.

[0048] In a specific example, video frames of the first quality type can be considered as video frames with clean document content, i.e., pure document content. Video frames of the second quality type can be considered as video frames without document content, i.e., non-document content. Video frames of the third quality type can be considered as video frames containing animated document content, i.e., dynamic document content. The fourth quality type can be considered as video frames with dirty document content, i.e., document content that has been scribbled on. In one feasible way of performing quality classification, quality classification can be carried out through a quality model. In the training phase, several video frames can be extracted from each document-class video among tens of thousands of document-class videos, and labeled manually according to the definitions of the above four quality types. Then, a classification network can be built based on any deep neural network (such as a residual neural network, etc.), using the data labeled above as training samples and a cross-entropy loss function for training. The quality model after training can then select key video frames with higher quality from each document-class video segment.

[0049] However, for some video frames, for example, video frames of the third quality type containing animated document content, since such video frames are intermediate frames of the animation effect and are diverse, it is difficult for training samples to cover all possible animation effects. Therefore, based on the trained quality model, further processing can be carried out to exclude invalid frames. This processing can include:

[0050] (1), if the current document-class video segment contains only one video frame with document content, then delete this video segment and do not output it;

[0051] (2), within the same document-class video segment, if a certain video frame has an obvious brightness difference from other video frames, that is, the brightness difference between this video frame and other video frames is greater than a preset brightness threshold, then this video frame is an invalid frame and is not selected as a key video frame;

[0052] (3), within the same document-class video segment, use a blur detection operator to detect the blur degree of each video frame. If a certain video frame is significantly blurrier than other video frames, that is, the blur degree difference between this video frame and other video frames is greater than a preset blur threshold, then this video frame is an invalid frame and is not selected as a key video frame;

[0053] (4), if a document-class video segment contains more than a certain number of video frames with document content, such as more than 5 PPT video frames, then the first video frame and the last video frame of this video segment are invalid frames and are not selected as key video frames of this video segment;

[0054] (5) For the remaining video frames excluding those in (1), (2), (3), and (4), according to the classification results of the quality model, select the video frames that are closest to "clean document content" (the first quality type) as key video frames.

[0055] Thus, the selected key video frames not only have high image quality but also can effectively represent the document content.

[0056] Step S206: Perform document localization on the key video frames of each extracted document-class video segment to determine the image area where the document content is located.

[0057] Through this step, the document content in each extracted key video frame can be localized, excluding the interference of non-document content such as portraits and chat boxes for subsequent correction.

[0058] In a feasible manner, document localization can be achieved through a detection model. Exemplarily, this detection model can be implemented using the Retina-Net detector architecture and the Swin-Transformer-Tiny backbone network. Retina-Net is a single-stage object detection model that can well detect dense and small-scale objects. Swin-Transformer is a deep learning model based on Transformer that can efficiently and accurately perform object detection in visual tasks. Swin-Transformer has four different scales, namely: Tiny (micro), Small (small), Base (basic), and Larger (large). Compared with other scales, the Tiny structure is more concise and has fewer model parameters.

[0059] In the training stage, based on the above video segmentation model and key video frame extraction algorithm, video frames can be extracted from tens of thousands of document-class videos and manually annotated. Use the annotated video frames as training samples to train the detection model. The Focal Loss function can be used to handle the problem of unbalanced positive and negative samples during training. The trained detection model can then well perform document content localization processing on the video frames.

[0060] Based on the localization of the document content, the position and image area of the document content in the video frame image can be determined.

[0061] Step S208: Correct the determined image area to obtain a pure document content image, and generate an editable document file based on the pure document content image.

[0062] By correcting the image area, non-document content in the image can be effectively removed, ensuring the purity, cleanliness, and clarity of the document in the obtained image.

[0063] In a feasible manner, correcting the determined image region may include at least one of the following: (1) performing sharpness enhancement processing on the determined image region; (2) determining whether there is an occluded document content region in the determined image region; if so, performing a first document content restoration process for removing occlusion on the image region; (3) determining whether there is a scribbled document content region in the determined image region; if so, performing a second document content restoration process for removing scribbles on the image region; (4) performing image normalization processing on the determined image region.

[0064] Among them, performing sharpness enhancement processing on the determined image region can be implemented as: performing contrast stretching processing on the determined image region to increase the contrast between the text and the background in the image region; and / or, performing high-frequency component enhancement and low-frequency component suppression processing on the determined image region. Through the above processing, the clarity of the image region corresponding to the document content can be made higher, or the image can be changed from blurred to high-definition.

[0065] Exemplarily, the CLAHE algorithm can be used first to stretch the contrast of the image region, making the text in the document content more prominent compared to the background; then, the high-frequency components of the image region are enhanced, and the low-frequency components of the image region are suppressed, making the text edges in the document content more prominent; the SwinIR deep neural network model is used to perform noise reduction processing on the image obtained in the above two steps to obtain the final output. Among them, the CLAHE algorithm is a histogram equalization algorithm, which can suppress noise while enhancing the image contrast.

[0066] The first document content restoration process for removing occlusion on the image region may include: determining the document class video segment to which the key video frame where the image region is located belongs; for the multiple video frames included in the determined document class video segment, performing foreground segmentation respectively to extract multiple document content image parts corresponding to the multiple video frames respectively; synthesizing the multiple document content image parts to generate a synthesized document content image; and performing denoising processing on the synthesized document content image. Through the above processing, the occluded region in the image region can be effectively restored, and clear document content can be obtained.

[0067] In a video, the occlusion of document content is often caused by the movement of the narrator, the waving of the arm, etc. However, in a video, a certain region in the document content cannot always be occluded. Therefore, the temporal information of the video can be used to restore the occluded region. Exemplarily, an image segmentation model can be used to implement the restoration of the occluded region, and the process may include:

[0068] (1) Train an image segmentation model in a conventional manner. Use the trained image segmentation model to perform foreground segmentation on video frames or video frames that have been standardized (cropped, size standardized, etc.). In this example, the key video frames are used to extract the pixels belonging to the document content.

[0069] (2) According to the foreground segmentation results of each video frame in the document - type video segment to which the key video frame in (1) belongs, synthesize the pixels of the document content of all the video frames in this document - type video segment, including the key video frame in (1), into the same new image, that is, the synthesized document content image.

[0070] (3) Based on the denoising model trained by SwinIR, perform denoising processing on the synthesized new image, that is, the synthesized document content image, to solve the ghosting problem caused by synthesis. Among them, SwinIR consists of three modules: shallow feature extraction, deep feature extraction, and high - quality image reconstruction module. The shallow feature extraction module uses convolutional layers to extract shallow features and directly transmits them to the reconstruction module to retain low - frequency information; the deep feature extraction module is mainly composed of the remaining Swin Transformer blocks (RSTB). Each Transformer block uses multiple Swin Transformers for local attention and cross - window interaction; and a convolutional layer is added at the end of the block for feature enhancement, and a residual connection is used to provide a shortcut for feature aggregation. Finally, shallow and deep features are fused in the reconstruction module for high - quality image reconstruction.

[0071] As described above, the occluded document content in the video frame can be effectively restored.

[0072] The second document content restoration process for removing scribbles from the image area can be implemented as follows: Use a denoising model for removing scribbles to perform the second document content restoration process on the image area. Among them, the denoising model is trained based on denoising training samples. The denoising training sample pair includes a pure document content sample image and a scribble sample image corresponding to the pure document content sample image. The scribble sample image is generated by synthesizing a scribble mask onto the pure document content sample image. In this way, the added content such as temporary scribbles, writing, etc. in the document content can be removed to restore the original document content.

[0073] When the narrator is explaining, they often make on - the - spot scribbles, writing, drawing, etc. on the document content, which affects the user experience of the recorded video. Therefore, in this example, these scribbles, writing, drawing, etc. (collectively referred to as scribbles in the embodiments of this application) can be restored, including:

[0074] (1) Use Retina-Net and Swin-Transformer-Tiny to build a detector network to accurately locate the scribbled area;

[0075] (2) Manually perform mask annotation on a number of "scribbles" and synthesize them onto an image of pure document content (such as a clean PPT image) (the pure document content image can be selected through a quality model), and construct a large number of (clean, scribbled) image pairs as training sample pairs;

[0076] (3) Use SwinIR to train a denoising model based on the above (clean, scribbled) training sample pairs, and then use the trained denoising model to restore the image area with scribbled document content in the key video frames to an image area of pure document content.

[0077] For the determined image area, image normalization processing can be performed. Based on the position information of the document content obtained by document positioning, operations such as cropping, magnifying to the standard size, automatic correction, and mirror image flipping can be performed to form a standardized image. It should be noted that this image normalization processing can be performed before the image sharpness enhancement processing, the first document content restoration processing, and the second document content restoration processing to improve the efficiency of these processes.

[0078] Through the above correction processing, a high-quality pure document content image can be obtained. Furthermore, based on this pure document content image, an editable document file can be generated, including but not limited to: PPT files, PDF files, WORD files, etc.

[0079] It should be noted that in addition to the document-based videos formed by explaining PPT files, there are also document-based videos formed by explaining PDF files and WORD files. Different from PPT files, PDF files and WORD files are usually scrolled from top to bottom for playback, rather than "page-turning" playback. Therefore, the document content in the two adjacent video frames is no longer a binary relationship of "the same PPT or different PPTs", but a relationship of "partially the same and partially different". In view of this situation, the embodiments of the present application adopt a splicing algorithm to splice a number of video frames containing pure document content to form a complete PDF file or WORD file.

[0080] Based on this, in a feasible manner for this situation, after obtaining the pure document content image, for the obtained multiple pure document content images with a time sequence relationship, the optical flow relationship between two adjacent pure document content images can be calculated respectively; according to the optical flow relationship, the pure document content images belonging to the same document page are merged.

[0081] For example, first, a deep optical flow model (FlowFormer) is used to calculate the optical flow relationship between two pure document content images, thereby calculating the same parts and different parts, and based on this, the two pure document content images are synthesized into one image; then SwinIR is used for denoising processing to solve the ghosting problem caused by synthesis, and the final pure document image for generating a PDF file or a WORD file is obtained.

[0082] On this basis, generating an editable document file based on the pure document content image may include: extracting content elements from the pure document content image to obtain the extracted content elements and the position information of the content elements in the pure document content image; generating an editable document file according to the content elements and their corresponding position information.

[0083] For example, a detector built with Retina-Net and Swin-Transformer-Tiny can be used to detect the possible picture elements and their regions in the document content, so as to extract the picture elements and obtain their position information; an OCR solution can be used to extract the text elements in the document content and obtain their position information; a detector + OCR is used to extract the table text and position information; finally, the picture elements, text elements, and table elements are synthesized into a PPT, PDF, or WORD file according to their position information in the pure document content image to form an editable document file.

[0084] Through this embodiment, first, each of the multiple document-class video segments is a single-document content video segment, that is, each video segment corresponds to only one document content. Thus, the subsequent extraction of video frames for the document content is more targeted and it is easier to extract the current document content. Second, when extracting key video frames for one document content, the extracted key video frames need to meet certain image quality standards so that the subsequent extracted document content can be clearer and have a higher imaging quality. Third, after extracting the key video frames, a document positioning operation for the document content is performed. Thus, other parts that may affect the document content, such as portraits, or human hands, or chat boxes, etc., which are non-document content, will be effectively removed, that is, the interference data affecting the document content can be removed. In addition, for the image region where the document content part located in each key video frame is located, image correction for obtaining pure document content will also be performed, including but not limited to removing marks, scribbles, etc. on the document content. Thus, the non-document content in the image can be effectively removed, ensuring the purity, cleanliness, and clarity of the document in the obtained image, greatly improving the quality of the document content image, and also making the editable document file generated based on this pure document content image clean, clear, with a high file quality, and convenient for subsequent use.

[0085] Refer toFigure 7 , which shows a schematic structural diagram of an electronic device according to an embodiment of the present application. The specific embodiments of the present application do not limit the specific implementation of the electronic device.

[0086] As Figure 7 shown, the electronic device may include: a processor 302, a communications interface 304, a memory 306, and a communication bus 308.

[0087] Among them:

[0088] The processor 302, the communications interface 304, and the memory 306 communicate with each other through the communication bus 308.

[0089] The communications interface 304 is used to communicate with other electronic devices or servers.

[0090] The processor 302 is used to execute the program 310, and specifically can execute the relevant steps in the above method embodiments.

[0091] Specifically, the program 310 may include program code, and the program code includes computer operation instructions.

[0092] The processor 302 may be a CPU, or a GPU (Graphic Processing Unit), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. One or more processors included in the intelligent device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0093] The memory 306 is used to store the program 310. The memory 306 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.

[0094] The program 310 may include multiple computer instructions. Specifically, the program 310 may cause the processor 302 to execute the operations corresponding to the methods described in any one of the foregoing multiple method embodiments through multiple computer instructions.

[0095] For the specific implementation of each step in Program 310, reference may be made to the corresponding descriptions in the corresponding steps and units in the foregoing method embodiments, and they have corresponding beneficial effects, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices and modules described above can refer to the corresponding process descriptions in the foregoing method embodiments, which will not be elaborated here.

[0096] The embodiments of the present application also provide a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method described in any one of the foregoing multiple method embodiments. The computer storage medium includes, but is not limited to: Compact Disc Read-Only Memory (CD-ROM), Random Access Memory (RAM), floppy disk, hard disk, magneto-optical disk, etc.

[0097] The embodiments of the present application also provide a computer program product, including computer instructions, and the computer instructions direct a computing device to perform the operations corresponding to any one of the foregoing multiple method embodiments.

[0098] In addition, it should be noted that the information related to users (including but not limited to user device information, user personal information, etc.) and data (including but not limited to sample data for training the model, data for processing, stored data, displayed data, etc.) involved in the embodiments of the present application are all information and data authorized by the users or fully authorized by all parties. And the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0099] It should be pointed out that according to the needs of implementation, each component / step described in the embodiments of the present application can be split into more components / steps, or two or more components / steps or partial operations of the components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present application.

[0100] The method according to the embodiments of the present application can be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through a network and to be stored in a local recording medium, so that the method described herein can be stored on such a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an Application Specific Integrated Circuit (ASIC) or a Field Programmable Gate Array (FPGA)) for software processing. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as a Random Access Memory (RAM), a Read-Only Memory (ROM), a flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown herein.

[0101] Those of ordinary skill in the art can realize that the units and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for a specific application, but such an implementation should not be considered to exceed the scope of the embodiments of the present application.

[0102] The above embodiments are only used to illustrate the embodiments of the present application, rather than to limit the embodiments of the present application. Those of ordinary skill in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present application, and the patent protection scope of the embodiments of the present application shall be defined by the claims.

Claims

1. A method for processing document - type video images, comprising: Obtain multiple document - type video segments, where each document - type video segment is a single - document content video segment; Extract key video frames that meet the image quality standard from each document - type video segment; Perform document localization on the key video frames of each extracted document - type video segment to determine the image area where the document content is located; Correct the determined image area to obtain a pure document content image, and generate an editable document file based on the pure document content image; Among them, the correction of the determined image area includes: judging whether there is an occluded document content area in the determined image area; if so, determining the document - type video segment to which the key video frame where the image area is located belongs; for the multiple video frames included in the determined document - type video segment, perform foreground segmentation respectively to extract multiple document content image parts corresponding to the multiple video frames; synthesize the multiple document content image parts to generate a synthesized document content image; perform denoising processing on the synthesized document content image.

2. The method according to claim 1, wherein, The correction of the determined image area further includes at least one of the following: Perform clarity enhancement processing on the determined image area; Judge whether there is a scribbled document content area in the determined image area; if so, perform a second document content restoration process for removing scribbles on the image area; Perform image normalization processing on the determined image area.

3. The method according to claim 2, wherein, The clarity enhancement processing of the determined image area includes: Perform contrast stretching processing on the determined image area to increase the contrast between text and background in the image area; and / or, perform high - frequency component enhancement and low - frequency component suppression processing on the determined image area.

4. The method according to claim 2, wherein, The second document content restoration process for removing scribbles on the image area includes: Use a denoising model for removing scribbles to perform a second document content restoration process for removing scribbles on the image area; Among them, the denoising model is trained based on denoising training samples, and the denoising training sample pair includes a pure document content sample image and a scribble sample image corresponding to the pure document content sample image; the scribble sample image is generated by synthesizing a scribble mask to the pure document content sample image.

5. The method according to any one of claims 1 - 4, wherein, The obtaining of multiple document - type video segments includes: Divide the document - type video into multiple segments in chronological order, perform parallel video frame extraction processing on each segment to obtain multiple candidate segments; Determine at least one group of video frames with the same document content from each candidate segment; Obtain multiple document - type video segments according to the multiple groups of video frames corresponding to the multiple candidate segments.

6. The method according to claim 5, wherein, The determining of at least one group of video frames with the same document content from each candidate segment includes: For each candidate segment, take two adjacent video frames in the candidate segment as a unit, and fuse the two video frames in the channel dimension; Perform semantic extraction processing based on the fused video frames, and classify whether the document content of the two video frames is the same according to the semantic extraction processing result; According to the classification results corresponding to multiple adjacent two video frames included in the candidate segment, the candidate segment is segmented into at least one group of video frames with the same document content.

7. The method according to any one of claims 1 - 4, wherein, Extracting key video frames that meet the image quality standard from each document-class video segment includes: Performing quality classification on multiple video frames included in each document-class video segment to obtain the quality types corresponding to the multiple video frames, where the quality types include: a first quality type for indicating that the video frame is pure document content, a second quality type for indicating that the video frame does not contain document content, a third quality type for indicating that the document content in the video frame is dynamic document content, and a fourth quality type for indicating that the document content in the video frame is painted. According to the quality types corresponding to the multiple video frames, extract key video frames that meet the image quality standard therefrom.

8. The method according to claim 7, wherein, The extracting key video frames that meet the image quality standard according to the quality types corresponding to the multiple video frames includes: Excluding video frames of the second quality type from the multiple video frames, and excluding invalid frames that cannot be key video frames among the video frames of the first quality type, the third quality type, or the fourth quality type. Among the remaining video frames of the multiple video frames, use the video frame with the highest similarity to the video frame of pure document content as the key video frame that meets the image quality standard.

9. The method according to claim 8, wherein, The excluding invalid frames that cannot be key video frames among the video frames of the first quality type, the third quality type, or the fourth quality type includes: For each frame among the video frames of the first quality type, the third quality type, or the fourth quality type, perform at least one of the following processes: If the brightness difference between the current video frame and other video frames in the video frame is greater than a preset brightness threshold, then confirm the current video frame as an invalid frame. If the blur degree between the current video frame and other video frames in the video frame is greater than a preset blur threshold, then confirm the current video frame as an invalid frame. If the number of the video frames is greater than a preset number threshold, then confirm the first video frame and the last video frame in the video frame as invalid frames.

10. The method according to any one of claims 1 - 4, wherein, After obtaining the image of pure document content, the method further includes: For multiple images of pure document content with a temporal relationship obtained, calculate the optical flow relationship between adjacent two images of pure document content respectively. According to the optical flow relationship, merge the images of pure document content belonging to the same document page.

11. The method according to any one of claims 1 - 4, wherein, Generating an editable document file based on the image of pure document content includes: Performing content element extraction on the image of pure document content to obtain the extracted content elements and the position information of the content elements in the image of pure document content. According to the content elements and the position information, generate an editable document file.

12. An electronic device, comprising: A processor, a memory, a communication interface, and a communication bus, the processor, the memory, and the communication interface complete communication with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the method according to any one of claims 1-11.

13. A computer storage medium, having stored thereon a computer program, which when executed by a processor, implements the method according to any one of claims 1 - 11.

Citation Information

Patent Citations

  • Document format conversion method and device, storage medium, equipment and program product

    CN115114229A