Document Format Conversion Method, Device, Storage Medium, Equipment and Program Product

The integration of OCR and instance segmentation techniques in PPT document conversion methods addresses the challenge of poor content extraction and editing capabilities, resulting in high-quality, editable PPT document reconstructions.

CN115114229BActive Publication Date: 2025-07-15TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210432095.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-22
Publication Date
2025-07-15
Estimated Expiration
2042-04-22

AI Technical Summary

Technical Problem

The prior art cannot effectively extract the picture content in the PPT document layout pictures, and the restoration effect is poor, resulting in insufficient editability and convenience of information acquisition of PPT files.

Method used

Optical character recognition (OCR) technology is used to identify text information and text locations in PPT documents, and the target area is extracted in combination with the instance segmentation model to generate target PPT documents with editable attributes.

Benefits of technology

It improves the restore degree and convenience of PPT information, and can effectively convert the picture content of the PPT document layout into an editable target PPT document.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115114229B_ABST
    Figure CN115114229B_ABST
Patent Text Reader

Abstract

The present application discloses a document format conversion method, device, storage medium, device and program product, which can be applied to various scenarios such as artificial intelligence, computer vision, image recognition, etc. The method includes: obtaining a to-be-recognized picture containing the layout of a PPT document, and performing character recognition processing on the to-be-recognized picture to obtain the text information in the to-be-recognized picture and the text position corresponding to the text information, and then performing region segmentation processing on the to-be-recognized picture to obtain the target region in the to-be-recognized picture, where the target region includes at least a text box region and an image region, and generating a target PPT document that matches the PPT document layout in the to-be-recognized picture according to the text information, text position and target region, where the content corresponding to the text information and the target region in the target PPT document has an editable attribute, which can effectively improve the PPT information restoration degree and enhance the convenience of PPT information acquisition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technologies, and particularly to a method, device, storage medium, equipment and program product for document format conversion. Background Art

[0002] In daily life or work, PPT documents are widely used. During the editing and production of PPT documents, it is often necessary to refer to and modify the picture material content with relevant PPT styles. Users may need to perform format conversion operations, such as converting the pictures on the PPT document layout into PPT documents.

[0003] In the current conversion methods, it is impossible to extract the pictures in the pictures on the PPT document layout, and the restoration effect of the PPT file obtained by restoring the content of the pictures on the PPT document layout is poor. Summary of the Invention

[0004] Embodiments of the present application provide a method, device, storage medium, equipment and program product for document format conversion, which can effectively convert the content in a to-be-recognized picture containing a PPT document layout into a target PPT document with editable attributes, improve the restoration degree of PPT information, and enhance the convenience of obtaining PPT information.

[0005] On the one hand, a method for document format conversion is provided. The method includes: obtaining a to-be-recognized picture containing a PPT document layout; performing character recognition processing on the to-be-recognized picture to obtain the text information in the to-be-recognized picture and the text position corresponding to the text information; performing region segmentation processing on the to-be-recognized picture to obtain the target region in the to-be-recognized picture; generating a target PPT document that matches the PPT document layout in the to-be-recognized picture according to the text information, the text position and the target region, wherein the text information and the content corresponding to the target region in the target PPT document have editable attributes.

[0006] On the other hand, a device for document format conversion is provided. The device includes:

[0007] An obtaining unit, configured to obtain a to-be-recognized picture containing a PPT document layout;

[0008] A recognition unit, configured to perform character recognition processing on the to-be-recognized picture to obtain the text information in the to-be-recognized picture and the text position corresponding to the text information;

[0009] A segmentation unit, configured to perform region segmentation processing on the to-be-recognized picture to obtain the target region in the to-be-recognized picture, where the target region includes at least a text box region and an image region;

[0010] A generating unit, configured to generate a target PPT document that matches the PPT document layout in the to-be-recognized picture according to the text information, the text position, and the target area, where the content corresponding to the text information and the target area in the target PPT document has an editable attribute.

[0011] On the other hand, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded by a processor to execute the steps in the document format conversion method described in any one of the foregoing embodiments.

[0012] On the other hand, a computer device is provided. The computer device includes a processor and a memory. A computer program is stored in the memory, and the processor is configured to execute the steps in the document format conversion method described in any one of the foregoing embodiments by calling the computer program stored in the memory.

[0013] On the other hand, a computer program product is provided, including computer instructions, and when the computer instructions are executed by a processor, the steps in the document format conversion method described in any one of the foregoing embodiments are implemented.

[0014] In the embodiment of the present application, a to-be-recognized picture containing a PPT document layout is obtained, and character recognition processing is performed on the to-be-recognized picture to obtain the text information in the to-be-recognized picture and the text position corresponding to the text information. Then, region segmentation processing is performed on the to-be-recognized picture to obtain the target area in the to-be-recognized picture, where the target area includes at least a text box area and an image area, and a target PPT document that matches the PPT document layout in the to-be-recognized picture is generated according to the text information, the text position, and the target area, where the content corresponding to the text information and the target area in the target PPT document has an editable attribute. In the embodiment of the present application, the text information in the to-be-recognized picture and the text position corresponding to the text information are recognized based on the optical character recognition (OCR) technology, and the target area in the to-be-recognized picture is obtained based on an instance segmentation model. Then, a target PPT document with an editable attribute that matches the PPT document layout in the to-be-recognized picture is generated according to the text information, the text position, and the target area, which can effectively convert the content in the to-be-recognized picture containing the PPT document layout into a target PPT document with an editable attribute, improve the PPT information restoration degree, and enhance the convenience of PPT information acquisition. Description of the Drawings

[0015] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0016] Figure 1 It is a schematic flowchart of the document format conversion method provided by the embodiment of the present application.

[0017] Figure 2 It is a schematic diagram of the picture to be recognized provided by the embodiment of the present application.

[0018] Figure 3 It is a schematic diagram of the first application scenario of the related technology provided by the embodiment of the present application.

[0019] Figure 4 It is a schematic diagram of the second process of the related technology provided by the embodiment of the present application.

[0020] Figure 5 It is a schematic diagram of the third process of the related technology provided by the embodiment of the present application.

[0021] Figure 6 It is a usage scenario diagram of the graphics device provided by the embodiment of the present application.

[0022] Figure 7 It is a schematic diagram of the first application scenario of the document format conversion method provided by the embodiment of the present application.

[0023] Figure 8 It is a schematic diagram of the second application scenario of the document format conversion method provided by the embodiment of the present application.

[0024] Figure 9 It is a schematic diagram of the third application scenario of the document format conversion method provided by the embodiment of the present application.

[0025] Figure 10 It is a schematic diagram of the fourth application scenario of the document format conversion method provided by the embodiment of the present application.

[0026] Figure 11 It is a schematic diagram of the fifth application scenario of the document format conversion method provided by the embodiment of the present application.

[0027] Figure 12 It is a schematic diagram of the structure of the document format conversion device provided by the embodiment of the present application.

[0028] Figure 13 It is a schematic diagram of the structure of the computer device provided by the embodiment of the present application. Specific embodiments

[0029] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0030] The embodiments of the present application provide a document format conversion method, apparatus, computer device, and storage medium. Specifically, the document format conversion method in the embodiments of the present application can be executed by a computer device, where the computer device can be a terminal or a server, etc. The terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart TV, a smart speaker, a wearable smart device, a smart vehicle terminal, etc. The terminal can also include a client, and the client can be a video client, a browser client, or an instant messaging client, etc. The server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0031] The embodiments of the present application can be applied to various scenarios such as artificial intelligence, computer vision, and image recognition.

[0032] First, some nouns or terms that appear in the process of describing the embodiments of the present application are explained as follows:

[0033] Artificial Intelligence (AI): It is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.

[0034] Computer Vision Technology (CV): Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as object recognition and measurement in machine vision, and further performing graphic processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision researches related theories and technologies, and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0035] Machine Learning (ML): It is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0036] Image recognition: It refers to the technology of using a computer to process, analyze, and understand images to identify various different patterns of targets and objects, and it is a practical application of deep learning algorithms.

[0037] PPT: It is a presentation program and is one of the components in the Microsoft Office system. A complete set of PPT files generally includes: title page, animation, PPT cover, preface, table of contents, transition pages, chart pages, picture pages, text pages, back cover, end animation, etc.; the materials used include: text, pictures, charts, animations, sounds, videos, etc.

[0038] Image segmentation: The process of subdividing a digital image into multiple image sub-regions (sets of pixels). The purpose of image segmentation is to simplify or change the representation form of the image, making it easier to understand and analyze. Image segmentation is further divided into semantic segmentation, instance segmentation, etc.

[0039] Semantic segmentation: A term in the field of deep learning. By classifying each pixel point in the image, the image is segmented into several regions with specific semantic categories.

[0040] Instance segmentation: A term in the field of deep learning. Different from semantic segmentation, it does not classify all pixels in an image. Instead, it only segments the objects of interest and also needs to find the bounding boxes of the objects, that is, determine the minimum bounding rectangle of the object area.

[0041] Optical Character Recognition (OCR): It refers to the process in which electronic devices (such as scanners or digital cameras) inspect the characters printed on paper, determine their shapes by detecting dark and bright patterns, and then translate the shapes into computer text using character recognition methods. That is, for printed characters, it uses an optical method to convert the text in a paper document into a black-and-white dot matrix image file, and through an OCR software, the text in the image is converted into a text format for further editing and processing by a word processing software.

[0042] Graphic: A basic structure in a PPT file. As Figure 6 shown in the graphic usage scenario diagram, it shows how to insert a graphic in a PPT, that is Figure 6 all the elements in the "Shape" button, including rectangles, rounded rectangles, triangles, rhombuses, parallelograms, regular hexagons, regular polygons, ellipses, stars, flags, etc.

[0043] Minimum bounding rectangle: It refers to the largest range of several two-dimensional shapes (such as points, lines, polygons) represented by two-dimensional coordinates, that is, a rectangle defined by the maximum abscissa, minimum abscissa, maximum ordinate, and minimum ordinate among the vertices of the given two-dimensional shapes. Such a rectangle contains the given two-dimensional shape and its sides are parallel to the coordinate axes.

[0044] Backbone: A term in the field of deep learning, mainly used in the definition of the algorithm network structure. It is a network used to extract data features. Backbone can be called the main network, the network used for feature extraction, representing a part of the network. Generally, it is used to extract image information at the front end to generate a feature map for use by the subsequent network.

[0045] The embodiments of this application propose a method based on instance segmentation, which can be used to extract the content in a to-be-recognized picture containing the layout of a PPT document. Specifically, it can extract the text, text boxes, images, tables, graphics, mathematical formulas, charts, flowcharts, etc. in the to-be-recognized picture, and then write the recognized content into the target PPT document to achieve editable attributes. The purpose is to convert the to-be-recognized picture containing the layout of a PPT document into a target PPT document.

[0046] In the embodiments of the present application, the text information of the image to be recognized and the text position corresponding to the text information can be recognized based on the OCR technology, and the target regions in the image to be recognized can be obtained based on the instance segmentation model. Then, according to the text information, the text position, and the target regions, a target PPT document with editable attributes that matches the PPT document layout in the image to be recognized can be generated, which can effectively convert the content in the image to be recognized containing the PPT document layout into a target PPT document with editable attributes, improve the restoration degree of PPT information, and enhance the convenience of PPT information acquisition. In the embodiments of the present application, based on the instance segmentation model, each target region in the image to be recognized containing the PPT document layout can be extracted at one time, and the edge information of the graphics region with different contours and shapes can be effectively obtained, providing favorable conditions for generating a target PPT document with editable attributes.

[0047] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the priority order of the embodiments.

[0048] Each embodiment of the present application provides a document format conversion method, which can be executed by a terminal or a server, or jointly executed by a terminal and a server; in the embodiments of the present application, the case where the document format conversion method is executed by a terminal is taken as an example for description.

[0049] Please refer to Figures 1 to 11 , Figure 1 which is a schematic flowchart of the document format conversion method provided by the embodiments of the present application, Figures 2 to 11 and are all schematic diagrams of related application scenarios provided by the embodiments of the present application. The method includes:

[0050] Step 110, obtain an image to be recognized containing a PPT document layout.

[0051] For example, the image to be recognized can be specifically understood as an image containing a PPT document layout with a recognition requirement. The image to be recognized can be an image uploaded by a user through a user terminal, or an image containing a PPT document layout obtained from a video file to be recognized.

[0052] For example, Figure 2 an image to be recognized 200 is given. Subsequently, the image to be recognized 200 can be used as input, and finally, the target PPT document corresponding to the image to be recognized 200 can be obtained. For example, the image to be recognized 200 contains three text paragraphs (text paragraph 211, text paragraph 212, text paragraph 213), three rectangular images (image 221, image 222, image 223), and a rounded rectangle region 231, and one text paragraph 213 is located in the rounded rectangle region 231.

[0053] For example, in the related art, text can be extracted based on the optical character recognition (OCR) technology. This solution mainly focuses on text extraction from images. After using the OCR technology to extract the text information in the image input by the user, the extracted text information is then written into a PPT (or PPTX) document to obtain a PPT file. As Figure 3 shown in the schematic diagram of the first application scenario in the related art, after using the OCR technology to extract the text information in the image to be recognized, only three pieces of text information in the image to be recognized are displayed in the obtained PPT file: "Cooperative Inquiry", "Through the analysis of the following pictures, let's discuss the following questions in groups:", and "Inquiry question: Which responsibilities are borne alone? Which responsibilities should have been borne by you but were borne by others instead." This text extraction solution based on the OCR technology focuses on extracting text information from images, and does not perform any restoration on the image area and graphics in the image, and there is no PPT background, so the restoration effect has a large gap from the image to be recognized.

[0054] For example, in another related art, the image to be recognized can be directly inserted into a PPT document. As Figure 4 shown in the schematic diagram of the second application scenario in the related art, this solution simply inserts the user-input image 200 directly into a PPT (or PPTX) file, but the elements in the image 400 cannot be edited, and this solution cannot recognize the text information in the image 400, resulting in a high user cost.

[0055] For example, in another related art, text boxes can be extracted based on the OCR method, and the picture is retained as the background. This solution mainly uses the OCR method to extract the text boxes in the picture, and then treats the area outside the text boxes as the background of the PPT file. As Figure 5 shown in the schematic diagram of the third application scenario in the related art, it can be seen from the figure that this solution extracts multiple text boxes (text box 511, text box 512, text box 513), and at the same time treats the other content in the image to be recognized as the PPT background. Although this solution detects the text boxes in the image to be recognized and makes the finally restored text information editable, it still cannot make the pictures and basic graphics in the PPT file editable.

[0056] Therefore, the embodiments of the present application propose a document format conversion method based on instance segmentation, which can extract the content in the image to be recognized containing the PPT document layout in subsequent steps. Specifically, it can extract the text, text boxes, images, tables, graphics, mathematical formulas, charts, flowcharts, etc. in the image to be recognized, and then write the recognized content into the target PPT document to achieve editable attributes. In order to achieve the purpose of converting the image to be recognized containing the PPT document layout into the target PPT document.

[0057] Optionally, the obtaining of the picture to be recognized includes: identifying image frames containing the layout of the PPT document from the video file to be recognized, and sorting the obtained multiple image frames according to the corresponding timestamps to form a set of pictures to be recognized.

[0058] For example, if it is necessary to obtain the PPT document information in the video file to be recognized, the video file to be recognized can be processed by frame division to obtain a series of multiple image frames containing the layout of the PPT document. Each image frame can be used as a picture to be recognized. Specifically, the video file to be recognized can be split into a series of image frames in the same time axis sequence, and this series of image frames forms a sequence picture library. For example, the processing of dividing the video file to be recognized into frames specifically may include: obtaining the total duration of the video file to be recognized, and then splitting the video file to be recognized into independent original image frames at preset time intervals. Among them, the smaller the preset time, the more original image frames are split from the video file to be recognized; the larger the preset time, the fewer original image frames are split from the video file. Among them, the more original image frames are split, the more image frames with high similarity will be, and the greater the similarity between adjacent image frames. Therefore, both the total duration of the video file to be recognized and the setting of the preset time used as the splitting condition in this step have an impact on the number of split image frames and the similarity between adjacent image frames. After obtaining the original image frames, the original image frames are recognized to obtain image frames containing the layout of the PPT document, and the obtained multiple image frames containing the layout of the PPT document are sorted according to the corresponding timestamps to form a set of pictures to be recognized.

[0059] For example, the image frame can also be a series of images containing the PPT document format screen obtained by photographing the video file to be recognized, and the series of photographed images are arranged in the order of timestamps to obtain a series of multiple image frames.

[0060] Step 120, perform character recognition processing on the picture to be recognized to obtain the text information in the picture to be recognized and the text position corresponding to the text information.

[0061] Optionally, the performing of character recognition processing on the picture to be recognized to obtain the text information in the picture to be recognized and the text position corresponding to the text information includes:

[0062] Performing character recognition processing on the picture to be recognized based on the optical character recognition (OCR) technology to locate and recognize the text information in the picture to be recognized, and determining the text information in the picture to be recognized and the text position corresponding to the text information.

[0063] For example, when performing character recognition processing on a picture to be recognized based on the optical character recognition (OCR) technology, the following processing flow may be included:

[0064] 2.1) Preprocess the picture to be recognized that contains text information to reduce useless information or interference information in the picture to be recognized, and perform preprocessing on the picture to be recognized that contains text information for subsequent feature extraction.

[0065] For example, the picture to be recognized can be subjected to processing such as binarization, denoising, skew correction, and line elimination.

[0066] For example, the binarization process is to set the grayscale value of the pixels on the picture to be recognized to 0 or 255 to convert the picture to be recognized into a binary image with a black-and-white visual effect, so as to distinguish the text from the background.

[0067] For example, the denoising process can be to remove the noise points in the picture to be recognized, or the picture to be recognized can be subjected to edge smoothing processing, or the picture to be recognized can be cropped to crop out the area that does not contain text information, etc.

[0068] For example, the skew correction process is that if the layout of the PPT document in the picture to be recognized is skewed, it may need to be skewed clockwise or counterclockwise by a few degrees to create a completely horizontal or vertical text line. For example, the line elimination process is to clean the non-symbol frames and lines in the picture to be recognized.

[0069] 2.2) Detect the text information in the preprocessed picture to be recognized. Common image detection algorithms can be used to frame the text area in the picture to be recognized to determine the text box area, that is, to determine the text position corresponding to the text information.

[0070] 2.3) Recognize the text information in the detected text box area through a text recognition algorithm to recognize the specific content of the text information. Specifically:

[0071] 2.31) Perform layout analysis on the preprocessed picture to be recognized to split and store the file to be recognized. For example, columns, paragraphs, headings, etc. are marked as blocks, which is particularly useful for layout analysis in multi-column layouts and tables.

[0072] 2.32) Perform character cutting on the image to be recognized after layout analysis. At this time, it is necessary to locate the cutting characters and the boundaries of the string, and then cut the string separately. After single-segment splitting, recognition is performed. When performing character switching processing, line character detection can be carried out to establish the shape baseline of words and characters, and words can be divided as needed. Script recognition can also be performed. In a multilingual document, the script may be converted at the word level. Therefore, before using relevant OCR to manage a specific script, script identification is crucial. Then, character isolation or "segmentation" processing can be carried out. For OCR characters, various characters linked to the image to be recognized should be split, and individual characters should be split into several artifact-based segments for linking. Then, normalization processing can also be carried out to normalize the aspect ratio and scale.

[0073] 3.33) Extract features from the image to be recognized after character cutting to extract the character features of the image to be recognized, providing a basis for subsequent character recognition. For example, character features can be defined by evaluating the lines and strokes of characters through feature detection algorithms, or the entire character features can be recognized through pattern recognition. For example, the image of a character can be converted into a binary matrix, where white pixels are 0 and black pixels are 1. Then, the distance formula is used to find the farthest distance from the center of the matrix to 1. Then, a radius of a circle is created and divided into finer-grained parts. At this stage, each segment is compared with a matrix database representing different font characters through an algorithm to determine the statistically most common character features.

[0074] 2.34) Perform character recognition based on the character features to obtain the character recognition result. For example, template rough classification and template fine matching can be used, and the feature vectors extracted from the current characters are used to recognize the characters in the feature template library.

[0075] 2.35) Typeset the character recognition result. Typeset the recognition result according to the original version and output the text information of the image to be recognized.

[0076] 2.36) Perform post-processing correction on the typeset recognition result. For example, the typeset recognition result can be corrected through the relationship of a specific language context.

[0077] Step 130, perform region segmentation processing on the image to be recognized to obtain the target region in the image to be recognized, where the target region includes at least a text box region and an image region.

[0078] Optionally, the target region may include at least one of a text box region, an image region, a graphics region, a table region, a mathematical formula region, a chart region, and a flowchart region.

[0079] Optionally, performing region segmentation processing on the image to be recognized to obtain the target region in the image to be recognized includes:

[0080] Performing region segmentation processing on the image to be recognized based on an instance segmentation model to obtain the target region in the image to be recognized.

[0081] Among them, the instance segmentation model is mainly used to obtain all target regions in the image to be recognized, mainly including text box regions, image regions, graphic regions, and table regions, and may also include mathematical formula regions, chart regions, flowchart regions, etc. Among them, as Figure 6 shown in the graphic usage scenario diagram, the graphic 610 includes multiple categories, such as rectangles, rounded rectangles, triangles, rhombuses, parallelograms, regular hexagons, regular polygons, ellipses, stars, flag shapes, etc. In addition, the image region may also include rectangular images, rounded rectangular images, elliptical images, etc.

[0082] Instance segmentation technology is a technical solution derived relative to image classification, object detection, and semantic segmentation. From Figure 7 a simple comparison of the four technical solutions is given, where Figure 7 (a) represents image classification. Through this image classification technical solution, only three categories of items including bottle, cup, and cube in the picture can be obtained; Figure 7 (b) represents object detection. This object detection technical solution can not only obtain Figure 7 the three categories of items including bottle, cup, and cube in (a), but also obtain the specific positions of the items; Figure 7 (c) represents semantic segmentation. This semantic segmentation technical solution can obtain regional information of different semantic information. In addition to obtaining the three categories of items in 7(a), it can obtain the regions where each category of items is distributed, and the regions of each category can be given in the same color, that is, directly classifying the pixels of the picture; Figure 7 (d) represents instance segmentation. Compared with the Figure 7 (c) solution, first, some unconcerned regions, such as the background region, can be removed. In addition, obvious regional segmentation boundaries can be provided for different categories of items. Even if there are two identical categories of cubes and there is an overlap, two different targets will be segmented. It can be seen that this instance segmentation technical solution can more conveniently obtain the range and position of each target instance.

[0083] Among them, according to the instance segmentation technology introduced above, it can be concluded that this instance segmentation technology solution is very applicable in the PPT area segmentation scenario because in the to-be-recognized pictures containing PPT document layouts, there are often cases of graphic overlaps or image overlaps. If only based on the semantic segmentation technology solution, each target area cannot be effectively distinguished. In addition, if the object detection technology solution is adopted, for area pictures similar to triangles and rounded rectangles, only the minimum bounding rectangle area can be obtained through object detection, and the extraction of the target contour cannot be realized, resulting in the inability to restore the accurate triangle area picture. Therefore, compared with the semantic segmentation technology solution and the object detection technology solution, the instance segmentation technology solution has a better segmentation effect on the target area.

[0084] For example, the instance segmentation model can be an instance segmentation algorithm based on CBNetv2. This algorithm is mainly used to construct a more efficient backbone network. Based on this algorithm, the embodiments of the present application can use an improved HTC algorithm (denoted as HTC++) framework to perform area segmentation on the to-be-recognized pictures containing PPT document layouts.

[0085] Figure 8 The structural schematic diagram of the instance segmentation model is shown. Among them, the instance segmentation model 800 can include a feature extraction module 801, a feature fusion module 802, an object classification module 803, an object localization module 804, an object contour extraction module 805, and a semantic segmentation module 806. For example, the feature extraction module 801 can adopt the solution of cbnetv2, and the subsequent feature fusion module 802 and the subsequent object classification module 803, object localization module 804, object contour extraction module 805, and semantic segmentation module 806 can be implemented based on the HTC++ algorithm framework. Input the to-be-recognized picture into the instance segmentation model 800, extract the feature vector of the to-be-recognized picture through the feature extraction module 801, and perform feature fusion processing through the feature fusion module 802 to obtain the fusion feature of the to-be-recognized picture, so as to integrate the high-level feature vector and the low-level feature vector of the to-be-recognized picture output by multiple backbone networks in the feature extraction module. Then, based on the object classification module 803, object localization module 804, object contour extraction module 805, and semantic segmentation module 806 in the instance segmentation model 800, process the fusion feature of the to-be-recognized picture to obtain the instance segmentation result, and determine the final area segmentation result according to the instance segmentation result and the confidence of each candidate area in the instance segmentation result. This area segmentation result contains all target areas.

[0086] For example, CBNetv2 incorporates existing pre-trained weights as the backbone of a detector. CBNetv2 directly enhances the representational ability of existing pre-trained models through a new fusion method. Without pre-training, it only needs to use the weights of an existing open-source pre-trained single backbone to initialize each assembled backbone of CBNetV2. The backbone network of CBNetV2 is divided into two parts: the Assisting Backbones and the Lead Backbone. The Lead Backbone contains only one backbone network, and the Assisting Backbones consist of multiple backbone networks. The output of each stage of the Assisting Backbones flows into the lower levels of its successor backbone (i.e., the next backbone network connected after it) as input. Finally, the features of the Lead Backbone will be input into the Neck and Detection Head for regression and classification prediction. Compared with simple deeper or wider networks, CBNetV2 integrates high-level and low-level features of multiple backbone networks, gradually expands the receptive field, and extends the receptive field to more efficient object detection.

[0087] The HTC (Hybrid Task Cascade) algorithm is a multi-task multi-stage hybrid cascade structure model. Its core idea is to integrate cascade and multi-tasking at each stage to process and improve the information flow, and further improve the accuracy by using spatial context. At each stage, HTC combines bbox regression and mask prediction in a multi-task manner. In addition, a direct connection is built between the mask branches of different stages: encoding the mask features of each stage and sending them to the next stage. For object detection, context information provides very important clues for object localization and classification. Therefore, HTC also additionally adopts a fully convolutional branch to perform segmentation. This branch not only encodes the context information from foreground instances but also encodes the information from background regions, further improving the prediction accuracy of bounding boxes and instance masks.

[0088] For example, each layer is connected to all subsequent layers to build huge features. For example, the embodiments of this application can adopt two backbones, combining all the high-level features of the previous backbone and connecting them to the low-level features of the subsequent backbone.

[0089] For example, the four modules (target classification module 803, target localization module 804, target contour extraction module 805, semantic segmentation module 806) can correspond to four heads in the HTC network and are four branches of a unified network. Among them, the target classification module 803 can be used to distinguish whether the detected target belongs to the target area; the target localization module 804 can be used to determine the specific position of the target area; the target contour extraction module 805, also known as the mask extraction module, can be used to determine the contour of the target area; the semantic segmentation module 806 can be used to classify each pixel point in the picture to be recognized, so as to segment the picture to be recognized into several regions with specific semantic categories. All of these four modules require supervision information to jointly supervise the training of the instance segmentation model 800. In the training data, it is first necessary to define a batch of sample pictures with region category labels such as table regions, text box regions, image regions, graphic regions, mathematical formula regions, chart regions, flowchart regions, etc. Then, based on the collected sample pictures, the entire instance segmentation model 800 is trained. The above four modules respectively correspond to four loss functions, and the loss functions are optimized through network training.

[0090] After the instance segmentation model 800 is trained, the trained instance segmentation model 800 is used to perform region segmentation processing on the picture to be recognized to obtain the instance segmentation result.

[0091] Among them, the obtained instance segmentation results can be sorted according to the confidence level to filter out candidate regions with a confidence level lower than the confidence level threshold, and the region segmentation result is obtained. The region segmentation result includes all target regions in the picture to be recognized. Among them, the instance segmentation model 800 will correspondingly output the confidence level of each candidate region for all the output candidate regions. The value range of the confidence level is 0-1. Generally, candidate regions with a confidence level lower than the confidence level threshold will be filtered out, and candidate regions with a confidence level greater than (or equal to) the preset threshold will be retained. The candidate regions with a confidence level greater than (or equal to) the confidence level threshold are determined as the target regions. For example, the value range of the confidence level threshold can be 0.5-0.85, and the specific confidence level threshold can be obtained based on testing different data sets.

[0092] For example, as Figure 9 shown, the region segmentation result obtained in the embodiment of the present application is given under the condition of Figure 2 as the input. The region segmentation result includes text box regions (text box region 911, text box region 912, text box region 913), image regions (image region 921, image region 922, image region 923), and graphic region 931.

[0093] For example, the instance segmentation model can also adopt other instance segmentation algorithms, such as using different backbones and different detection and segmentation heads. For example, the instance segmentation network for remote sensing images (cascade mask rcnn) is an improved network obtained by adding a mask branch to the cascade faster rcnn network and can be used to implement the instance segmentation task.

[0094] For example, more different categories of target regions can also be designed according to actual needs. The specific categories of target regions exemplified in the embodiments of the present application do not limit the embodiments of the present application.

[0095] Step 140: Generate a target PPT document that matches the PPT document layout in the to-be-recognized picture according to the text information, the text position, and the target region, where the content corresponding to the text information and the target region in the target PPT document has an editable attribute.

[0096] Among them, the content corresponding to the target region may include the region border of the target region, the content of the target region itself, the relevant plug-ins called in the target PPT document and matching the target region, and the content written in and displayed in the matching relevant plug-ins, etc.

[0097] For example, if the target region is a text region, the content corresponding to the target region may include the region border of the text box region and the text information displayed within the text box region.

[0098] For example, if the target region is an image region, the content corresponding to the target region may include the region border of the image region and the image displayed within the image region.

[0099] For example, if the target region is a graphics region, the content corresponding to the target region may include the region border displayed in the target PPT document by the called target graphics plug-in and the content displayed within the target graphics plug-in.

[0100] For example, if the target region is a table region, the content corresponding to the target region may include the region border displayed in the target PPT document by the called table plug-in and the table content displayed within the table plug-in.

[0101] For example, if the target region is a mathematical formula region, the content corresponding to the target region may include the region border displayed in the target PPT document by the called mathematical formula plug-in and the formula content displayed within the mathematical formula plug-in.

[0102] For example, if the target area is a chart area, the content corresponding to the target area may include the area border displayed by the invoked chart plugin in the target PPT document and the chart content displayed within the chart.

[0103] For example, if the target area is a flowchart area, the content corresponding to the target area may include the area border displayed by the invoked flowchart plugin in the target PPT document and the flowchart content displayed within the flowchart.

[0104] For example, after obtaining the ORC recognition results (text information and text positions) and the region segmentation results (all target regions in the image to be recognized), a target PPT document that matches the PPT document layout in the image to be recognized is generated. For example, Figure 10 the restored target PPT document 1000 shown in the figure is obtained. The target PPT document 1000 includes text box areas (text box area 1011, text box area 1012, text box area 1013) with editable attributes, image areas (image area 1021, image area 1022, image area 1023), and a graphics area 1031. Among them, each text box area also contains corresponding text information. The text box area 1011 corresponds to the restoration of the text paragraph 211 in the image to be recognized, the text box area 1012 corresponds to the restoration of the text paragraph 212 in the image to be recognized, and the text box area 1013 corresponds to the restoration of the text paragraph 213 in the image to be recognized. Among them, each image area contains a corresponding image. The image area 1021 corresponds to the restoration of the image 221 in the image to be recognized, the image area 1022 corresponds to the restoration of the image 222 in the image to be recognized, and the image area 1023 corresponds to the restoration of the image 223 in the image to be recognized. Among them, the graphics area 1031 restores the rounded rectangle area 231 in the image to be recognized, and the text box area 1013 is located within the graphics area 1031. From the restoration results, it can be seen that the target PPT document 1000 restored based on the embodiments of the present application can not only restore the text boxes, but also restore the images on the opposite side, and at the same time restore the graphics, and all can be made editable.

[0105] Optionally, the generating a target PPT document that matches the initial PPT document layout in the image to be recognized according to the text information, the text positions, and the target area includes:

[0106] Performing position matching between the text box area in the target area and the text positions corresponding to the text information, so as to locate the text box area to a first target position corresponding to the text positions in the target PPT document, and writing the text information into the text box area displayed at the first target position in the target PPT document;

[0107] Crop the image area in the target area, and insert the cropped image area into the target PPT document after scaling it according to a preset ratio.

[0108] For example, match the position of the text box area in the region segmentation result with the text position recognized in the OCR recognition result, so as to locate the text box area to the first target position corresponding to the text position in the target PPT document, and then write the text information recognized in the OCR recognition result into the text box area displayed at the first target position in the target PPT document, so as to generate editable text information in the target PPT document. Among them, the first target position can be determined according to the pixel position of the text box area in the picture to be recognized and the size ratio between the picture to be recognized and the target PPT document. Compared with Figure 3 the corresponding solution that only uses OCR technology to recognize text information, in the embodiment of the present application, the text information in the target PPT document is generated by combining the OCR recognition result and the text box area in the region segmentation result, and the reduction degree of the text box area is higher. It can not only restore editable text information, but also restore the text position in the picture to be recognized to the target PPT document.

[0109] For example, the image area can also be cropped along the contour of the image area, and then inserted into the target PPT document at the second target position corresponding to the position of the image area in the target PPT document after scaling it according to a preset ratio, so as to generate an editable image in the target PPT document. Among them, the second target position can be determined according to the pixel position of the image area in the picture to be recognized and the size ratio between the picture to be recognized and the target PPT document. Compared with Figure 4 the corresponding solution of directly inserting the picture to be recognized into the PPT document, or compared with Figure 5 the corresponding solution of using other content outside the text box of the picture to be recognized as the PPT background, in the embodiment of the present application, the cropped image area is inserted into the target PPT document after scaling it according to a preset ratio, and the reduction degree of the image area in the picture to be recognized is higher, and the image generated corresponding to the image area in the target PPT document has an editable attribute.

[0110] Optionally, if the target area further includes a graphics area, call a target graphics plug-in matching the graphics area from the target PPT document, and write the target graphics plug-in into the target PPT document.

[0111] For example, determine the target graphics plugin to be called according to the category of the graphics area, and then directly call the target graphics plugin in the target PPT document to write the target graphics plugin corresponding to all graphics areas to the third target position corresponding to the position of the graphics area in the target PPT document, so as to generate an editable graphics in the target PPT document. Among them, the third target position can be determined according to the pixel position of the graphics area in the picture to be recognized and the size ratio between the picture to be recognized and the target PPT document. Compared with Figure 4 the solution of directly inserting the picture to be recognized into the PPT document, or compared with Figure 5 the solution of using other content outside the text box of the picture to be recognized as the PPT background, in the embodiment of the present application, the target graphics plugin corresponding to all graphics areas is called and written into the target PPT document, so that the reduction degree of the graphics area in the picture to be recognized is higher, and the generated graphics in the target PPT document has an editable attribute.

[0112] Optionally, if the target area further includes a table area, then crop the table area, identify the table content in the cropped table area, call the table plugin matching the table area from the target PPT document, and write the table content into the target PPT document according to the position of the table area and the table plugin.

[0113] For example, crop along the contour of the table area according to the position of the table area, and send the cropped table area to the table recognition module to identify the table content in the cropped table area. Then call the table plugin matching the table area in the target PPT document, insert a table at the fourth target position corresponding to the position of the table area in the target PPT document, and call the writing module to directly write the table content into the table using the pptx library in the python code, so as to generate editable table content in the target PPT document. Among them, the fourth target position can be determined according to the pixel position of the table area in the picture to be recognized and the size ratio between the picture to be recognized and the target PPT document. In the embodiment of the present application, by identifying the table content of the table area, calling the table plugin corresponding to the table area and writing the table content into the target PPT document, the reduction degree of the table area in the picture to be recognized is higher, and the generated table content in the target PPT document has an editable attribute.

[0114] Optionally, if the target area further includes a mathematical formula area, then identify the formula content in the mathematical formula area, call the formula editor matching the mathematical formula area from the target PPT document, and write the formula content into the target PPT document according to the position of the mathematical formula area and the formula editor.

[0115] For example, to identify the formula content in the mathematical formula area, where the formula content includes formula symbols and mathematical values, a formula editor that matches the mathematical formula area in the target PPT document can be called according to the formula content, and the formula symbols and mathematical values in the formula content can be written into the formula editor at the fifth target position corresponding to the position of the mathematical formula area in the target PPT document, so as to generate an editable mathematical formula in the target PPT document. Among them, the fifth target position can be determined according to the pixel position of the mathematical formula area in the picture to be recognized and the size ratio between the picture to be recognized and the target PPT document. By recognizing the mathematical content of the mathematical formula area in the embodiments of the present application, calling the formula editor corresponding to the mathematical formula area and writing the formula content into the target PPT document, the reduction degree of the mathematical formula area in the picture to be recognized is higher, and the generated formula content in the target PPT document has an editable attribute.

[0116] Optionally, if the target area further includes a chart area, then identify the chart content in the chart area, call a chart plug-in that matches the chart area from the target PPT document, and write the chart content into the target PPT document according to the position of the chart area and the chart plug-in.

[0117] For example, if the chart content is statistical chart content, a chart plug-in that matches the chart area in the target PPT document can be called, and the chart content can be written into the chart plug-in at the sixth target position corresponding to the position of the chart area in the target PPT document, so as to generate editable chart content in the target PPT document. Among them, the sixth target position can be determined according to the pixel position of the chart area in the picture to be recognized and the size ratio between the picture to be recognized and the target PPT document. By recognizing the chart content of the chart area in the embodiments of the present application, calling the chart plug-in corresponding to the chart area and writing the chart content into the target PPT document, the reduction degree of the chart area in the picture to be recognized is higher, and the generated chart content in the target PPT document has an editable attribute.

[0118] Optionally, if the target area further includes a flowchart area, then identify the flowchart content in the flowchart area, call a flowchart plug-in that matches the flowchart area from the target PPT document, and write the flowchart content into the target PPT document according to the position of the flowchart area and the flowchart plug-in.

[0119] For example, if the target area also includes a flowchart area, first identify the flowchart content in the flowchart area. The flowchart content may include a flowchart block and the flowchart text within the flowchart block. Then, according to the flowchart block, call a flowchart plug-in that matches the flowchart area from the target PPT document, and write the flowchart text into the flowchart plug-in to obtain a corresponding target flowchart. Then, insert the editable target flowchart at a seventh target position corresponding to the position of the flowchart area in the target PPT document. Among them, the seventh target position can be determined according to the pixel position of the flowchart area in the to-be-recognized picture and the size ratio between the to-be-recognized picture and the target PPT document. By identifying the chart content of the flowchart area in this embodiment of the present application, calling the flowchart plug-in corresponding to the flowchart area and writing the flowchart content into the target PPT document, the reduction degree of the flowchart area in the to-be-recognized picture is higher, and the generated flowchart content in the target PPT document has an editable attribute.

[0120] Optionally, the method further includes: extracting a background image of the to-be-recognized picture according to the target area; and inserting the background image into the target PPT document.

[0121] For example, since it is necessary to obtain the background of the target PPT document, all the target area to-be-recognized pictures obtained by region segmentation need to be input into a background extraction module to extract the background image of the to-be-recognized picture. After the background image is extracted, insert the background image into the target PPT document. Among them, all the target areas input into the background extraction module can be used as the foreground information of the to-be-recognized picture, and the background image in the to-be-recognized picture is determined according to the foreground information. In this embodiment of the present application, the recognized background image through the target area is the closest to the background in the to-be-recognized picture, and the background in the to-be-recognized picture can be highly restored, and the generated background image in the target PPT document has an editable attribute.

[0122] Optionally, the generating a target PPT document that matches the PPT document layout in the to-be-recognized picture according to the text information, the text position, and the target area further includes:

[0123] Generating a target PPT document corresponding to each to-be-recognized picture in the to-be-recognized picture set according to the text information, the text position, and the target area corresponding to each to-be-recognized picture in the to-be-recognized picture set;

[0124] Performing document merging on the target PPT documents corresponding to each to-be-recognized picture in the to-be-recognized picture set to obtain a target PPT file corresponding to the to-be-recognized video file.

[0125] For example, sort the obtained multiple image frames according to the corresponding timestamps to form a set of pictures to be recognized. This set of pictures to be recognized can be used to individually identify the target PPT document corresponding to each picture to be recognized, traverse and process each picture to be recognized in the set of pictures to be recognized, identify the text information and the corresponding text position of each picture to be recognized in the set of pictures to be recognized based on the OCR technology, and obtain the target area in each picture to be recognized in the set of pictures to be recognized based on the instance segmentation model. Then, according to the text information, the text position, and the target area, generate a target PPT document with editable attributes that matches the PPT document layout in the picture to be recognized, obtain multiple target PPT documents, and arrange the multiple target PPT documents in the order in the set of pictures to be recognized to obtain the target PPT file corresponding to the video file to be recognized. In this way, a target PPT file with editable attributes corresponding to the video file to be recognized can be obtained, so that the user can collect relevant PPT file information in a timely manner, improve the efficiency of information acquisition, highly restore the PPT file in the video file to be recognized, and make the collected target PPT file information have editable attributes for the convenience of the user.

[0126] For example, for better understanding of the document format conversion method in the embodiments of the present application, reference can be made to Figure 11The shown process schematic diagram inputs the to-be-recognized picture containing the layout of the PPT document, performs character recognition processing on the to-be-recognized picture through the ORC model to obtain all the text information in the to-be-recognized picture and the text positions corresponding to the text information; and performs region segmentation processing on the to-be-recognized picture through the instance segmentation model to obtain the region segmentation result, which includes all the target regions in the to-be-recognized picture, and generates a target PPT document that matches the PPT document layout in the to-be-recognized picture according to the text information, text positions, and target regions. Among them, the text information and the content corresponding to the target region in the target PPT document have editable attributes. For example, the target region includes at least one of a text box region, an image region, a graphics region, a table region, a mathematical formula region, a chart region, and a flowchart region; if the target region includes a text region, the text box region in the target region is position-matched with the text position corresponding to the text information to locate the text box region to the first target position corresponding to the text position in the target PPT document, and the text information is written into the text box region displayed at the first target position in the target PPT document; if the target region further includes an image region, the image region is cropped, and the cropped image region is scaled according to a preset ratio and inserted into the target PPT document; if the target region further includes a graphics region, a target graphics plug-in that matches the graphics region is called from the target PPT document, and the target graphics plug-in is written into the target PPT document; if the target region further includes a table region, the table region is cropped, and a table recognition module is called to recognize the table content in the cropped table region, a table plug-in that matches the table region is called from the target PPT document, and according to the position of the table region and the called table plug-in in the target PPT document, a writing module is called to write the table content into the target PPT document; if the target region further includes a mathematical formula region, the formula content in the mathematical formula region is recognized, a formula editor that matches the mathematical formula region is called from the target PPT document, and according to the position of the mathematical formula region and the called formula editor in the target PPT document, the formula content is written into the target PPT document; if the target region further includes a chart region, the chart content in the chart region is recognized, a chart plug-in that matches the chart region is called from the target PPT document, and according to the position of the chart region and the called chart plug-in in the target PPT document, the chart content is written into the target PPT document; if the target region further includes a flowchart region, the flowchart content in the flowchart region is recognized, a flowchart plug-in that matches the flowchart region is called from the target PPT document, and according to the position of the flowchart region and the called flowchart plug-in in the target PPT document, the flowchart content is written into the target PPT document; the background image of the to-be-recognized picture can also be extracted according to all the target regions in the region segmentation result and inserted into the target PPT document.

[0127] Any combination of the above technical solutions can form an alternative embodiment of the present application, which will not be elaborated one by one here.

[0128] In an embodiment of the present application, a to-be-recognized picture containing the layout of a PPT document is obtained, and character recognition processing is performed on the to-be-recognized picture to obtain the text information in the to-be-recognized picture and the text position corresponding to the text information. Then, region segmentation processing is performed on the to-be-recognized picture to obtain the target region in the to-be-recognized picture, where the target region at least includes a text box region and an image region. And based on the text information, the text position, and the target region, a target PPT document matching the PPT document layout in the to-be-recognized picture is generated, where the content corresponding to the text information and the target region in the target PPT document has an editable attribute. In the embodiment of the present application, the text information in the to-be-recognized picture and the text position corresponding to the text information are recognized based on the optical character recognition (OCR) technology, and the target region in the to-be-recognized picture is obtained based on an instance segmentation model. Then, based on the text information, the text position, and the target region, a target PPT document with an editable attribute and matching the PPT document layout in the to-be-recognized picture is generated, which can effectively convert the content in the to-be-recognized picture containing the PPT document layout into a target PPT document with an editable attribute, improve the restoration degree of PPT information, and enhance the convenience of PPT information acquisition.

[0129] To facilitate better implementation of the document format conversion method in the embodiment of the present application, an embodiment of the present application further provides a document format conversion device. Please refer to Figure 12 , Figure 12 which is a schematic structural diagram of the document format conversion device provided in the embodiment of the present application. Among them, the document format conversion device 1200 may include:

[0130] An acquisition unit 1210, configured to acquire a to-be-recognized picture containing the layout of a PPT document;

[0131] A recognition unit 1220, configured to perform character recognition processing on the to-be-recognized picture to obtain the text information in the to-be-recognized picture and the text position corresponding to the text information;

[0132] A segmentation unit 1230, configured to perform region segmentation processing on the to-be-recognized picture to obtain the target region in the to-be-recognized picture, where the target region at least includes a text box region and an image region;

[0133] A generation unit 1240, configured to generate a target PPT document matching the PPT document layout in the to-be-recognized picture based on the text information, the text position, and the target region, where the content corresponding to the text information and the target region in the target PPT document has an editable attribute.

[0134] Optionally, the splitting unit 1230 can be used to perform region splitting processing on the to-be-recognized picture based on an instance segmentation model to obtain a target region in the to-be-recognized picture, where the target region includes at least a text box region and an image region.

[0135] Optionally, the generating unit 1240 can be used to perform position matching between the text box region in the target region and the text position corresponding to the text information, so as to locate the text box region to a first target position corresponding to the text position in the target PPT document, and write the text information into the text box region displayed at the first target position in the target PPT document;

[0136] Crop the image region in the target region, and insert the cropped and scaled image region into the target PPT document according to a preset ratio.

[0137] Optionally, the generating unit 1240 can also be used to, if the target region further includes a graphics region, call a target graphics plug-in matching the graphics region from the target PPT document, and write the target graphics plug-in into the target PPT document.

[0138] Optionally, the generating unit 1240 can also be used to, if the target region further includes a table region, crop the table region, identify the table content in the cropped table region, call a table plug-in matching the table region from the target PPT document, and write the table content into the target PPT document according to the position of the table region and the table plug-in.

[0139] Optionally, the generating unit 1240 can also be used to, if the target region further includes a mathematical formula region, identify the formula content in the mathematical formula region, call a formula editor matching the mathematical formula region from the target PPT document, and write the formula content into the target PPT document according to the position of the mathematical formula region and the formula editor.

[0140] Optionally, the generating unit 1240 can also be used to, if the target region further includes a chart region, identify the chart content in the chart region, call a chart plug-in matching the chart region from the target PPT document, and write the chart content into the target PPT document according to the position of the chart region and the chart plug-in.

[0141] Optionally, the generating unit 1240 may further be configured to: if the target area further includes a flowchart area, identify the flowchart content in the flowchart area, call a flowchart plug-in matching the flowchart area from the target PPT document, and write the flowchart content into the target PPT document according to the position of the flowchart area and the flowchart plug-in.

[0142] Optionally, the generating unit 1240 may further be configured to: extract the background image of the to-be-recognized picture according to the target area; insert the background image into the target PPT document.

[0143] Optionally, the obtaining unit 1210 may be configured to: identify an image frame containing the PPT document layout from the to-be-recognized video file, and sort the obtained multiple image frames according to the corresponding timestamps to form a to-be-recognized picture set.

[0144] Optionally, the generating unit 1240 may further be configured to: generate a target PPT document corresponding to each to-be-recognized picture in the to-be-recognized picture set according to the text information, the text position, and the target area corresponding to each to-be-recognized picture in the to-be-recognized picture set;

[0145] Merge the target PPT documents corresponding to each to-be-recognized picture in the to-be-recognized picture set to obtain a target PPT file corresponding to the to-be-recognized video file.

[0146] Optionally, the recognizing unit 1220 is configured to: perform character recognition processing on the to-be-recognized picture based on the optical character recognition (OCR) technology to locate and recognize the text information in the to-be-recognized picture, and determine the text information in the to-be-recognized picture and the text position corresponding to the text information.

[0147] Each unit in the above document format conversion device may be implemented in whole or in part by software, hardware, and their combination. Each of the above units may be embedded in the processor of the computer device in a hardware form or be independent of the processor, or may be stored in the memory of the computer device in a software form so that the processor can call and execute the operations corresponding to each of the above units.

[0148] The document format conversion device 1200 may be integrated in a terminal or a server that has a memory and is equipped with a processor and has computing capabilities, or the document format conversion device 1200 is the terminal or the server.

[0149] Optionally, the present application further provides a computer device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0150] As Figure 13 shown Figure 13 Figure 13 is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device may be a terminal. The computer device 1300 includes a processor 1301 having one or more processing cores, a memory 1302 having one or more computer-readable storage media, and a computer program stored on the memory 1302 and executable on the processor. Among them, the processor 1301 is electrically connected to the memory 1302. Those skilled in the art can understand that the computer device structure shown in the figure does not constitute a limitation on the computer device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0151] The processor 1301 is the control center of the computer device 1300, connecting various parts of the entire computer device 1300 through various interfaces and lines. By running or loading software programs and / or modules stored in the memory 1302, and calling data stored in the memory 1302, it executes various functions of the computer device 1300 and processes data, thereby performing overall processing on the computer device 1300.

[0152] In the embodiment of the present application, the processor 1301 in the computer device 1300 will load the instructions corresponding to the processes of one or more application programs into the memory 1302 according to the following steps, and the processor 1301 will run the application programs stored in the memory 1302 to implement various functions:

[0153] Obtain a to-be-recognized picture containing the layout of a PPT document; perform character recognition processing on the to-be-recognized picture to obtain the text information in the to-be-recognized picture and the text position corresponding to the text information; perform region segmentation processing on the to-be-recognized picture to obtain the target region in the to-be-recognized picture, where the target region at least includes a text box region and an image region; generate a target PPT document that matches the PPT document layout in the to-be-recognized picture according to the text information, the text position, and the target region, where the content corresponding to the text information and the target region in the target PPT document has an editable attribute.

[0154] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, and details are not described herein again.

[0155] Optionally, as Figure 13As shown, the computer device 1300 further includes: a touch display screen 1303, a radio frequency circuit 1304, an audio circuit 1305, an input unit 1306, and a power supply 1307. Among them, the processor 1301 is electrically connected to the touch display screen 1303, the radio frequency circuit 1304, the audio circuit 1305, the input unit 1306, and the power supply 1307 respectively. Those skilled in the art can understand that Figure 13 the computer device structure shown in

[0156] does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements. The touch display screen 1303 can be used to display a graphical user interface and receive operation instructions generated by a user acting on the graphical user interface. The touch display screen 1303 may include a display panel and a touch panel. Among them, the display panel can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the computer device. These graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. The touch panel can be used to collect touch operations of the user on or near it (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel), and generate corresponding operation instructions, and the operation instructions execute the corresponding program. Optionally, the touch panel can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch orientation of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into touch point coordinates, and then sends it to the processor 1301, and can receive commands sent by the processor 1301 and execute them. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it transmits it to the processor 1301 to determine the type of touch event. Subsequently, the processor 1301 provides a corresponding visual output on the display panel according to the type of touch event. In the embodiments of the present application, the touch panel and the display panel can be integrated into the touch display screen 1303 to implement input and output functions. However, in some embodiments, the touch panel and the touch panel can be implemented as two independent components to implement input and output functions. That is, the touch display screen 1303 can also be used as part of the input unit 1306 to implement the input function.

[0157] The radio frequency circuit 1304 can be used to transmit and receive radio frequency signals to establish wireless communication with a network device or other computer devices through wireless communication, and transmit and receive signals with the network device or other computer devices.

[0158] The audio circuit 1305 can be used to provide an audio interface between the user and the computer device through a speaker and a microphone. The audio circuit 1305 can convert the received audio data into an electrical signal and transmit it to the speaker, which converts it into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 1305 and then converted into audio data. After the audio data is output to the processor 1301 for processing, it is sent through the radio frequency circuit 1304 to, for example, another computer device, or the audio data is output to the memory 1302 for further processing. The audio circuit 1305 may also include an earphone jack to provide communication between a peripheral earphone and the computer device.

[0159] The input unit 1306 can be used to receive input digital, character information or object feature information (such as fingerprint, iris, facial information, etc.), and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0160] The power supply 1307 is used to supply power to each component of the computer device 1300. Optionally, the power supply 1307 can be logically connected to the processor 1301 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 1307 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0161] Although Figure 13 not shown in the figure, the computer device 1300 may also include a camera, a sensor, a Wi-Fi module, a Bluetooth module, etc., which will not be elaborated here.

[0162] This application also provides a computer-readable storage medium for storing a computer program. The computer-readable storage medium can be applied to a computer device, and the computer program enables the computer device to execute the corresponding process in the document format conversion method in the embodiments of this application. For the sake of brevity, it will not be elaborated here.

[0163] This application also provides a computer program product. The computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, enabling the computer device to execute the corresponding process in the document format conversion method in the embodiments of this application. For the sake of brevity, it will not be elaborated here.

[0164] The present application also provides a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the corresponding processes in the document format conversion method in the embodiments of the present application. For the sake of brevity, details are not described herein again.

[0165] It should be understood that the processor in the embodiments of the present application may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method embodiments can be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0166] It can be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include but not be limited to these and any other suitable types of memory.

[0167] It should be understood that the above memory is by way of example but not limitation. For example, the memory in the embodiments of the present application can also be a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synch link dynamic random access memory (SLDRAM), and a direct rambus random access memory (Direct Rambus RAM), etc. That is to say, the memory in the embodiments of the present application is intended to include but not be limited to these and any other suitable types of memory.

[0168] Those of ordinary skill in the art will realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0169] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0170] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0171] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0172] In addition, the functional units in the embodiments of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0173] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer or a server) to execute all or part of the steps of the methods described in the various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0174] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A document format conversion method, characterized in that The method includes: Obtaining a to-be-recognized picture containing the layout of a PPT document; Performing character recognition processing on the to-be-recognized picture to obtain the text information in the to-be-recognized picture and the text position corresponding to the text information; Performing region segmentation processing on the to-be-recognized picture based on an instance segmentation model to obtain the target regions in the to-be-recognized picture, where the target regions at least include text box regions and image regions, and also include at least one of a graphics region, a table region, a mathematical formula region, a chart region, and a flowchart region. The instance segmentation model includes a feature extraction module, a feature fusion module, a target classification module, a target localization module, a target contour extraction module, and a semantic segmentation module. Among them, the steps of the region segmentation processing include: inputting the to-be-recognized picture into the instance segmentation model, and extracting the feature vector of the to-be-recognized picture through the feature extraction module; performing feature fusion processing on the feature vector of the to-be-recognized picture through the feature fusion module to obtain the fused feature of the to-be-recognized picture; processing the fused feature of the to-be-recognized picture based on the target classification module, the target localization module, the target contour extraction module, and the semantic segmentation module in the instance segmentation model to obtain an instance segmentation result; determining the final region segmentation result according to the instance segmentation result and the confidence of each candidate region in the instance segmentation result, and the final region segmentation result includes all the target regions in the to-be-recognized picture; Generating a target PPT document that matches the PPT document layout in the to-be-recognized picture according to the text information, the text position, and the target regions, where the content corresponding to the text information and the target regions in the target PPT document has an editable attribute. The content corresponding to the target regions includes the region border of the target region, the content of the target region itself, the relevant plug-ins called in the target PPT document that match the target regions, and the content written into and displayed in the matching relevant plug-ins; Extracting the background image of the to-be-recognized picture according to the target regions; Inserting the background image into the target PPT document.

2. The document format conversion method according to claim 1, wherein The generating a target PPT document that matches the initial PPT document layout in the to-be-recognized picture according to the text information, the text position, and the target regions includes: Performing position matching between the text box region in the target region and the text position corresponding to the text information, so as to locate the text box region to the first target position corresponding to the text position in the target PPT document, and writing the text information into the text box region displayed at the first target position in the target PPT document; Cropping the image region in the target region, and inserting the cropped and scaled image region into the target PPT document according to a preset ratio.

3. The document format conversion method according to claim 2, wherein Generating a target PPT document that matches the initial PPT document layout in the to-be-recognized picture according to the text information, the text position, and the target area further includes: If the target area further includes a graphics area, calling a target graphics plug-in that matches the graphics area from the target PPT document and writing the target graphics plug-in into the target PPT document.

4. The document format conversion method according to claim 2, wherein Generating a target PPT document that matches the initial PPT document layout in the to-be-recognized picture according to the text information, the text position, and the target area further includes: If the target area further includes a table area, cropping the table area, recognizing the table content in the cropped table area, calling a table plug-in that matches the table area from the target PPT document, and writing the table content into the target PPT document according to the position of the table area and the table plug-in.

5. The document format conversion method according to claim 2, characterized in that, Generating a target PPT document that matches the initial PPT document layout in the to-be-recognized picture according to the text information, the text position, and the target area further includes: If the target area further includes a mathematical formula area, recognizing the formula content in the mathematical formula area, calling a formula editor that matches the mathematical formula area from the target PPT document, and writing the formula content into the target PPT document according to the position of the mathematical formula area and the formula editor.

6. The document format conversion method according to claim 2, wherein Generating a target PPT document that matches the initial PPT document layout in the to-be-recognized picture according to the text information, the text position, and the target area further includes: If the target area further includes a chart area, recognizing the chart content in the chart area, calling a chart plug-in that matches the chart area from the target PPT document, and writing the chart content into the target PPT document according to the position of the chart area and the chart plug-in.

7. The document format conversion method according to claim 2, wherein Generating a target PPT document that matches the initial PPT document layout in the to-be-recognized picture according to the text information, the text position, and the target area further includes: If the target area further includes a flowchart area, recognizing the flowchart content in the flowchart area, calling a flowchart plug-in that matches the flowchart area from the target PPT document, and writing the flowchart content into the target PPT document according to the position of the flowchart area and the flowchart plug-in.

8. The document format conversion method according to any one of claims 1-7, characterized in that Obtaining the to-be-recognized picture includes: Identifying image frames containing the PPT document layout from the to-be-recognized video file, sorting the obtained multiple image frames according to the corresponding timestamps to form a to-be-recognized picture set.

9. The document format conversion method according to claim 8, wherein Generating a target PPT document that matches the PPT document layout in the to-be-recognized picture according to the text information, the text position, and the target area further includes: Generate a target PPT document corresponding to each to-be-recognized picture in the to-be-recognized picture set according to the text information, the text position, and the target area corresponding to each to-be-recognized picture in the to-be-recognized picture set; Merge the target PPT documents corresponding to each to-be-recognized picture in the to-be-recognized picture set to obtain a target PPT file corresponding to the to-be-recognized video file.

10. The document format conversion method according to any one of claims 1-7, characterized in that The character recognition processing of the to-be-recognized picture to obtain the text information in the to-be-recognized picture and the text position corresponding to the text information includes: Performing character recognition processing on the to-be-recognized picture based on the optical character recognition (OCR) technology to locate and recognize the text information in the to-be-recognized picture, and determining the text information in the to-be-recognized picture and the text position corresponding to the text information.

11. A document format conversion device, characterized in that, The device includes: An acquisition unit, configured to acquire a to-be-recognized picture containing the layout of a PPT document; A recognition unit, configured to perform character recognition processing on the to-be-recognized picture to obtain the text information in the to-be-recognized picture and the text position corresponding to the text information; A segmentation unit, configured to perform region segmentation processing on the to-be-recognized picture based on an instance segmentation model to obtain the target area in the to-be-recognized picture, where the target area at least includes a text box area and an image area, and further includes at least one of a graphics area, a table area, a mathematical formula area, a chart area, and a flowchart area. The instance segmentation model includes a feature extraction module, a feature fusion module, a target classification module, a target localization module, a target contour extraction module, and a semantic segmentation module. The steps of the region segmentation processing include: inputting the to-be-recognized picture into the instance segmentation model, and extracting a feature vector of the to-be-recognized picture through the feature extraction module; performing feature fusion processing on the feature vector of the to-be-recognized picture through the feature fusion module to obtain a fusion feature of the to-be-recognized picture, so as to integrate the high-level feature vector and the low-level feature vector of the to-be-recognized picture output by multiple backbone networks in the feature extraction module; processing the fusion feature of the to-be-recognized picture based on the target classification module, the target localization module, the target contour extraction module, and the semantic segmentation module in the instance segmentation model to obtain an instance segmentation result; and determining a final region segmentation result according to the instance segmentation result and the confidence of each candidate region in the instance segmentation result, where the final region segmentation result includes all target areas in the to-be-recognized picture; A generating unit, configured to generate a target PPT document that matches the PPT document layout in the to-be-recognized picture according to the text information, the text position, and the target area, where the content corresponding to the text information and the target area in the target PPT document has an editable attribute, and the content corresponding to the target area includes the area border of the target area, the own content of the target area, the relevant plug-ins called in the target PPT document and matching the target area, and the content written into and displayed in the matching relevant plug-ins; The generating unit is further configured to extract the background image of the to-be-recognized picture according to the target area and insert the background image into the target PPT document.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the steps in the document format conversion method according to any one of claims 1-10.

13. A computer device, characterized in that, The computer device includes a processor and a memory. The memory stores a computer program, and the processor is configured to execute the steps in the document format conversion method according to any one of claims 1-10 by calling the computer program stored in the memory.

14. A computer program product, comprising computer instructions, characterized in that, When the computer instruction is executed by the processor, the steps in the document format conversion method according to any one of claims 1-10 are implemented.

Citation Information

Patent Citations

  • Method for extracting PPT file information from video file and related equipment

    CN110414352A

  • Document image recognition method and device, electronic equipment and computer readable medium

    CN114187448A