Slide restoration method based on background reconstruction and reverse injection of time series
Through the slide restoration method based on the generation adversarial network and convolutional neural network, the problem of time-consuming and labor-intensive slide restoration and shooting jitter is solved, and fast and accurate slide restoration is achieved, improving learning and office efficiency.
Patent Information
- Application Number
- CN202210605531.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-31
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-05-31
AI Technical Summary
The existing slide restoration methods are time-consuming and labor-intensive, and problems such as jitter, tilt, and occlusion are prone to occur when shooting through a mobile terminal, which affects the restoration effect.
The image clarity is enhanced based on the generative adversarial network, combined with object detection and bilinear interpolation to correct edge contours, and the image features and recurrent neural networks are extracted using convolutional neural networks to learn timing features, and background template reverse injection is carried out through improved object detection algorithms to achieve graphic and text separation and text recognition.
Quickly and accurately restore background information, text information and picture information from slide pictures, improve learning and office efficiency, and realize the transformation from unstructured data to structured data.
Smart Images

Figure CN115239574B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing, and particularly relates to a method for restoring slides based on background reconstruction and reverse injection of time series. Background Art
[0002] In daily study, work, and communication, slides have gradually become a tool for sharing ideas in daily work due to their simplicity, strong expressiveness, and powerful functions. In real life, due to reasons such as slide file transmission and ownership, it is impossible to directly obtain the original slides from the speaker. Many audiences will quickly record the content on the slides through mobile phones, mobile terminals, or cameras at exhibitions, lectures, or in the classroom. After obtaining a series of photos, the user needs to copy all the photos to the computer, then manually crop each slide, extract the picture part, adjust it to the appropriate size of the slide, and then paste it into the designated position of the slide file, and manually edit the text part and fill it into the designated position of the slide.
[0003] This process is time-consuming and laborious. Especially when the user fails to organize in time and needs to view the recorded content and open the pictures one by one, the file management is rather messy. At the same time, when shooting through a mobile terminal, there will be situations such as shaking, tilting, and occlusion, which will affect the subsequent slide restoration effect.
[0004] In summary, the existing slide restoration methods have obvious inconveniences and advantages in actual use. There is an urgent need for an efficient and fast automatic slide restoration method that can restore text, pictures, and backgrounds and improve learning and work efficiency. Summary of the Invention
[0005] Aiming at the problems existing in the prior art, the present invention provides a method for restoring slides based on background reconstruction and reverse injection of time series, which can conveniently restore background information, text information, and picture information from slide pictures, improve learning and office efficiency, and realize the transformation from unstructured data to structured data.
[0006] To achieve the above technical objectives, the present invention adopts the following technical solutions: A method for restoring slides based on background reconstruction and reverse injection of time series, specifically including the following steps:
[0007] Step 1: Enhance the clarity of the slide picture file through an image enhancement algorithm based on a generative adversarial network.
[0008] Step 2: Obtain the edge contour of the slide picture through a slide contour detection algorithm based on object detection, and then correct the edge contour of the slide through a bilinear interpolation algorithm to obtain a corrected slide contour.
[0009] Step 3: The slide correction contour extracts the image features of the slide through a convolutional neural network, and then learns the temporal features of the slide through a recurrent neural network to extract the background template of the slide;
[0010] Step 4: By improving the object detection algorithm, the extracted background template is injected backward, and the slide picture is analyzed for text, and the text area and the picture area are divided;
[0011] Step 5: Apply text recognition technology to the text area to recognize the text;
[0012] Step 6: Save the position information and the picture of the picture area;
[0013] Step 7: In the extracted background template of the slide, paste the picture at the original picture position and write the text at the original text position to obtain the slide document.
[0014] Further, the specific process of Step 1 is as follows: The slide picture file is input into the generation module in the generative adversarial network, and the original blurred data in the slide picture is automatically processed into clear data.
[0015] Further, Step 2 includes the following sub-steps:
[0016] Step 21: Pass the slide picture with enhanced clarity through the object detection algorithm to extract the four vertex coordinates of the slide picture, recorded as (x1, y1), (x2, y2), (x3, y3), (x4, y4);
[0017] Step 22: According to the recorded coordinate position information of the four vertices of the slide picture, correct the edge contour of the slide picture through the bilinear interpolation algorithm, and convert the coordinates of the four vertices of the slide picture into rectangular coordinates, recorded as (xx1, yy1), (xx2, yy2), (xx3, yy3), (xx4, yy4).
[0018] Further, in Step 3, the convolutional neural network and the recurrent neural network are nested networks, and the output end of the convolutional neural network is connected to the input end of the recurrent neural network.
[0019] Further, the improved object detection model of the improved object detection algorithm is to add an input branch on the basis of the object detection model.
[0020] Further, Step 4 includes the following sub-steps:
[0021] Step 41: Input the slide picture and the background template into the improved object detection model respectively;
[0022] Step 42: Extract the slide image features F_ppt from the slide image through the convolutional neural network in the improved object detection model according to the slide image weight α, and extract the background template features F_template from the background template through the convolutional neural network in the improved object detection model according to the background template weight β;
[0023] Step 43: Perform weighted processing on the slide image features F_ppt and the background template features F_template to obtain the weighted feature F_final = α * F_ppt + β * F_template;
[0024] Step 44: Input the weighted feature into the pyramid network for scale fusion, and then extract the type and the position of the detection box through the classification branch and the detection branch respectively to obtain the text area and the picture area.
[0025] Further, the slide image weight α and the background template weight β are obtained by training the improved object detection model: Input the slide image and the background template into the improved object detection model, initialize the slide image weight α and the background template weight β to 0.5 respectively, use α and β as parameters to train the improved object detection model, perform backpropagation of the reverse gradient through the loss function corresponding to object detection, update the slide image weight α and the background template weight β until the loss function corresponding to object detection is minimized, complete the training of the improved object detection model, and output the updated slide image weight α and the background template weight β.
[0026] Compared with the prior art, the present invention has the following beneficial effects: The slide restoration method based on time series background reconstruction and reverse injection of the present invention uses a nested network of convolutional neural network and recurrent neural network. By inputting the slide image, the slide background template is extracted; at the same time, the improved object detection model with reverse injection of the background template is used. By inputting the slide image and the background template, during the process of object detection, the background template information can be referred to, and more accurate text areas and picture areas can be obtained. Through the slide restoration method based on time series background reconstruction and reverse injection of the present invention, the slide file can be quickly and accurately restored from the slide image file, improving the learning and office efficiency and realizing the transformation from unstructured data to structured data. Brief Description of the Drawings
[0027] Figure 1 It is a flowchart of the slide restoration method based on time series background reconstruction and reverse injection of the present invention;
[0028] Figure 2 It is a schematic diagram of the background template extraction process in the present invention;
[0029] Figure 3This is the flowchart of the improved object detection algorithm for template reverse injection in the present invention. Detailed implementation manners
[0030] The technical solution of the present invention will be further explained below with reference to the accompanying drawings.
[0031] As Figure 1 This is the flowchart of the slide restoration method based on time series background reconstruction and reverse injection in the present invention. The slide restoration method specifically includes the following steps:
[0032] Step 1: Enhance the clarity of the slide picture file through the image enhancement algorithm based on the generative adversarial network. Specifically, input the slide picture file into the generation module in the generative adversarial network, and automatically process the originally blurred data in the slide picture into clear data.
[0033] Step 2: Pass the slide picture with enhanced clarity through the slide contour detection algorithm based on object detection to obtain the edge contour of the slide picture, and then correct the edge contour of the slide through the bilinear interpolation algorithm to obtain the corrected slide contour. Through this step, the situation of picture tilt and distortion caused by the shooting angle or jitter can be corrected, and a basically positive slide contour can be obtained, providing better picture quality for subsequent slide extraction and slide recognition, and improving the slide restoration effect. Specifically, it includes the following sub-steps:
[0034] Step 21: Pass the slide picture with enhanced clarity through the object detection algorithm to extract the four vertex coordinates of the slide picture, recorded as (x1, y1), (x2, y2), (x3, y3), (x4, y4);
[0035] Step 22: According to the coordinate position information of the four vertices of the recorded slide picture, correct the edge contour of the slide picture through the bilinear interpolation algorithm, and convert the coordinates of the four vertices of the slide picture into rectangular coordinates, recorded as (xx1, yy1), (xx2, yy2), (xx3, yy3), (xx4, yy4).
[0036] Step 3: As Figure 2 , the corrected slide contour extracts the image features of the slide through the convolutional neural network, and then learns the temporal features of the slide through the recurrent neural network to extract the background template of the slide. Compared with the traditional pure convolutional network, by adding the recurrent neural network, the network can learn the temporal features and associate the front and back features, so as to better restore the slide background template. In the present invention, the convolutional neural network and the recurrent neural network are nested networks, and the output end of the convolutional neural network is connected to the input end of the recurrent neural network.
[0037] Step 4: By improving the object detection algorithm and injecting the extracted background template in reverse, perform discourse analysis on the slide images, divide the text regions and image regions, and achieve text-image separation; in the present invention, the improved object detection model of the improved object detection algorithm is to add an input branch on the basis of the object detection model. By improving the object detection algorithm and adding an input branch, the object detection model can obtain more features. By using the background template as a reference for comparison, the object detection model can better distinguish text and images and achieve text-image division. For example Figure 3 , specifically including the following sub-steps:
[0038] Step 41: Input the slide image and the background template into the improved object detection model respectively, and set the slide image weight α and the background template weight β;
[0039] Step 42: Pass the slide image through the convolutional neural network in the improved object detection model to extract the slide image feature F_ppt according to the slide image weight α, and pass the background template through the convolutional neural network in the improved object detection model to extract the background template feature F_template according to the background template weight β;
[0040] Step 43: Perform weighted processing on the slide image feature F_ppt and the background template feature F_template to obtain the weighted feature F_final = α * F_ppt + β * F_template;
[0041] Step 44: Input the weighted feature into the pyramid network for scale fusion to enhance its multi-scale detection performance, and then extract the type and detection box position through the classification branch and the detection branch respectively to obtain the text region and the image region.
[0042] In the present invention, the slide image weight α and the background template weight β are obtained by training the improved object detection model: input the slide image and the background template into the improved object detection model, initialize the slide image weight α and the background template weight β to 0.5 respectively, use α and β as parameters to train the improved object detection model, perform backpropagation of the reverse gradient through the loss function corresponding to object detection, update the slide image weight α and the background template weight β until the loss function corresponding to object detection is minimized, complete the training of the improved object detection model, and output the updated slide image weight α and the background template weight β.
[0043] Step 5: Adopt text recognition technology for the text region to recognize the text;
[0044] Step 6: Save the position information and the image of the image region;
[0045] Step 7: In the background template of the extracted slide, paste the picture at the original picture position and write the text at the original text position to obtain the slide document, thereby realizing the restoration of the slide picture.
[0046] The slide restoration method based on background reconstruction and reverse injection of time series of the present invention can quickly and accurately restore the slide file from the slide picture file, with good restoration effect. At the same time, it improves learning and office efficiency and realizes the transformation from unstructured data to structured data.
[0047] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above in the preferred implementation manner, it is not used to limit the present invention. Any person skilled in the art can make some changes or modifications to equivalent embodiments with equivalent changes by using the above-disclosed technical content within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention and is based on the technical essence of the present invention, any simple modification, equivalent replacement, and improvement of the above embodiments are still within the protection scope of the technical solution of the present invention.
Claims
1. A slide restoration method based on background reconstruction and reverse injection of time series, characterized in that, Specifically, it includes the following steps: Step 1: Enhance the clarity of the slide picture file through an image enhancement algorithm based on a generative adversarial network. Step 2: Pass the slide picture with enhanced clarity through a slide contour detection algorithm based on object detection to obtain the edge contour of the slide picture, and then correct the edge contour of the slide through a bilinear interpolation algorithm to obtain the corrected slide contour. Step 3: The corrected slide contour extracts the image features of the slide through a convolutional neural network, and then learns the temporal features of the slide through a recurrent neural network to extract the background template of the slide. Step 4: Through an improved object detection algorithm, inject the extracted background template backward to perform discourse analysis on the slide picture and divide the text area and the picture area; it includes the following sub-steps: Step 41: Input the slide picture and the background template into the improved object detection model respectively. Step 42: Pass the slide picture through the convolutional neural network in the improved object detection model to extract the slide picture feature F_ppt according to the slide picture weight α, and pass the background template through the convolutional neural network in the improved object detection model to extract the background template feature F_template according to the background template weight β. Step 43: Perform weighted processing on the slide picture feature F_ppt and the background template feature F_template to obtain the weighted feature F_final = α*F_ppt + β*F_template. Step 44: Input the weighted feature into a pyramid network for scale fusion, and then extract the type and the detection box position through the classification branch and the detection branch respectively to obtain the text area and the picture area. Step 5: Adopt an optical character recognition technology for the text area to recognize the text. Step 6: Save the position information and the picture of the picture area. Step 7: In the extracted background template of the slide, paste the picture at the original picture position and write the text at the original text position to obtain the slide document.
2. The slide restoration method based on time series background reconstruction and reverse injection according to claim 1, characterized in that The specific process of Step 1 is as follows: Input the slide picture file into the generation module in the generative adversarial network, and automatically process the original blurred data in the slide picture into clear data.
3. The slide restoration method based on time series background reconstruction and reverse injection according to claim 1, characterized in that Step 2 includes the following sub-steps: Step 21: Pass the slide picture with enhanced clarity through an object detection algorithm to extract the four vertex coordinates of the slide picture, recorded as (x1, y1), (x2, y2), (x3, y3), (x4, y4). Step 22: According to the recorded coordinate position information of the four vertices of the slide picture, correct the edge contour of the slide picture through a bilinear interpolation algorithm, and convert the coordinates of the four vertices of the slide picture into rectangular coordinates, recorded as (xx1, yy1), (xx2, yy2), (xx3, yy3), (xx4, yy4).
4. The slide restoration method based on time series background reconstruction and reverse injection according to claim 1, characterized in that In Step 3, the convolutional neural network and the recurrent neural network are nested networks, and the output end of the convolutional neural network is connected to the input end of the recurrent neural network.
5. The slide restoration method based on time series background reconstruction and reverse injection according to claim 1, characterized in that The improved object detection model of the improved object detection algorithm is to add an input branch on the basis of the object detection model.
6. The slide restoration method based on time series background reconstruction and reverse injection according to claim 1, characterized in that The weights α of the slide image and β of the background template are obtained by training to improve the object detection model: input the slide image and the background template into the improved object detection model, initialize the weights α of the slide image and β of the background template to 0.5 respectively, use α and β as parameters to train the improved object detection model, perform backpropagation of the gradient through the loss function corresponding to object detection, update the weights α of the slide image and β of the background template until the loss function corresponding to object detection is minimized, complete the training of the improved object detection model, and output the updated weights α of the slide image and β of the background template.
Citation Information
Patent Citations
Character detection method for fusing character area edge information in character image
CN110738207A
Presentation file generation method and device, equipment and medium
CN111753108A