Image Processing Method, Apparatus, Device, and Storage Medium
By identifying and stitching the document materials in the image, the problem of difficulty for the listener to take notes is solved, and the effect of efficiently organizing PPT or courseware documents is achieved.
Patent Information
- Application Number
- CN202011065951.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2040-09-30
AI Technical Summary
When using PPT or courseware to give a speech, it is difficult for the listener to record notes in a timely manner, and the existing methods of recording videos or taking photos and sorting documents are inefficient.
By identifying the image content, intercepting images containing document materials, performing perspective correction, cropping and stitching, and outputting splicing files in electronic document format.
It improves the efficiency of sorting PPT or courseware documents and reduces the user's time to sort out.
Smart Images

Figure CN114359920B_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to the field of image technology, and particularly relates to an image processing method, apparatus, device, and storage medium. Background Art
[0002] With the development of technology, currently, in meetings, trainings, and teachings, the method of applying other document materials such as PPTs or courseware is very popular. Using the method of documents such as PPTs or courseware for lectures brings convenience to the lecturer, and can avoid the low efficiency of real-time writing on a whiteboard or blackboard during the lecture. However, it brings inconvenience to the listeners. Since the method of using documents such as PPTs or courseware saves the real-time writing time of the lecturer, the lecture speed will be relatively fast, and the listeners will not have enough time to take notes.
[0003] Now, most listeners use the method of recording videos or taking photos to record the content of documents such as PPTs or courseware, and after the lecture, they organize the PPTs or courseware and other documents. This method is inefficient. Summary of the Invention
[0004] In view of the above-mentioned defects or deficiencies in the prior art, it is desirable to provide an image processing method, apparatus, device, and storage medium.
[0005] In a first aspect, the present application provides an image processing method, which includes:
[0006] Identifying the image content of N images;
[0007] When it is recognized that the image content of the N images contains document materials, intercepting the document material images from the N images to obtain M intercepted images;
[0008] Stitching the M intercepted images and outputting the stitched file in an electronic document format;
[0009] Wherein, N is a positive integer, and M is a positive integer less than or equal to N.
[0010] In one embodiment, the image is a video frame image;
[0011] Before identifying the image content of the N images, it further includes:
[0012] Obtaining the marked points recorded in the target video;
[0013] Determining the video frame images corresponding to the marked points in the target video according to the marked points.
[0014] In one embodiment, before obtaining the marked points recorded in the target video, it further includes:
[0015] During the recording of the target video or during the playback of the target video, receive a marking input on the target video;
[0016] In response to the marking input, mark a marking point on the video frame image corresponding to the target video;
[0017] Wherein, one video frame image corresponds to each marking point.
[0018] In one embodiment, splicing the intercepted images includes:
[0019] Obtain the playback timing of the document material images corresponding to the M intercepted images in the target video;
[0020] Determine the first splicing order of the M intercepted images according to the playback timing;
[0021] Splice the M intercepted images according to the first splicing order.
[0022] In one embodiment, splicing the intercepted images includes:
[0023] Determine the document page numbers of the document material images corresponding to the M intercepted images;
[0024] Determine the second splicing order of the M intercepted images according to the document page numbers;
[0025] Splice the M intercepted images according to the second splicing order.
[0026] In one embodiment, in the step of intercepting document material images from N images to obtain M intercepted images:
[0027] In the case where there are identical document material images among the document material images intercepted from the N images, use one of the identical document material images as one intercepted image.
[0028] In one embodiment, when it is recognized that there is an included angle between any boundary of the document material included in the image content and the boundary of the image where the document material is located, and the included angle is greater than the included angle threshold,
[0029] Intercepting document material images from N images includes:
[0030] Perform perspective correction and cropping on the document material.
[0031] In one embodiment, the electronic document format includes any one of presentation file format, PDF format, rich text format, word format, and text editing system document format.
[0032] In one embodiment, the document material includes any one of PPT documents and courseware documents.
[0033] In a second aspect, the present application provides an image processing apparatus, which includes:
[0034] An identification module, configured to identify the image content of N images;
[0035] A cropping module, configured to crop document material images from the N images to obtain M cropped images when it is identified that the image content of the N images contains document materials;
[0036] An output module, configured to splice the M cropped images and output the spliced file in an electronic document format.
[0037] In a third aspect, the present application provides a device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the image processing method as in the first aspect.
[0038] In a fourth aspect, the present application provides a readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the image processing method as in the first aspect.
[0039] The technical solution provided by the embodiments of the present application effectively reduces the time for organizing documents such as PPTs or courseware and improves efficiency by cropping document material images from images containing document materials, splicing the cropped document material images, and outputting the spliced file in an electronic document format. Description of the Drawings
[0040] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, objectives, and advantages of the present application will become more apparent:
[0041] Figure 1 It is a flowchart of the image processing method provided by the embodiments of the present invention;
[0042] Figure 2 It is a structural diagram of the image processing apparatus provided by the embodiments of the present invention;
[0043] Figure 3 It is a structural diagram of an electronic device provided by the embodiments of the present invention. Detailed Embodiments
[0044] The present application will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention and are not intended to limit the invention. Additionally, it should be noted that for the sake of description, only parts related to the invention are shown in the drawings.
[0045] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.
[0046] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described can be implemented in an order other than those illustrated or described here.
[0047] In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or modules does not necessarily have to be limited to those steps or modules clearly listed, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products or devices.
[0048] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will detail this application with reference to the drawings and in conjunction with the embodiments.
[0049] Currently, in meetings, trainings, and teachings, the use of PPT or other document materials such as courseware is very popular. Using PPT or other documents for presentations brings convenience to the presenter, avoiding the low efficiency of real-time writing on a whiteboard or blackboard during the presentation. Since the time for the presenter's real-time writing is saved when using PPT or other documents, the presentation speed will be relatively fast, and the listeners will not have enough time to take notes.
[0050] Now, most listeners use the method of recording videos or taking photos to record the content of documents such as PPT or courseware, and then organize the PPT or courseware after the presentation. This method is less efficient.
[0051] Based on the above problems, this application hopes to propose an image processing method that is highly efficient and has a high user satisfaction when organizing document materials such as PPT or courseware recorded by the method of recording videos or taking photos.
[0052] The above method can be applied to a terminal device equipped with a camera. The terminal device can be a mobile phone, a tablet computer, a laptop computer, a smart helmet, smart glasses, a smartwatch, etc.
[0053] It should be noted that for the image processing method provided in the embodiments of the present invention, the execution subject may be an image processing device, and the image processing device may be implemented as part or all of a terminal device through software, hardware, or a combination of software and hardware. In the following method embodiments, the execution subject is taken as a terminal device as an example for description.
[0054] Referring to Figure 1 , which shows a schematic flowchart of an image processing method provided according to an embodiment of the present application.
[0055] As Figure 1 shown, an image processing method may include:
[0056] S110. Recognize the image content of N images.
[0057] Specifically, the image may be a video frame image (i.e., the frame image corresponding to a certain frame in the video), or a picture image (such as a photo taken by a camera, a screenshot image, etc.). The image may be directly obtained from the terminal device that records the video or picture, or obtained from a storage device that stores the recorded video or picture, or obtained through downloading, etc. Here, the form and acquisition method of the image are not limited.
[0058] The recognition of the image content of the image can be achieved by training a neural network. It can also be recognized by other means.
[0059] If the image is a picture image, the obtained picture image can be directly input into the neural network model to complete the recognition of the image content of the image.
[0060] If the image is a video frame image, the obtained video needs to be processed first to obtain the video frame image.
[0061] In one embodiment, when the image is a video frame image, before recognizing the image content of N images, the method further includes:
[0062] Obtain the marked points recorded in the target video;
[0063] Determine the video frame image corresponding to the marked points in the target video according to the marked points.
[0064] Specifically, the target video is a video with marked points marked and recorded in a video recorded by a user, or a stored video, or a downloaded video, etc. Among them, the marked points marked and recorded can be input by the user or input by a terminal device, etc. The number of marked points is N, and N is a positive integer. The video frame image can be any frame image in the image. In this embodiment, the video frame image is the video frame image corresponding to the marked point in the target video. Marking N marked points corresponds to N video frame images.
[0065] In one embodiment, before obtaining the marked points marked and recorded in the target video, it further includes:
[0066] During the recording process of the target video or during the playing process of the target video, receive the mark input on the target video;
[0067] In response to the mark input, mark the marked points in the corresponding video frame image of the target video;
[0068] Among them, one video frame image corresponds to each marked point.
[0069] Specifically, when the user records or plays a video, mark the marked points on the video while recording or playing the video according to actual needs. When marking the marked points, it can be set to automatically mark at preset time intervals, and the preset time interval can be set according to actual needs. It can be understood that if the preset time interval is set too large, that is, marking a marked point at a long interval, it may miss marking the images containing document materials in the image content; if the preset time interval is set too small, that is, marking a marked point at a short interval, the images containing the same document materials in the image content may be marked multiple times, and there will be more images to be recognized when recognizing the images, which takes a long time. The preset time interval can be set according to the learning and training of the neural network model.
[0070] When marking the marked points, it can also be that the user manually marks the marked points while recording or playing the video and turning the pages of the PPT or courseware, etc. during the speech.
[0071] When marking the marked points, an algorithm for judging whether the document materials contained in the image content of adjacent frames of images change can also be used for real-time judgment. If it is judged that the document materials contained in the image content of adjacent frames of images change, the marked points can be automatically marked, or a pop-up window can be used to ask the user whether to mark. The user can choose whether to mark the marked points according to actual needs. It should be noted that other methods can also be used to mark the marked points on the video, and no limitation is made here.
[0072] After obtaining the target video and the marked points marked in the target video, the video frame images marked can be determined according to the marked points in the target video. When identifying the image content of the images, all the marked video frame images can be input into a neural network model to complete the determination of whether the image content of the video frame images contains document materials.
[0073] S120. When it is recognized that the image content of N images contains document materials, cut out the document material images from the N images to obtain M cut-out images.
[0074] Specifically, when it is recognized that the image content of an image contains document materials, due to environmental factors during recording, the image may be too dark or overexposed. At this time, the too dark or overexposed image needs to be processed to the normal brightness range first. This processing can adopt existing technologies and will not be elaborated here. Optionally, the document materials can include any one of PPT documents, courseware documents, etc.
[0075] For the processed images, detect the boundaries of the document materials in the images according to the boundary recognition technology, and crop the images according to the detected boundaries. It can be understood that in order to make the cropped document material images beautiful, when cropping the images according to the detected boundaries, the boundaries on all four sides can be extended outward (that is, the left boundary extends to the left, the right boundary extends to the right, the upper boundary extends upward, and the lower boundary extends downward) by a preset length. The preset lengths extended in these four directions can be equal or unequal and can be set according to actual needs.
[0076] In one embodiment, in the step of cutting out the document material images from N images to obtain M cut-out images:
[0077] In the case where there are identical document material images among the document material images cut out from the N images, one of the identical document material images is used as one cut-out image.
[0078] Specifically, since there may be identical document material images among the document material images cut out from the N images, it is possible to determine whether there are identical document materials in the image content of the cut-out images. If there are, one cut-out image corresponding to one of the identical document materials is retained, and the cut-out images corresponding to the remaining copies of the identical document materials are all excluded.
[0079] When determining whether there are identical document materials in the image content of the cut-out images, a comparison algorithm for the text in the images can be used to compare the document materials contained in all the cut-out images.
[0080] As described above, since there may be cases where the same document image exists in the intercepted document image data, the number of intercepted images obtained may be less than or equal to the number of images. That is, the number M of the obtained screenshot images is a positive integer less than or equal to the number N of images.
[0081] When recording a video, it is generally not recorded directly facing the screen. That is, the document data included in the recorded video is usually inclined (here, inclined means that any boundary of the document data has an angle with the corresponding boundary of the image where the document data is located, and the angle is greater than the angle threshold). Therefore, when intercepting document data, it is necessary to use the method of tilt detection and correction to process it. That is, first perform tilt detection on the document data. If the document data is inclined, it is necessary to correct the document data first. Commonly used tilt detection methods include: text line-based detection method, projection profile analysis method, Hough transform method, etc.
[0082] In one embodiment, when it is recognized that there is an angle between any boundary of the document data included in the image content and the corresponding boundary of the image where the document data is located, and the angle is greater than the angle threshold, intercepting the document data in the image includes: performing perspective correction and cropping on the document data.
[0083] Specifically, performing perspective correction and cropping on the document data, that is, correcting the angle between all boundaries of the document data and the corresponding boundary of the image where the document data is located to within the angle threshold. Photoshop technology can be used, or other technologies such as distorted document image restoration technology can also be used, which are not limited here.
[0084] The angle threshold can be set according to actual needs. Exemplarily, the angle threshold can be set to 5°.
[0085] S130. Stitch the intercepted images and output the stitched file in the format of an electronic document.
[0086] Specifically, the intercepted images are document images intercepted from the image. Stitching the intercepted images may include stitching the screenshot images to obtain a stitched image, and outputting the stitched image in the format of an electronic document. It may also include inputting the intercepted images into a word document, PPT document, PDF document, etc., stitching the intercepted images in any of the above documents, or taking each intercepted image as a page in any of the above documents, and then uniformly outputting it as a stitched file in the format of an electronic document. After outputting the stitched file in the format of an electronic document, the path for saving the file can be sent to the user, and the saved file can be found in the file manager.
[0087] The electronic document format can be set according to the actual needs of users. Optionally, the electronic document format can include any one of the presentation file format, PDF (Portable Document Format) format, Rich Text Format (RTF) format, word format, and Word Processing System (WPS) format. It can also include Excel workbook format, web page format, MHT file format, and other image-displayable formats.
[0088] It can be understood that when a speaker is giving a speech, it often happens that they jump back to previously presented documents such as PPTs or courseware. In this case, the video frame images marked by the person recording the video or the photos taken by the person taking pictures may contain the same content as the video frame images or photos corresponding to the previous marked points. If all the intercepted images are directly spliced, the page numbers in the spliced image may not correspond to the page numbers of the original PPT or courseware, etc., and may contain duplicate content. Therefore, when splicing, it is necessary to sort the intercepted images.
[0089] In one embodiment, splicing the intercepted images includes:
[0090] Obtain the playback timing of the document material images corresponding to M intercepted images in the target video;
[0091] Determine the first splicing order of M intercepted images according to the playback timing;
[0092] Splice M intercepted images according to the first splicing order.
[0093] Specifically, the playback timing of the document material images in the target video is related to the time when the marking points are marked in the target video. The playback timing corresponding to the previously marked marking points is earlier, and the playback timing corresponding to the later marked marking points is later. That is, the playback timing of the document material images in the target image is the time order of marking the marking points.
[0094] The first splicing order is the display order of the intercepted images in the output spliced file. This first splicing order is consistent with the playback timing and is the time order when the marking points are marked. Splice M intercepted images according to this first splicing order.
[0095] In one embodiment, splicing the intercepted images includes:
[0096] Determine the document page numbers of the document material images corresponding to M intercepted images;
[0097] Determine the second splicing order of M intercepted images according to the document page numbers;
[0098] Splice M intercepted images according to the second splicing order.
[0099] Specifically, usually in document materials such as PPT or courseware, the page number can be set at the top, bottom, left, middle, right, etc. of the page number. Detect the possible positions of the page number in the intercepted image, determine the page number of the document material image, and determine the document page numbers of the M intercepted images according to the page number of the document material image.
[0100] The second splicing order is the display order of the intercepted images in the output spliced file, and this second splicing order is consistent with the page number order of document materials such as PPT or courseware. Splice M intercepted images according to this second splicing order.
[0101] In the embodiments of the present application, when it is recognized that the image content of N images contains document materials, intercept the document material images from the N images to obtain M intercepted images, splice the M intercepted images, and output the spliced file in an electronic document format, which can reduce the time for users to organize documents such as PPT or courseware and improve efficiency.
[0102] The following takes the example of recording a tag (mark) video to illustrate the image processing method proposed in the embodiments of the present application.
[0103] The user records a tag video with a camera (manually marking tags while recording). After recording, open the tag video album on the mobile phone, and the viewing mark entry can be displayed. Click the viewing mark entry, and the viewing mark expands to the time in the video corresponding to the video frame image of each mark point. Identify each video frame image corresponding to the mark point to determine whether the video frame image contains file materials. When it is recognized that the video frame image contains document materials, an export file button can be displayed on the mobile phone album interface. Click the export file button to intercept and splice the files in the video, perform perspective correction and cropping on the pages that need to be corrected. During the splicing process, the document material order can be determined based on information such as the page and time of the video frame image, and it can be judged whether there are the same document materials. If so, perform duplicate removal processing. The document materials in the video are exported and stored in a PDF format file, and the storage path of the file can also be prompted to the user, and the file can be found in the file manager.
[0104] As Figure 2 is a schematic structural diagram of an image processing device 200 provided by an embodiment of the present application. As Figure 2 shown, this device can implement the method as Figure 1 shown, and this device may include:
[0105] An identification module 210, configured to identify the image content of N images;
[0106] The cropping module 220 is configured to crop the document images from the N images to obtain M cropped images when it is recognized that the image content of the N images contains document materials;
[0107] The output module 230 is configured to splice the M cropped images and output the spliced file in the format of an electronic document.
[0108] Optionally, the image is a video frame image, and the apparatus further includes:
[0109] The first acquisition module is configured to acquire a target video and the marked points marked in the target video;
[0110] The determination module is configured to determine the video frame image corresponding to the marked point in the target video according to the marked point.
[0111] Optionally, the apparatus further includes:
[0112] The input receiving module is configured to receive the marking input on the target video during the recording process or the playing process of the target video;
[0113] The response module is configured to mark the marked point on the video frame image corresponding to the marked input in the target video in response to the marked input;
[0114] Wherein, one video frame image corresponds to each marked point.
[0115] Optionally, the output module 230 is further configured to:
[0116] Obtain the playing time sequence of the document images corresponding to the M cropped images in the target video;
[0117] Determine the first splicing order of the M cropped images according to the playing time sequence;
[0118] Splice the M cropped images according to the first splicing order.
[0119] Optionally, the output module 230 is further configured to:
[0120] Determine the document page numbers of the document images corresponding to the M cropped images;
[0121] Determine the second splicing order of the M cropped images according to the document page numbers;
[0122] Splice the M cropped images according to the second splicing order.
[0123] Optionally, the cropping module 220 is further configured to:
[0124] In the case where there are identical document images among the document images cropped from the N images, use one of the identical document images as one cropped image.
[0125] Optionally, when there is an angle between any boundary of the document material included in the recognized image content and the boundary corresponding to the image where the document material is located, and the angle is greater than the angle threshold, the cropping module 220 is further configured to:
[0126] Perform perspective correction cropping on the document material.
[0127] Optionally, the electronic document format includes any one of presentation file format, PDF format, rich text format, word format, and text editing system document format.
[0128] Optionally, the document material includes any one of PPT documents and courseware documents.
[0129] The image processing device provided in this embodiment can execute the embodiments of the above method, and its implementation principle and technical effects are similar, which will not be elaborated here.
[0130] Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 3 shown, it shows a schematic structural diagram of an electronic device 300 suitable for implementing the embodiments of the present application.
[0131] As Figure 3 shown, the electronic device 300 includes a central processing unit (CPU) 301, which can execute various appropriate actions and processes according to the program stored in the read-only memory (ROM) 302 or the program loaded from the storage section 308 into the random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the device 300 are also stored. The CPU 301, ROM 302, and RAM 303 are connected to each other through a bus 304. The input / output (I / O) interface 306 is also connected to the bus 304.
[0132] The following components are connected to the I / O interface 305: an input section 306 including a keyboard, a mouse, etc.; an output section 307 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. The drive 310 is also connected to the I / O interface 306 as required. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as required, so that the computer program read from it can be installed into the storage section 308 as required.
[0133] In particular, according to the embodiments of the present disclosure, as referred to above Figure 1The described process can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product that includes a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing the above-described image processing method. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 309, and / or installed from the removable medium 311.
[0134] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code, and the foregoing module, segment of a program, or portion of code includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0135] The units or modules described in the embodiments of the present application can be implemented in software or in hardware. The described units or modules can also be provided in a processor. The names of these units or modules do not, in some cases, constitute a limitation on the units or modules themselves.
[0136] As another aspect, the present application also provides a storage medium, which can be the storage medium included in the foregoing device in the above embodiments; or can exist separately and be not assembled into the device. The storage medium stores one or more programs, and the foregoing programs are used by one or more processors to perform the image processing method described in the present application.
[0137] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, technical solutions formed by mutually replacing the above features with (but not limited to) technical features having similar functions disclosed in the present application.
Claims
1. An image processing method, characterized in that, The method includes: Identifying the image content of N images; the images are at least one of video frame images or picture images; When it is recognized that the image content of the N images contains document materials, intercepting document material images from the N images to obtain M intercepted images; Stitching the M intercepted images and outputting the stitched file in an electronic document format; Wherein, N is a positive integer, and M is a positive integer less than or equal to N; When the image is a video frame image, before identifying the image content of the N images, the method further includes: Obtaining the marked points recorded in the target video; according to the marked points, determining the video frame images corresponding to the marked points in the target video; the target video is a video with marked points recorded in a video recorded by the user, or a stored video, or a downloaded video, etc., wherein each marked point corresponds to a video frame image; Before obtaining the marked points recorded in the target video, it further includes: Receiving a marking input on the target video during the recording process or the playing process of the target video; In response to the marking input, marking marked points on the corresponding video frame images in the target video; Wherein, the generation process of the marking input includes: Obtaining the change situation of the document materials included in the image content of adjacent video frame images, and if the document materials included in the image content of the adjacent video frame images change, generating the marking input by an automatic marking method or a method of generating a pop-up window to ask the user.
2. The method according to claim 1, wherein The stitching of the M intercepted images includes: Obtaining the playing time sequence of the document material images corresponding to the M intercepted images in the target video; Determining the first stitching order of the M intercepted images according to the playing time sequence; Stitching the M intercepted images according to the first stitching order.
3. The method according to claim 1, wherein The stitching of the M intercepted images includes: Determining the document page numbers of the document material images corresponding to the M intercepted images; Determining the second stitching order of the M intercepted images according to the document page numbers; Stitching the M intercepted images according to the second stitching order.
4. The method according to claim 1, wherein In the step of intercepting document material images from the N images to obtain M intercepted images: In the case where there are identical document material images among the document material images intercepted from the N images, taking one of the identical document material images as one of the intercepted images.
5. The method according to claim 1, wherein When it is recognized that there is an included angle between any boundary of the document materials included in the image content and the boundary corresponding to the image where the document materials are located, and the included angle is greater than the included angle threshold, The intercepting of the document material images from the N images includes: Performing perspective correction and cropping on the document materials.
6. The method according to claim 1, wherein The electronic document format includes any one of presentation file format, PDF format, rich text format, word format, and text editing system document format.
7. The method according to claim 1, characterized in that The document materials include any one of PPT documents and courseware documents.
8. An image processing apparatus, characterized in that, The device includes: An identification module for identifying the image content of N images; the images are at least one of video frame images or picture images; A cropping module for cropping document material images from the N images to obtain M cropped images when it is recognized that the image content of the N images contains document materials; An output module for splicing the M cropped images and outputting a spliced file in an electronic document format; When the images are video frame images, before identifying the image content of the N images, the apparatus is further configured to: Obtain the marked points marked in the target video; determine the video frame images corresponding to the marked points in the target video according to the marked points; the target video is a video with marked points marked in a video recorded by a user, or a stored video, or a downloaded video, etc., where each marked point corresponds to a video frame image; Before obtaining the marked points marked in the target video, it further includes: Receiving a marking input on the target video during the recording process or the playing process of the target video; In response to the marking input, marking marked points on the corresponding video frame images in the target video; Wherein, the generation process of the marking input includes: Obtaining the change situation of the document materials contained in the image content of adjacent video frame images, and if the document materials contained in the image content of the adjacent video frame images change, generating the marking input by an automatic marking method or a method of generating a pop-up window to ask the user.
9. An apparatus, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the image processing method according to any one of claims 1-7.
10. A readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the image processing method according to any one of claims 1-7.
Citation Information
Patent Citations
Video note generation method and device, storage medium and computer device
CN110381382A
Method for extracting PPT file information from video file and related equipment
CN110414352A
Video tagging system
US20130051754A1