Document generation method and device
By processing and grouping multiple images, the system automatically identifies and generates documents, solving the problem of low efficiency in manually organizing images to generate documents in existing technologies, and achieving efficient document generation.
Patent Information
- Application Number
- CN202511146239.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-25
AI Technical Summary
In existing technologies, users need to manually organize images to generate documents, which is inefficient.
By acquiring multiple images and performing image correction, cropping, composition, and filtering, similar image groups are identified and representative images are selected to automatically generate documents.
It enables automatic image organization and document generation, significantly reducing user complexity and improving document generation efficiency.
Smart Images

Figure CN121010673A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of communication, and specifically relate to a document generation method and device. BACKGROUND
[0002] In daily work and life, a user usually needs to arrange photographed images into a document for later viewing or sharing. For example, for images containing document content photographed in a meeting or a class, the user usually needs to arrange them into a document in the form of a presentation, for later viewing. For example, for images photographed in a concert, the user usually needs to arrange them into a document in the form of a nine-square grid, for sharing.
[0003] In the prior art, the user usually needs to manually arrange images, for example, manually selects the best images from photographed images, and manually arranges the selected best images to form a document. This document generation method is inefficient. SUMMARY
[0004] An object of embodiments of the present application is to provide a document generation method and device, which improves the efficiency of document generation.
[0005] In a first aspect, embodiments of the present application provide a document generation method, which includes: obtaining a plurality of first images collected in a target scene; performing first processing on the plurality of first images to obtain a plurality of second images, wherein the first processing includes at least one of image correction processing, image edge cutting processing, composition processing, filter processing, color enhancement processing, noise suppression processing, color cast correction processing, and sharpening processing; identifying similar image groups in the plurality of second images, and displaying a representative image of each similar image group, the representative image of each similar image group being determined based on a target element of images in the each similar image group, the target element including at least one of quality and content; and in a case where a first input of a user is received, generating a document according to the representative image of each similar image group.
[0006] In a second aspect, an embodiment of the present application provides a document generation apparatus, the apparatus comprising: an acquisition unit configured to acquire a plurality of first images captured in a target scene; a processing unit configured to perform first processing on the plurality of first images to obtain a plurality of second images, wherein the first processing comprises at least one of image correction processing, image edge cropping processing, composition processing, filter processing, color enhancement processing, noise suppression processing, color cast correction processing, and sharpening processing; a display unit configured to identify a similar image group in the plurality of second images, and display a representative image of each similar image group, the representative image of each similar image group being determined based on a target element of images in the each similar image group, the target element comprising at least one of quality and content; and a generation unit configured to, in a case where a first input of a user is received, generate a document according to the representative image of each similar image group.
[0007] In a third aspect, an embodiment of the present application provides an electronic device, the electronic device comprising a processor and a memory, the memory storing programs or instructions executable on the processor, the programs or instructions being executed by the processor to implement the steps of the method according to the first aspect.
[0008] In a fourth aspect, an embodiment of the present application provides a readable storage medium, the readable storage medium storing a computer program, the computer program being executed by a processor to implement the steps of the method according to the first aspect.
[0009] In a fifth aspect, an embodiment of the present application provides a chip, the chip comprising a processor and a communication interface, the communication interface being coupled to the processor, the processor being configured to execute programs or instructions to implement the method according to the first aspect.
[0010] In a sixth aspect, an embodiment of the present application provides a computer program product, the program product being stored in a storage medium, the program product being executed by at least one processor to implement the method according to the first aspect.
[0011] In the embodiments of the present application, first, a plurality of first images captured in a target scene are acquired; then, first processing is performed on the plurality of first images to obtain a plurality of second images; then, a similar image group in the plurality of second images is identified, and a representative image of each similar image group is displayed, the representative image of each similar image group being determined based on a target element of images in the each similar image group; and then, in a case where a first input of a user is received, a document is generated according to the representative image of each similar image group. In the above process, by automatically performing image first processing, image grouping, representative image selection, and document generation in sequence, automatic arrangement of images and automatic generation of a document are achieved, the complexity of user operation is significantly reduced, and the efficiency of document generation is improved. BRIEF DESCRIPTION OF DRAWINGS
[0012] Figure 1 is a flowchart of a document generation method provided by an embodiment of the present application;
[0013] Figure 2 is a schematic diagram of an application scenario of the document generation method provided by an embodiment of the present application;
[0014] Figure 3 is a schematic diagram of an application scenario of the document generation method provided by an embodiment of the present application;
[0015] Figure 4 is a schematic diagram of an application scenario of the document generation method provided by an embodiment of the present application;
[0016] Figure 5 is a schematic diagram of an application scenario of the document generation method provided by an embodiment of the present application;
[0017] Figure 6 is a schematic diagram of an application scenario of the document generation method provided by an embodiment of the present application;
[0018] Figure 7 is a schematic diagram of an application scenario of the document generation method provided by an embodiment of the present application;
[0019] Figure 8A is a schematic diagram of an application scenario of the document generation method provided by an embodiment of the present application;
[0020] Figure 8B is a schematic diagram of an application scenario of the document generation method provided by an embodiment of the present application;
[0021] Figure 8C is a schematic diagram of an application scenario of the document generation method provided by an embodiment of the present application;
[0022] Figure 9 is a structural schematic diagram of a document generation apparatus provided by an embodiment of the present application;
[0023] Figure 10 is a structural schematic diagram of an electronic device provided by an embodiment of the present application;
[0024] Figure 11 is a hardware structural schematic diagram of an electronic device suitable for implementing an embodiment of the present application. DETAILED DESCRIPTION
[0025] The technical solutions in the embodiments of the present application will be clearly described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art belong to the scope of protection of the present application.
[0026] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0027] The document generation method and apparatus provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0028] Please refer to Figure 1 This document illustrates one of the flowcharts of the document generation method provided in this application embodiment. The document generation method provided in this application embodiment can be applied to electronic devices with camera functionality. In practice, the aforementioned electronic devices can be smartphones, tablets, laptops, wearable devices, etc.
[0029] The document generation method provided in this application includes the following steps:
[0030] Step 101: Acquire multiple first images captured in the target scene.
[0031] In this embodiment, the target scene can refer to a shooting scene with specific semantic features, such as, but not limited to, a meeting scene, a classroom scene, a lecture scene, a concert scene, a banquet scene, a scenic spot scene, etc.
[0032] In meeting, classroom, and lecture scenarios, if the first image captured is a document image, these scenarios can be referred to as document capture scenarios. The document image refers to an image containing document content. This document content can include, but is not limited to, electronic document content and printed document content. Electronic document content can further include PPT (PowerPoint) content, Word document content, PDF (Portable Document Format) content, etc., which will not be listed here.
[0033] In this embodiment, the first image can be captured in a normal shooting mode, and at this time, the first image is an image that has not been processed. The first image can also be captured in a super-clear document shooting mode. In the super-clear document shooting mode, the electronic device can automatically identify the edges of a document such as a PPT, a screen, or paper, and automatically perform edge cropping, correction, and the like after shooting to obtain a rectangular image, and save the original image and the corrected image. At this time, the first image can be the corrected image. The original image can be provided for the user to view and edit again.
[0034] In this embodiment, during image capturing, scene detection can be performed based on a shooting preview screen or a captured image to determine whether the current shooting scene is a target scene. If the current shooting scene is the target scene, a plurality of first images captured in the current shooting scene can be obtained. The plurality of first images can be images captured at the same geographic location, and the time interval between the plurality of first images should be less than a threshold, and the image content of the plurality of first images should have the characteristics of the same shooting object in the target scene, for example, both contain paper documents, or both contain screens displaying electronic document content.
[0035] For example, in a conference scene, the user continuously captures a plurality of slide photos, and each slide photo in the plurality of slide photos can be determined to be a first image. If the plurality of slide photos contain one other photo, for example, a conference room seating chart, the photo can be ignored because it does not contain a slide, and only the slide photos are extracted.
[0036] In some optional implementations of this embodiment, the plurality of first images captured in the target scene can be further obtained by the following steps:
[0037] In a first step, in a case where it is detected that a plurality of continuous image capturing operations are performed and the current shooting scene is the target scene, second information is displayed, and the second information is used to inquire whether the user performs image arrangement.
[0038] In a second step, in a case where a sixth input of the user to the second information is received, a plurality of first images captured in the target scene are obtained. The sixth input can be used to return confirmation information to determine that image arrangement is automatically performed. The sixth input can be a touch input, a voice instruction, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual use requirements, and embodiments of this application are not limited. The specific gesture can be any one of a single-click gesture, a sliding gesture, a dragging gesture, a pressure recognition gesture, a long-press gesture, an area change gesture, a double-press gesture, and a double-click gesture. The click input can be a single-click input, a double-click input, or any number of click inputs, and can also be a long-press input or a short-press input.
[0039] As an example, refer to Figure 2 After the user takes multiple photos, it is detected that the current shooting scene is a conference scene or the like, and a second information as shown by a label 201, "It is detected that you are in a conference scene. Do you want to merge the document images into the XX folder and merge the pictures of the same page of PPT into a continuous shooting picture?" can be displayed. If the user clicks the "Confirm" control as shown by a label 202, multiple first images collected in the current shooting scene can be obtained.
[0040] By displaying the second information in the case that the continuous multiple image collection operations are detected and the current shooting scene is the target scene, the scene that needs to be arranged can be automatically recognized, and the second information can be avoided from being mistakenly prompted in the continuous shooting process in the non-target scene, thereby improving the pertinence of information display. By delaying the image obtaining action until the user explicitly confirms, the images can be avoided from being automatically collected without the consent of the user, thereby improving the degree of self-control of the user on data processing.
[0041] In some optional implementations of the embodiment, a virtual folder can also be created according to time and place, and the obtained multiple first images can be stored in the virtual folder, thereby facilitating unified management, as shown by a label 301 in Figure 3 The name of the virtual folder can be generated by document content recognition through artificial intelligence technology, or can be named or edited by the user.
[0042] In some optional implementations of the embodiment, if the user does not want to keep the virtual folder, the virtual folder can be deleted in the interface shown in FIG. 3 by long pressing the virtual folder or the like, thereby reducing interference.
[0043] In some optional implementations of the embodiment, after the virtual folder is created, if the user does not operate the virtual folder within a specified time length, it means that the user does not need to arrange the images in the virtual folder, and the virtual folder can be automatically deleted, thereby reducing interference. The specified time length can be pre-set by the user according to use habits, or automatically set as a default value, for example, one week, and the like, which is not limited herein.
[0044] In step 102, first processing is performed on the multiple first images to obtain multiple second images.
[0045] In the embodiment, the first processing can include, but is not limited to, at least one of the following: image correction processing, image edge cutting processing, composition processing, filter processing, color enhancement processing, noise suppression processing, color cast correction processing, sharpening processing, and the like. Each of the multiple first images can be subjected to the first processing to obtain multiple second images corresponding to the multiple first images.
[0046] Specifically, the first processing manner can be determined according to the target scene. As an example, if the target scene is a document shooting scene such as a meeting scene, a classroom scene, a lecture scene, etc., the first processing can include image correction processing, image edge cutting processing, etc. As another example, if the target scene is a non-document shooting scene such as a concert scene, a banquet scene, a scenic spot scene, etc., the first processing can include but is not limited to composition processing, filter processing, etc.
[0047] In actual application, the electronic device can identify the type of each first image in the album in the background. If a certain first image is a PPT image, a blackboard writing image, etc., and is not shot by the super clear document mode of the camera, the first image can be automatically subjected to image correction processing, etc. by the super document mode of the album application to obtain a second image, and the storage and display of the image can be performed in multiple ways. The super document mode is similar to the super clear document mode of the camera application and is used in the first processing stage of the image.
[0048] As an example, the original first image can be displayed in a normal folder of the album, and the first processed second image can be displayed in a virtual folder.
[0049] As another example, the first processed second image can be displayed in both the normal folder and the virtual folder. If the user is not satisfied with the first processing effect, the second image can be restored to the first image by performing a specific gesture input on the second image.
[0050] As a further example, the original first image can be displayed in both the normal folder and the virtual folder. After the user opens the first image and performs a specific gesture input on the first image, the corresponding second image can be displayed.
[0051] In step 103, similar image groups in the plurality of second images are identified, and a representative image of each similar image group is determined, the representative image of each similar image group being determined based on a target element of the images in each similar image group.
[0052] In this embodiment, the similar image group refers to an image subset obtained by clustering based on image content similarity in the image set corresponding to the plurality of second images. The electronic device can identify the similarity between the second images in the background, and the second images with high similarity can be summarized into the same similar image group.
[0053] As an example, the first image is an image of a PPT document shown by a speaker, and the second image is an image generated after image correction processing on the first image. Since the content of the same page in the PPT document is usually played by the speaker line by line or region by region, it is easy for the user to take multiple first images corresponding to the same page in the PPT document but with incomplete content. Accordingly, multiple second images corresponding to the same page in the PPT document but with incomplete content can be generated. The electronic device can identify the document content in the images in the background, determine the relationship between the second images, and aggregate the second images corresponding to the same page in the PPT document into the same similar image group.
[0054] As another example, in a conference scenario, the content of the PPT document shown by the speaker can be blocked by the front row personnel or the speaker, resulting in the user taking multiple first images for the same page in the PPT document. In addition, due to the user's operation habit, operation error, or poor imaging effect, the user can take multiple first images for the same page in the PPT document. Accordingly, multiple second images corresponding to the content of the same page in the PPT document can be generated. The electronic device can identify the image similarity in the background, determine the similarity between the second images, and aggregate the second images corresponding to the same page in the PPT document into the same similar image group.
[0055] In this embodiment, each similar image group can include at least one second image, and one second image can be used as a representative image of the similar image group. The representative image in each similar image group can be determined based on a target element in the similar image group, and the target element includes at least one of the following: quality, content.
[0056] As an example, for a document image, for each similar image group, the second image with the most complete document content in the similar image group can be used as the representative image of the similar image group. In the case where the second image with the most complete document content includes at least two, the image with the highest image clarity or the least occlusion can be further selected from the at least two second images with the most complete document content as the representative image of the similar image group.
[0057] In some optional implementations of this embodiment, the target scenario includes a document shooting scenario. On this basis, the representative image of each similar image group can be determined through the following sub-steps:
[0058] For each similar image group, if the similar image group includes at least two second images, the clarity of each second image in the similar image group, the completeness of the document content, and the occlusion area are determined. Based on the clarity, the completeness of the document content, and the occlusion area, the representative image in each similar image group is determined.
[0059] For each similar image group, if the similar image group includes a second image, the second image in the similar image group is determined as the representative image.
[0060] The clarity can be determined by a clarity detection model trained based on an artificial intelligence algorithm, or can be determined by a Laplacian operator, and the like, which is not limited herein. The completeness of the document content can be determined by a ratio of a text box area to an image area through an optical character recognition (OCR) technology. The area of the occlusion region can be determined by a ratio of an occlusion region pixel ratio to a whole image area through a human body or foreground segmentation model.
[0061] Specifically, the second image with the highest completeness of the document content can be selected as the representative image first. If the second image with the highest completeness of the document content includes at least two images, the second image with the highest clarity is further selected from the at least two images as the representative image. If the second image with the highest clarity still includes at least two images, the second image with the smallest area of the occlusion region is further selected from the at least two images as the representative image. If the second image with the smallest area of the occlusion region still includes at least two images, a second image is randomly selected from the at least two images as the representative image.
[0062] In this way, the second image with the optimal target element in each similar image group can be selected as the representative image, so as to improve the quality of the generated document.
[0063] At step 104, in a case where the first input of the user is received, a document is generated according to the representative image of each similar image group.
[0064] In this embodiment, the first input can be used to trigger the generation of the document. The first input can be a touch input, or a voice instruction, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual use requirements, and the present embodiment is not limited herein. The specific gesture in the present embodiment can be any one of a single-click gesture, a sliding gesture, a dragging gesture, a pressure recognition gesture, a long-press gesture, an area change gesture, a double-press gesture, and a double-click gesture. The click input in the present embodiment can be a single-click input, a double-click input, or any number of click inputs, and can also be a long-press input or a short-press input.
[0065] In this embodiment, the document is a carrier of recorded information or data, which can be text, image, chart, sound, video or a combination thereof, and is intended to convey information or data. In this embodiment, the format of the document can be determined according to the target scenario. For example, if the target scenario is a document image collection scenario such as a meeting scenario, a classroom scenario, a lecture scenario, etc., the generated document can be a PDF document, a PPT document, etc.; if the target scenario is a concert scenario, a scenic spot scenario, a banquet scenario, etc., the generated document can be an image format document with a nine-square structure layout, etc. The document can include a representative image of each similar image group, can include a thumbnail of a non-representative image in each similar image group, and can include text information generated based on image content, which is not limited here. The document can support sharing.
[0066] As an example, see Figure 4 The virtual folder includes a control for triggering document generation, as shown by reference numeral 401. When the user clicks the control, the determined representative images can be laid out, and a PPT or PDF document containing the representative images can be generated based on the layout result. It can be understood that before clicking the control 401, the user can manually check or uncheck part of the representative images from the virtual folder; all representative images can be checked at one key; the representative images can be replaced by non-representative images, etc., which is not limited here.
[0067] The method provided by the above embodiments of the present application first acquires a plurality of first images collected in a target scenario; then performs first processing on the plurality of first images to obtain a plurality of second images; then identifies similar image groups in the plurality of second images and displays representative images of each similar image group, the representative images of each similar image group being determined based on target elements of images in each similar image group; and generates a document based on the representative images of each similar image group in the case that a first input of a user is received. In the above process, by automatically performing image first processing, image grouping, representative image selection, document generation, etc. in turn, automatic arrangement of images and automatic generation of documents are realized, the complexity of user operation is significantly reduced, and the document generation efficiency is improved.
[0068] In some optional embodiments, after step 103 is performed, the following steps can also be performed to update the representative images:
[0069] Step S11, display the representative images of each similar image group and hide the non-representative images in each similar image group.
[0070] As an example, see Figure 4For PPT document images, for each page in the PPT document, the second image containing the content of the page but not complete can be hidden to display only the image of the page with the most complete content. For PPT document images with the same content taken multiple times, the image with the clearest and least occlusion is displayed, and the area image is hidden.
[0071] Folding and hiding the non-representative images can greatly simplify the user's browsing interface while maintaining the integrity of the information. The user does not need to look at all the images one by one, but only needs to focus on the representative images to quickly understand the main content, reducing information redundancy and browsing time.
[0072] Step S12, in the case of receiving the fifth input of the user, determining a target representative image in the retained representative images, and displaying at least one non-representative image in the similar image group to which the target representative image belongs, wherein the at least one non-representative image is marked with the location of the missing area of the document content.
[0073] The fifth input can be used to select the target representative image and trigger the display of at least one non-representative image in the similar image group to which the target representative image belongs. The fifth input can be a touch input, a voice instruction, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual use requirements, and the embodiments of the present application are not limited.
[0074] The non-representative image refers to other images in the similar image group except the representative image. These images, although not selected as representative images in the initial screening, may contain some important information missing from the representative image or have better visual effects, and therefore can be displayed and selected when the user needs. The non-representative image can be displayed in the form of a thumbnail, and the display method and display style thereof are not limited by the embodiments of the present application.
[0075] The missing area of the document content refers to the area in the image where the document content is not completely presented due to the shooting angle, occlusion, etc. These areas may contain key text, chart, and other information, which are crucial for the user to fully understand the document content. By identifying and marking these missing areas, the user can quickly locate and view them.
[0076] As an example, the user long-presses a representative image A in the virtual folder, and the electronic device identifies that the similar image group to which it belongs contains two non-representative images B and C, so it can display images B and C to the user in the form of thumbnails. At the same time, the electronic device finds through algorithm analysis that there are missing areas of the document content in images B and C compared to image A. When the user browses one of images B or C, a rectangular box can be used to mark the missing area of the document content, as shown in Figure 5The target representative image is replaced by the target non-representative image. As shown in FIG. 5, the target representative image is replaced by the target non-representative image as shown in FIG. 5. By marking the missing area of the document content, the user can intuitively understand the advantages and disadvantages of each non-representative image, which facilitates the user to make more accurate selection.
[0077] Step S13, in the case of receiving the sixth input of the user, determining a target non-representative image in the at least one non-representative image, and replacing the target representative image with the target non-representative image.
[0078] The sixth input can be used to select the target non-representative image and trigger the replacement of the target representative image. The sixth input can be a touch input, or a voice instruction, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual use requirements, and the embodiments of the present application are not limited.
[0079] As an example, the user can select the target non-representative image by clicking, dragging, or other operations, and after detecting that the user selects the non-representative image, the target representative image can be replaced by the target non-representative image for display. Further, the user can also choose to clean up the images that are not displayed to reduce the invalid picture space occupation.
[0080] By allowing the user to select a new representative image and replace the display, the user's individual needs and actual use scenarios can be better met. The user can select the most suitable image as the representative image according to his own judgment and needs, thereby improving the quality and accuracy of the representative image.
[0081] In some optional embodiments, step 104 can further include:
[0082] Step S21, displaying a document preview page, and the document preview page includes a representative image of each similar image group.
[0083] The document preview page is a page for presenting the sorted images for the user to view and edit. The document preview page provides various editing tools and options, such as replacement, insertion, deletion, sorting, and image editing functions, to meet the user's needs for adjusting the document content.
[0084] Optionally, the representative images can be sorted based on the collection time of the representative images of each similar image group, and then the document preview page is displayed according to the sorting result.
[0085] Optionally, for the document shooting scene, the page number information in the representative image of each similar image group can be automatically identified, the representative images are sorted according to the page number information, and then the document preview page is displayed according to the sorting result.
[0086] Optionally, for the document shooting scenario, the document preview page can also be displayed according to the following steps: first, the representative images of each similar image group are sorted based on the image collection time, to obtain a first sorting result; then, the page number information is obtained from the representative images of each similar image group; after that, the representative images of each similar image group are sorted based on the page number information, to obtain a second sorting result; then, in the case that the first sorting result and the second sorting result do not match, the first information is displayed, the first information is used to inquire whether the user adjusts the image order; finally, in the case that the fourth input of the user to the first information is received, the document preview page is displayed based on the second sorting result.
[0087] As an example, in the case that the first sorting result and the second sorting result do not match, the first information "there is a page number exception, do you want to adjust the order" can be displayed. If the user confirms, the order of the representative images can be set according to the page number information in the representative images, and then the document preview page containing the representative images is displayed.
[0088] By sorting the images based on the page number information, the image arrangement order that is more consistent with the document content logic can be generated. This sorting method can ensure the coherence and integrity of the document content, so that the user can obtain information in the correct order when viewing the document.
[0089] Further, the order of the images in the document preview page can be adjusted. Referring to Figure 6 , the "sorting" control 601 can be displayed in the document preview page. After the user clicks the control, the "smart page number", "import order", "shooting order" and other sorting options can be displayed. Among them, the "smart page number" refers to sorting according to the page number information in each image in the document preview page. The "import order" refers to sorting according to the selected order of each image in the document preview page. The "shooting order" refers to sorting according to the collection time of each image in the document preview page. By allowing the user to select the image sorting method, the flexibility of setting the image order can be improved.
[0090] Further, referring to Figure 6The document preview page can also display a "fold similar pictures" control 602. By default, the document preview page can only display representative images in each similar image group. If the determination of the similar image group and the selection of the representative image are not performed, i.e., step 103 is not performed, or the user manually checks one or more non-representative images for document generation, the non-representative images can also be displayed in the document preview page. At this time, if the user clicks the "fold similar pictures" control 602, the determination of the similar image group and the selection of the representative image can be performed, so as to hide the non-representative images to reduce interference. Alternatively, if the same or similar representative images exist in the execution result of step 103, when the user clicks the "fold similar pictures" control 602, the representative images can be further screened, and redundant images among the representative images can be hidden.
[0091] Further, the user can also manually hide similar images in the document preview page. For example, the user can long press image A in the document preview page and drag it to the area where image B is located, so that the two images overlap. At this time, image B can be hidden. At the same time, image A and image B can be further compared. If it is determined based on the target element of the image that image B is better than image A, a prompt information "The covered picture, the information is more complete, do you want to display it?" can be displayed. If the user confirms, image B is displayed and image A is hidden. In addition, referring to Figure 7 If it is determined based on the content of the image that the similarity of image A and image B is lower than a threshold, a prompt information "There is no similarity between the two, do you want to merge them?" can be displayed to prevent user misoperation.
[0092] In step S22, in a case where a second input of the user to the document preview page is received, an editing process is performed on the document preview page to update the document preview page, wherein the editing process includes at least one of the following: replacing at least one representative image in the document preview page with a non-representative image, inserting at least one non-representative image in the document preview page, deleting at least one image in the document preview page, adjusting the order of the images in the document preview page, and editing at least one image in the document preview page.
[0093] The second input can be used to trigger the editing of the document preview page. The second input can be a touch input, a voice instruction, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual use requirements, and embodiments of the present application are not limited.
[0094] Specifically, referring to Figure 8AFor each image in the document preview page, if the image has a corresponding hidden image, a control as shown by reference numeral 801 can be displayed below the image. Here, if the image is a representative image of a group of similar images, the corresponding hidden image is a non-representative image of the group of similar images. If the image is manually set by the user, the associated image can be the image manually folded by the user. For example, the user long-presses an image A in the document preview page and drags it to the area where image B is located, so that the two images overlap, and the corresponding hidden image of image A includes image B. For another example, if the user manually replaces a representative image C in a group of similar images with a non-representative image D, the corresponding hidden image of non-representative image D includes representative image C.
[0095] After the user clicks the control 801 below a certain image, a thumbnail of the hidden image corresponding to the control can be displayed, as shown by reference numeral 802. Figure 8B For the convenience of detail comparison, the thumbnail can also be indicated to be enlarged. The user can browse different hidden images by left and right swiping.
[0096] Further, the displayed images in the document preview page can be replaced. For example, after the user performs a certain operation on a thumbnail, such as long-pressing, a "use current image" control as shown by reference numeral 803 can be displayed. The user can switch the displayed image by clicking the control. Through this process, a non-representative image can be replaced by a representative image for display. Figure 8B
[0097] Further, a new image can be inserted in the document preview page. For example, if the user wants to keep a hidden image, the user can long-press the thumbnail corresponding to the hidden image and drag it to a certain position in the document preview page to insert the hidden image at the position in the document. Through this process, a non-representative image can be inserted in the document preview page.
[0098] Further, an image in the document preview page can be deleted or reselected. For example, as shown in Figure 8C For a certain image in the document preview page, after the user performs a certain gesture such as right swiping, a "delete" control 804 and a "reselect" control 805 can be displayed. If the user clicks the "delete" control 804, the image can be deleted. If the user clicks the "reselect" control 805, thumbnails of the hidden images corresponding to the image can be displayed, and the user can select one or more hidden images to be displayed. It can be understood that if the image does not have a corresponding hidden image, for example, a certain page of PPT only has one image, the "reselect" control 805 is not displayed.
[0099] Further, at least one image in the document preview page can be edited. For example, when a user performs a specific operation such as clicking on an image in the document preview page, the image can be opened and enter an editing state. The user can perform operations such as cropping, straightening, adding filters, denoising, brightness adjustment, contrast adjustment, and the like on the image.
[0100] At step S23, in a case where a third input of the user is received, a document is generated based on the updated document preview page. Specifically, the third input can be used to trigger the generation of the document. The third input can be a touch input, a voice instruction, or a specific gesture input by the user, or other possible inputs, which can be determined according to actual use requirements, and embodiments of the present application are not limited thereto. After the document preview page is updated, the user can perform a gesture operation or a voice instruction to generate and export the document.
[0101] Through the above process, the user can deeply customize and optimize the content of the document. The user can flexibly adjust the content of the document according to actual requirements, ensure that the generated document meets the specifications and individual requirements, and improve the accuracy and efficiency of document organization.
[0102] It should be noted that the document generation method provided by the embodiments of the present application can be executed by a document generation device. In the embodiments of the present application, the document generation device is taken as an example to illustrate the document generation device provided by the embodiments of the present application.
[0103] As shown in Figure 9 The document generation device 900 of the present embodiment includes: an acquisition unit 901 configured to acquire a plurality of first images collected in a target scene; a processing unit 902 configured to perform first processing on the plurality of first images to obtain a plurality of second images, wherein the first processing includes at least one of image straightening processing, image edge cutting processing, composition processing, filter processing, color enhancement processing, noise suppression processing, color cast correction processing, and sharpening processing; a determination unit 903 configured to identify similar image groups in the plurality of second images and display representative images of each similar image group, wherein the representative images of each similar image group are determined based on target elements of images in each similar image group, and the target elements include at least one of quality and content; and a generation unit 904 configured to generate a document based on the representative images of each similar image group in a case where a first input of the user is received.
[0104] In some optional implementation of the present embodiment, the generation unit 904 is further configured to: display a document preview page, the document preview page including the representative image of each similar image group; and in response to receiving a second input of the user on the document preview page, perform editing processing on the document preview page to update the document preview page, wherein the editing processing includes at least one of the following: replacing at least one representative image in the document preview page with a non-representative image, inserting at least one non-representative image in the document preview page, deleting at least one image in the document preview page, adjusting the order of images in the document preview page, and editing at least one image in the document preview page; and generating the document based on the updated document preview page. Through the above process, the user can deeply customize and optimize the content of the document. The user can flexibly adjust the content of the document according to actual needs, ensure that the generated document meets the specification and individual requirements, and improve the accuracy and efficiency of document arrangement.
[0105] In some optional implementation of the present embodiment, the generation unit 904 is further configured to: sort the representative images of each similar image group based on image acquisition time to obtain a first sorting result; obtain page number information from the representative images of each similar image group; sort the representative images of each similar image group based on the page number information to obtain a second sorting result; and in response to the first sorting result not matching the second sorting result, display first information for querying whether the user adjusts the image order; and in response to receiving a fourth input of the user on the first information, display the document preview page based on the second sorting result. By sorting the images based on the page number information, a more logical image arrangement order of the document content can be generated. This sorting method can ensure the coherence and integrity of the document content, so that the user can obtain information in the correct order when viewing the document.
[0106] In some optional implementation of the present embodiment, the target scene includes a document shooting scene; and the determination unit 903 is further configured to: for each similar image group, if the similar image group includes at least two second images, determine the clarity of each second image in the similar image group, the completeness of the document content, and the area of the occlusion region; determine the representative image in each similar image group based on the clarity, the completeness of the document content, and the area of the occlusion region; and for each similar image group, if the similar image group includes one second image, determine the second image in the similar image group as the representative image. Through the above method, the second image with the optimal target element in each similar image group can be selected as the representative image, thereby improving the quality of the generated document.
[0107] In some optional implementations of the embodiment, the apparatus further includes an updating unit, and the updating unit is further configured to: display the representative image of each similar image group and hide the non-representative images in each similar image group; in a case where a fifth input of the user is received, determine a target representative image in the retained representative images, and display at least one non-representative image in the similar image group to which the target representative image belongs, wherein the at least one non-representative image is marked with a position of a document content missing area; and in a case where a sixth input of the user is received, determine a target non-representative image in the at least one non-representative image, and replace the target representative image with the target non-representative image. By allowing the user to select a new representative image and replace the display, the personalized needs and actual use scenarios of the user can be better met. The user can determine and select the most suitable image as the representative image according to his / her own judgment and needs, thereby improving the quality and accuracy of the representative image.
[0108] In some optional implementations of the embodiment, the acquisition unit 901 is further configured to: in a case where the continuous multiple times of image acquisition operations are detected and the current shooting scene is the target scene, display second information, the second information being used to inquire whether the user performs image arrangement; and in a case where a sixth input of the user on the second information is received, acquire the multiple first images collected in the target scene. By displaying the second information in a case where the continuous multiple times of image acquisition operations are detected and the current shooting scene is the target scene, the scene that needs to perform image arrangement can be automatically identified, and the second information can be prevented from being mistakenly prompted in the continuous shooting process of the non-target scene, thereby improving the pertinence of information display. By delaying the image acquisition action until after the user explicitly confirms, the images can be prevented from being automatically collected without the consent of the user, thereby improving the degree of self-control of the user over data processing.
[0109] The apparatus provided by the above embodiment of the present application first acquires the multiple first images collected in the target scene, then performs the first processing on the multiple first images to obtain the multiple second images, and then identifies the similar image groups in the multiple second images and displays the representative image of each similar image group, the representative image of each similar image group being determined based on the target element of the images in each similar image group, so that in a case where a first input of the user is received, a document is generated according to the representative image of each similar image group. In the above process, by automatically performing the image first processing, image grouping, representative image selection, and document generation in sequence, the automatic arrangement of images and the automatic generation of documents are realized, the complexity of user operation is significantly reduced, and the document generation efficiency is improved.
[0110] The document generation apparatus in the embodiments of the present application can be an electronic device, or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or other devices than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an Augmented Reality (AR) / Virtual Reality (VR) device, a robot, a wearable device, an Ultra-Mobile Personal Computer (UMPC), a netbook, a Personal Digital Assistant (PDA), or the like, and can also be a server, a Network Attached Storage (NAS), a Personal Computer (PC), a Television (TV), a cash register, a self-service machine, or the like. The embodiments of the present application are not limited in this regard.
[0111] The document generation apparatus in the embodiments of the present application can be an apparatus with an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems, and the embodiments of the present application are not limited in this regard.
[0112] The document generation apparatus provided in the embodiments of the present application can implement the method embodiments, and each process of the method embodiments is not repeated here. Figure 1 The method embodiments implement each process, and each process is not repeated here.
[0113] Optionally, as shown in Figure 10 The embodiments of the present application also provide an electronic device 1000, which includes a processor 1001 and a memory 1002. The memory 1002 stores programs or instructions that can run on the processor 1001. When the programs or instructions are executed by the processor 1001, each step of the above document generation method embodiments is implemented, and the same technical effects are achieved. Each step is not repeated here.
[0114] It should be noted that the electronic device in the embodiments of the present application includes the mobile electronic device and the non-mobile electronic device.
[0115] Figure 11 A hardware structure diagram of an electronic device according to an embodiment of the present application is shown.
[0116] The electronic device 1100 includes, but is not limited to, a radio frequency unit 1101, a network module 1102, an audio output unit 1103, an input unit 1104, a sensor 1105, a display unit 1106, a user input unit 1107, an interface unit 1108, a memory 1109, and a processor 1110, etc.
[0117] Those skilled in the art can understand that the electronic device 1100 can also include a power supply (such as a battery) for supplying power to each component, and the power supply can be logically connected to the processor 1110 through a power management system, so that the power management system can realize the functions of managing charging, discharging, and power consumption management, etc. Figure 11 The electronic device structure shown in the figure does not constitute a limitation on the electronic device, and the electronic device can include more or fewer components than the figure, or combine certain components, or different component arrangements, which are not described here.
[0118] The processor 1110 is configured to acquire a plurality of first images collected in a target scene; perform first processing on the plurality of first images to obtain a plurality of second images, wherein the first processing includes at least one of image correction processing, image edge cutting processing, composition processing, filter processing, color enhancement processing, noise suppression processing, color cast correction processing, and sharpening processing; identify similar image groups in the plurality of second images, and display a representative image of each similar image group, wherein the representative image of each similar image group is determined based on a target element of images in the each similar image group, and the target element includes at least one of quality and content; and in a case where a first input of a user is received, generate a document according to the representative image of each similar image group.
[0119] By sequentially and automatically performing image first processing, image grouping, representative image selection, and document generation, automatic arrangement of images and automatic generation of documents are realized, which significantly reduces the complexity of user operations and improves the efficiency of document generation.
[0120] In some optional implementations of the embodiment, the processor 1110 is further configured to display, through the display unit 1106, a document preview page, the document preview page including representative images of each of the similar image groups; in a case where a second input of the user on the document preview page is received through the user input unit 1107, perform editing processing on the document preview page to update the document preview page, the editing processing including at least one of the following: replacing at least one representative image in the document preview page with a non-representative image, inserting at least one non-representative image in the document preview page, deleting at least one image in the document preview page, adjusting the order of images in the document preview page, editing at least one image in the document preview page; and generating the document based on the updated document preview page. Through the above process, the user can deeply customize and optimize the content of the document. The user can flexibly adjust the content of the document according to actual needs, ensure that the generated document meets the specification and individual requirements, and improve the accuracy and efficiency of document arrangement.
[0121] In some optional implementations of the embodiment, the processor 1110 is further configured to sort the representative images of each of the similar image groups based on image acquisition time to obtain a first sorting result; obtain page number information from the representative images of each of the similar image groups; sort the representative images of each of the similar image groups based on the page number information to obtain a second sorting result; in a case where the first sorting result and the second sorting result do not match, display first information, the first information being used to inquire whether the user adjusts the image order; and in a case where a fourth input of the user on the first information is received through the user input unit 1107, display the document preview page through the display unit 1106 based on the second sorting result. By sorting the images based on the page number information, a more logical image arrangement order of the document content can be generated. This sorting method can ensure the coherence and integrity of the document content, so that the user can obtain information in the correct order when viewing the document.
[0122] In some optional implementations of the present embodiment, the target scene includes a document shooting scene; the processor 1110 is further configured to, for each similar image group, if the similar image group includes at least two second images, determine the sharpness of each second image in the similar image group, the completeness of the document content, and the area of the occlusion region; determine a representative image in each similar image group based on the sharpness, the completeness of the document content, and the area of the occlusion region; and for each similar image group, if the similar image group includes one second image, determine the second image in the similar image group as the representative image. In this way, the second image with the optimal target element in each similar image group can be selected as the representative image, thereby improving the quality of the generated document.
[0123] In some optional implementations of the present embodiment, the processor 1110 is further configured to display, by the display unit 1106, the representative image of each similar image group and hide the non-representative images in each similar image group; in a case where a fifth input of a user is received by the user input unit 1107, determine a target representative image from the retained representative images, and display, by the display unit 1106, at least one non-representative image in the similar image group to which the target representative image belongs, wherein the at least one non-representative image is marked with the position of the missing region of the document content; and in a case where a sixth input of a user is received by the user input unit 1107, determine a target non-representative image from the at least one non-representative image, and replace the target representative image with the target non-representative image. By allowing the user to select a new representative image and replace the display, the personalized needs and actual use scenarios of the user can be better met. The user can set the most suitable image as the representative image according to his / her own judgment and needs, thereby improving the quality and accuracy of the representative image.
[0124] In some optional implementations of the present embodiment, the processor 1110 is further configured to, in a case where a continuous multiple-time image capturing operation is detected and the current shooting scene is the target scene, display, by the display unit 1106, second information for querying whether the user wants to perform image arrangement; and in a case where a sixth input of a user on the second information is received by the user input unit 1107, acquire the multiple first images captured in the target scene. By displaying the second information in a case where a continuous multiple-time image capturing operation is detected and the current shooting scene is the target scene, the scene that needs to perform image arrangement can be automatically identified, and the second information can be prevented from being mistakenly prompted in a continuous shooting process in a non-target scene, thereby improving the pertinence of information display. By delaying the image acquisition action until after the user explicitly confirms, the images can be prevented from being automatically collected without the consent of the user, thereby improving the degree of self-control of the user over data processing.
[0125] It should be understood that in the embodiments of the present application, the input unit 1104 can include a graphics processor (GPU) 11041 and a microphone 11042. The graphics processor 11041 processes image data of a still picture or a video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 1106 can include a display panel 11061, which can be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1107 includes at least one of a touch panel 11071 and other input devices 11072. The touch panel 11071 is also referred to as a touch screen. The touch panel 11071 can include two parts of a touch detection device and a touch controller. The other input devices 11072 can include, but are not limited to, a physical keyboard, function keys (such as volume control keys, on-off keys, etc.), a trackball, a mouse, a joystick, and the like, which will not be described here.
[0126] The memory 1109 can be used to store software programs and various data. The memory 1109 can mainly include a first storage area storing programs or instructions and a second storage area storing data, wherein the first storage area can store an operating system, application programs or instructions required by at least one function (such as a sound playing function, an image playing function, etc.), and the like. In addition, the memory 1109 can include a volatile memory or a non-volatile memory, or the memory 1109 can include both volatile and non-volatile memories. The non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 1109 in the embodiments of the present application includes but is not limited to these and any other suitable types of memories.
[0127] The processor 1110 can include one or more processing units; optionally, the processor 1110 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and an application program, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 1110.
[0128] The embodiments of the present application also provide a readable storage medium, the readable storage medium stores programs or instructions, the programs or instructions are executed by a processor to realize various processes of the above-mentioned document generation method embodiments, and the same technical effects can be achieved. To avoid repetition, details are not described here.
[0129] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes a computer readable storage medium, such as a computer readable only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0130] The embodiment of the present application further provides a chip, which comprises a processor and a communication interface, the communication interface is coupled with the processor, the processor is used for running programs or instructions to realize the processes of the above document generation method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein.
[0131] It should be understood that the chip mentioned in the embodiment of the present application can also be referred to as a system level chip, a system chip, a chip system or a system on chip, etc.
[0132] The embodiment of the present application provides a computer program product, which is stored in a storage medium, and is executed by at least one processor to realize the processes of the above document generation method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein.
[0133] It should be noted that in this document, the term "comprising" or "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the method and device in the embodiment of the present application is not limited to the order of performing the functions as shown or discussed, but can also include performing the functions in a substantially simultaneous manner or in the opposite order, for example, the described method can be performed in an order different from that described, and various steps can also be added, omitted or combined. In addition, the features described with reference to certain examples can be combined in other examples.
[0134] Through the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned example methods can be realized by means of software and a necessary general hardware platform, and of course, can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a computer software product in essence or in the form of a part that contributes to the prior art, which is stored in a storage medium (such as a ROM / RAM, a magnetic disk, or an optical disk) and includes a plurality of instructions for causing a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application.
[0135] The embodiments of the present application are described above in combination with the drawings, but the present application is not limited to the above-mentioned specific embodiments, and the above-mentioned specific embodiments are only illustrative and not restrictive. Those skilled in the art can make many forms under the inspiration of the present application without departing from the scope of the present application and the scope protected by the claims.
Claims
1. A document generation method, characterized in that, The method includes: Acquire multiple first images captured in the target scene; Perform a first process on the plurality of first images to obtain a plurality of second images; Identify similar image groups among the multiple second images, and determine a representative image for each similar image group, wherein the representative image for each similar image group is determined based on the target elements of the images in each similar image group; Upon receiving the user's initial input, a document is generated based on the representative image of each similar image group.
2. The method according to claim 1, characterized in that, The step of generating a document based on a representative image of each similar image group includes: Display a document preview page, which includes a representative image for each of the similar image groups; Upon receiving a second input from a user on the document preview page, the document preview page is edited to update it. The editing process includes at least one of the following: replacing at least one representative image on the document preview page with a non-representative image, inserting at least one non-representative image into the document preview page, deleting at least one image on the document preview page, adjusting the order of the images on the document preview page, and editing at least one image on the document preview page. Upon receiving a third input from the user, the document is generated based on the updated document preview page.
3. The method according to claim 2, characterized in that, The document preview page includes: Based on the image acquisition time, the representative images of each similar image group are sorted to obtain a first sorting result; Page number information is obtained from the representative images of each similar image group; Based on the page number information, the representative images of each similar image group are sorted to obtain a second sorting result; If the first sorting result does not match the second sorting result, the first information is displayed, which asks the user whether to adjust the image order. Upon receiving a fourth input from the user regarding the first information, the document preview page is displayed based on the second sorting result.
4. The method according to claim 1, characterized in that, The target scene includes a document shooting scene; determining the representative image for each similar image group includes: For each similar image group, if the similar image group includes at least two second images, then determine the sharpness, document content integrity, and occlusion area of each second image in each similar image group; based on the sharpness, document content integrity, and occlusion area, determine the representative image in each similar image group; For each group of similar images, if the group of similar images includes a second image, then the second image in the group of similar images is determined as the representative image.
5. The method according to claim 1, characterized in that, After determining the representative image for each similar image group, the method further includes: Display the representative image of each similar image group and hide the non-representative images in each similar image group; Upon receiving the user's fifth input, a target representative image is determined from the retained representative images, and at least one non-representative image from the similar image group to which the target representative image belongs is displayed, wherein the location of the missing document content area is marked in the at least one non-representative image; Upon receiving a sixth input from the user, a target non-representative image is determined among the at least one non-representative image, and the target representative image is replaced with the target non-representative image.
6. A document generation device, characterized in that, The device includes: The acquisition unit is used to acquire multiple first images captured in the target scene; A processing unit is configured to perform a first process on the plurality of first images to obtain a plurality of second images; A determining unit is configured to identify similar image groups among the plurality of second images and display a representative image for each similar image group, wherein the representative image for each similar image group is determined based on the target elements of the images in each similar image group; The generation unit is configured to generate a document based on a representative image of each similar image group upon receiving a first input from the user.
7. The apparatus according to claim 6, characterized in that, The generation unit is further configured to: Display a document preview page, which includes a representative image for each of the similar image groups; Upon receiving a second input from a user on the document preview page, the document preview page is edited to update it. The editing process includes at least one of the following: replacing at least one representative image on the document preview page with a non-representative image, inserting at least one non-representative image into the document preview page, deleting at least one image on the document preview page, adjusting the order of the images on the document preview page, and editing at least one image on the document preview page. The document is generated based on the updated document preview page.
8. The apparatus according to claim 7, characterized in that, The generation unit is further configured to: Based on the image acquisition time, the representative images of each similar image group are sorted to obtain a first sorting result; Page number information is obtained from the representative images of each similar image group; Based on the page number information, the representative images of each similar image group are sorted to obtain a second sorting result; If the first sorting result does not match the second sorting result, the first information is displayed, which asks the user whether to adjust the image order. Upon receiving a fourth input from the user regarding the first information, the document preview page is displayed based on the second sorting result.
9. The apparatus according to claim 6, characterized in that, The target scene includes a document shooting scene; the determining unit is further configured to: For each similar image group, if the similar image group includes at least two second images, then determine the sharpness, document content integrity, and occlusion area of each second image in each similar image group; based on the sharpness, document content integrity, and occlusion area, determine the representative image in each similar image group; For each group of similar images, if the group of similar images includes a second image, then the second image in the group of similar images is determined as the representative image.
10. The apparatus according to claim 6, characterized in that, The device also includes an updating unit, which is further used for: Display the representative image of each similar image group and hide the non-representative images in each similar image group; Upon receiving the user's fifth input, a target representative image is determined from the retained representative images, and at least one non-representative image from the similar image group to which the target representative image belongs is displayed, wherein the location of the missing document content area is marked in the at least one non-representative image; Upon receiving a sixth input from the user, a target non-representative image is determined among the at least one non-representative image, and the target representative image is replaced with the target non-representative image.