Document generation method and device and electronic equipment

By adding image processing functions to the album, automatically extracting and generating documents, the problem of cumbersome user operations in the prior art is solved and the efficiency of document generation is improved.

CN120407050APending Publication Date: 2025-08-01VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510566432.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

In the prior art, when users generate documents, they need to select images one by one, press and copy text and paste them. The operation is cumbersome, resulting in low document generation efficiency.

Method used

It provides an album with image processing function. Users only need to add images to the album, use the album's image processing function to automatically extract text and generate documents, reducing user manual operations.

Benefits of technology

Simplifies the document generation process, improves document generation efficiency, and reduces the steps of users to manually extract and paste copy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407050A_ABST
    Figure CN120407050A_ABST
Patent Text Reader

Abstract

The invention discloses a document generation method and device and electronic equipment, and belongs to the technical field of artificial intelligence. The method can comprise the steps that first input of adding a first image in a first photo album by a user is received; in response to the first input, processing the first image according to a first image processing function corresponding to the first photo album to obtain image information; and generating a document according to the image information, the document including information related to the content of the first image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, and particularly relates to a document generation method, apparatus, and electronic device. Background Art

[0002] With the development of artificial intelligence technology, users can view the images downloaded, taken, and screen-captured by the user through the photo album provided by the electronic device. When the user wants to obtain the content in the image, the user usually selects the image from the photo album, enters the editing interface of the image, long-presses the image in the editing interface, copies the text in the image, and then pastes the extracted text into the document specified by the user to organize it into the document required by the user.

[0003] However, the user operation in this document generation process is cumbersome, resulting in low document generation efficiency. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a document generation method, apparatus, and electronic device, which can simplify the user operation in the document generation process and improve the document generation efficiency.

[0005] In a first aspect, the embodiments of this application provide a document generation method, including:

[0006] Receiving a first input from the user to add a first image to a first photo album;

[0007] In response to the first input, processing the first image according to the first image processing function corresponding to the first photo album to obtain image information;

[0008] Generating a document according to the image information, where the document includes information related to the content of the first image.

[0009] In a second aspect, the embodiments of this application provide a document generation apparatus, including:

[0010] A receiving module, configured to receive a first input from the user to add a first image to a first photo album;

[0011] A processing module, configured to process the first image according to the first image processing function corresponding to the first photo album in response to the first input to obtain image information;

[0012] A generating module, configured to generate a document according to the image information, where the document includes information related to the content of the first image.

[0013] In a third aspect, the embodiments of this application provide an electronic device, which includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the document generation method as shown in the first aspect are implemented.

[0014] Fourthly, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the document generation method as shown in the first aspect are implemented.

[0015] Fifthly, an embodiment of the present application provides a chip, which includes a processor and a display interface. The display interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the steps of the document generation method as shown in the first aspect.

[0016] Sixthly, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the steps of the document generation method as shown in the first aspect.

[0017] In the embodiment of the present application, in the case of receiving a first input for adding a first image to a first album, the first image can be processed according to the first image processing function corresponding to the first album to obtain image information; and a document including information related to the content of the first image can be generated according to the image information. In this way, the user only needs to add the image to be processed to the first album, and can use the first image processing function corresponding to the first album to process the first image to obtain the image information for generating a document, without the user manually extracting the text in the image, nor manually copying and pasting it into a specified document, reducing the user operations for generating a document based on the text in the image and improving the document generation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a flowchart of the document generation method provided by some embodiments of the present application;

[0019] Figure 2 It is a schematic diagram of the interface of the document generation method provided by some embodiments of the present application;

[0020] Figure 3 It is a schematic diagram of the processing logic of the translation function of the document generation method provided by some embodiments of the present application;

[0021] Figure 4 It is a schematic diagram of the interface of the document generation method provided by some embodiments of the present application;

[0022] Figure 5 It is a schematic diagram of the interface of the document generation method provided by some embodiments of the present application;

[0023] Figure 6 It is a schematic diagram of the interface of the document generation method provided by some embodiments of the present application;

[0024] Figure 7Interface schematic diagram of the document generation method provided for some embodiments of the present application;

[0025] Figure 8 Interface schematic diagram of the document generation method provided for some embodiments of the present application;

[0026] Figure 9 Interface schematic diagram of the document generation method provided for some embodiments of the present application;

[0027] Figure 10 Interface schematic diagram of the document generation method provided for some embodiments of the present application;

[0028] Figure 11 Interface schematic diagram of the document generation method provided for some embodiments of the present application;

[0029] Figure 12 Structural schematic diagram of a document generation device provided for embodiments of the present application;

[0030] Figure 13 Structural schematic diagram of an electronic device provided for embodiments of the present application;

[0031] Figure 14 Hardware structural schematic diagram of an electronic device provided for embodiments of the present application. Detailed implementation manners

[0032] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0033] The terms "first", "second", etc. in the description and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order different from those illustrated or described herein, and the objects distinguished by "first", "second", etc. generally belong to the same category, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0034] With the development of electronic device technology, users are becoming increasingly dependent on the functions provided by electronic devices, such as phone calls, videos, photo albums, etc., to meet the needs of users' work and life. For the photo album function, users can save the images taken, the images downloaded from the browser, and the user screenshots in the photo album for viewing.

[0035] Among them, user screenshots are used in a very large number of scenarios. For example, in the scenario where parents check the students' mastery of knowledge, as parents, they often go online to search for relevant materials on knowledge points. When screening materials, parents will save the materials by taking screenshots, and then organize the content in the screenshots into a document. Another example is in the scenario of saving famous quotes when reading e-books. When a user wants to save a certain text in an e-book while reading, the user can save it by taking a screenshot, and then extract the text in the screenshot to generate reading notes.

[0036] However, when recognizing and extracting text from images, users can only select images one by one, which brings a lot of trouble to users. Specifically, users need to select the images to be processed from the photo album of the electronic device, enter the editing interface of the image, long-press the image in the editing interface, copy the text in the image, paste the extracted text into the document specified by the user, and then, select the next image to be processed from the photo album again, and so on, to organize it into the document required by the user. In this way, the user operation in the process of generating this document is cumbersome, resulting in a low document generation efficiency.

[0037] To solve the problems in the related technology, the embodiments of the present application provide a document generation method, device, equipment and storage medium, which can create a first photo album with a first image processing function. In this way, users only need to add the images to be processed to the first photo album, and they can use the first image processing function corresponding to the first photo album to process the first image to obtain the image information for generating a document. There is no need for users to manually extract the text in the image, nor do they need to manually copy and paste it into the specified document, reducing the user operation for generating a document based on the text in the image and improving the document generation efficiency. If there are multiple images that the user needs to process, they can be added to the first photo album in batches. In this way, through the first image processing function corresponding to the first photo album, multiple first images are processed in batches. Without the need for users to operate one by one, a document can be generated through the image information of multiple first images, which can not only simplify the user operation in the process of generating a document, but also improve the document generation efficiency.

[0038] The following combines the attached Figures 1 to 14 , and through specific embodiments and their application scenarios, the document generation method provided by the embodiments of the present application is described in detail.

[0039] First, in combination with Figure 1A document generation method provided in an embodiment of the present application is described in detail.

[0040] Figure 1 A flowchart of a document generation method provided for some embodiments of the present application.

[0041] like Figure 1 As shown, the document generation method provided in the embodiment of the present application can be applied to electronic devices. Based on this, the document generation method can include steps 110 to 130, as shown below.

[0042] It should be noted that the document generation method provided in the embodiments of the present application can be executed by electronic devices such as mobile phones, tablet computers, laptop computers, PDAs, and in-vehicle electronic devices. In some embodiments of the present application, the document generation method provided in the embodiments of the present application is described by taking an electronic device as the execution subject to execute the document generation method.

[0043] Step 110, receiving a first input from a user to add a first image to a first album; Step 120, in response to the first input, processing the first image according to a first image processing function corresponding to the first album to obtain image information, where the image information includes information related to the content of the first image; Step 130, generating a document based on the image information.

[0044] For example, Figure 2 As shown, the first image may include image 1 and image 2, which may be images taken by the user or images downloaded from a browser. The first image processing function corresponding to the first album may be a translation function. Based on this, when the user moves image 1 and image 2 to the first album, the user may receive a first input to add the first image 201 to the first album. Figure 3 As shown, according to the translation function, the English "The moon is so beautiful tonight." in the first image 201 is translated into Chinese "今夜月亮好美." At this time, the image information may include "今夜月亮好美", "The moon is so beautiful tonight." and "今夜月亮好美". Figure 4 As shown, based on this, according to "The moon is so beautiful tonight." and "The moon is so beautiful tonight", generate documents.

[0045] In this way, the user only needs to add the image to be processed to the first album, and can use the first image processing function corresponding to the first album to process the first image to obtain the image information for generating a document. The user cannot manually extract the text in the image, nor does the user need to manually copy and paste it into the specified document, reducing the user operations for generating a document based on the text in the image and improving the document generation efficiency.

[0046] It should be noted that the number of the first images in the embodiments of the present application can be at least one. When the first image includes at least two first images, multiple first images can be processed batch by batch according to the first image processing function corresponding to the first album. Without the need for the user to operate one by one, a document can be generated through the image information of multiple first images, which can not only simplify the user operations in the document generation process, but also improve the document generation efficiency.

[0047] The following will detail the technical terms that appear in the embodiments of the present application.

[0048] The image processing function refers to the functional characteristics possessed by the album, and the functional characteristics can be specified by the user. The image processing function in the embodiments of the present application can refer to the first image processing function and the candidate processing image in the embodiments of the present application. Among them, the first image processing function can include at least one of the following: translation function, multi-modal information extraction function.

[0049] A document refers to an electronic document, which is an information carrier stored and presented in digital form, and is a carrier for recording, storing, and transmitting information. It can contain various forms such as text, pictures, tables, and charts, and will be saved as an electronic file or a printed matter. The document can be a document of various office software, and can also be a data medium such as a note or a notepad and the data recorded thereon. Among them, the documents of various office software include but are not limited to word processing documents, spreadsheet documents, and presentation documents.

[0050] The translation function refers to the function of converting the content of one language, language variety, or language variant into another language or language variant. In the embodiments of the present application, it is to translate the first language such as English into the second language such as Chinese.

[0051] A language variety refers to the type of language, which is used to describe a specific language or a group of languages classified according to certain criteria. For example, Chinese, English, Japanese, etc. can all be called independent language varieties.

[0052] The multi-modal information extraction function refers to the function of extracting information in multiple different data forms from an image. These data forms can include images, texts, voices, videos, etc.

[0053] Modal information refers to the information of one data form extracted from multiple different data forms in an image, such as graphic information, text information, etc.

[0054] The document storage space refers to the place for storing data. The document storage space includes the space inside the electronic device for storing data, and may also include the storage space of the cloud service of the application program in the electronic device. The document storage space in the embodiments of the present application may include a first document storage space and a second document storage space for storing documents.

[0055] The above steps will be described in detail as follows.

[0056] First, regarding step 110, in some embodiments of the present application, before step 110, the embodiments of the present application provide a creation process of the first album. Based on this, the document generation method may further include steps 210 to 240.

[0057] Step 210, in the case of displaying the album interface of the album application, receive a fifth input from the user to the new function album control in the album interface.

[0058] Exemplarily, as Figure 5 shown, display the album interface 50 of the album application. The album interface 50 includes a new function album control 51. The user can click on the new function album control 51 to create a new album with image processing functions.

[0059] Step 220, in response to the fifth input, display a function list, where the function list includes at least two candidate image processing functions.

[0060] Exemplarily, as Figure 6 shown, when receiving the fifth input from the user to the new function album control 51, a function list 60 can be displayed. The function list 60 includes two candidate image processing functions, such as a translation function and a multimodal information extraction function.

[0061] Step 230, receive a sixth input from the user to the first image processing function among at least two candidate image processing functions.

[0062] Exemplarily, continue to refer to Figure 6 , and the user can select the first image processing function required by the user, such as the translation function, from the translation function and the multimodal information extraction function according to their own needs.

[0063] Step 240, in response to the sixth input, create a new first album and associate the first album with the first image processing function.

[0064] Exemplarily, as Figure 7 shown, a new first album can be created in the album application and the first album is associated with the translation function.

[0065] In this way, a first photo album with a first image processing function can be created according to the user's image processing needs, so that the first images can be processed in batches through the first image processing function corresponding to the first photo album to obtain image information for generating documents. Without the need for the user to operate one by one, the document can be generated through the image information of multiple first images, which can not only simplify the user operations in the document generation process, but also improve the document generation efficiency.

[0066] It should be noted that if the first image processing function selected by the user in step 230 is at least two image processing functions, for example, the first image processing function is a translation function and a multimodal information extraction function, then the number of newly created first albums can be two, i.e., a first album corresponding to the translation function and a first album corresponding to the multimodal information extraction function can be created in the album application; or, a single first album can be created, i.e., a first album having both the translation function and the multimodal information extraction function can be created in the album application.

[0067] Secondly, regarding step 120, the first image processing function in the embodiment of the present application may include a translation function or a multimodal information extraction function. The process of obtaining image information is described in detail below based on different first image processing functions.

[0068] In some embodiments of the present application, the first image processing function includes a translation function. Based on this, step 120 may specifically include step 1201 and step 1202 .

[0069] Step 1201 : Using a translation function, translate a first text in a first language in a first image to obtain a second text in a second language.

[0070] For example, if the first language is an application, the first text is "The moon is so beautiful tonight.", and the second language is Chinese, the first text in the first language can be translated through the translation function to obtain "The moon is so beautiful tonight."

[0071] It should be noted that the second language may be Chinese by default, or multiple languages may be provided for the user to select, so that the second text is translated into the language required by the user.

[0072] Step 1202: Generate image information based on the second text, where the image information includes the second text, or the first text and the second text.

[0073] For example, “The moon is so beautiful tonight” may be determined as image information, and “The moon is so beautiful tonight.” and “The moon is so beautiful tonight” may also be determined as image information.

[0074] Thus, the first image processing function corresponding to the first photo album can be used to process the first image to obtain the image information for generating a document. The user cannot manually extract the text in the image, nor does the user need to manually copy and paste it into a specified document, reducing the user operations for generating a document based on the text in the image and improving the document generation efficiency.

[0075] In some embodiments of the present application, the first image processing function includes a multimodal information extraction function. Based on this, step 120 may specifically include step 1203 and step 1204.

[0076] Step 1203: Extract information from the first image according to the multimodal information extraction function to obtain modal information.

[0077] Exemplarily, still taking the first image in Figure 2 as Image 201 for example, the text “The moonisso beautifultonight.” is included in Image 201. Through the multimodal information extraction function, this Image 201 can be recognized and the text “Themoonis so beautifultonight.” can be extracted.

[0078] Step 1204: Generate image information according to the modal information.

[0079] Exemplarily, the text “The moonis so beautifultonight.” can be determined as the image information.

[0080] It should be noted that if the first image includes both text and an image, the text and the image in the first image can be extracted separately, and the text and the image can be determined as the image information.

[0081] Thus, the first image processing function corresponding to the first photo album can be used to process the first image to obtain the image information for generating a document. The user cannot manually extract the text in the image, nor does the user need to manually copy and paste it into a specified document, reducing the user operations for generating a document based on the text in the image and improving the document generation efficiency.

[0082] In some embodiments, if the user wants to both translate and extract the first text, the first image can be added separately to the first album with a translation function and the first album with a multimodal information extraction function. At this time, the image information obtained for the same image under different functions can be determined as the image information of the first image. For example, the image information obtained for the first image through the translation function and the image information obtained for the first image through the multimodal information extraction function are both used as the image information of the first image, so that the image information obtained for the first image through the translation function and the image information obtained for the first image through the multimodal information extraction function can be jointly displayed in the same document for the user to view.

[0083] Then, referring to step 130, in some embodiments of the present application, before step 130, it can be determined whether there is a first document associated with the first album. If it is determined that there is a first document associated with the first album, the first document can be updated according to the image information. Conversely, if it is determined that there is no first document associated with the first album, a new second document can be created to add the image information to the second document, so as to obtain the document required by the user. Based on this, step 130 can specifically include step 1301 or step 1302.

[0084] Step 1301, in the case where the first album is associated with the first document, add the image information to the first document to obtain a document.

[0085] Exemplarily, as Figure 8 shown, if the electronic device includes a first document associated with the first album, that is, the first album includes a second image other than the first image, and the first document includes the image information of the second image, such as the text "The spring scenery is just right, and the gentle breeze is not dry". At this time, "The moon is so beautiful tonight." and "The moon is so beautiful tonight" can be added to the first document. Specifically, it can be added after the text "The spring scenery is just right, and the gentle breeze is not dry".

[0086] Step 1302, in the case where the first album is not associated with the first document, create a second document and add the image information to the second document to obtain a document, where the document name of the document is determined by the album name of the first album.

[0087] Exemplarily, in the case where the first album is not associated with the first document, a second document can be created. The document name of the second document can be determined by the album name of the first album. Referring to Figure 4 , the document name of the second document can be "Document of the First Album". At this time, "The moon is so beautiful tonight." and "The moon is so beautiful tonight" can be added to the second document.

[0088] Thus, the user only needs to add the image to be processed to the first album, and can utilize the first image processing function corresponding to the first album to process the first image to obtain the image information for generating a document. There is no need for the user to manually extract the text in the image, nor to manually copy and paste it into a specified document, reducing the user operations for generating a document based on the text in the image and improving the document generation efficiency.

[0089] Here, it should be noted that if the image processing function of the first album includes a translation function and a multimodal information extraction function, the first document can include two documents. One document is used to store and present the image information after being processed by the translation function, and the other document is used to store and present the image information after being processed by the multimodal information extraction function; alternatively, the first document can be one document, and in different regions of this document, the image information after being processed by the translation function and the image information after being processed by the multimodal information extraction function are stored and presented.

[0090] In some embodiments of the present application, in step 1302, before creating a new blank second document, a function of providing the user with a choice of document format can also be provided. Based on this, the document generation method can further include steps 1303 to 1305.

[0091] Step 1303, display at least two document format options. Among them, the document format options can include the document formats provided by various office software documents, including but not limited to the word document format, Excel document format options, and pdf document format.

[0092] Exemplarily, as Figure 9 shown, after the user adds the first image to the first album, such as Figure 2 shown, after adding image 201 to the first album, 3 document format options are displayed, such as the word document format option 909, the pdf document format option 910, and the Excel document format option 911.

[0093] Step 1304, receive a second input from the user for the target document format option among at least two document format options.

[0094] Exemplarily, still referring to Figure 9 , if the user wants to generate a word document, then, the word document format option 909 can be clicked.

[0095] Step 1305, in response to the second input, create a second document, where the document format of the second document is the document format indicated by the target document format option.

[0096] Exemplarily, the document format of the second document is word.

[0097] Accordingly, options for the user to select the format of the document can be provided. That is, when the user selects the document format, a document with that document format can be generated, without the user separately selecting the document format for storing the document, thereby improving the efficiency of generating the document.

[0098] In some embodiments of the present application, the embodiments of the present application further provide a function of determining the location for storing the document. Based on this, after step 130, the method may further include step 1401 or step 1402.

[0099] Step 1401: When the first album is associated with the first document storage space, store the document in the first document storage space.

[0100] Step 1402: When the first album is not associated with the first document storage space, store the document in the second document storage space specified by the user.

[0101] Accordingly, a function of allowing the user to select the storage location of the document can be provided, facilitating the user to store the document in the document storage space indicated by the user and facilitating the user to view the document.

[0102] In some embodiments of the present application, the embodiments of the present application further provide a function of sharing the document. Based on this, after step 130, the method may further include steps 1501 to 1504.

[0103] Step 1501: Receive a third input from the user for the document.

[0104] Exemplarily, as Figure 10 shown, the interface for displaying the document may include a document sharing control 1001, and the user can click on the document sharing control 1001 to share the document.

[0105] Step 1502: In response to the third input, display at least two sharing application options.

[0106] Exemplarily, as Figure 11 shown, in response to the third input from the user for the document sharing control 1001, display at least two sharing application options, such as option 1101 for application A and option 1102 for application B.

[0107] Step 1503: Receive a fourth input from the user for the target sharing application option among the at least two sharing application options.

[0108] Exemplarily, as Figure 11 shown, if the user wants to share the document with application B, then the user can click on option 1102 for receiving application B.

[0109] Step 1504: In response to a fourth input, send a document to a target sharing application, where the target sharing application is the application indicated by the target sharing application option.

[0110] Exemplarily, share the document with Application B so that the user can view the document on Application B.

[0111] Thus, options for applications to which the document can be forwarded can be provided to the user. That is, when the user selects an application to receive the document, a document is generated and sent to the application selected by the user or a contact in the application, without the user having to switch to the interface of the application or the contact in the application to select sharing the document, thereby improving the efficiency of sharing the document.

[0112] In the embodiment of the present application, the execution subject of the document generation method may be a document generation device. In the embodiment of the present application, taking the document generation device executing the document generation as an example, the device of the document generation method provided in the embodiment of the present application is described.

[0113] The present application also provides a document generation device. Specifically, it is described in detail in combination with Figure 12 for detailed description.

[0114] Figure 12 FIG. is a schematic structural diagram of a document generation device provided in an embodiment of the present application.

[0115] As Figure 12 shown, the document generation device 120 may be applied to an electronic device. Specifically, the document generation device 120 may include:

[0116] A receiving module 1201, configured to receive a first input of a user adding a first image to a first album;

[0117] A processing module 1202, configured to, in response to the first input, process the first image according to a first image processing function corresponding to the first album to obtain image information;

[0118] A generating module 1203, configured to generate a document according to the image information, where the document includes information related to the content of the first image.

[0119] The document generation device 120 in the embodiment of the present application is described in detail below, as specifically shown below.

[0120] In some embodiments of the present application, the processing module 1202 may specifically be configured to, when the first image processing function includes a translation function, translate a first text in a first language in the first image according to the translation function to obtain a second text in a second language;

[0121] Generate image information according to the second text, where the image information includes the second text, or the first text and the second text.

[0122] In some embodiments of the present application, the processing module 1202 may specifically be configured to, when the first image processing function includes a multimodal information extraction function, extract information from the first image according to the multimodal information extraction function to obtain modal information;

[0123] Generate image information according to the modal information.

[0124] In some embodiments of the present application, the generating module 1203 may specifically be configured to, when the first album is associated with the first document, add the image information to the first document to obtain a document;

[0125] When the first album is not associated with the first document, create a second document and add the image information to the second document to obtain a document, where the document name of the document is determined by the album name of the first album.

[0126] In some embodiments of the present application, the document generation device 120 may further include a display module for displaying at least two document format options;

[0127] The receiving module 1201 may further be configured to receive a second input from the user for a target document format option among the at least two document format options;

[0128] The generating device 120 may further include a creating module for creating a second document in response to the second input, where the document format of the second document is the document format indicated by the target document format option.

[0129] In some embodiments of the present application, the document generation device 120 may further include a storage module for storing the document in the first document storage space when the first album is associated with the first document storage space;

[0130] When the first album is not associated with the first document storage space, store the document in a second document storage space specified by the user.

[0131] In some embodiments of the present application, the receiving module 1201 may further be configured to receive a third input from the user for the document;

[0132] The document generation device 120 may further include a display module for displaying at least two sharing application options in response to the third input;

[0133] The receiving module 1201 may further be configured to receive a fourth input from the user for a target sharing application option among the at least two sharing application options;

[0134] The document generation device 120 may further include a sending module, configured to send a document to a target sharing application in response to a fourth input, where the target sharing application is the application indicated by the target sharing application option.

[0135] In some embodiments of the present application, the receiving module 1201 may further be configured to, when the album interface of the album application is displayed, receive a fifth input from the user to a new function album control in the album interface;

[0136] The document generation device 120 may further include a display module, configured to display a function list in response to the fifth input, where the function list includes at least two candidate image processing functions;

[0137] The receiving module 1201 may further be configured to receive a sixth input from the user to a first image processing function among the at least two candidate image processing functions;

[0138] The document generation device 120 may further include a creation module, configured to create a first album in response to the sixth input and associate the first album with the first image processing function.

[0139] The document generation device in the embodiments of the present application may be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than terminals. Exemplarily, the electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. It may also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0140] The document generation device in the embodiments of the present application may be a device with an operating system. The operating system may be an Android operating system, an IOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0141] The device collaboration device provided by the embodiments of the present application can achieveFigures 1 to 11 For the sake of avoiding repetition, the processes implemented by the embodiments of the document generation method shown above achieve the same technical effects and will not be elaborated here.

[0142] Based on this, the document generation apparatus provided by the embodiments of the present application can, when receiving a first input for a user to add a first image to a first photo album, process the first image according to the first image processing function corresponding to the first photo album to obtain image information, where the image information includes information related to the content of the first image; and generate a document according to the image information. In this way, the user only needs to add the image to be processed to the first photo album, and can use the first image processing function corresponding to the first photo album to process the first image to obtain the image information for generating a document, without the user manually extracting the text in the image or manually copying and pasting it into a specified document, reducing the user operations for generating a document based on the text in the image and improving the document generation efficiency.

[0143] Optionally, as Figure 13 shown, the embodiments of the present application further provide an electronic device 130, including a processor 1301 and a memory 1302. A program or instruction that can run on the processor 1301 is stored on the memory 1302. When the program or instruction is executed by the processor 1301, it implements each step of the above-mentioned document generation method embodiments and can achieve the same technical effects. For the sake of avoiding repetition, it will not be elaborated here.

[0144] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0145] Figure 14 FIG. is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application.

[0146] The electronic device 1400 includes, but is not limited to: a radio frequency unit 1401, a network module 1402, an audio output unit 1403, an input unit 1404, a sensor 1405, a display unit 1406, a user input unit 1407, an interface unit 1408, a memory 1409, a processor 1410, and other components.

[0147] Those skilled in the art can understand that the electronic device 1400 may further include a power source (such as a battery) for supplying power to each component. The power source can be logically connected to the processor 1410 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 14 The structure of the electronic device shown in FIG. does not limit the electronic device. The electronic device may include more or fewer components than shown, or combine some components, or have different component arrangements, which will not be elaborated here.

[0148] Among them, in the embodiments of the present application, the user input unit 1407 is used to receive a first input from the user to add a first image to the first album. The processor 1410 is configured to, in response to the first input, process the first image according to the first image processing function corresponding to the first album to obtain image information. The processor 1410 is configured to generate a document according to the image information, where the document includes information related to the content of the first image.

[0149] The electronic device 1400 will be described in detail below, as follows.

[0150] In some embodiments of the present application, the processor 1410 may specifically be configured to, when the first image processing function includes a translation function, translate the first text in the first language in the first image according to the translation function to obtain a second text in a second language;

[0151] Generate image information according to the second text, where the image information includes the second text, or the first text and the second text.

[0152] In some embodiments of the present application, the processor 1410 may specifically be configured to, when the first image processing function includes a multi-modal information extraction function, extract information from the first image according to the multi-modal information extraction function to obtain modal information;

[0153] Generate image information according to the modal information.

[0154] In some embodiments of the present application, the processor 1410 may specifically be configured to, when the first album is associated with a first document, add the image information to the first document to obtain a document;

[0155] When the first album is not associated with the first document, create a second document and add the image information to the second document to obtain a document, where the document name of the document is determined by the album name of the first album.

[0156] In some embodiments of the present application, the display unit 1406 is configured to display at least two document format options;

[0157] The user input unit 1407 may also be configured to receive a second input from the user for a target document format option among the at least two document format options;

[0158] The processor 1410 may also be configured to, in response to the second input, create a second document, where the document format of the second document is the document format indicated by the target document format option.

[0159] In some embodiments of the present application, the memory 1409 is configured to store the document in the first document storage space when the first album is associated with the first document storage space;

[0160] When the first photo album is not associated with the first document storage space, store the document in the second document storage space specified by the user.

[0161] In some embodiments of the present application, the user input unit 1407 is configured to receive a third input of a document from the user;

[0162] The display unit 1406 is configured to display at least two sharing application options in response to the third input;

[0163] The receiving module 1201 may also be configured to receive a fourth input of the user for a target sharing application option among the at least two sharing application options;

[0164] The network module 1402 is configured to send the document to the target sharing application in response to the fourth input, where the target sharing application is the application indicated by the target sharing application option.

[0165] In some embodiments of the present application, the user input unit 1407 may also be configured to receive a fifth input of the user for a new function photo album control in the photo album interface when the photo album application's photo album interface is displayed;

[0166] The display unit 1406 is configured to display a function list in response to the fifth input, and the function list includes at least two candidate image processing functions;

[0167] The user input unit 1407 may also be configured to receive a sixth input of the user for a first image processing function among the at least two candidate image processing functions;

[0168] The processor 1410 is configured to create a first photo album in response to the sixth input and associate the first photo album with the first image processing function.

[0169] It should be understood that the input unit 1404 may include a Graphics Processing Unit (GPU) 14041 and a microphone 14042. The graphics processor 14041 processes the image data of static images or videos obtained by an image capture device (such as a camera) in the video capture mode or the image capture mode. The display unit 1406 may include a display panel, and the display panel may be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 1407 includes at least one of a touch panel 14071 and other input devices 14072. The touch panel 14071 is also called a touch screen. The touch panel 14071 may include two parts: a touch detection device and a touch display. The other input devices 14072 may include, but are not limited to, a physical keyboard, function keys (such as volume display keys, switch keys, etc.), a trackball, a mouse, a joystick, which will not be elaborated here.

[0170] The memory 1409 can be used to store software programs and various data. The memory 1409 mainly includes a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area can store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 1409 can include a volatile memory or a non-volatile memory, or the memory 1409 can include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDRSDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 1409 in the embodiments of the present application includes, but is not limited to, these and any other suitable types of memories.

[0171] The processor 1410 can include one or more processing units; optionally, the processor 1410 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless display signals, such as a baseband processor. It can be understood that the above-mentioned modem processor may not be integrated into the processor 1410.

[0172] The embodiments of the present application also provide a readable storage medium. A program or instruction is stored on the readable storage medium. When the program or instruction is executed by a processor, it implements each process of the above-mentioned embodiment of the document generation method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0173] Among them, the processor is the processor in the electronic device in the above embodiments. Among them, the readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disks or optical discs, etc.

[0174] In addition, an embodiment of the present application further provides a chip, which includes a processor and a display interface. The display interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above embodiment of the document generation method, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0175] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0176] An embodiment of the present application provides a computer program product. The program product is stored in a storage medium and is executed by at least one processor to implement each process of the above embodiment of the document generation method, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0177] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.

[0178] In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed. It may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0179] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of various embodiments of the present application.

[0180] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

Claims

1. A document generation method, characterized in that, Including: Receiving a first input from a user to add a first image to a first photo album; In response to the first input, processing the first image according to a first image processing function to obtain image information; Generating a document according to the image information, where the document includes information related to the content of the first image.

2. The method according to claim 1, characterized in that, The first image processing function includes a translation function; the processing the first image according to the first image processing function to obtain image information includes: Translating the first text in the first language in the first image according to the translation function to obtain a second text in a second language; Generating the image information according to the second text, where the image information includes the second text, or the first text and the second text.

3. The method according to claim 1, characterized in that, The first image processing function includes a multi-modal information extraction function; the processing the first image according to the first image processing function to obtain image information includes: Extracting information from the first image according to the multi-modal information extraction function to obtain modal information; Generating the image information according to the modal information.

4. The method according to claim 1, wherein The generating the document according to the image information includes: In the case where a first document is associated with the first photo album, adding the image information to the first document to obtain the document; In the case where the first photo album is not associated with the first document, creating a second document and adding the image information to the second document to obtain the document, where the document name of the document is determined by the photo album name of the first photo album.

5. The method according to claim 4, characterized in that, Before creating the second document, the method further includes: Displaying at least two document format options; Receiving a second input from the user for a target document format option among the at least two document format options; In response to the second input, creating a second document, where the document format of the second document is the document format indicated by the target document format option.

6. The method according to claim 1, characterized in that The method further includes: In the case where a first document storage space is associated with the first photo album, storing the document in the first document storage space; In the case where the first photo album is not associated with the first document storage space, storing the document in a second document storage space specified by the user.

7. The method according to claim 1, characterized in that, The method further includes: Receiving a third input from the user for the document; In response to the third input, displaying at least two sharing application options; Receiving a fourth input from the user for a target sharing application option among the at least two sharing application options; In response to the fourth input, sending the document to a target sharing application, where the target sharing application is the application indicated by the target sharing application option.

8. The method according to claim 1, wherein Before receiving the first input from the user to add a first image to a first photo album, the method further includes: In the case of displaying the photo album interface of the photo album application, receiving a fifth input from the user for a new function photo album control in the photo album interface; In response to the fifth input, displaying a function list, where the function list includes at least two candidate image processing functions; Receiving a sixth input from the user for a first image processing function among the at least two candidate image processing functions; In response to the sixth input, create the first photo album and associate the first photo album with the first image processing function.

9. A document generation device, characterized in that, Comprising: a receiving module, configured to receive a first input from a user for adding a first image to the first photo album; a processing module, configured to, in response to the first input, process the first image according to the first image processing function to obtain image information; a generating module, configured to generate a document according to the image information, where the document includes information related to the content of the first image.

10. An electronic device, characterized in that, Comprising: a processor, a memory, and a program or instruction stored on the memory and executable on the processor, where when the program or instruction is executed by the processor, the steps of the document generation method according to any one of claims 1-8 are implemented.