Method and apparatus for processing document, and device and medium

By generating structured data and combining it with descriptive data in text and graphical formats, the problem that machine learning models cannot fully reflect the document summary is solved, resulting in more accurate and richer document descriptions.

WO2026044623A1PCT designated stage Publication Date: 2026-03-05BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/115649
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-29
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Existing machine learning models cannot fully and accurately reflect the summary content of multimodal documents, especially the information in graphical data is not effectively utilized.

Method used

By generating structured data and combining it with descriptive data in text and graphics formats, and using language models and graphics models to process the text and graphics components respectively, descriptive data that includes multiple aspects of text and graphics is generated.

Benefits of technology

It provides a more comprehensive and accurate document summary, helping users better understand the document content and improving the accuracy and richness of the described data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024115649_05032026_PF_FP_ABST
    Figure CN2024115649_05032026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a method and apparatus for processing a document, and a device and a medium. The method comprises: receiving a user request, which is used for acquiring description data of a document; on the basis of the document, generating structured data; and on the basis of the structured data, determining the description data of the document, wherein the description data comprises a portion represented in a text format and a portion represented in a graphical format. By using an exemplary implementation of the present disclosure and in this way, description data can be determined in a multi-modal manner, thereby facilitating better understanding of the content of a document by a user.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatus, devices, and media for processing documents Technical Field

[0001] Exemplary implementations of this disclosure generally relate to document processing, and more particularly to methods, apparatus, devices, and computer-readable storage media for extracting descriptive data from documents. Background Technology

[0002] With the rapid development of computer technology, machine learning models (e.g., language models) can provide users with various services for processing documents. For example, an application can receive user requests to process documents (e.g., extracting descriptive data including a summary of the document's content, etc.). Documents may include text data, and some documents may include data in multiple modalities (e.g., text, images, etc.). However, the descriptive data provided by a language model only includes text data and cannot comprehensively and accurately reflect the summary content of the document. Therefore, a more efficient way to process documents is desired.

[0003] Summary of the Invention

[0004] In a first aspect of this disclosure, a method for processing a document is provided. In this method, a user request for obtaining descriptive data of a document is received; structured data is generated based on the document; and descriptive data of the document is determined based on the structured data, the descriptive data including portions represented in text format and portions represented in graphical format.

[0005] In a second aspect of this disclosure, an apparatus for processing a document is provided. The apparatus includes: a receiving module configured to receive a user request for obtaining descriptive data of a document; a generating module configured to generate structured data based on the document; and a determining module configured to determine descriptive data of the document based on the structured data, the descriptive data including portions represented in text format and portions represented in graphic format.

[0006] In a third aspect of this disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to a first aspect of this disclosure when executed by the at least one processing unit.

[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, causes the processor to implement the method according to a first aspect of this disclosure.

[0008] In a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method according to a first aspect of this disclosure.

[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the implementation of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] In the following detailed description, the above and other features, advantages, and aspects of the various implementations of this disclosure will become more apparent, taken in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0011] Figure 1 shows a block diagram of an application environment according to an exemplary implementation of the present disclosure;

[0012] Figure 2 shows a block diagram for processing documents according to some implementations of this disclosure;

[0013] Figure 3 shows a block diagram of generating structured data according to some implementations of this disclosure;

[0014] Figure 4 shows a block diagram illustrating the extraction of graphical descriptions according to some implementations of this disclosure;

[0015] Figure 5 shows a block diagram of structured data according to some implementations of this disclosure;

[0016] Figure 6 shows a block diagram of obtaining descriptive data according to some implementations of this disclosure;

[0017] Figure 7 shows a block diagram of obtaining descriptive data according to some implementations of this disclosure;

[0018] Figure 8 shows a block diagram of a user page according to some implementations of this disclosure;

[0019] Figure 9 shows a flowchart of a method for processing documents according to some implementations of this disclosure;

[0020] Figure 10 shows a block diagram of an apparatus for processing documents according to some implementations of the present disclosure; and

[0021] Figure 11 shows a block diagram of a device capable of implementing various implementations of the present disclosure. Detailed Implementation

[0022] Implementations of this disclosure will now be described in more detail with reference to the accompanying drawings. While some implementations of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the implementations set forth herein. Rather, these implementations are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and implementations of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0023] In the description of the implementation methods disclosed herein, the term "comprising" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". Other explicit and implicit definitions may also be included below. As used herein, the term "model" can represent the relationships between various data. For example, the aforementioned relationships can be obtained based on various currently known and / or future-developed technical solutions.

[0024] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0025] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained.

[0026] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0027] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, for example, via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose whether to "agree" or "disagree" to provide personal information to the electronic device.

[0028] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0029] The term "in response to" as used herein refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of subsequent actions performed in response to such event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is met. For example, in some cases, subsequent actions may be performed immediately upon the occurrence of the event or the fulfillment of the condition; while in others, they may be performed some time after the occurrence of the event or the fulfillment of the condition.

[0030] Example Environment

[0031] With the rapid development of computer technology, machine learning models (e.g., language models) can provide various services to users. For example, an application can obtain user requests from users to process documents (e.g., extract descriptive data including a summary of the document's content, etc.). Referring to Figure 1, which illustrates a block diagram 100 of an application environment according to an exemplary implementation of this disclosure, as shown in Figure 1, a user can enter a user request in area 140 of a user page 110, click control 142 to submit the user request, and receive a response from the application.

[0032] For example, a user can submit a user request 120 and a document 122 to be processed. The application will then analyze the specific content of document 122 and provide a response 130 that includes the main content of the document. The document may include text data, and alternatively and / or additionally, it may include data in multiple modalities (e.g., text, images, etc.). However, the descriptive data provided by the language model only includes text data and cannot fully and accurately reflect the summary content of the document. Therefore, a more efficient way to process the document is desired.

[0033] Summary of document processing

[0034] To at least partially address the shortcomings of the prior art, a method for processing documents is proposed according to an exemplary implementation of this disclosure. Referring to FIG2, which describes an outline of an exemplary implementation of this disclosure, FIG2 illustrates a block diagram 200 for processing documents according to some implementations of this disclosure. As shown in FIG2, a user request for obtaining descriptive data of document 210 can be received. Here, document 210 may include text data. Alternatively and / or additionally, document 210 may be represented in various formats, and document 210 may include text data 212 (e.g., text represented in various natural languages) and graphic data 214 (e.g., other data with visual effects besides text data, such as images, tables, etc.). Here, graphic data includes at least one graphic representation. Further, the document may include a local document or a web page document.

[0035] Structured data can be generated from documents; furthermore, descriptive data of the document can be determined based on the structured data, which includes parts represented in text format and parts represented in graphic format. Specifically, text data can be extracted from the document to generate corresponding structured data. Furthermore, structured data can be used to generate descriptive data that includes graphic data.

[0036] Alternatively and / or additionally, when the document includes graphic data, a graphic description of at least one graphic representation can be extracted. Here, the graphic description is represented in text format and may include various descriptive aspects of the graphic representation, such as text identified from the graphic representation, the image content of the graphic representation, etc. Structured data 220 can be generated based on the text data and the graphic description of the graphic representation. Here, the structured data 220 is represented in text format. Further, description data 230 of document 210 can be determined based on the structured data 220, the description data 230 including graphic representations (e.g., graphic representations 232 and 234, etc.).

[0037] Using the exemplary implementation of this disclosure, regardless of whether the document is represented in a multimodal manner, the generated descriptive data can include a variety of information, including text and graphics, thus better describing the summary content of the document. In this way, descriptive data can be determined in a multimodal manner, thereby helping users to better understand the document content.

[0038] Detailed process of document processing

[0039] Having described an outline of some implementations according to this disclosure, further details for processing documents are provided below. According to some implementations of this disclosure, generating structured data based on a document includes: generating structured data based on the text data in response to determining that the document includes text data represented in text format; and generating graphical data based on the text content. According to some implementations of this disclosure, text data 212 may include one or more text fragments, which can be extracted from the document to generate structured data. Here, text fragments may correspond to different text units, such as pages, paragraphs, sentences, or have a predetermined length (e.g., 200 words, or other numerical values), etc.

[0040] According to some implementations of this disclosure, structured data can be input into a machine learning model to generate corresponding descriptive data. Here, the machine learning model can be implemented in various ways; a language model can be used to generate the text portion of the descriptive data, and a graph model (e.g., a text-to-image model, a text-to-formula model, etc.) can be used to generate the graphical portion of the descriptive data. For example, a language model can be used to summarize the document's overview information, and the overview information can be input into a graph model, with prompts constructed to invoke the graph model to generate the corresponding graphical representation. Specifically, assuming the overview information includes: Step 1: ...; Step 2: ...; and Step 3, the prompts could include "Please generate a flowchart according to the steps in the overview information, where each step corresponds to a block diagram in the flowchart." In this case, regardless of whether the original document includes graphical formats, descriptive data involving both text and graphical formats can be generated, thereby providing a more comprehensive and richer document overview.

[0041] According to some implementations of this disclosure, the document may further include graphic data 214, and the graphic data 214 may include one or more graphic representations. See Figure 3 for more information, which shows a block diagram 300 for generating structured data according to some implementations of this disclosure. As shown in Figure 3, the document 210 may include text fragments 310, 312, 314, etc. Here, text fragments may correspond to different text units, such as pages, paragraphs, sentences, etc. For ease of description, only paragraphs are used as examples of text fragments in the following description of the document processing process.

[0042] According to some implementations of this disclosure, in the process of generating structured data based on a document, in response to determining that the document includes graphic data represented in a graphic format, at least one graphic representation is extracted from the graphic data; a graphic description of at least one graphic representation is extracted from the document, the graphic description being represented in text format; and the structured data is updated based on the graphic description.

[0043] Here, graphical representations can include various types, such as, but not limited to: still image types, table types, formula types, canvas types, animated image types, and video types. Still image types can include illustrations in a document, which can be represented using different image formats. Table types can include text tables, graphic tables, etc., and can be represented in image format or a dedicated table format. Formula types can include various formulas and can be represented in image format or the format of a dedicated formula editor. Canvas types can include various drawn patterns or diagrams and can be represented in image format or a dedicated canvas format. Animated image types and video types can include multiple images.

[0044] Using some implementations of this disclosure, the generated descriptive data can include multifaceted information. In this way, the descriptive data can be determined in a multimodal manner, thereby helping users better understand the document content. For ease of description, the following description uses only static image types as examples of graphical representation to illustrate more details of some implementations according to this disclosure. Specifically, graphical descriptions can be inserted into structured data determined based on various text fragments; that is, graphical descriptions can be used to update the structured data.

[0045] According to some implementations of this disclosure, during the generation of structured data, the position of the graphic representation in the document can be determined; a graphic description can be inserted at the corresponding graphic position in the text data to form structured data. As shown in Figure 3, text fragments 310, 312, graphic representation 320, and 314 are arranged in sequence. At this time, a graphic description 322 in text format can be extracted from the graphic representation 320, and structured data 220 is generated according to the order of text fragments 310, 312, graphic description 322, and text fragment 314. In other words, the graphic representation 320 can be replaced by the graphic description 322.

[0046] Using some implementation methods of this disclosure, the generated structured data 220 only involves text-formatted data and does not include graphical data. Therefore, it can be input into a language model to extract corresponding descriptive data. In this way, the powerful processing capabilities of the language model can be utilized, thereby improving the accuracy of obtaining descriptive data.

[0047] According to some implementations of this disclosure, the graphic description of a graphic representation can be extracted based on at least one of the following: performing text recognition on the graphic representation to determine the graphic description; performing image recognition on the graphic representation to determine the graphic description; performing context recognition on the graphic representation to determine the graphic description; or performing speech recognition on the graphic representation to determine the graphic description.

[0048] See Figure 4 for further details, which illustrates a block diagram 400 of extracting a graphic representation according to some implementations of this disclosure. As shown in Figure 4, text recognition 410 can be performed in various ways. For example, an optical character recognition (OCR) process can be used to extract text information from the graphic representation 320. Alternatively and / or additionally, a text recognition model can be used to extract text information. The graphic representation 320 can be processed using an image recognition process 420; for example, an image-to-text model can be used to determine the specific content in the image and extract text information. The context recognition process 430 can be used to extract text information from the context of the graphic representation 320 (e.g., paragraphs referencing the graphic representation 320, the title of the graphic representation 320, paragraphs before and after the graphic representation 320 in a document, etc.).

[0049] Alternatively and / or additionally, when the graphic representation 320 involves a moving image type and / or a video type, the processing procedures described above can be performed for each video frame. Furthermore, if the video-type graphic representation 320 includes speech data, a speech recognition 440 process can also be performed to extract text information.

[0050] Furthermore, the text information extracted in the above process can be integrated and used to generate a graphical description 322. At this point, the graphical description 322 only involves text formatting and can be inserted into the structured data 220 as input data for a language model. Using some implementations of this disclosure, rich information from the graphical representation 320 can be extracted from multiple aspects, thereby facilitating the generation of more accurate descriptive data.

[0051] According to some implementations of this disclosure, the graphic description may further include identification information of the graphic representation (e.g., name, identifier, etc.), thereby facilitating the differentiation of the graphic representation corresponding to the graphic description. For example, the identification information of each graphic representation can be determined according to the order in which the graphic representations appear in the document. A mapping relationship between a graphic representation and an address used to access that graphic representation can be stored to insert the graphic representation into the description data. As another example, the identification information can be determined using the name of each graphic representation.

[0052] Specifically, for image-type graphic representations, the corresponding graphic description can be: "IMAGE01, Image processing procedure, including: Step 1, Step 2, Step 3…". Here, "IMAGE01" is the identification information, and the subsequent text is the extracted descriptive information; the two parts together constitute the graphic description. For example, for table-type graphic representations, the corresponding graphic description can be: "TABLE01, Experimental data, Comparison of processing performance of various image processing methods…".

[0053] According to some implementations of this disclosure, a uniform label {Graphic}{ / Graphic} can be used to indicate a graphic representation; alternatively and / or additionally, special labels {Image}{ / Image}, etc., can be used to represent different types of graphic representations. In this case, the graphic description can be represented as: "{Image}IMAGE01, image processing procedure, including: step 1, step 2, step 3…{ / Image}", and "{Table}TABLE01, experimental data, comparison of processing performance of various image processing methods…{ / Table}". It should be understood that the above structure is merely illustrative, and other data structures can be used to store the graphic description 322.

[0054] Referring to Figure 5 for more information on generating structured data, Figure 5 shows a block diagram 500 of structured data according to some implementations of this disclosure. As shown in Figure 5, structured data 220 can be generated in the order of text fragment 310, text fragment 312, graphic description 322, and text fragment 314. In other words, the graphic representation 320 in the document can be replaced by the graphic description 322. In this case, the structured data 220 only involves text formatting, while also including the text description corresponding to the graphic representation 320. Using some implementations of this disclosure, the structured data 220 can be input into a language model, and the powerful text processing capabilities of the language model can be used to extract descriptive data from the document.

[0055] According to some implementations of this disclosure, in the process of determining the descriptive data of a document based on structured data, intermediate data associated with the structured data can be determined according to a predetermined model, the intermediate data including the descriptive data of the document represented in text format. Furthermore, in response to determining that the intermediate data includes identification information corresponding to a graphical representation, the relevant portions of the identification information are updated using the graphical representation to determine the descriptive data.

[0056] Referring to Figure 6 for further information, Figure 6 illustrates a block diagram 600 of acquiring descriptive data according to some implementations of this disclosure. A document can be processed and structured data 220 can be generated according to the methods described above. Assuming the document includes multiple paragraphs and Image 1 (identified as IMAGE01) and Table 2 (identified as TABLE02), the generated structured data can include the text of the multiple paragraphs and graphical descriptions corresponding to Image 1 and Table 2, respectively. Cue words 610 can be determined, and the structured data 220 can be input into a predetermined model (e.g., a language-based model 620).

[0057] Here, prompt 610 can specify the task expected to be performed by the predefined model. For example, prompt 610 can be expressed as: Please extract the summary of the structured data, Please summarize the main content of the structured data, etc. Alternatively and / or additionally, the model 620 can be controlled to handle the processing capability related to graphical representations. For example, prompt 610 can be added with: "When the content includes <<{Image}{ / Image}>> or <<{Table}{ / Table}>> fields, the corresponding graphical representation fields can be displayed in an appropriate position in the response."

[0058] At this point, model 620 can process structured data 220 and generate corresponding intermediate data 630. Intermediate data 630 may include summary portions of various text fragments in the document, such as portions 640 and 642. Alternatively and / or additionally, intermediate data 630 may include portions corresponding to summaries of graphical representations. For example, portion 650 corresponds to image 1, and portion 652 corresponds to table 2. Specifically, portion 650 may include the identification information "IMAGE01" for image 1, and portion 652 may include the identification information "TABLE01" for table 1.

[0059] At this time, the relevant part of the identification information can be updated using a graphical representation to determine the description data 230. Specifically, if the intermediate data 630 is detected to include a graphical representation of an image type represented by "IMAGE01", the part 650 can be updated using the graphical representation "Image 1"; if the intermediate data 630 is detected to include a graphical representation of a table type represented by "TABLE02", the part 652 can be updated using the graphical representation (Table 2). At this time, the description data 230 may include a summary of text data, and also includes Image 1 (i.e., graphical representation 232) and Table 2 (i.e., graphical representation 234).

[0060] According to some implementations of this disclosure, in the process of updating relevant parts of the identification information using a graphical representation, at least one of the following can be added to the location of the identification information: a graphical representation, or an access address for accessing the graphical representation. Specifically, part 650 can be updated directly using the graphical representation 232; alternatively and / or additionally, access address 660 can be queried and the access address URL01 corresponding to IMAGE01 can be determined, and part 650 can be updated using that access address. Similar processing can be performed for graphical representations of table type. Specifically, part 652 can be updated directly using the graphical representation 234; alternatively and / or additionally, access address 660 can be queried and the access address URL02 corresponding to TABLE02 can be determined, and part 652 can be updated using that access address.

[0061] According to some implementations of this disclosure, during the process of updating relevant parts of the identification information using graphical representation, it is possible to further verify whether the context of the identification information in the intermediate data matches the graphical description. If a match is determined, the relevant parts of the identification information are updated using graphical representation.

[0062] According to some implementations of this disclosure, in the process of performing context recognition on the graphical representation to determine the graphical description, the text portion in the document that references the graphical representation is identified; and semantic recognition is performed on the text portion to determine the image description. For example, expressions such as "as shown in Figure 1" or "in Figure 1" can be searched in the document to determine which paragraph in the document references the graphical representation. Further, the semantic meaning of the found paragraph can be analyzed to determine the image description. For example, suppose Figure 1 is a flowchart, and the document contains: "As shown in Figure 1, this diagram illustrates three steps: Step 1: ...; Step 2: ...; Step 3: ...". In this case, the image description could include, for example, "It illustrates three steps: Step 1: ...; Step 2: ...; Step 3: ...".

[0063] It should be understood that some documents may contain multiple similar figures, and certain paragraphs may describe each figure separately. In this way, paragraphs that better match the content of the figures can be found within the document, thus determining the corresponding image description. Alternatively and / or additionally, multiple paragraphs in a document may describe the same figure. For example, one paragraph in the document may describe step 1 in the figure, and another paragraph may describe step 2. In this way, all paragraphs related to the figures can be found separately within the document, thereby generating a more complete image description.

[0064] It should be understood that the intermediate data 630 generated by model 620 may sometimes not include text content related to a certain graphic representation. Even if the intermediate data 630 includes the identification information of the graphic representation, it is not necessary to insert the graphic representation. Suppose the document includes image 1A, image 1B, and table 2, and the generated intermediate data only involves image 1A and table 2 but not image 1B. In this case, image 1A and table 2 can be inserted only at the corresponding positions in the intermediate data. Furthermore, the identification information corresponding to image 1B can be removed from the intermediate data.

[0065] It should be understood that the predetermined model here can be a language model fine-tuned using dedicated reference data, and the predetermined model describes the relationship between the reference structured data of the reference document and the reference descriptive data of the reference document. The reference descriptive data is represented in text format and includes reference identification information represented by reference graphics in the reference document. Reference sample data (also known as training data, labeled data, etc.) can be collected and used to fine-tune the language model.

[0066] Specifically, documents (e.g., papers, news articles, emails, etc.) including text and graphic data can be collected. The corresponding structured data and labeled data can be generated according to the predetermined format described above. Specifically, the format of the labeled data can be similar to intermediate data 630, and the labeled data can include a summary portion corresponding to the text data and a summary portion corresponding to the graphic data (represented in text format, e.g., {IMAGE}IMAGE01...{ / IMAGE}, etc.).

[0067] Model 620 can be fine-tuned to reduce its loss. Specifically, the loss function can be constructed based on factors such as recall, precision, and false positive rate for image-type graphical representations, and / or error rates of individual fields in table-type tables, so that model 620 can more accurately describe the relationship between the reference structured data and the reference intermediate data. Utilizing some implementations of this disclosure, the accuracy of the model can be improved, thereby generating more informative and visually appealing descriptive data based on the intermediate data output by the model in a more comprehensible manner.

[0068] It should be understood that the above description of the process for determining descriptive data uses only image and table types as examples of graphical representations. Alternatively and / or additionally, graphical representations of formula and canvas types in documents can be processed in a similar manner. Alternatively and / or additionally, for moving image types and video types that include multiple image frames, each image frame can be processed in a similar manner to generate the corresponding descriptive data.

[0069] Specifically, in response to determining that the graphical representation comprises multiple images (e.g., the type of the graphical representation is a moving image type and / or a video type), multiple keyframes can be extracted from the graphical representation. Processing can be performed on at least one of the multiple keyframes to determine the corresponding graphical description. Utilizing some implementations of this disclosure, richer graphical information can be extracted from the document, resulting in output descriptive data that is more comprehensive and easier for users to understand.

[0070] Referring to Figure 7 for further information, Figure 7 illustrates a block diagram 700 of acquiring descriptive data according to some implementations of this disclosure. As shown in Figure 7, assuming the graphic data 710 is of video type, corresponding keyframes 720, ..., and 722 can be extracted. Further, graphic descriptions of some or all of the keyframes can be extracted to generate corresponding structured data 730. At this time, the structured data 730 may include graphic descriptions corresponding to one or more keyframes. The generated descriptive data 740 may include multiple graphic representations from a single graphic data 710, such as image 742 corresponding to keyframe 720, and image 744 corresponding to keyframe 722, etc.

[0071] According to some implementations of this disclosure, descriptive data can be provided as a response to a user request. For example, descriptive data can be provided on user page 110 as shown in Figure 1. In this case, the descriptive data can present the main content of the document in a graphic and textual format. According to some implementations of this disclosure, the user can specify whether to include graphical representations in the descriptive data. For example, an "Enable / Disable" control can be provided on user page 110 to specify whether to include graphical representations in the descriptive data. Alternatively and / or additionally, the user can specify whether to enable this feature in natural language in the user request.

[0072] According to some implementations of this disclosure, attribute information of the terminal device used to execute the method can be obtained; based on the attribute information of the terminal device, presentation attributes for presenting the graphical representation can be determined; and the graphical representation in the description data can be presented according to the presentation attributes. Here, the presentation attributes may include at least one of the following: format, size, and position. Specifically, the image presentation format can be determined based on the image decoding standard supported by the terminal device. Alternatively and / or additionally, the image size and position can be determined based on the resolution of the terminal device, and so on. For example, if the terminal device has a low resolution, the image size can be appropriately reduced and the image can be presented, and so on. In this way, the description data can be presented in a format and layout more suitable for the terminal device.

[0073] According to some implementations of this disclosure, users can specify whether to include graphical data in the description data. For example, a control can be provided on the user page to specify whether to include graphical data, or the user can specify whether to include graphical data in the user request. In this case, in response to determining that the user has specified that the description data should include graphical data, the above-described procedure can be performed. If it is determined that the user has not specified that the description data should include graphical data, the procedure can be performed in a conventional manner, for example, the document can be processed directly using a language model to provide the description data.

[0074] According to some implementations of this disclosure, a document can be presented in a first display area; and descriptive data and interactive controls for submitting user requests can be presented in a second display area. Figure 8 shows a block diagram 800 of a user page according to some implementations of this disclosure. As shown in Figure 8, the content of the original document 812 can be provided in display area 810, and descriptive data 230 and various interactive controls can be provided in display area 820. In this way, the user can further interact with the descriptive data 230, for example, by selecting a part of the descriptive data to present the corresponding content in the original document 812. Specifically, the user can click on the graphic representation 232, at which point the original document 812 will scroll to the part corresponding to the graphic representation 232, for example, presenting the original paragraph including the graphic representation 232. In this way, it can help the user discover the correspondence between the original document and the descriptive data, thereby better understanding the content of the original document.

[0075] Using the exemplary implementation of this disclosure, the generated descriptive data can include a variety of information, including text and graphics, thus better describing the summary content of the document. In this way, the descriptive data can be determined in a multimodal manner, thereby helping users to better understand the document content.

[0076] Example process

[0077] Figure 9 shows a flowchart of a method 900 for processing a document according to some implementations of this disclosure. At block 910, a user request for obtaining descriptive data of the document is received. At block 920, structured data is generated based on the document. At block 930, based on the structured data, descriptive data of the document is determined, the descriptive data including portions represented in text format and portions represented in graphical format.

[0078] According to some implementations of this disclosure, generating structured data based on a document includes: generating structured data based on the text data in response to determining that the document includes text data represented in text format; and generating graphical data based on text content.

[0079] According to some implementations of this disclosure, generating structured data based on a document includes: in response to determining that the document includes graphic data represented in a graphic format, extracting at least one graphic representation from the graphic data; extracting a graphic description of the at least one graphic representation from the document, the graphic description being represented in text format; and updating the structured data based on the graphic description.

[0080] According to some implementations of this disclosure, generating structured data based on a document further includes: determining the position of the graphic representation in the document; and inserting a graphic description at the corresponding position in the text data to form structured data.

[0081] According to some implementations of this disclosure, the graphical description extracted from the graphical representation includes at least one of the following: performing text recognition on the graphical representation to determine the graphical description; performing image recognition on the graphical representation to determine the graphical description; performing context recognition on the graphical representation to determine the graphical description; or performing speech recognition on the graphical representation to determine the graphical description.

[0082] According to some implementations of this disclosure, performing contextual recognition on a graphic representation to determine a graphic description includes: identifying the text portion in a document that references the graphic representation; and performing semantic recognition on the text portion to determine the image description.

[0083] According to some implementations of this disclosure, determining the description data of a document based on structured data includes: determining intermediate data associated with the structured data according to a predetermined model, the intermediate data including the description data of the document in text format; and in response to determining that the intermediate data includes identification information corresponding to a graphical representation, updating relevant portions of the identification information using the graphical representation to determine the description data.

[0084] According to some implementations of this disclosure, a predetermined model describes the relationship between the reference structured data of the reference document and the reference intermediate data of the reference document, wherein the reference intermediate data is represented in text format.

[0085] According to some implementations of this disclosure, updating relevant parts of the identification information using a graphical representation includes adding at least one of the following to the location of the identification information: a graphical representation, or an access address for accessing the graphical representation.

[0086] According to some implementations of this disclosure, updating the relevant part of the identification information using a graphical representation further includes: in response to determining that the context of the identification information in the intermediate data matches the graphical description, updating the relevant part of the identification information using a graphical representation.

[0087] According to some implementations of this disclosure, the method further includes: obtaining attribute information of a terminal device used to perform the method; determining presentation attributes for presenting a graphical representation based on the attribute information of the terminal device, the presentation attributes including at least one of the following: format, size, position; and presenting the graphical representation in the description data according to the presentation attributes.

[0088] According to some implementations of this disclosure, the method further includes: in response to determining that user-specified descriptive data includes graphical data, executing the method.

[0089] According to some implementations of this disclosure, the method further includes: presenting a document in a first display area; and presenting descriptive data and interactive controls for submitting user requests in a second display area.

[0090] According to some implementations of this disclosure, the graphical representation includes at least one of the following types: static image type, table type, formula type, canvas type, dynamic image type, and video type.

[0091] Example devices and equipment

[0092] Figure 10 shows a block diagram of an apparatus 1000 for processing documents according to some implementations of the present disclosure. The apparatus includes: a receiving module 1010 configured to receive a user request for obtaining descriptive data of a document; a generating module 1020 configured to generate structured data based on the document; and a determining module 1030 configured to determine descriptive data of the document based on the structured data, the descriptive data including portions represented in text format and portions represented in graphical format.

[0093] According to some implementations of this disclosure, the generation module is further configured to: generate structured data based on the text data in response to determining that the document includes text data represented in text format; and generate graphic data based on the text content.

[0094] According to some implementations of this disclosure, the generation module is further configured to: extract at least one graphical representation from the graphical data in response to determining that the document includes graphical data represented in graphical format; extract a graphical description of the at least one graphical representation from the document, the graphical description being represented in text format; and update the structured data based on the graphical description.

[0095] According to some implementations of this disclosure, the generation module is further configured to: determine the position of the graphic representation in the document; and insert a graphic description at the corresponding graphic position in the text data to form structured data.

[0096] According to some implementations of this disclosure, the generation module is configured to: perform text recognition on the graphic representation to determine the graphic description; perform image recognition on the graphic representation to determine the graphic description; perform context recognition on the graphic representation to determine the graphic description; or perform speech recognition on the graphic representation to determine the graphic description.

[0097] According to some implementations of this disclosure, the generation module is configured to: determine the text portion of a document that represents a referenced graphic; and perform semantic recognition on the text portion in order to determine the image description.

[0098] According to some implementations of this disclosure, the determining module is further configured to: determine intermediate data associated with the structured data based on a predetermined model, the intermediate data including descriptive data of a document represented in text format; and, in response to determining that the intermediate data includes identification information corresponding to a graphical representation, update relevant portions of the identification information using the graphical representation to determine the descriptive data.

[0099] According to some implementations of this disclosure, a predetermined model describes the relationship between the reference structured data of the reference document and the reference intermediate data of the reference document, wherein the reference intermediate data is represented in text format.

[0100] According to some implementations of this disclosure, the determining module is further configured to: add at least one of the following to the location of the identification information: a graphical representation, or an access address for accessing the graphical representation.

[0101] According to some implementations of this disclosure, the determining module is further configured to: update relevant portions of the identifying information using a graphical representation in response to determining that the context of the identifying information in the intermediate data matches the graphical description.

[0102] According to some implementations of this disclosure, the apparatus further includes a presentation module configured to: acquire attribute information of a terminal device for performing a method; determine presentation attributes for presenting a graphical representation based on the attribute information of the terminal device, the presentation attributes including at least one of the following: format, size, position; and present the graphical representation in the description data according to the presentation attributes.

[0103] According to some implementations of this disclosure, the apparatus further includes: a calling module configured to call the apparatus in response to determining that user-specified descriptive data includes graphical data.

[0104] According to some implementations of this disclosure, the presentation module is further configured to: present a document in a first display area; and present descriptive data and interactive controls for submitting user requests in a second display area.

[0105] According to some implementations of this disclosure, the graphical representation includes at least one of the following types: static image type, table type, formula type, canvas type, dynamic image type, and video type.

[0106] Figure 11 shows a block diagram of a device 1100 capable of implementing various implementations of the present disclosure. It should be understood that the computing device 1100 shown in Figure 11 is merely exemplary and should not constitute any limitation on the functionality and scope of the implementations described herein. The computing device 1100 shown in Figure 11 can be used to implement the methods described above.

[0107] As shown in Figure 11, computing device 1100 is in the form of a general-purpose computing device. Components of computing device 1100 may include, but are not limited to, one or more processors or processing units 1110, memory 1120, storage device 1130, one or more communication units 1140, one or more input devices 1150, and one or more output devices 1160. Processing unit 1110 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 1120. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of computing device 1100.

[0108] Computing device 1100 typically includes multiple computer storage media. Such media can be any available media accessible to computing device 1100, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 1120 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 1130 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within computing device 1100.

[0109] The computing device 1100 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG11, disk drives for reading or writing from removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading or writing from removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. The memory 1120 may include a computer program product 1125 having one or more program modules configured to perform various methods or actions of various implementations of this disclosure.

[0110] The communication unit 1140 enables communication with other computing devices via a communication medium. Additionally, the components of the computing device 1100 can function as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device 1100 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or another network node.

[0111] Input device 1150 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 1160 can be one or more output devices, such as a monitor, speaker, printer, etc. Computing device 1100 can also communicate as needed with one or more external devices (not shown) via communication unit 1140. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with computing device 1100, or with any device that enables computing device 1100 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0112] According to exemplary implementations of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is provided that stores a computer program thereon, which, when executed by a processor, implements the methods described above.

[0113] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0114] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0115] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0116] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0117] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for processing documents, comprising: Receive user requests for descriptive data of a document; Structured data is generated based on the document; as well as Based on the structured data, the description data of the document is determined, the description data including a portion represented in text format and a portion represented in graphic format.

2. The method according to claim 1, wherein generating the structured data based on the document comprises: In response to determining that the document includes text data represented in the text format, the structured data is generated based on the text data; The graphic data is generated based on the text content.

3. The method according to claim 1, wherein generating the structured data based on the document comprises: In response to determining that the document includes graphic data represented in a graphic format, at least one graphic representation is extracted from the graphic data; Extract a graphical description of the at least one graphical representation from the document, the graphical description being represented in text format; as well as The structured data is updated based on the graphical description.

4. The method of claim 3, wherein generating the structured data based on the document further comprises: Determine the position of the graphic representation within the document; as well as The graphic description is inserted at the graphic position corresponding to the position in the text data to form the structured data.

5. The method of claim 3, wherein the graphical description extracted from the graphical representation includes at least one of the following: Perform text recognition on the graphic representation to determine the graphic description; Perform image recognition on the graphic representation to determine the graphic description; Perform context recognition on the graphic representation to determine the graphic description; or Speech recognition is performed on the graphic representation to determine the graphic description.

6. The method of claim 5, wherein performing the context recognition for the graphical representation to determine the graphical description comprises: Identify the text portions in the document that reference the graphic representation; as well as Semantic recognition is performed on the text portion in order to determine the image description.

7. The method of claim 1, wherein determining the description data of the document based on the structured data comprises: Based on the structured data, intermediate data associated with the structured data is determined according to a predetermined model, the intermediate data including the description data of the document in text format; as well as In response to determining that the intermediate data includes identification information corresponding to the graphical representation, the relevant portion of the identification information is updated using the graphical representation to determine the descriptive data.

8. The method of claim 7, wherein the predetermined model describes the relationship between reference structured data of the reference document and reference intermediate data of the reference document, the reference intermediate data being represented in the text format.

9. The method of claim 7, wherein updating the relevant portion of the identification information using the graphical representation includes: Add at least one of the following to the location of the identification information: the graphic representation, or an access address for accessing the graphic representation.

10. The method of claim 7, wherein updating the relevant portion of the identification information using the graphical representation further comprises: In response to determining that the context of the identification information in the intermediate data matches the graphical description, the relevant portion of the identification information is updated using the graphical representation.

11. The method of claim 1, further comprising: Obtain attribute information of the terminal device used to execute the method; Based on the attribute information of the terminal device, a presentation attribute for presenting the graphic representation is determined, and the presentation attribute includes at least one of the following: format, size, and position; as well as The graphical representation of the descriptive data is presented according to the presentation attributes.

12. The method of claim 1, further comprising: The method is executed in response to determining that the user-specified description data includes graphical data.

13. The method of claim 1, further comprising: The document is displayed in the first display area; as well as The description data and interactive controls for submitting user requests are presented in the second display area.

14. The method of claim 1, wherein the graphical representation includes at least any of the following types: static image type, table type, formula type, canvas type, dynamic image type, and video type.

15. An apparatus for processing documents, comprising: The receiving module is configured to receive user requests for obtaining descriptive data of a document; A generation module is configured to generate structured data based on the document; as well as A determination module is configured to determine the description data of the document based on the structured data, the description data including portions represented in text format and portions represented in graphic format.

16. An electronic device comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 4 when executed by the at least one processing unit.

17. A computer-readable storage medium having a computer program stored thereon, the computer program causing the processor to implement the method according to any one of claims 1 to 14 when executed by a processor.

18. A computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Knowledge graph establishment method, device and system based on textbook

    CN116090560A

  • Industrial document-oriented multi-modal information extraction method and system

    CN116796288A

  • Information display method and device, electronic equipment and storage medium

    CN117370586A

  • Flow chart generation method and device and electronic equipment

    CN117745226A

  • Generating Alternative Descriptions for Images

    US20140146053A1