Document reading comprehension support method

The document reading support method using a language model efficiently generates accurate summaries by classifying and summarizing selected document parts within a token limit, addressing the limitations of conventional methods.

JP2025079326APending Publication Date: 2025-05-21SEMICON ENERGY LAB CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024193034
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-09
Filing Date
2024-11-01
Publication Date
2025-05-21

AI Technical Summary

Technical Problem

Conventional document comprehension methods, such as displaying specification and drawings side by side, can be time-consuming and may not provide accurate understanding, especially when using transformer-based language models with memory constraints, and the quality of extracted sentences varies with the extractor's knowledge and experience.

Method used

A document reading support method using a language model that classifies documents into chapters or paragraphs, allows selection of specific parts, and generates summaries within a specified token limit, highlighting or linking to original text for accuracy.

Benefits of technology

Enables efficient and accurate document summarization by ensuring summaries are within a specified token count, maintaining relevance to the user's purpose, and providing quality control through highlighting or linking to original text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025079326000001_ABST
    Figure 2025079326000001_ABST
Patent Text Reader

Abstract

To provide a document reading comprehension support method using a language model.SOLUTION: A document reading comprehension support method includes the steps of: displaying segmented documents, accepting selection of a part of the documents, inputting the part and an instruction sentence to summarize the part to a language model, determining whether a token of the part is less than or equal to a predetermine value, and acquiring the summary of the part determined to be less than or equal to the predetermined value.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a language model, and in particular to a document comprehension support method using a generative AI model.

[0002] The above technical field is one embodiment of the present invention, and the present invention is not limited to the above technical field. Other embodiments of the present invention can include, for example, a semiconductor device, a display device, a light-emitting device, a power storage device, a memory device, an electronic device, a lighting device, an input device (e.g., a touch sensor), an input / output device (e.g., a touch panel), a driving method thereof, or a manufacturing method thereof. [Background technology]

[0003] When reading a document, it is necessary to understand the contents accurately. Also, reading a document requires an overview that is in line with the reader's purpose. However, people sometimes interpret the document by arbitrarily stringing together words, which can result in an inaccurate understanding of the document. In the case of a long document, there is a problem that it takes a long time for people to read it. In addition, drawings are often attached to documents related to patents (typically patent specifications, publications, or patent bulletins, which are called patent documents), but since the drawings and the documents explaining the drawings are separated, it is necessary to understand them while comparing them, and information on the technical content may be missing. Patent Document 1 discloses a system that displays the contents of the specification and the drawings side by side to enable efficient viewing of patent documents.

[0004] ChatGPT is an example of a large-scale interactive language model. The large-scale language model (LLM) used in ChatGPT includes GPT-3 (Generative Pre-trained Transformer 3) and GPT-4 (Generative Pre-trained Transformer 4). [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Publication No. 8-339380 Summary of the Invention [Problem to be solved by the invention]

[0006] Conventional document comprehension support methods may not provide sufficient support in some cases. For example, a document comprehension support method such as that in Patent Document 1, which only displays the contents of the specification and drawings side by side, may take a long time to read and may not allow accurate understanding.

[0007] It is possible to try to generate a document summary using a language model such as a conversational generation model, but if the language model being used is based on a transformer architecture, there is an upper limit to the number of characters that can be input depending on the equipment and memory constraints used, making it difficult to read the entire document and generate a summary. In addition, the task of extracting sentences from a document so that the number of characters is appropriate (the task of preparing extracted sentences) is laborious. Furthermore, the quality of the extracted sentences varies depending on the knowledge and experience of the extractor. Furthermore, while document comprehension support according to the user's purpose is desired, if the extracted sentences are not related to the topic, the processing is very inefficient.

[0008] The present invention has been made in view of the above problems, and an object of one embodiment of the present invention is to provide a novel document reading support method. Another embodiment of the present invention is a document reading support method using a language model, and an object of one embodiment of the present invention is to provide a document reading support method that makes it possible to obtain an appropriate prompt. Another embodiment of the present invention is a document reading support method using a language model, and an object of one embodiment of the present invention is to provide a document reading support method that makes it possible to obtain an accurate answer sentence.

[0009] The present invention does not necessarily have to solve all of these problems. Furthermore, the description of these problems does not preclude the existence of other problems of the present invention. For example, problems other than these can be extracted from the description of the specification, drawings, and claims. [Means for solving the problem]

[0010] In consideration of the above problems, a document comprehension support method is provided, which includes the steps of: displaying a classified document (referred to as a first document to distinguish it from other documents); accepting a selection of a part of the first document (referred to as a second document to distinguish it from other documents); inputting the second document and an instruction sentence to summarize the second document into a language model; determining whether the tokens of the second document are below a specified value; and obtaining a summary of the second document determined to be below the specified value.

[0011] Another aspect of the present invention is a document comprehension support method including the steps of: displaying a document divided into multiple chapters including at least chapter 1, accepting a selection of chapter 1, inputting chapter 1 and a directive to summarize chapter 1 into a language model, determining whether the tokens of chapter 1 are equal to or less than a specified value, and acquiring a summary of chapter 1 determined to be equal to or less than the specified value. Note that the document has multiple chapters, and any one of the multiple chapters is referred to as chapter 1.

[0012] Another aspect of the present invention is a document comprehension support method comprising the steps of: displaying one or more drawings and a segmented document; accepting a selection of at least one drawing; collecting sentences related to the selected drawing from the document; inputting the collected sentences and instruction sentences for summarizing the collected sentences into a language model; determining whether tokens of the collected sentences are below a specified value; and obtaining summaries of the collected sentences determined to be below the specified value.

[0013] Another aspect of the present invention is a document reading support method comprising the steps of: displaying a document that contains one or more words and is segmented; searching for words contained in the document and collecting paragraphs from the document in which the words are used; inputting the collected paragraphs and instructions for summarizing the collected paragraphs into a language model; determining whether tokens of the collected paragraphs are below a specified value; and obtaining summaries of the collected paragraphs determined to be below the specified value.

[0014] In the present invention, a word may be accompanied by a symbol or a number.

[0015] In the present invention, the language of the document may be a first language, the instructional text may be a second language, and the summary generated by the language model may be in the first language.

[0016] In the present invention, the summary may be translated from a second language into a first language.

[0017] In the present invention, it is preferable that the method further comprises the step of displaying the summary generated by the language model, and in the displayed summary, a first word that is not used in the document is highlighted.

[0018] In the present invention, it is preferable that the method further comprises a step of displaying the summary generated by the language model, and that when a second word used in the document is selected in the displayed summary, a sentence of the document containing the second word is displayed.

[0019] In the present invention, it is preferable that in the step of determining whether the number of tokens is equal to or less than a specified value, an alert is displayed if the number of tokens is greater than the specified value. Effect of the Invention

[0020] According to one aspect of the present invention, a novel document reading support method can be provided. According to another aspect of the present invention, a document reading support method using a language model that makes it possible to select an appropriate prompt can be provided. According to another aspect of the present invention, a document reading support method using a language model that makes it possible to obtain an accurate answer sentence can be provided.

[0021] The present invention does not necessarily have to have all of these effects. Furthermore, the description of these effects does not preclude the existence of other effects of the present invention. For example, effects other than these can be extracted from the description of the specification, drawings, and claims. [Brief description of the drawings]

[0022] [Figure 1] FIG. 1 is a diagram showing an example of a system configuration that enables the document reading support method of the present invention to be implemented. [Diagram 2] FIG. 2 is a diagram showing an example of the configuration of an information processing device that can implement the document reading support method of the present invention. [Diagram 3] FIG. 3 is a flowchart showing an example of the document reading assistance method of the present invention. [Figure 4] 4(A) and 4(B) are diagrams showing an example of transition of a screen of an information terminal used in the document reading support method of the present invention. [Diagram 5] FIG. 5 is a diagram showing an example of a system that makes it possible to implement the document reading assistance method of the present invention. [Figure 6] FIG. 6 is a diagram showing a specific example of the configuration of a system that can implement the document reading assistance method of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0023] The following describes the embodiments of the present invention with reference to the drawings. However, it will be easily understood by those skilled in the art that the present invention can be modified in various ways without departing from the spirit of the invention. Therefore, the present invention should not be interpreted as being limited to the contents of the description of the embodiments shown below.

[0024] In the drawings, the position, size, range, etc. of each component may not accurately represent the actual components. Therefore, the position, size, range, etc. of each component are not necessarily limited to the position, size, range, etc. disclosed in the drawings.

[0025] In this specification, the language model is based on the transformer architecture and is additionally trained to become a conversational (also called dialogue-based) model. In other words, a conversational generative model corresponds to a lower level concept of a language model. A language model is also generally called a large-scale language model.

[0026] In this specification, the words "first" and "second" are used for the convenience of understanding the technical content or to identify each component. Therefore, the words "first" and "second" do not limit the number of each component. Furthermore, the words "first" and "second" do not limit the order of each component. Furthermore, the words "first" and "second" or identifying symbols used in this specification may not match the words or identifying symbols in the claims.

[0027] In this specification, a document refers to a written expression of human intent using characters or symbols. A document may be, for example, a book, a patent document, or a paper. A document may be divided into multiple chapters or multiple paragraphs. A chapter may also be composed of multiple paragraphs.

[0028] (Embodiment) In this embodiment, a configuration example of an information processing system capable of implementing document reading support according to one embodiment of the present invention will be described with reference to Fig. 1. An information processing system capable of implementing document reading support may be called a reading support system.

[0029] <Example of information processing system configuration> The information processing system of this embodiment preferably includes a first information processing device 10, a second information processing device 40, a first information terminal 20a, a second information terminal 20b, a third information terminal 20c, and a fourth information terminal 20d, as shown in Fig. 1. The first information terminal 20a, the second information terminal 20b, the third information terminal 20c, and the fourth information terminal 20d are collectively referred to as a plurality of information terminals 20.

[0030] 1, a first information processing device 10 is connected to a plurality of information terminals 20 via a network 31. The first information processing device 10 is also connected to a second information processing device 40 via a network 30.

[0031] The first information terminal 20a to the fourth information terminal 20d are each an information terminal operated by a user of the document reading support system, and may also be called a client computer. In FIG. 1, as an example, the first information terminal 20a is shown as a desktop computer, the second information terminal 20b is shown as a notebook computer, the third information terminal 20c is shown as a smartphone, and the fourth information terminal 20d is shown as a tablet computer. Note that the fourth information terminal 20d is one in which the housing 21 having an input unit (typically a keyboard) can be separated.

[0032] Next, in this embodiment, a configuration example of a first information processing device or the like that can implement a document reading support system according to one embodiment of the present invention will be described with reference to FIG.

[0033] Configuration example of first information processing device 10 As shown in Fig. 2, a first information processing device 10 according to an embodiment of the present invention includes an input unit 110, a storage unit 120, a processing unit 130, an output unit 140, and a transmission path 150. In addition to the first information processing device 10, Fig. 2 also shows a first information terminal 20a and a second information processing device 40, and arrows indicate transmission and reception of data.

[0034] [Input section 110] The input unit 110 can receive data from outside the first information processing device 10. For example, the input unit 110 can receive data from the first information terminal 20a. The input unit 110 can also receive data from the second information processing device 40.

[0035] The input unit 110 can supply the received data to one or both of the storage unit 120 and the processing unit 130 via a transmission path 150 .

[0036] [Storage section 120] The storage unit 120 has a function of storing programs and the like executed by the processing unit 130. The storage unit 120 may also have a function of storing data (e.g., calculation results, analysis results, inference results, and the like) generated by the processing unit 130. The storage unit 120 may also have a function of storing data and the like accepted by the input unit 110.

[0037] The storage unit 120 may also have a database. In the database, document data described below can be stored and managed. The first information processing device 10 may have a database separate from the storage unit 120. Specifically, the first information processing device 10 may have a function of retrieving data from a database that exists outside the storage unit 120, outside the first information processing device 10, or outside the information processing system. The first information processing device 10 may also have a function of retrieving data from both a database inside the first information processing device 10, that is, a database owned by the first information processing device 10 itself, and a database that exists outside.

[0038] The storage unit 120 has at least one of a volatile memory and a non-volatile memory. Examples of the volatile memory include a dynamic random access memory (DRAM) and a static random access memory (SRAM). Examples of the non-volatile memory include a resistive random access memory (ReRAM), a phase change random access memory (PRAM), a ferroelectric random access memory (FeRAM), a magnetoresistive random access memory (MRAM), and a flash memory. The storage unit 120 can be configured with a SiLSI (a circuit using silicon transistors).

[0039] The storage unit 120 may also include at least one of NOSRAM (registered trademark) and DOSRAM (registered trademark). The storage unit 120 may also include a recording media drive. Examples of the recording media drive include a hard disk drive (HDD) and a solid state drive (SSD).

[0040] NOSRAM is an abbreviation of "Nonvolatile Oxide Semiconductor Random Access Memory (RAM)". NOSRAM is a memory that uses transistors (also called OS transistors) that use metal oxides in the channel formation region as memory cells that are two-transistor (2T) type or three-transistor (3T) type gain cells. The current that flows between the source and drain in the off state, that is, the leakage current, is extremely small. NOSRAM can be used as a nonvolatile memory by retaining a charge according to data in the memory cell using its extremely small leakage current characteristic. In particular, NOSRAM can read the stored data without destroying it (non-destructive read), so it is suitable for calculation processing that repeats a large number of data read operations only. NOSRAM can be configured by stacking memory cells. Since stacking increases the data capacity, it can be used as a large-scale cache memory, main memory, or storage memory to achieve high performance.

[0041] DOSRAM is an abbreviation for "Dynamic Oxide Semiconductor RAM" and refers to RAM with one transistor (1T) one capacitance (1C) type memory cells. DOSRAM is a DRAM formed using OS transistors, and is a memory that temporarily stores information sent from the outside. DOSRAM is a memory that takes advantage of the small off-current of OS transistors.

[0042] In this specification and the like, a metal oxide is an oxide of a metal in a broad sense. Metal oxides are classified into oxide insulators, oxide conductors (including transparent oxide conductors), oxide semiconductors (also referred to as oxide semiconductors or simply OS), and the like. For example, when a metal oxide is used for a semiconductor layer of a transistor, the metal oxide may be referred to as an oxide semiconductor.

[0043] The metal oxide in the channel formation region preferably contains indium (In), that is, indium oxide. An OS transistor using a metal oxide containing indium for a channel formation region has high carrier mobility (electron mobility). The metal oxide in the channel formation region is preferably an oxide semiconductor containing an element M described below instead of or in addition to In. The element M is preferably at least one of aluminum (Al), gallium (Ga), and tin (Sn). Other elements applicable to the element M include boron (B), silicon (Si), titanium (Ti), iron (Fe), nickel (Ni), germanium (Ge), yttrium (Y), zirconium (Zr), molybdenum (Mo), lanthanum (La), cerium (Ce), neodymium (Nd), hafnium (Hf), tantalum (Ta), and tungsten (W). In the metal oxide, a plurality of elements listed as the element M may be combined. The element M is an element having a high bond energy with oxygen, and the bond energy with oxygen is higher than the bond energy between oxygen and indium. The metal oxide in the channel formation region is preferably a metal oxide containing zinc (Zn) instead of or in addition to In. Metal oxides containing zinc may be easily crystallized.

[0044] The metal oxide contained in the channel formation region is not limited to the above-mentioned elements, typically, metal oxides containing indium. The metal oxide contained in the channel formation region may be, for example, metal oxides containing zinc but not indium, such as zinc tin oxide and gallium tin oxide, metal oxides containing gallium, and metal oxides containing tin.

[0045] [Processing section 130] The processing unit 130 has a function of performing processes such as calculation, analysis, and inference using data supplied from one or both of the input unit 110 and the storage unit 120. The processing unit 130 can supply generated data (e.g., calculation results, analysis results, inference results) to one or both of the storage unit 120 and the output unit 140.

[0046] The processing unit 130 has a function of acquiring data from the storage unit 120. The processing unit 130 may also have a function of recording or registering data in the storage unit 120.

[0047] The processing unit 130 may have, for example, an arithmetic circuit. The processing unit 130 may have, for example, a central processing unit (CPU). The CPU has an arithmetic unit, a primary cache memory, a secondary cache memory, and the like. The processing unit 130 may also have a GPU (Graphics Processing Unit). The GPU has an arithmetic unit, a primary cache memory, a secondary cache memory, and the like. The CPU or GPU may have one or both of an OS transistor and a transistor having silicon in a channel formation region (Si transistor).

[0048] The processing unit 130 may have a register and a main memory in addition to the CPU or the GPU. The register and the main memory may be said to be owned by the CPU. The register and the main memory may be said to be owned by the GPU. The main memory is capable of sending and receiving data to and from a secondary cache or the like. The main memory has at least one of a volatile memory such as a RAM (Random Access Memory) and a non-volatile memory such as a ROM (Read Only Memory). The main memory may also have at least one of the above-mentioned NOSRAM and DOSRAM. The main memory may have one or both of an OS transistor and a transistor having silicon in a channel formation region (Si transistor).

[0049] The RAM may be, for example, a DRAM or an SRAM, and a memory space is virtually allocated and used as a working space for the processing unit 130. The operating system, application programs, program modules, program data, lookup tables, and the like stored in the storage unit 120 are loaded into the RAM for execution. The data, programs, and program modules loaded into the RAM are each operated by accessing the processing unit 130.

[0050] ROM can store BIOS (Basic Input / Output System) and firmware that do not require rewriting. Examples of ROM include mask ROM, OTPROM (One Time Programmable Read Only Memory), EPROM (Erasable Programmable Read Only Memory), etc. Examples of EPROM include UV-EPROM (Ultra-Violet Erasable Programmable Read Only Memory), which allows the erasure of stored data by exposure to ultraviolet light, EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory, etc.

[0051] The processing unit 130 may have a microprocessor such as a DSP (Digital Signal Processor). Since the DSP is specialized for digital signal processing, it is preferable to mount it to control peripheral circuits such as a CPU. The microprocessor may be realized by a PLD (Programmable Logic Device) that operates on hardware such as an FPGA (Field Programmable Gate Array) or an FPAA (Field Programmable Analog Array). The processing unit 130 may also have a quantum processor. The processing unit 130 can perform various data processing and program control by interpreting and executing instructions from various programs using the processor. Programs that can be executed by the processor are stored in at least one of the memory area of ​​the processor and the storage unit 120.

[0052] The processing unit 130 preferably has an OS transistor. Since the off-current of the OS transistor is extremely small, by using the OS transistor as a switch for holding the charge (data) flowing into the capacitive element, the data can be held for a long period of time. By using this characteristic in at least one of the register and the cache memory of the processing unit 130, the processing unit can be operated only when necessary, and in other cases, the processing unit can be turned off by saving the information of the immediately previous processing in the capacitive element. In other words, the OS transistor enables normally-off computing, and the power consumption of the information processing system can be reduced.

[0053] By using a CPU or the like capable of high speed operation in the processing unit 130, AI can be used for some of the processing executed by the first information processing device 10. The first information processing device 10 may be provided with an artificial neural network (ANN, hereinafter also simply referred to as a neural network) to enable processing using AI. Since a neural network is realized by a circuit (hardware) or a program (software), the first information processing device 10 may have the above circuit or the above program.

[0054] In this specification, the term "neural network" refers to a general model that imitates the neural circuit network of a living organism, determines the connection strength between neurons through learning, and has problem-solving capabilities. A neural network has an input layer where information is input, an output layer where information is output, and an intermediate layer (hidden layer) between the input layer and the output layer, and optimizes the weights for the input data to obtain a correct output result.

[0055] In this specification and the like, when discussing neural networks, determining neurons and their weight coefficients from existing information may be referred to as "learning."

[0056] In this specification and the like, constructing a neural network using weighting coefficients obtained by learning and deriving a new conclusion therefrom may be referred to as "inference."

[0057] [Output section 140] The output unit 140 can output the calculation result in the processing unit 130 to the outside of the first information processing device 10. For example, the output unit 140 can transmit data to the second information processing device 40. In addition, the output unit 140 can transmit data to a plurality of information terminals 20.

[0058] [Transmission Line 150] The transmission path 150 has a function of transmitting data. Data can be transmitted and received between the input unit 110, the storage unit 120, the processing unit 130, and the output unit 140 via the transmission path 150.

[0059] Configuration example of second information processing device 40 The second information processing device 40 can process the received data and transmit the processing result. For example, the second information processing device 40 can perform processing such as calculation using the data received from the first information processing device 10. In addition, the second information processing device 40 can transmit the processing result to the first information processing device 10. This can reduce the calculation load on the first information processing device 10.

[0060] The second information processing device 40 can perform processing using a natural language processing model using AI. For example, the second information processing device 40 can perform processing using a natural language processing model using AI, such as BERT (Bidirectional Encoder Representations from Transformers) and T5 (Text-to-Text Transfer Transformer).

[0061] In addition, the second information processing device 40 can perform processing using a model (such as a document generation model or a dialogue model) that uses a large-scale language model. It is preferable that the generation of the summary sentence in step S151 described later is processed using a model that uses a large-scale language model. For example, the processing can be performed using a large-scale language model such as GPT-3, GPT-3.5, GPT-4, LaMDA (Language Model for Dialogue Applications), PaLM (Pathways Language Model), or Llama2.

[0062] Furthermore, the second information processing device 40 can execute processing using a general-purpose language processing model capable of performing various natural language processing tasks.

[0063] In addition, the document reading support service provider does not necessarily have to own the second information processing device 40. For example, the service provider can use part of the service provided by another business operator using the second information processing device 40.

[0064] Configuration Example of First Information Terminal 20a The first information terminal 20a can receive data input by a user, and can provide the user with data output by the information processing system according to an aspect of the present invention.

[0065] Furthermore, the first information terminal 20a can transmit data received from a user to the first information processing device 10. Furthermore, the first information terminal 20a can provide data received from the first information processing device 10 to the user.

[0066] Furthermore, the first information terminal 20a can transmit data generated based on data received from a user to the first information processing device 10. Furthermore, the first information terminal 20a can provide data generated based on data received from the first information processing device 10 to the user.

[0067] For example, dedicated application software or a web browser is installed on the first information terminal 20a. A user can access the first information processing device 10 via either of them. This allows the user to enjoy a service using the information processing system according to one embodiment of the present invention, for example, using a computer with a lower processing capacity than the first information processing device 10.

[0068] The first information terminal 20a can also be called a client computer, etc. Furthermore, each of the multiple information terminals 20 is an information terminal operated by a user.

[0069] "Network 30" The network 30 connects the first information processing device 10 and the second information processing device 40. This allows input data and processed data to be transmitted and received between the two devices. In addition, the load related to information processing can be distributed.

[0070] In this embodiment, the network 30 is mainly a computer network that is larger than the network 31. For example, a global network can be used for the network 30. Specifically, the Internet, which is the foundation of the World Wide Web (WWW), can be used.

[0071] "Network 31" The network 31 connects a plurality of information terminals 20 and the first information processing device 10. This enables data to be transmitted and received between the two. Also, the load related to information processing can be distributed. Also, a service provider can provide a user with a service using an information processing method according to one aspect of the present invention via the network 31, for example.

[0072] For example, a local network can be used for the network 31. Also, an intranet or an extranet can be used for the network 31. Also, a personal area network (PAN), a local area network (LAN), a campus area network (CAN), a metropolitan area network (MAN), a wide area network (WAN), a global area network (GAN), etc. can be used for the network 31.

[0073] When performing wireless communication, communication standards such as the fourth generation mobile communication system (4G), the fifth generation mobile communication system (5G), and the sixth generation mobile communication system (6G), or specifications standardized by IEEE such as Wi-Fi (registered trademark) and Bluetooth (registered trademark), can be used as communication protocols or communication technologies.

[0074] When a provider of a service using the document reading support method according to one embodiment of the present invention and a user who enjoys the service belong to the same organization such as a company, it is preferable that data transmission between the multiple information terminals 20 and the first information processing device 10 is performed, for example, using a network 31 constructed within the organization. This makes it possible to transmit and receive data between the multiple information terminals 20 and the second information processing device 40 more securely than when data is transmitted via the Internet. It is also possible to prevent confidential information within the organization from leaking to the outside. Alternatively, data transmission and reception between the multiple information terminals 20 and the first information processing device 10 may be performed using a network 30 (for example, the Internet).

[0075] A more specific configuration example of the document reading support system in this embodiment will be described with reference to FIG. 6. The server computer 201 has an input unit, a storage unit, a processing unit, an output unit, and a transmission path, and is also called a backend. The processing unit of the server computer 201 includes a text data supply unit 203 and a drawing data supply unit 204 shown in FIG. 6. The text data supply unit 203 has a function of structuring the text data of the document to be read and a function of supplying the structured data to the display data processing unit 205. Structuring includes division by paragraph numbers of a sentence, division by chapters including multiple paragraph numbers, and the like. For example, in the case of a patent document, division into chapters such as in the first embodiment and the second embodiment can be mentioned. Furthermore, since the first embodiment includes multiple paragraph numbers, division by chapters including these paragraph numbers can also be used. The text data of the divided sentences can be stored in the storage unit. The drawing data supply unit 204 supplies image data of the drawings included in the document to the display data processing unit 205. For example, data in which image data is linked to each drawing number can be used as the image data. The display data processing unit 205 constructs relationship information between the drawing and the text so that the user can display the drawing and the corresponding text side by side on the screen. Specifically, the display data processing unit 205 performs processing to embed HTML tags in the text data and the image data of the drawing. For example, the image data of the drawing A can be displayed with a link to the paragraph X that mentions the drawing embedded as a tag, and the paragraph X that mentions the drawing A can be displayed with a link to the drawing A embedded as a tag. The display data processing unit 205 can recognize the divided chapters and form buttons for instructing the language model for the chapters. Furthermore, the display data processing unit 205 can obtain instruction sentences for the language model from the prompt data supply unit 206. The data created in the display data processing unit 205 is preferably sent to the individual terminal 301 operated by the user as display data in HTML format for display.

[0076] The individual terminal 301 has an input unit, a storage unit, a processing unit, an output unit, and a transmission path, and is also called a front end, and corresponds to a terminal operated by a user. The individual terminal 301 processes the display data received from the server computer 201 by the display processing unit 302 to display or operate text, or to display or operate drawings, and displays the data on the display device 304. The display device 304 includes a panel unit that enables display. The display processing unit 302 is, for example, a browser. The prompt processing unit 303 allows the user to specify a chapter as an object to be summarized via the display device 304, or to edit a prompt. The communication processing unit 305 transmits and receives a prompt to and from the language model 404, and feeds back the received data to the display processing unit 302, so that an answer from the language model can be displayed on the display device 304. The answer received by the individual terminal 301 may be transmitted to the server computer 201, where further processing may be performed. The language model 404 may be installed in a cloud environment so as to be available via the Internet or a communication line, or may be installed in the server computer 201. Although not shown, the language model 404 can also send and receive prompts to and from the individual terminal 301 via the server computer 201 .

[0077] The document reading support system in this embodiment may have a function that allows the user to perform a text search on the text displayed on the display device 304. In this case, the search processing unit 207 provided in the server computer 201 receives a text search command from the individual terminal 301 and reconstructs the text data. The reconstructed text data is sent to the display processing unit 302 of the individual terminal 301.

[0078] <Example of method> In this embodiment, a method for supporting document reading comprehension, which is one aspect of the present invention, will be described. Fig. 3 shows an example of a flowchart relating to the method for supporting document reading comprehension, which is one aspect of the present invention. Fig. 4 shows an example of the transition of the screen of the first information terminal 20a in the document reading comprehension support method. Fig. 5 explains information processing performed by the first information terminal 20a, the first information processing device 10, and the second information processing device 40 with reference to the step numbers of the flowchart in Fig. 3.

[0079] The document reading support method according to one embodiment of the present invention is started, and as shown in step S101 of FIG. 3, the user selects a document for which he / she wishes to receive document reading support. The document in the document reading support method according to one embodiment of the present invention is preferably a patent document. The input related to step S101 is performed by the first information terminal 20a, and as shown in FIG. 4(A), the screen 50 of the first information terminal 20a preferably has a text box 61 as a display corresponding to step S101. Also, the main information processing related to step S101 is preferably executed by the first information processing device 10 as shown in FIG. 5. The main information processing includes the processing in which the input unit 110 accepts data from the first information terminal 20a, the data is supplied to one or both of the storage unit 120 and the processing unit 130 via the transmission path 150, and at least the processing unit 130 selects a document based on the data. The document can be selected from a database in the storage unit 120 or a database outside the first information processing device 10.

[0080] Note that the screen 50 may be provided with a list button 61a for displaying a list of multiple documents instead of or in addition to the text box 61. FIG. 4(A) shows the screen 50 with the list button 61a provided in addition to the text box 61. In this case, the user can select a document from the list. The list can be stored in a database in the storage unit 120 or in a database external to the first information processing device 10. The document selected from the list is transmitted to the first information terminal 20a via the output unit 140.

[0081] Moreover, in the screen 50, instead of or in addition to the text box 61, a search button 61b for executing a search for a plurality of documents may be provided. FIG. 4(A) shows the screen 50 provided with the search button 61b in addition to the text box 61. When the search button 61b is selected, a box for inputting a search keyword is displayed, or a box for inputting a search keyword may be displayed on a new screen different from the screen 50. The user can select a document using the search results. Information processing related to the search can be executed by the processing unit 130 of the first information processing device 10. In addition, in order to enable the search, the data related to the document may include at least text data. The text data can be acquired from OCR (Optical Character Recognition) of the document. Furthermore, a search can be performed using an identification number or the like associated with the document. When the document is a patent document, the application number, publication number, registration number, or other number can be used as the identification number. Other numbers include a management number of a company. The text information and the identification number can be stored using a database of the storage unit 120 or a database outside the first information processing device 10.

[0082] Here, the language of the document will be described. The document is written in English, Japanese, and other languages. The language of the document is not limited in any way, but in the natural language processing model using AI of the second information processing device 40, the language of the document is preferably English. In this specification, the languages ​​are described as a first language and a second language using ordinal numbers to distinguish them from each other. In addition to the document written in the first language, the document written in the second language can be stored in a database held by the storage unit 120 or a database external to the first information processing device 10.

[0083] A user's desired language can be registered as a setting of the information processing system. When the user's desired language is a first language and a second language is used in the natural language processing model using AI of the second information processing device 40, a translation function may be provided in the information processing system. The translation function can be executed by the processing unit 130, by processing using AI of the first information processing device 10, or by processing using a natural language processing model using AI of the second information processing device 40. Such a translation function can improve the convenience of the information processing system.

[0084] Next, as shown in step S111 of FIG. 3, the document selected by the user is displayed in a categorized manner. Note that the document may contain categorized information in advance, and the document may be displayed using the categorized information. For example, if the document is a patent document, paragraph numbers, chapters (corresponding to each item such as the effects of the invention), etc. may be used as the categorized information. It is preferable that the main information processing related to step S111 is executed by the first information processing device 10 as shown in FIG. 5. As the main information processing, the categorized information is transmitted to the first information terminal 20a via the output unit 140 together with the data of the document previously selected. The result of step S111 is displayed on the first information terminal 20a. As shown in FIG. 4(A), the screen 50 displays the document in a categorized manner into chapters 51 and sentences 52 as a display corresponding to step S111. If the document is a patent document, the chapter 51 displays each heading such as the effects of the invention, and the sentence 52 displays a sentence including a paragraph number. Furthermore, if the document contains drawings, the screen 50 displays the drawings 53 alongside the chapters 51 and sentences 52. On the screen 50, the chapter 51, the sentence 52, and the drawing 53 are displayed in this order from the left, but the layout is not limited to this. Depending on the user's settings, it is also possible to display the drawing 53, the chapter 51, and the sentence 52 in this order from the left.

[0085] The chapter 51 and the sentence 52 are linked to each other using the storage unit 120 or the processing unit 130. Furthermore, the sentence 52 and the drawing 53 are linked to each other using the storage unit 120 or the processing unit 130. Of course, the drawing 53 and the chapter 51 may be linked to each other using the storage unit 120 or the processing unit 130. By using the linked information, when the chapter 51 is selected on the screen 50, the linked sentence 52 can be displayed. Furthermore, when the drawing 53 is selected on the screen 50, the linked sentence 52 can be displayed. In the sentence 52, the figure number can be highlighted, and when the figure number is selected, the linked drawing 53 is displayed. By using the chapter 51, the sentence 52, and the drawing 53 arranged on the screen 50 in this way, document reading support can be enjoyed.

[0086] Next, while checking the screen 50, the user selects a portion of the document as shown in step S121 of FIG. 3. The selected portion can be used as the extracted sentence. Specifically, the user selects a specific chapter 51 as the extracted sentence. Or the user selects a specific sentence 52 as the extracted sentence. Or the user selects a specific drawing 53, and selects sentence 52 describing the selected drawing 53 as the extracted sentence. Such an extraction operation is preferable because it can be performed in a short time and there is no variation in quality. In addition, it is preferable that the main information processing related to step S121 is executed by the first information processing device 10 as shown in FIG. 5. The selected portion of the document is included in the prompt.

[0087] Here, the tokens of the extracted sentence will be explained. The tokens of a specific chapter 51 are clearly fewer than the tokens of the entire document and often satisfy the specified value or less, so they are preferable as an extracted sentence. The tokens of a specific sentence 52 are clearly fewer than the tokens of the entire document and often satisfy the specified value or less, so they are preferable as an extracted sentence. The tokens of sentence 52 describing a selected drawing 53 are also clearly fewer than the tokens of the entire document and often satisfy the specified value or less, so they are preferable as an extracted sentence. Such extracted sentences are in line with the user's purpose, and processing using the extracted sentences is efficient.

[0088] With regard to step S121, input for selection is performed on the first information terminal 20a. The screen 50 of the first information terminal 20a has selection buttons 62 as a display corresponding to step S121. In the screen 50 of Fig. 4(A), the selection buttons 62 are provided for each chapter and are arranged adjacent to the chapters 51. Although the arrangement of the selection buttons 62 is not limited to this, such an arrangement of the selection buttons 62 is preferable because it makes it easy to operate to select a specific chapter 51.

[0089] <Variation 1> A modified example of step S121 in FIG. 3 will be described. In step S121, after selecting one drawing from drawings 53, there is a method of collecting paragraphs in which the selected drawing number is described, and making this the extracted sentence. That is, sentences 52 related to the drawing can be collected from a document and made the extracted sentence. If there are multiple collected paragraphs, the sentences 52 may be collected in the order in which they appear in the document. Also, the sentences 52 may be collected with a priority according to the frequency of appearance of the selected drawing number. For example, when a drawing A2 is selected, a paragraph in which the drawing A2 is used can be selected as a part of the document. In this case, if there is a paragraph in which the drawing A2 is used many times according to the purpose of the user, it is preferable to collect the paragraph so that the priority of the paragraph is high. Instead of sentence 52, a chapter (for example, embodiment 1) including a paragraph using the drawing number may be selected as the extracted sentence. According to modified example 1, a summary sentence corresponding to the selected drawing 53 can be obtained in step S161 in FIG. 3. In this case, it is preferable to use a prompt such as instruction sentence 2 described later. By using this method, it is possible to efficiently know what the document explains about the specified drawing.

[0090] <Variation 2> A modified example of step S121 in FIG. 3 will be described. In step S121, there is a method of searching for a word used in a document, and then selecting a paragraph in which the word is used. For example, a text search can be performed using the character string "third insulating film" to collect paragraphs in which the third insulating film is used. In addition, in patent documents, a code such as A or a number such as 100 is often attached to the third insulating film. Instead of the step of searching for a word, a search using the code or number may be performed. The collected paragraph can be selected as a part of the document, and a summary sentence regarding the third insulating film can be obtained in step S161 in FIG. 3. Instead of a paragraph, a chapter (for example, embodiment 1) containing a paragraph using the word may be selected as an extracted sentence. Since insulating films exhibit various functions, the function of the insulating film of interest can be accurately understood by checking the summary sentence in step S161. In this case, it is recommended to use a prompt such as instruction sentence 2 described later. This usage makes it possible to efficiently know what the document explains about the specified word.

[0091] Next, as shown in step S131 of FIG. 3, the user inputs the extracted sentence (a part of the selected sentence) and the instruction sentence (an instruction sentence for summarizing the part) to the language model. In step S131, the input is performed at the first information terminal 20a, and the instruction sentence is received by the second information processing device 40 as shown in FIG. 5. In step S131, the instruction sentence may be input to the second information processing device 40 via the first information processing device 10. When inputting, it is preferable to switch from the screen 50 of FIG. 4(A) to the screen 50a of FIG. 4(B), and the screen 50a has a text box 63 as a display corresponding to step S131. The user inputs the instruction sentence, for example, "summarize the selected part of the document", into the text box 63. It is preferable to use a fixed phrase as the instruction sentence. For example, it is preferable that the user selects the instruction content from a list or the like, and a fixed phrase corresponding to the instruction content is input into the text box 63. When the instruction sentence is input as a fixed phrase, it is easy to infer the token. In addition, instructions can be provided that are inferred from the usage history of the information processing system. By using standard phrases or inferred phrases as instructions, the instructions can be of a certain quality even if the user has little knowledge or experience.

[0092] The user can edit the text box 63. It is preferable to edit it depending on how the part of the document is selected. For example, when chapter 51 is selected, instruction 1 "a general summary of the selected range" may be entered, and when sentence 52 describing drawing 53 is selected, instruction 2 "a summary of the description of drawing 53 in the selected range" may be entered in addition to instruction 1. Since the summary sentences as the answer sentences may differ between instruction 1 and instruction 2, the user may consider the instruction in order to obtain a summary sentence suitable for document comprehension support.

[0093] Here, the language of the instruction sentence will be explained. In the natural language processing model using AI of the second information processing device 40, it is preferable to use English prompts. Therefore, it is preferable to write the instruction sentence in English. It is possible to provide a function for translating the instruction sentence, but since the instruction sentence is often a short sentence and a fixed phrase can also be used for the instruction sentence, the convenience of the system is not impaired even if the above-mentioned translation function is not provided. It is preferable that the document and the instruction sentence are in a common language, but the language of the document and the language of the instruction sentence may be different.

[0094] Thereafter, the user selects the execute button 64 on the screen 50a in Fig. 4(B) The main information processing related to this step up to the selection of the execute button 64 is preferably executed by the second information processing device 40 as shown in Fig. 5 .

[0095] Next, as shown in step S141 of Fig. 3, it is determined whether the tokens of a part of the selected document are equal to or less than a specified value. Also, in step S141, it may be determined whether the tokens of a part of the selected sentence and a prompt including an instruction sentence are equal to or less than a specified value. It is preferable that the main information processing related to step S141 is executed by the second information processing device 40 as shown in Fig. 5 without user input. Therefore, a display corresponding to step S141 is not necessary on the screen 50.

[0096] If the number of tokens is greater than the specified value (if No in the figure), an alert may be displayed as shown in step S142 in FIG. 3. The main information processing related to step S142 is preferably executed by the second information processing device 40 as shown in FIG. 5. The execution result may be displayed as an alert on the screen 50 of the first information terminal 20a. The alert display may include tokens that are the specified value in addition to the tokens of the selected document. For example, the content "The maximum context length of this system is 8000 tokens. However, the extracted sentence has become 12500 tokens. Please reduce the length of the extracted sentence." may be displayed on the screen 50. An alert containing specific content is preferable because it makes it easier for the user to understand the error. In this case, the process returns to step S121 as shown in FIG. 3, and the user can select a part of the document again.

[0097] If the token is equal to or smaller than the specified value (YES in the figure), a summary sentence is generated as an answer sentence using a language model or the like, as shown in step S151 of Figure 3. Specifically, based on a prompt having an extracted sentence and an instruction sentence, the language model or the like generates a summary sentence as an answer sentence. Regarding step S151, it is preferable that main information processing is executed by the second information processing device 40.

[0098] Thereafter, as shown in step S161 of Fig. 3, the selected part of the summary is acquired. With regard to step S161, it is preferable that the summary is displayed on the first information terminal 20a. The screen 50a of Fig. 4(B) has a text box 68 for displaying the summary as a display corresponding to step S161. Furthermore, if it takes some time until the summary is displayed, the screen 50a of Fig. 4(B) may display the time. The time may be a predicted waiting time, or may be the time elapsed since the execute button 64 was pressed.

[0099] Here, the language of the summary sentence will be explained. In the natural language processing model using the AI ​​of the second information processing device 40, it is preferable to use an English prompt. Therefore, it is preferable that the summary sentence, which is the answer sentence, is also written in English. For example, if the user's desired language is not English, it is preferable to provide a translation function to the document reading support system. Specifically, when the summary sentence is in English, a translation function for translating it into Japanese may be added to the document reading support system. Such a translation function can improve the convenience of the system. In addition, it is preferable to use GPT-3 instead of GPT-4 as the language model for executing the translation function. This is because translation of the summary sentence does not require a huge amount of information processing. In the document reading support system, which is one aspect of the present invention, it is preferable that the document, the instruction sentence, and the answer sentence are in a common language, but the document may be in a first language, the instruction sentence in a second language, and the answer sentence in the first language. The second language corresponds to a translation of the first language.

[0100] The document reading assistance method according to one aspect of the present invention can end after step S161.

[0101] <Additional feature 1> An additional function 1 for displaying a summary will be described. Words included in the summary displayed in the text box 68 can be highlighted. Words that are not used in the document can be highlighted. The summary generated by the language model may contain hallucinations and may be of low quality. Hallucinations are often words that are not used in the document, so highlighting the words not used in the document in the summary makes it easier for the user to check the quality of the summary.

[0102] <Additional function 2> Additional function 2 for displaying the summary will now be described. Words contained in the summary displayed in text box 68 can be linked to words used in the document. Furthermore, by selecting a linked word, chapter 51 or sentence 52 containing that word can be displayed on screen 50. By comparing the summary with the contents displayed in chapter 51 or sentence 52, the user can check the quality of the summary.

[0103] The document reading support method having the above-mentioned steps can provide sufficient reading support. Furthermore, in the document reading support method, an extracted sentence can be obtained by selecting a chapter, sentence, or figure displayed side by side on the screen 50. Therefore, even a user with no knowledge or experience can obtain an extracted sentence with little variation in quality in a short time. Furthermore, since an extracted sentence according to the user's topic or purpose can be obtained, the quality of the summary sentence is improved. Moreover, by making the instruction sentence a language suitable for the language model and providing a translation function into the language desired by the user, an accurate summary sentence can be obtained.

[0104] According to this embodiment, a novel document reading support method can be provided. According to this embodiment, a document reading support method using a language model that makes it possible to select an appropriate prompt can be provided. According to this embodiment, a document reading support method using a language model that makes it possible to obtain an accurate answer sentence can be provided. [Explanation of symbols]

[0105] 10 First information processing device 20 Information terminal 20a First information terminal 20b Second information terminal 20c Third information terminal 20d 4th information terminal 21 Case 30 Network 31 Network 40 Second information processing device 50 screens 50a screen Chapter 51 52 sentences 53 Drawings 61 Text Box 61a List button 61b Search button 62 Selection button 63 Text Box 64 Run button 68 Text Box 110 Input section 120 Storage section 130 Processing section 140 Output section 150 Transmission Line 201 Server Computer 203 Text Data Supply Unit 204 Drawing Data Supply Department 205 Display data processing unit 206 Prompt Data Supply Department 207 Search processing unit 301 Individual terminal 302 Display processing unit 303 Prompt Processing Unit 304 Display devices, Display devices 305 Communication Processing Unit 404 Language Model

Claims

1. displaying the segmented first document; accepting a selection of a second document that is a portion of the first document; inputting the second document and a directive to summarize the second document into a language model; determining whether the tokens of the second document are less than or equal to a predetermined value; obtaining a summary of the second document determined to be equal to or less than the specified value.

2. displaying a document divided into a plurality of chapters including at least a first chapter; accepting a selection of the first chapter; inputting the first chapter and a directive to summarize the first chapter into a language model; determining whether the tokens of the first chapter are equal to or smaller than a predetermined value; and acquiring a summary of the first chapter which is determined to be equal to or smaller than the specified value.

3. displaying one or more drawings and classified documents; accepting a selection of at least one of said drawings; collecting sentences from the document relating to the selected drawings; inputting the collected sentences and instructions for summarizing the collected sentences into a language model; determining whether the collected sentence tokens are equal to or smaller than a specified value; and acquiring summaries of the collected sentences that are determined to be equal to or smaller than the specified value.

4. displaying a document that includes the one or more words and that has been segmented; searching for words contained in the document and collecting paragraphs in which the words are used from the document; inputting the collected paragraphs and instructions for summarizing the collected paragraphs into a language model; determining whether the collected paragraph tokens are equal to or less than a specified value; and obtaining summaries of the collected paragraphs determined to be equal to or smaller than the specified value.

5. 5. The document reading support method according to claim 4, wherein the words are accompanied by symbols or numbers.

6. In any one of claims 1 to 5, A document reading support method, wherein the language of the document is a first language, the instructional sentence is a second language, and the summary generated by the language model is in the first language.

7. 7. The document reading support method according to claim 6, wherein the summary is translated from the second language to the first language.

8. In any one of claims 1 to 5, displaying the summary generated by the language model; In the displayed summary, a first word not used in the document is highlighted.

9. In any one of claims 1 to 5, displaying the summary generated by the language model; A document reading assistance method, wherein when a second word used in the document is selected in the displayed summary, a sentence of the document including the second word is displayed.

10. In any one of claims 1 to 5, In the step of determining whether the number of tokens is equal to or less than the specified value, an alert is displayed if the number of tokens is greater than the specified value.

Citation Information

Patent Citations

  • Information retrieval processing method

    JP1996339380A