Search support method
The search support method addresses the burden of checking multiple gazettes by using a language model to cluster and prioritize search results, ensuring efficient content verification and reducing the risk of missing important documents.
Patent Information
- Application Number
- JP2024210128
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-08
- Filing Date
- 2024-12-03
- Publication Date
- 2025-06-19
AI Technical Summary
Users face a burden when the number of retrieved gazettes increases, as it takes time to check their contents, and there is a risk of missing important gazettes by changing search terms.
A search support method that utilizes a language model to receive patent documents and a first citation group, extract summaries, and cluster similarities to divide the citation group into subsets, allowing for the deletion of unnecessary subsets based on labels and differences.
This method assists users in efficiently verifying the content of retrieved gazettes, reducing the burden of checking multiple documents and minimizing the risk of missing important information.
Smart Images

Figure 2025092458000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a language model, particularly a search assistance method using a generative AI model.
[0002] The above technical field is one aspect of the present invention, and the present invention is not limited to the above technical field. Another aspect of the present invention can include, for example, a semiconductor device, a display device, a light-emitting device, a power storage device, a storage device, an electronic device, a lighting device, an input device (e.g., a touch sensor), an input / output device (e.g., a touch panel), their driving methods, or their manufacturing methods.
Background Art
[0003] ChatGPT can be cited as a service using generative AI (Artificial Intelligence). Large Language Models (LLMs) such as GPT-3 (Generative Pre-trained Transformer 3) or GPT-4 (Generative Pre-trained Transformer 4, registered trademark) are used in the large-scale language model utilized by ChatGPT.
[0004] As the performance of generative AI improves, the scope of services using generative AI is expanding. For example, in Patent Document 1, a document search assistance method is proposed that includes a display processing unit that displays an interactive guidance screen to the user according to the user's operation history when searching for gazettes from a patent document database.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] Even if a user can find accurate search terms, when the number of retrieved gazettes increases, it takes time to check the contents, which becomes a burden on the user. Patent Document 1 describes this as a disadvantage when the number of retrieved gazettes increases, and further narrows down the search after changing the search terms. However, there is a risk that important gazettes may be missed from the search results by changing the search terms.
[0007] The present invention has been made in view of the above problems, and an aspect of the present invention aims to provide a novel search support method as one of the problems. Another aspect of the present invention aims to provide a search support method for assisting in checking the contents of gazettes retrieved by a search.
[0008] The present invention does not necessarily need to solve all of these problems. Also, the description of these problems does not prevent the existence of other problems of the present invention. For example, it is possible to extract other problems from the description regarding the specification, drawings, and claims.
Means for Solving the Problems
[0009] In view of the above problems, an aspect of the present invention includes steps of receiving patent documents and a first citation group, obtaining a first summary extracted from the patent documents, and obtaining a plurality of second summaries extracted from each document belonging to the first citation group, inputting a first instruction sentence for outputting similarities with the first summary to a language model for each of the plurality of second summaries, clustering the similarities to divide the first citation group into two or more second citation groups and obtaining a label for each second citation group, and deleting at least one second citation group from the two or more second citation groups based on the label, which is a search support method.
[0010] Another aspect of the present invention is a search support method including steps of: receiving patent documents and a first set of cited examples; obtaining a first summary extracted from the patent documents and a plurality of second summaries extracted from each document belonging to the first set of cited examples; inputting, to a language model, a first instruction sentence for outputting a first similarity with the first summary for each of the plurality of second summaries; clustering the similarities to divide the first set of cited examples into two or more second sets of cited examples and obtaining a label for each second set of cited examples; inputting, to the language model, a second instruction sentence for outputting a difference from the first summary for each of the second sets of cited examples; and deleting at least one second set of cited examples from the two or more second sets of cited examples based on the label and the difference.
[0011] Another aspect of the present invention is a search support method including steps of: receiving patent documents and a first set of cited examples; obtaining a first summary extracted from the patent documents and a plurality of second summaries extracted from each document belonging to the first set of cited examples; inputting, to a language model, a first instruction sentence for outputting a first similarity with the first summary for each of the plurality of second summaries; clustering the first similarities to divide the first set of cited examples into two or more second sets of cited examples and obtaining a first label for each second set of cited examples; inputting, to the language model, a second instruction sentence for outputting a difference from the first summary for each of the second sets of cited examples; when the user cannot determine a second set of cited examples to be deleted based on the label, extracting second similarities; clustering the second similarities to divide the first set of cited examples into two or more third sets of cited examples and obtaining a second label for each third set of cited examples; and deleting at least one third set of cited examples from the two or more third sets of cited examples based on the second label.
[0012] Another aspect of the present invention includes steps of receiving patent documents and a first set of cited examples, obtaining a first summary extracted from the patent documents, and obtaining a plurality of second summaries extracted from each document belonging to the first set of cited examples, inputting a first instruction statement for outputting similarities with the first summary to a language model for each of the plurality of second summaries, clustering the similarities to divide the first set of cited examples into two or more second sets of cited examples and obtaining a label for each second set of cited examples, inputting a second instruction statement for outputting first differences with the first summary to the language model for each of the second sets of cited examples, extracting second differences when the user cannot determine a second set of cited examples to be deleted based on the label, and deleting at least one second set of cited examples from the two or more second sets of cited examples based on the label and the second differences. This is a search support method.
Advantages of the Invention
[0013] According to one aspect of the present invention, a novel search support method can be provided. Also, according to one aspect of the present invention, a search support method can be provided that provides assistance for verifying the content of the gazettes hit in the search.
[0014] The present invention does not necessarily have to have all of these effects. Also, the description of these effects does not prevent the existence of other effects of the present invention. For example, it is possible to extract other effects from the description regarding the specification, drawings, and claims.
Brief Description of the Drawings
[0015]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Mode for Carrying Out the Invention
[0016] Embodiments of the present invention will be described with reference to the drawings. However, it will be easily understood by those skilled in the art that the present invention can be variously modified without departing from the spirit thereof. Therefore, the present invention is not construed as being limited to the contents of the embodiments described below.
[0017] In these drawings and the like, the positions, sizes, and ranges of the respective components may not accurately represent the actual components. For this reason, the positions, sizes, and ranges of the respective components are not necessarily limited to the positions, sizes, and ranges disclosed in the drawings.
[0018] In this specification and the like, the words "first" and "second" may be used for convenience in understanding the technical content or for identifying each component. For this reason, the words "first" and "second" do not limit the number of each component. Also, the words "first" and "second" do not limit the order of each component. Further, the words "first" and "second" or the identification codes used in this specification may not match the words or identification codes in the claims of this patent.
[0019] In this specification and the like, a document refers to something that expresses the intention of a person using characters or symbols. A document includes a state in which it is divided into a plurality of chapters or a plurality of paragraphs. Also, one chapter includes a state composed of a plurality of paragraphs. Further, in this specification and the like, a document that becomes a reference material for research is called a document, and the document includes publications such as academic papers or magazines.
[0020] In this specification and the like, the gazette includes the application form required for a patent application, the specification attached to the application form, and a document corresponding to the claims attached to the application form, and further includes bibliographic data of the application form, text data of the specification and the claims. In this specification and the like, the text data may be obtained from OCR (Optical Character Recognition) of a document file. The gazette may include the drawings attached to the application form, and further the gazette may include one or more selected from mathematical formulas, chemical formulas, and tables in the specification. Also, it is assumed that one or more numbers selected from the application number, publication number, and registration number assigned by the Patent Office are associated with the gazette.
[0021] In this specification and the like, the document related to a patent includes a patent specification other than the gazette and a document corresponding to the claims. The document related to a patent is typically a draft before filing or an application form related to a patent. The application form may be a document created by the inventor. Also, the application form may be a document modified by a person other than the inventor, typically an intellectual property official. The document related to a patent may be said to be a document that can be prepared before filing with the Patent Office or before the Patent Office publishes it.
[0022] In this specification and the like, the gazette and the document related to a patent are collectively called patent documents. The document corresponding to the patent documents includes something that can be called a document.
[0023] In this specification and the like, a cited example basically refers to a document that has become publicly known before the filing of a patent, although there are some differences in the eligible conditions depending on the country. The document corresponding to the cited example includes something that can be called a document. In this specification and the like, two or more cited examples are called a "group of cited examples".
[0024] In this specification and the like, the language constituting the document is not limited in any way. The language may be Japanese or a foreign language (such as English, Chinese, Korean, German, etc.).
[0025] In this specification and the like, the language model is based on the Transformer architecture and is additionally trained to be a dialogue (also called conversation) model. Also, a typical language model is an LLM. An LLM is specialized in a text generation function that processes based on given text data, and generative AI has not only a text generation function but also an image generation function that processes based on image data.
[0026] (Embodiment 1) In this embodiment, a configuration example of a system (referred to as a search support system) that enables the implementation of search support, which is one aspect of the present invention, will be described with reference to FIGS. 1 and 2.
[0027] <Configuration Example 1 of Search Support System> The search support system of this embodiment preferably has a configuration including a first information processing device 10, a second information processing device 40, and an information terminal 20 as shown in FIG. 1. As shown in FIG. 1, the information terminal 20 is connected to the first information processing device 10 via a network 31. Also, the first information processing device 10 is connected to the second information processing device 40 via a network 30. In the search support system, the network 31 in FIG. 1 can be replaced with the network 30. That is, the first information processing device 10, the second information processing device 40, and the information terminal 20 may all be connected by the network 30.
[0028] <Configuration Example 2 of Search Support System> The search support system according to this embodiment may be configured to include a second information processing device 40 and an information terminal 20 without using the first information processing device 10 as shown in FIG. 2. As shown in FIG. 2, the information terminal 20 is connected to the second information processing device 40 via a network 30. The configurations of the first information processing device 10 can be provided in the second information processing device 40. Also, the configurations of the first information processing device 10 may be provided in the information terminal 20. By providing each configuration in the second information processing device 40 or the information terminal 20, the first information processing device 10 can be omitted. In the search support system, the network 30 in FIG. 2 can be replaced with a network 31. The network 30 and the network 31 will be described later.
[0029] 《Configuration Example of Information Terminal》 In Configuration Example 1 and Configuration Example 2 of the search support system, the information terminal 20 is operated by a user and can also be referred to as a client computer or the like. In FIGS. 1 and 2, as an example, a desktop computer is shown, but a notebook computer, a smartphone, or a tablet computer may be used as the information terminal 20. A notebook computer can fold a housing having an input unit (typically a keyboard) and overlap it with the main body. Also, some notebook computers can separate the housing from the main body. Since a notebook computer, a smartphone, or a tablet computer is easy to carry, it is suitable when the user uses the search support system, which is an aspect of the present invention, from a location outside the office.
[0030] 《Configuration Example of the First Information Processing Device 10》 Next, a configuration example of the first information processing device 10 will be described with reference to FIG. 3.
[0031] As shown in FIG. 3, the first information processing apparatus 10 includes an input unit 110, a storage unit 120, a processing unit 130, an output unit 140, a transmission path 150, and a display 160. In FIG. 3, in addition to the first information processing apparatus 10, an information terminal 20 and a second information processing apparatus 40 are shown, and the arrows indicate the transmission and reception of data. The input unit and the output unit may be collectively referred to as a communication unit. Through the communication unit, the first information processing apparatus 10 can transmit and receive data to and from the outside. Note that in the first information processing apparatus 10, if it is not necessary to display to the user of the search support system, the display 160 may not be provided.
[0032] [Input unit 110] The input unit 110 enables the first information processing apparatus 10 to have a function of receiving external data. For example, the input unit 110 can receive data from the information terminal 20. Also, the input unit 110 can receive data from the second information processing apparatus 40. The input unit 110 can also receive data from other terminals or information processing apparatuses. Examples of the data include document data corresponding to patent documents or document data corresponding to each of the citation groups. The document data includes text data, and may also include image data such as drawings. Further, the document data may include one or more number data selected from the application number, publication number, and registration number.
[0033] The input unit 110 can supply the received data to one or more selected from the storage unit 120, the processing unit 130, and the display 160 via the transmission path 150.
[0034] [Storage unit 120] The storage unit 120 enables the first information processing apparatus 10 to have a storage function. The storage unit 120 is a memory area and can store programs and / or data. Representative examples of the programs are the programs executed by the processing unit 130. The data includes data generated by the processing unit 130 (e.g., calculation results, analysis results, inference results). Also, the data includes the data received by the input unit 110.
[0035] In addition, the storage unit 120 may have a database. As the database, there is a database of patent documents. As the database, there is a database of cited example groups. As the database, there is a database of user operation histories (including search histories). The storage unit 120 has a function of storing the database and can further manage it. Managing includes appropriately deleting unnecessary data.
[0036] The database may not be possessed by the storage unit 120. Also, in addition to the storage unit 120, a database existing outside the first information processing apparatus 10 can be used. The databases existing outside include databases in which patent management operators in each country accumulate gazettes.
[0037] The storage unit 120 has at least one of a volatile memory and a non-volatile memory. Examples of the volatile memory include DRAM (Dynamic Random Access Memory) and SRAM (Static Random Access Memory). Examples of the non-volatile memory include ReRAM (Resistive Random Access Memory, also referred to as a resistive change type memory), PRAM (Phase change Random Access Memory), FeRAM (Ferroelectric Random Access Memory), MRAM (Magnetoresistive Random Access Memory, also referred to as a magnetic resistance type memory), and flash memory. The storage unit 120 can be configured by a Si LSI (circuit using silicon transistors).
[0038] Further, the memory unit 120 may have at least one of NOSRAM (registered trademark) and DOSRAM (registered trademark). Also, the memory unit 120 may have a recording media drive. Examples of the recording media drive include a hard disk drive (HDD) and a solid state drive (SSD).
[0039] NOSRAM is an abbreviation of "Nonvolatile Oxide Semiconductor Random Access Memory (RAM)". NOSRAM refers to a memory in which the memory cell is a gain cell of a 2-transistor (2T) type or a 3-transistor (3T) type, and a transistor using a metal oxide in the channel formation region (also referred to as an OS transistor) is used as the transistor. The OS transistor has an extremely small current flowing between the source and the drain in the off state, that is, a leakage current. NOSRAM can be used as a non-volatile memory by holding charges corresponding to data in the memory cell using the characteristic of extremely small leakage current. In particular, since NOSRAM can read the stored data without destroying it (non-destructive readout), it is suitable for arithmetic processing that repeatedly performs a large number of only data read operations. NOSRAM can be provided by stacking memory cells. By stacking them, the data capacity increases, so high performance can be achieved by using it as a large-scale cache memory, main memory, or storage memory.
[0040] DOSRAM is an abbreviation of "Dynamic Oxide Semiconductor RAM" and refers to a RAM having a 1-transistor (1T) 1-capacity (1C) type memory cell. DOSRAM is a DRAM formed using an OS transistor, and DOSRAM is a memory that temporarily stores information sent from the outside. DOSRAM is a memory that utilizes the small off-current of the OS transistor.
[0041] In this specification and the like, a metal oxide is an oxide of a metal in a broad sense. Metal oxides are classified into oxide insulators, oxide conductors (including transparent oxide conductors), oxide semiconductors (also referred to as Oxide Semiconductor or simply OS), and the like. For example, when a metal oxide is used for the semiconductor layer of a transistor, the metal oxide may be referred to as an oxide semiconductor.
[0042] The metal oxide included in the channel formation region preferably contains indium (In). An OS transistor using a metal oxide containing indium for the channel formation region has high carrier mobility (electron mobility). Further, the metal oxide included in the channel formation region is preferably an oxide semiconductor containing an element M described below instead of or in addition to In. The element M is preferably at least one of aluminum (Al), gallium (Ga), and tin (Sn). Other elements applicable to the element M include boron (B), silicon (Si), titanium (Ti), iron (Fe), nickel (Ni), germanium (Ge), yttrium (Y), zirconium (Zr), molybdenum (Mo), lanthanum (La), cerium (Ce), neodymium (Nd), hafnium (Hf), tantalum (Ta), and tungsten (W). In the metal oxide, a plurality of the elements listed as the element M may be combined. The element M is an element having a high binding energy with oxygen, and is an element having a binding energy with the oxygen higher than the binding energy between oxygen and indium. Further, the metal oxide included in the channel formation region is preferably a metal oxide containing zinc (Zn) instead of or in addition to In. A metal oxide containing zinc may be likely to crystallize.
[0043] The metal oxide included in the channel formation region is not limited to the metal oxide containing the above-described elements, typically indium. The metal oxide included in the channel formation region is preferably a metal oxide not containing indium, such as zinc tin oxide and gallium tin oxide, and typically, a metal oxide containing zinc, a metal oxide containing gallium, or a metal oxide containing tin can be used.
[0044] [Processing unit 130] The processing unit 130 enables the first information processing apparatus 10 to have functions for performing processes such as calculation, analysis, and inference. Typically, data is supplied from one or both of the input unit 110 and the storage unit 120, and processes such as calculation, analysis, and inference can be performed using the supplied data. Also, the data can be acquired from the storage unit 120, and processes such as calculation, analysis, and inference can be performed using the acquired data.
[0045] The processing unit 130 can supply data (for example, calculation results, analysis results, inference results) generated by itself to one or both of the storage unit 120 and the output unit 140. For example, the data can be supplied to the storage unit 120 via the transmission path 150.
[0046] Also, when the first information processing apparatus 10 has the display 160, it is preferable that the processing unit 130 also has a function of constructing display data. The processing unit 130 preferably constructs display data so that the layout is easy for the user to view on the display 160. Of course, the layout according to the user's settings can also be used on the display 160.
[0047] The processing unit 130 has at least an arithmetic circuit. As the arithmetic circuit, for example, a central processing unit (CPU) can be provided. The CPU has an arithmetic unit, a primary cache memory, a secondary cache memory, and the like. Also, in addition to or instead of the CPU, the processing unit 130 may have a GPU (Graphics Processing Unit). The GPU has an arithmetic unit, a primary cache memory, a secondary cache memory, and the like. One or both of an OS transistor (a transistor using an oxide semiconductor layer as a channel formation region) and an Si transistor (a transistor using a semiconductor layer having silicon as a channel formation region) can be provided in a switch or the like provided in the CPU or the GPU.
[0048] The processing unit 130 may have registers and a main memory in addition to the CPU. The registers and the main memory may also be those of the CPU. The main memory can transmit and receive data to and from a secondary cache or the like. The main memory has at least one of a volatile memory such as a RAM and a non-volatile memory such as a ROM (Read Only Memory). Also, the main memory may have at least one of a NOSRAM and a DOSRAM. The main memory can have one or both of an OS transistor and an Si transistor. Note that if the CPU in this paragraph is replaced with a GPU, the configurations of the registers and the main memory can be understood.
[0049] Examples of the RAM include a DRAM or an SRAM. It is also possible to virtually allocate a memory space as a work space of the processing unit 130 and use it. The operating system, application programs, program modules, program data, look-up tables, etc. stored in the storage unit 120 are loaded into the RAM immediately before execution. The operating system, application programs, program modules, program data, and look-up tables loaded into the RAM can each be accessed from the processing unit 130.
[0050] The ROM can store systems that do not require rewriting, such as BIOS (Basic Input / Output System) firmware. Examples of ROM include mask ROM, OTPROM (One Time Programmable Read Only Memory), EPROM (Erasable Programmable Read Only Memory), etc. Examples of EPROM include UV-EPROM (Ultra-Violet Erasable Programmable Read Only Memory) that enables erasure of stored data by ultraviolet irradiation, EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory, etc.
[0051] In addition to the CPU or GPU, the processing unit 130 may have a microprocessor such as a DSP (Digital Signal Processor). Since the DSP is specialized for digital signal processing, it is preferably installed to control the peripheral circuits of the CPU or GPU. The microprocessor may be configured by a PLD (Programmable Logic Device) that operates on hardware such as an FPGA (Field Programmable Gate Array) or FPAA (Field Programmable Analog Array). Also, the processing unit 130 may have a quantum processor. The processing unit 130 can interpret instructions from various programs by a processor such as a quantum processor and execute various data processing and program controls. Programs executable by the processor are stored in at least one of the memory area of the processor and the storage unit 120.
[0052] Since the off-current of the above-described OS transistor is extremely small, by using the OS transistor as a switch for holding the charge (data) flowing into the capacitive element, the data retention period can be ensured over a long period. By using this characteristic for at least one of the register and the cache memory included in the processing unit 130, the processing unit is operated only when necessary, and in other cases, the information of the previous processing is saved in the storage element, so that the signal input or power supply to the processing unit 130 can be turned off. That is, with the OS transistor, normally-off computing becomes possible, and the power consumption of the search support system can be reduced.
[0053] Furthermore, by using a CPU or the like that can operate at high speed for the processing unit 130, some of the processes executed by the first information processing apparatus 10 can be processes using AI. The first information processing apparatus 10 preferably includes an artificial neural network (ANN: Artificial Neural Network, hereinafter also simply referred to as a neural network) in order to enable processing using AI. Since the neural network is realized by a circuit (hardware) or a program (software), the first information processing apparatus 10 preferably has the above circuit or the above program in addition to a CPU that can operate at high speed.
[0054] In this specification and the like, the neural network refers to a model in general that mimics the neural circuit network of a living organism, determines the connection strength between neurons by learning, and has problem-solving ability. The neural network has an input layer to which information is input, an output layer from which information is output, and an intermediate layer (hidden layer) between the input layer and the output layer, and optimizes the weights for the data input to obtain a correct output result.
[0055] In this specification and the like, when describing the neural network, determining the weight coefficient between neurons from existing information may be referred to as "learning".
[0056] In this specification and the like, when a neural network is configured using weight coefficients obtained by learning and new conclusions are derived therefrom, it may be referred to as "inference".
[0057] [Output unit 140] The output unit 140 enables the first information processing device 10 to have a function of outputting calculation results and the like to the outside. For example, the output unit 140 can output the calculation results and the like in the processing unit 130 to the outside of the first information processing device 10. Examples of the outside include one or more selected from the second information processing device 40 and the information terminal 20.
[0058] [Transmission path 150] The transmission path 150 has a function of transmitting data. The transmission and reception of data between the input unit 110, the storage unit 120, the processing unit 130, the output unit 140, and the display 160 can be performed via the transmission path 150.
[0059] 《Configuration example of the second information processing device 40》 Next, a configuration example of the second information processing device 40 will be described.
[0060] The second information processing device 40 can process the received data and transmit the result of the processing. For example, using the data received from the first information processing device 10, processing such as calculation can be performed. Also, the second information processing device 40 can transmit the result of the processing to the first information processing device 10. Thereby, the calculation burden on the first information processing device 10 can be reduced.
[0061] The second information processing device 40 can perform processing using a natural language processing model using a generative AI. For example, processing using natural language processing models such as BERT (Bidirectional Encoder Representations from Transformers) and T5 (Text-to-Text Transfer Transformer) can be executed.
[0062] In addition, the second information processing device 40 can perform processing using a model (such as a document generation model or a dialogue model) that utilizes a large language model. For example, processing can be executed using large language models such as GPT-3, GPT-3.5, GPT-4, LaMDA (Language Model for Dialogue Applications), PaLM (Pathways Language Model), Llama2, etc.
[0063] In addition, the second information processing device 40 can execute processing using a general-purpose language processing model capable of performing various natural language processing tasks.
[0064] In the search support system, the provider of the search support service does not necessarily have to own the second information processing device 40 in-house. For example, the service provider can utilize a part of the service provided by another operator or the like using the second information processing device 40.
[0065] 《Configuration Example of Information Terminal 20》 The information terminal 20 has a function for the user to input data. In addition, the information terminal 20 can provide the data output by the search support system to the user. That is, in the search support system, it is preferable that the user operates the information terminal 20 and the first information processing device 10 and the second information processing device 40 do not operate. Security can be improved in such a form.
[0066] For example, when the provider of a service using a search support system and the user who enjoys the service belong to the same organization such as the same company, the data transmission and reception between the information terminal 20 and the first information processing device 10 are preferably performed using, for example, the network 31 constructed within the organization. Thereby, data can be transmitted and received between the information terminal 20 and the second information processing device 10 more securely than when performed via the Internet. Also, leakage of confidential information within the organization to the outside can be prevented. Alternatively, the data transmission and reception between the information terminal 20 and the first information processing device 10 may be performed using the network 30 (for example, the Internet).
[0067] It is preferable that, for example, dedicated application software or a web browser or the like is installed in the information terminal 20. The user can also access the first information processing device 10 via the application software or the web browser. Thereby, the user can enjoy the service using the search support system using the information terminal 20 having a lower processing capacity than the first information processing device 10.
[0068] 《Network 30》 The network 30 is an example of a network connecting the first information processing device 10 and the second information processing device 40. Thereby, transmission and reception of the input data and the processed data are possible between the two. Also, the load related to information processing can be dispersed.
[0069] For example, a global network can be used as the network 30. Specifically, the Internet, which is the basis of the World Wide Web (WWW), can be used as the global network. The network 30 is preferably a larger-scale computer network than the network 31.
[0070] 《Network 31》 Network 31 is an example of a network that connects information terminal 20 and the first information processing device 10. As a result, data can be transmitted and received between the two, and the load related to information processing can be dispersed. In addition, a service provider can provide a service using a search support system to a user via network 31, for example.
[0071] For example, a local network can be used as network 31. Also, an intranet or an extranet can be used as network 31. Further, a PAN (Personal Area Network), a LAN (Local Area Network), a CAN (Campus Area Network), a MAN (Metropolitan Area Network), a WAN (Wide Area Network), a GAN (Global Area Network), etc. can be used as network 31.
[0072] When performing wireless communication, as a communication protocol or communication technology, communication standards such as the fourth-generation mobile communication system (4G), the fifth-generation mobile communication system (5G), the sixth-generation mobile communication system (6G), or specifications standardized by the IEEE such as Wi-Fi (registered trademark), Bluetooth (registered trademark), etc. can be used.
[0073] <Example 1 Regarding the Method> In the present embodiment, a search support method, which is one aspect of the present invention, will be described separately for each step. It is desirable that the order of each step described later be executed as described, but the order of each step in this search support method is not necessarily limited.
[0074] FIG. 4 is an example showing each step related to a search support method, which is one aspect of the present invention, in a flowchart. Note that the search support method, which is one aspect of the present invention, is started after the user logs in. The operation history (including the search history) of the user can be associated with the login information.
[0075] <Step S101> As shown in step S101 of FIG. 4, the search support system receives patent documents for which the user desires search support. Further, in step S101, the search support system receives a group of cited examples for which the user desires search support. The group of cited examples for which search support is desired is referred to as the first group of cited examples. The order of reception is not particularly limited, and it is possible to receive patent documents and the first group of cited examples in the same step.
[0076] The first group of cited examples can be prepared by the user. For example, the first group of cited examples can be a group of cited examples collected by the user using a search tool attached to the search support system. Also, the first group of cited examples can be a group of cited examples collected by the user using a search tool other than the search support system. The first group of cited examples can also be prepared by the search support system. For example, the first group of cited examples can be a group of cited examples associated with the user's login information.
[0077] The user's input operation corresponding to step S101 is performed on the information terminal 20. FIG. 5 shows an example of a screen 50 that can be displayed on the information terminal 20. Since the screen 50 can be set by the user, the layout of the screen 50 is not limited to that shown in FIG. 5.
[0078] The screen 50 preferably has an input box 61 for inputting patent documents corresponding to step S101. Also, the content of the patent document input to the input box 61 can be displayed in the box 65.
[0079] Also, the screen 50 preferably has a text box 71 for inputting the first group of cited examples corresponding to step S101. For example, after inputting the link destination of the folder in which the first group of cited examples is stored into the text box 71, the first group of cited examples can be uploaded to the search support system. The screen 50 may newly have a box for displaying the content of each document belonging to the first group of cited examples, or for displaying a list of the first group of cited examples, etc.
[0080] The main information processing related to step S101 is preferably executed by the first information processing apparatus 10. Examples of the main information processing include a process in which the input unit 110 receives each piece of data from the information terminal 20, and each piece of data is supplied to one or both of the storage unit 120 and the processing unit 130 via the transmission path 150.
[0081] <Additional Function 1: List Button> The additional function 1 of the screen 50 will be described. Although not shown in FIG. 5, the screen 50 may be provided with a list button for list-displaying a plurality of patent documents instead of or in addition to the input box 61. Although not shown in FIG. 5, the screen 50 may be provided with a list button for list-displaying the first cited example group instead of or in addition to the text box 71.
[0082] As the specification of the above-described list button, it is preferable that when the user selects the list button, two or more patent documents are list-displayed. The user can select a patent document for which search support is desired from the list. Note that if the patent documents are replaced with the first cited example group, the specification of the list button for list-displaying the first cited example group can be understood.
[0083] The above-described list can be stored in the database included in the storage unit 120 or a database external to the first information processing apparatus 10. In this case, the patent document selected from the list is transmitted to the information terminal 20 via the output unit 140, and the user can check the content in the box 65 or the like. Note that if the patent documents are replaced with the first cited example group, the specification for checking the content of any one of the documents belonging to the first cited example group can be understood. In this case, it is preferable that the screen 50 newly has a box for displaying the content of any one of the documents belonging to the first cited example group.
[0084] <Additional Function 2: Search Button> The additional function 2 of screen 50 will be described. Although not shown in FIG. 5, the screen 50 may be provided with a search button for a plurality of patent documents stored in the database instead of, or in addition to, the input box 61. Also, although not shown in FIG. 5, the screen 50 may be provided with a search button for executing a search for a group of cited examples stored in the database instead of, or in addition to, the text box 71.
[0085] As the specification of the above-described search button, it is preferable that a box for inputting a search term enabling a general patent search be displayed. The box for inputting the search term may be configured to be displayed on a new screen different from the screen 50. With such an additional function, the user can select a patent document based on the results of a general patent search. Note that if the patent document is read as the first group of cited examples, the specification of the search button for the first group of cited examples can be understood.
[0086] The above-described search can be mainly executed by the processing unit 130 of the first information processing apparatus 10. This search can utilize the database possessed by the storage unit 120 or a database external to the first information processing apparatus 10, similar to the above list display. The search results are transmitted to the information terminal 20 via the output unit 140. The user can confirm the search results, and the patent document selected from the search results can be checked for its content in the box 65. Note that if the patent document is read as the first group of cited examples, the content of any one of the documents belonging to the first group of cited examples can be checked. In this case, the screen 50 may newly have a box for displaying the content of any one of the documents belonging to the first group of cited examples.
[0087] The search support system can be provided with both the additional function 1 and the additional function 2.
[0088] <Step S111> Next, as shown in step S111 of FIG. 4, the search support system obtains an outline (sometimes referred to as an abstract) extracted from a patent document. Also in step S111, each outline extracted from each document belonging to the first cited example group is obtained. Step S111 is called an extraction operation. The order of the extraction operation is not particularly limited, and it is possible to process the extraction operations for the patent document and the first cited example group in the same step.
[0089] As an extraction operation, the content itself described in a specific chapter of a patent document can be extracted as an outline. The specific chapter can be specified by the user. Typically, the specific chapter includes an abstract, problems, claims, examples, or embodiments of the patent document. Also, if the patent document is replaced with each document belonging to the first cited example group, the extraction operation for each document belonging to the first cited example group can be understood. Also, each document belonging to the first cited example group may not have an abstract, problems, claims, examples, or embodiments, etc., but the user can appropriately select a chapter corresponding to the abstract, problems, claims, examples, or embodiments, etc. Note that it is preferable that the chapter specified in the patent document and the chapter specified in each document belonging to the first cited example group disclose equivalent content. For example, if the chapters have the same title or substantially the same title, it can be determined that they are chapters that disclose equivalent content. By such a specification, the variation in the extraction results can be suppressed. In this way, the search support system can obtain each outline.
[0090] Since the processing of step S111 can be automatically advanced by the search support system, the display screen 50 does not require a display corresponding to this step. Of course, a display corresponding to step S111 may be performed on the display screen 50. Also, although not shown, it is preferable to display the waiting time for the processing of this step on the display screen 50.
[0091] The main information processing related to step S111 is preferably executed by the first information processing device 10. As the main information processing, there is a process in which the processing unit 130 generates an operation result, an analysis result, or an inference result.
[0092] <Variation of Step S111> The extraction operation of Step S111 may use generative AI. In this case, the patent document may be referred to as the original text. Also, each document belonging to the first citation group may be referred to as the original text. Using the original text and the instruction text (listing patent documents, etc., such as "Please create an abstract of the following documents."), as prompts, input them into a language model, typically an LLM, to extract these abstracts. When extracting the abstract, the user can specify a paragraph or chapter in the patent document. In this case, the search support system can set the specified paragraph or chapter as the original text in the instruction text. "Setting" includes listing the specified paragraph or chapter in the instruction text. Another example of the instruction text may be "Please create an abstract of the following specified chapter."
[0093] In addition, the user can also set keywords for each of the patent documents or each document belonging to the first citation group. When extracting the abstract and keywords are set, the search support system can create an instruction text containing the specified keywords. As a specific example of the instruction text, when there are keywords, it is preferable to use "Please create an abstract regarding the 'keyword' in the following document." Note that it is preferable for the search support system to prepare one or more instruction texts as fixed-form texts and for the user to select the instruction text.
[0094] In Step S111 using generative AI, a step for the user to confirm the abstract can be added. For confirmation, it is preferable for the search support system to newly provide a box on Screen 50 to display the abstract of the patent document. Similarly, it is preferable to newly provide a box on Screen 50 to display the abstract of each document belonging to the first citation group.
[0095] Even when using generative AI, since step S111 can be automatically processed by the search support system, the screen 50 does not require a display corresponding to this step. Of course, a display corresponding to step S111 may be provided on the screen 50. Also, although not shown, it is preferable to display the waiting time for the processing of this step on the screen 50.
[0096] Although not shown, it is preferable that the main information processing related to step S111 using generative AI is executed by the second information processing device 40. Using generative AI improves the accuracy of the outline.
[0097] <Step S113> Next, in step S113 of FIG. 4, for each document summary belonging to the first citation group, similarities (which may be referred to as matching points) with the patent document summary are obtained using generative AI. The generative AI that can be used is the same as the variant of step S111. By using the same generative AI, historical information related to step S111 can be added, so that high-accuracy similarities can be obtained in step S113. Of course, different generative AI may be used in step S113 from that in step S111.
[0098] In this step, it is preferable that the original text be the abstract of the patent documents obtained in the previous step and the abstract of each document belonging to the first group of cited examples. In this step, the method of outputting similarities can be selected by the user. The instruction text only needs to instruct that similarities to each other be extracted. Typically, it can be something like "Please extract the similarities between the abstract of the patent documents and the abstract of each document belonging to the first group of cited examples." In the search support system, it is preferable to prepare one or more candidates for the instruction text as fixed texts, and for the user to select the instruction text. Also, the user can be made to output similarities after specifying keywords. For example, based on the abstract of the patent documents, keywords specified by the user are prepared, and similarities regarding the "keywords" can also be obtained by saying "Please extract the similarities regarding the 'keywords' for the abstract of each document belonging to the first group of cited examples." This is called keyword similarity. Also, based on the abstract of the patent documents, a configuration specified by the user is prepared, and similarities can also be obtained by saying "Please output the parts that match the configuration of the abstract of the patent documents for the configuration of the abstract of each document belonging to the first group of cited examples." This is called configuration similarity. Such a method of outputting similarities can also be proposed by the search support system in association with the user's login information. Note that the prompt including the above instruction text is input to a language model, typically an LLM.
[0099] Although the method of obtaining similarities has been described above, for the abstract of each document belonging to the first group of cited examples, the differences from the abstract of the patent documents may be obtained using generative AI. However, for the labeling in the steps described later, labeling based on similarities is more preferable in the search support system than labeling based on differences. This is because the abstract of the patent documents should be close to the abstract of each document belonging to the first group of cited examples, and it is considered that the processing becomes more efficient when labels are extracted from similar abstracts.
[0100] As shown in FIG. 5, the screen 50 preferably has a button 90 for starting the process of step S113 as a display corresponding to step S113. Further, the screen 50 preferably has a box 91 for displaying an instruction text. Further, the screen 50 preferably has a box 92 for displaying similarities as a display corresponding to step S113. Since the similarities are obtained for each document belonging to the first citation group, the similarities in the box 92 may be displayed in association with the corresponding citations. For example, the box 92 may list a plurality of similarities, and when any one of the similarities is selected, the corresponding citation may be displayed. A user who contacts such a screen 50 can confirm the instruction text in the box 91, confirm each similarity in the box 92, and further confirm the content of the patent document in the box 65.
[0101] The main information processing regarding step S113 is preferably executed by the second information processing device 40.
[0102] <Step S114> Next, step S114 in FIG. 4 vectorizes and clusters each similarity between each document belonging to the first citation group and the patent document, and divides the first citation group into two or more second citation groups. Typically, clustering may be performed using k-means or DBSCAN, etc., and based on the similarity degree of each similarity, the first citation group can be divided into two or more second citation groups.
[0103] Various methods can be used for the method of vectorizing a document. For example, Bag-of-Words, BERT (Bidirectional Encoder Representations from Transformer), etc. can be used.
[0104] Also, as a means for calculating the similarity degree of similarities between documents, the number of occurrences of words may be used. As a method of vectorizing a document using the number of occurrences of words, TF-IDF (Term Frequency-Inverse Document Frequency), etc. can be used.
[0105] Moreover, even if the distributed representation of words is used instead of the number of occurrences of words, the similarity of similarities between documents can be calculated. As a method for vectorizing documents using the distributed representation of words, for example, Word2vec, Doc2Vec, Sent2Vec, etc. can be used.
[0106] <Additional Function 3: Labeling> Furthermore, in this step, it is preferable to obtain a label for each second citation group. The label is preferably a word calculated from the similarity of documents. Therefore, typically, TF-IDF or the like can be used to obtain the label. Also, a generative AI may be used to extract the label. Specifically, it is preferable to input a prompt including appropriate instructions into a language model, typically an LLM, with the abstract of each document belonging to the second citation group as the original text to extract the label. Specifically, it can be "Create appropriate labels representing each of the following second citation groups using the following multiple documents belonging to the second citation group as the original text. Citation Group 1, Citation Group 2, Citation Group n" (where n is the number of groups belonging to the second citation group). In extracting the label, it is not necessary to target all the documents belonging to the second citation group in the instruction. It is advisable to select an arbitrary document from one of the second citation groups, preferably 2 or more and 10 or fewer documents. In this specification, etc., selecting an arbitrary document from the citation group is called sampling.
[0107] Note that in the search support system, each document belonging to the second citation group with label A is allowed to belong to the second citation group with label B.
[0108] Corresponding to step S114, as shown in FIG. 5, the screen 50 has a box 73 for displaying one second citation group and a box 74 for displaying another second citation group. There may be two or more boxes for displaying the second citation group. The label can be displayed on the screen 50 in the box 73 and the box 74. Of course, a new box for the label may be provided on the screen 50.
[0109] The main information processing related to step S114 may be executed by the second information processing apparatus 40.
[0110] <Step S115> Next, in step S115 of FIG. 4, unnecessary citation groups are deleted. Specifically, the user can determine unnecessary second citation groups based on labels or the like. When the search support system receives the result, it is deleted from the search results.
[0111] Corresponding to step S115, as shown in FIG. 5, the screen 50 preferably has a delete button. Specifically, it preferably has a delete button 96 corresponding to one second citation group and a delete button 97 corresponding to another second citation group.
[0112] The main information processing related to step S115 may be executed by the first information processing apparatus 10. Examples of the main information processing include the input unit 110 receiving each data from the information terminal 20, and each data being supplied to one or both of the storage unit 120 and the processing unit 130 via the transmission path 150.
[0113] The search support method according to one aspect of the present invention can end after step S115.
[0114] By such a search support method according to one aspect of the present invention, support can be received for confirming the content of citations, and the user can efficiently and appropriately reduce the citations, thereby reducing the burden. The greater the number of documents belonging to the first citation group, the more remarkable the effect of the search support method according to one aspect of the present invention.
[0115] <Example 2 related to the method> A method different from the above Example 1 will be described with reference to FIG. 6. Note that Example 2 is the same as Example 1 from step S101 to step S114.
[0116] <Step S116> Next, step S116 in FIG. 6 extracts the differences between the outlines of each document belonging to the second cited example group and the outline of the patent document using a generative AI. That is, in this step, a prompt including the instruction sentence corresponding to the above is input to the language model. Also, step S116 in FIG. 6 can extract the differences between the label corresponding to the second cited example group and the outline of the patent document. Each difference may be displayed in box 93 or box 94 on screen 50.
[0117] The main information processing regarding step S116 is preferably executed by the second information processing device 40.
[0118] <Step S115> Next, in step S115 of FIG. 5, unnecessary cited example groups are deleted. Specifically, the user can determine an unnecessary second cited example group based on the differences in addition to the labels. When the search support system receives the result, it is deleted from the search results. Note that this step can refer to the description of step S115 described in Example 1. By going through such steps, the cited examples can be reduced accurately.
[0119] The search support method according to one aspect of the present invention can end after step S115.
[0120] With such a search support method according to one aspect of the present invention, support can be received for checking the content of the cited examples, and the user can efficiently and appropriately reduce the cited examples, thereby reducing the burden. The greater the number of documents belonging to the first cited example group, the more remarkable the effect of the search support method according to one aspect of the present invention.
[0121] <Example 3 regarding the method> A method different from the above Examples 1 and 2 will be described with reference to FIG. 7. Note that Example 3 is the same as Example 2 up to step S116, and the search support system that has gone through step S116 has obtained similarities and differences.
[0122] <Step S119> Step S119 in FIG. 7 is a step that is preferably added when the user cannot determine unnecessary reference examples. If it cannot be determined in step S119 (in the figure, in the case of "no"), for example, it is possible to return to step S113 to newly extract similarities.
[0123] When returning to step S113, if the previously obtained similarities include two or more features, on screen 50, these features may be listed. The user who touches screen 50 can delete inappropriate features among the listed features. The search support system can accept the deletion, add conditions to ignore inappropriate features, and then execute step S113. As a result, appropriate similarities can be newly obtained. Then, execute step S114 to cluster similarities different from the previous ones and divide the first reference example group into a new second reference example group (referred to as the third reference example group), and new labels can also be obtained. Next, by executing step S116, new differences can also be obtained.
[0124] <Step S115> Next, in step S115 of FIG. 7, the user can determine an unnecessary second reference example group based on a label different from the previous label. When the search support system accepts the result, the unnecessary second reference example group is deleted from the search results. Note that this step can refer to the description of step S115 described in Example 1. By going through such steps, unnecessary reference examples can be reduced with higher accuracy.
[0125] <Variant example of step S119> If it cannot be determined in step S119 (in the case of "no" in the figure), it is possible to return to step S116 in order to newly extract the differences. When returning to step S116, if the previously obtained differences include two or more features, on screen 50, these features may be listed item by item. The user who touches screen 50 can delete inappropriate features among the itemized features. The search support system accepts the deletion, adds a condition to ignore the inappropriate features, and then can execute step S116. As a result, appropriate differences can be newly obtained.
[0126] After that, in step S115, the user can determine an unnecessary second group of cited examples based on differences different from the previous differences and the previous label. When the search support system accepts the result, the unnecessary second group of cited examples is deleted from the search results. Note that this step can refer to the description of step S115 described in Example 1. By going through such steps, the cited examples can be reduced with higher accuracy.
[0127] The search support method according to one aspect of the present invention can end after step S115.
[0128] By such a search support method according to one aspect of the present invention, support can be received for confirming the content of the cited examples, and the user can efficiently and appropriately reduce the cited examples, thereby reducing the burden. The greater the number of documents belonging to the first group of cited examples, the more remarkable the effect of the search support method according to one aspect of the present invention.
[0129] According to the present embodiment, a novel search support method can be provided.
Description of Reference Numerals
[0130] 10 First information processing apparatus 20 Information terminal 30 Network 31 Network 40 Second information processing apparatus 50 Screen 61 Input box 65 Box 71 Text Box 73 Box 74 Box 90 Button 91 Box 92 Box 93 Box 94 Box 96 Delete Button 97 Delete Button
Claims
1. receiving a patent document and a first set of references; obtaining a first abstract extracted from the patent document and obtaining a plurality of second abstracts extracted from respective documents in the first set of cited references; inputting a first instruction sentence to a language model, the first instruction sentence causing a similarity between each of the plurality of second summaries and the first summaries to be output; clustering the similarities to divide the first set of references into two or more second sets of references, and obtaining a label for each of the second sets of references; removing at least one second reference from the two or more second references based on the label; The search support method includes:
2. receiving a patent document and a first set of references; obtaining a first abstract extracted from the patent document and obtaining a plurality of second abstracts extracted from respective documents in the first set of cited references; inputting a first instruction sentence to a language model, the first instruction sentence causing a similarity between each of the plurality of second summaries and the first summaries to be output; clustering the similarities to divide the first set of references into two or more second sets of references, and obtaining a label for each of the second sets of references; inputting a second instruction sentence to the language model, the second instruction sentence causing a difference between the first summary and each of the second citations to be output; removing at least one second set of references from the two or more second sets of references based on the label and the difference; The search support method includes:
3. receiving a patent document and a first set of references; obtaining a first abstract extracted from the patent document and obtaining a plurality of second abstracts extracted from respective documents in the first set of cited references; inputting a first instruction sentence to a language model to output a first similarity between the first summary and each of the plurality of second summaries; clustering the first similarities to divide the first set of references into two or more second sets of references, and obtaining a first label for each of the second sets of references; inputting a second instruction sentence to the language model, the second instruction sentence causing a difference between the first summary and each of the second citations to be output; extracting a second set of similarities if the user cannot determine a second set of citations to be deleted based on the labels; clustering the second similarities to divide the first set of references into two or more third sets of references, and obtaining a second label for each of the third sets of references; removing at least one third set of references from the two or more third set of references based on the second label; The search support method includes:
4. receiving a patent document and a first set of references; obtaining a first abstract extracted from the patent document and obtaining a plurality of second abstracts extracted from respective documents in the first set of cited references; inputting a first instruction sentence to a language model, the first instruction sentence causing a similarity between each of the plurality of second summaries and the first summaries to be output; clustering the similarities to divide the first set of references into two or more second sets of references, and obtaining a label for each of the second sets of references; inputting a second instruction to the language model to output a first difference between the first summary and each of the second citations; extracting a second set of differences when the user cannot determine a second set of citations to be deleted based on the labels; removing at least one second set of references from the two or more second sets of references based on the label and the second difference; The search support method includes:
Citation Information
Patent Citations
Document search support system, document search support method, and document search support program
JP2021072009A