Programs, methods, information processing devices, graph-type data structures
Patent Information
- Application Number
- JP2026001167
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-01-07
- Publication Date
- 2026-09-09
AI Technical Summary
【0010】 本開示によれば、判断の根拠となる多様な情報源を総合的に検討することを支援することができる。
Smart Images

Figure 2026144981000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a program, a method, an information processing apparatus, and a graph-type data structure. [Background Art]
[0002] Various stakeholders including general consumers and business operators carry out procedures based on laws. Users in various positions, such as experts like attorneys and employees of business companies, conduct surveys of documents and the like through searching and collect judgment materials in order to make decisions from a legal perspective. Then, judgments on specific events, such as matters to be stipulated in terms of service and evaluation of the validity of customer attraction measures implemented in conjunction with service provision, are made from a legal perspective. Therefore, surveys that serve as a premise for judgments are important.
[0003] Patent Document 1 relates to document searching, and points out that "A document corpus including legal documents, patent documents, medical journals, and the like is searched using a query expression. ... In many cases, a user can create a plurality of search queries when investigating a specific topic. However, it may be difficult for the user to efficiently determine which search query provides the most relevant search results and how completely the search query covers a specific topic. Accordingly, many users do not trust their own document corpus search, and may believe that the generated search results have low reliability or are not sufficiently complete".
[0004] Patent Document 1 focuses on presenting search results as described above, and sets as a problem that "another method for graphically displaying electronic document search to improve the electronic document search experience is needed".
[0005] Patent Document 1 describes "generating a Venn diagram including a first circle representing a first set of documents and a second circle representing a second set of documents, and displaying it on a graphic display device," "the first circle overlaps the second circle in an overlapping area representing common electronic documents present in the first set of documents and the second set of documents," "the sizes of the first circle and the second circle reflect the number of electronic documents in the first set of documents and the second set of documents, respectively," "the first circle overlaps the second circle in an overlapping area representing common electronic documents present in the first set of documents and the second set of documents," and "in response to user input, indicating the separation of the first circle and the second circle on a graphic display device, and generating a first visualization chart from the first circle and a second visualization chart from the second circle." [Prior art documents] [Patent Documents]
[0006] [Patent Document 1] Japanese Patent Publication No. 2017-010580 [Overview of the project] [Problems that the invention aims to solve]
[0007] On the other hand, in order to make appropriate judgments in legal practice, for example, it is necessary to comprehensively consider a variety of information sources that serve as the basis for the judgment, such as relevant laws and regulations, similar precedents, and explanations in specialized books and guidelines.
[0008] Therefore, there is a need for technology that supports the comprehensive examination of diverse sources of information that form the basis of judgments. [Means for solving the problem]
[0009] According to one embodiment shown in this disclosure, a program is provided for operating a computer having one or more computer processors. The storage unit is configured to store data in the form of a graph-type data structure in which reference relationships are defined between the parts that represent the content of each piece of content for a plurality of pieces of content to be viewed. The program causes one or more computer processors to perform the following steps: receiving input of a question from a user; searching for a plurality of pieces of content based on the content of the question; selecting a second piece of content that has a reference relationship defined with the first piece of content based on the graph-type data structure data for one or more first pieces of content that are the search results; generating a prompt that includes the content of the question, the searched first piece of content, and the selected second piece of content, and includes instructions that generate an answer to the content of the question by referring to the first piece of content and the second piece of content; providing the generated prompt to an information processing system that performs language processing, thereby obtaining an answer output by the information processing system that performs language processing; and presenting the obtained answer to the user. [Effects of the Invention]
[0010] This disclosure can help in comprehensively considering a variety of sources of information that form the basis of a decision. [Brief explanation of the drawing]
[0011] [Figure 1] Figure 1 shows the configuration of System 1. [Figure 2] Figure 2 shows the configuration of server 20. [Figure 3] Figure 3 shows the configuration of terminal 10. [Figure 4] Figure 4 shows the data structure of the user database 211. [Figure 5] Figure 5 shows the data structure of the legal content database 212. [Figure 6]Figure 6 shows the data structure of the content usage history database 213. [Figure 7] Figure 7 shows the data structure of graph-type data structure 214. [Figure 8] Figure 8 shows the process flow for creating a graph database using information from multiple legal content sources. [Figure 9] Figure 9 shows the processing flow using a graph database related to laws. [Figure 10] Figure 10 is an example of an operation screen that displays the results of generating answers to user questions while referring to a graph-type data structure 214 related to the law. [Figure 11] Figure 11 shows the process flow for identifying the relevant laws and regulations in response to the user's question and presenting those laws and regulations to the user. [Figure 12] Figure 12 shows an example of an operation screen that presents the user with the relevant laws and regulations identified in response to the user's question. [Figure 13] Figure 13 illustrates the process flow for generating a database that identifies legal issues from information on multiple legally related content. [Figure 14] Figure 14 shows the process flow for searching for content related to legal issues that address a user's question. [Figure 15] Figure 15 shows an example of an operation screen that responds to a user's question by searching various databases based on the corresponding legal issues. [Figure 16] Figure 16 shows the process flow for identifying the relevant laws and regulations in the content that are identified as the basis for the answers to the questions, and outputting the latest laws and regulations. [Figure 17] Figure 17 shows an example screen of a user interface for viewing content, which displays both the content itself and a table of contents summarizing it. [Figure 18]FIG. 18 is a screen example of a screen that displays content details and other content related to the displayed content on an operation screen for browsing content. [Figure 19] FIG. 19 is a screen example of a screen that displays laws and regulations and information on law amendments as other content related to the displayed content on an operation screen for browsing content. [Figure 20] FIG. 20 is a diagram showing the flow of processing for displaying a part of other content related to the currently displayed content while presenting an operation screen for browsing content. [Figure 21] FIG. 21 is a diagram showing the flow of processing for generating an answer by determining whether to use a first mode for generating an answer at high speed or a second mode for generating a detailed answer at low speed as an answer generation mode in accordance with the content of a question. [Figure 22] FIG. 22 is a screen example of a screen that displays the result of determining whether to answer in the first mode (high-speed mode) or the second mode (detailed answer mode) at a stage before analyzing the content of an input question and accepting an operation to transmit the question. [Figure 23] FIG. 23 is a screen example of a screen that displays the determination result of an answer generation mode on an operation screen during a period from when a question is transmitted until an answer that also uses an external information processing system is generated. [Figure 24] FIG. 24 is a screen example that displays the determination result of an answer generation mode on a screen displaying an answer generated to a question. [Figure 25] FIG. 25 is a diagram showing the flow of processing for accepting a re-question that is an in-depth inquiry regarding a question input by a user. [Figure 26] FIG. 26 is a screen example of a screen for accepting a re-question that is an in-depth inquiry regarding a question input by a user. [Figure 27] FIG. 27 is a diagram showing the flow of processing for displaying a list of content related to an answer while presenting the answer to a question, and changing a display mode of the list according to a user's operation. [Figure 28]Figure 28 is an example of a screen that displays a list of content related to a question while presenting the answer, in a first manner that facilitates understanding the overall picture of the content. [Figure 29] Figure 29 is an example screen of a second mode in which a list of content related to a question is displayed, with the amount of information in a portion of each content increased to assist in selecting content to view in more detail. [Figure 30] Figure 30 shows a graph-type data structure for each type of content: laws, precedents, books, and guidelines. [Figure 31] Figure 31 shows the data structure of the abbreviation dictionary database 215, which defines the rules for abbreviating laws and regulations. [Figure 32] Figure 32 shows the process flow for identifying sections in the content that refer to laws and regulations, depending on the various forms of legal notation, defining reference relationships with the information of the laws and regulations which have a graph-type data structure, and generating a graph-type data structure. [Figure 33] Figure 33 shows the process flow for generating answers to questions by searching for various types of content in response to a question, extracting relevant content by referring to a graph-type data structure for the search results, and using this as the basis for generating an answer to the question. [Modes for carrying out the invention]
[0012] Embodiments of the present disclosure will be described below with reference to the drawings. In the following description, identical parts are denoted by the same reference numerals. Their names and functions are also the same. Therefore, detailed descriptions of them will not be repeated.
[0013] <Outline of the Embodiment> In the first embodiment, we will describe a method for generating a data structure that defines reference relationships between legal content based on wording (names of laws, etc.) contained in laws, precedents, legal books, and legal guidelines, and a method for generating answers to user questions using the graph-type data structure generated thereby.
[0014] <1.1 System Configuration Diagram> Figure 1 shows the configuration of System 1.
[0015] System 1, shown in Figure 1, includes a server 20, a user terminal 10, a case law search service server 91, a legal search service server 92, a book browsing service server 93, an information media service server 94, an artificial intelligence (large-scale language model) service server 95 (hereinafter sometimes referred to as "large-scale language model service server 95"), a business operator server 96, an SNS server 97, and a user terminal 10A. These devices communicate with each other via a network 80.
[0016] In the illustrated example, terminals used by users of the service provided by server 20 are shown as terminal 10, terminal 10A, etc., and each user operates their own terminal.
[0017] In this embodiment, each device (terminal device, server, etc.) can also be considered as an information processing device. That is, the collection of each device can be considered as a single "information processing device," and System 1 may be formed as a collection of multiple devices. The way in which the multiple functions required to realize System 1 according to this embodiment are distributed to one or more hardware can be appropriately determined in view of the processing capacity of each hardware and / or the specifications required for System 1.
[0018] Terminal 10 is a device operated by the user. In this embodiment, the user who conducts legal research, makes judgments, etc., operates Terminal 10. Terminal 10 and Terminal 10A have similar functional configurations. Terminal 10 is implemented, for example, as follows. • Handheld mobile devices such as smartphones and tablets • Stationary PCs (Personal Computers), Laptop PCs • Wearable devices worn by the user (watch-type, glasses-type, etc.) Terminal 10 includes a communication interface (IF) 12, an input device 13, an output device 14, memory 15, storage 16, and a processor 19.
[0019] The communication interface 12 is an interface for inputting and outputting signals so that terminal 10 can communicate with an external device.
[0020] The input device 13 is a device for receiving input operations from the user (for example, a touch panel, touchpad, pointing device such as a mouse, keyboard, etc.).
[0021] The output device 14 is a device (such as a display or speaker) for presenting information to the user.
[0022] Memory 15 is for temporarily storing programs and data processed by programs, etc., and is a volatile memory such as DRAM (Dynamic Random Access Memory).
[0023] Storage 16 is for storing data, and can be, for example, flash memory or an HDD (Hard Disk Drive).
[0024] The processor 19 is hardware for executing the instruction set described in the program, and consists of an arithmetic unit, registers, peripheral circuits, etc.
[0025] Server 20 is a device for providing a service to users that accepts legal questions and answers them by presenting multiple pieces of relevant legal content in response to the questions. In this embodiment, Server 20 accepts legal questions from the user of Terminal 10 in a chat format via free text or voice input, and outputs answers to the received questions. In generating answers to such questions, Server 20 may also have Server 95 of an artificial intelligence (large-scale language model) service generate the answers, and provides the service by responding to the user with the generated results.
[0026] In this embodiment, the server 20 provides services to the following types of users. • Users who perform legal duties and provide answers to legal consultations, similar to the legal departments of business companies. • Users who do not necessarily have dedicated legal staff, such as business divisions within a company. • Users who provide legal consultations to clients as a service, such as law firms. Server 20 may accept requests for legal consultations from clients to experts by matching users who provide legal consultation services as professionals, such as law firms, with users who request legal consultations from experts such as lawyers, such as businesses or individuals. For example, the content of the consultation may be posted on a bulletin board that is accessible to third parties, and the expert's response may also be made public, or the content of the consultation may be kept confidential and not disclosed to third parties, allowing the consultation to take place privately between the client and the expert. Server 20 may also store such client consultation content and expert response content.
[0027] The server 20 includes a communication interface 22, an input / output interface 23, memory 25, storage 26, and a processor 29.
[0028] Communication IF22 is an interface for inputting and outputting signals so that the server 20 can communicate with external devices.
[0029] Input / Output IF23 functions as an interface between an input device for receiving user input operations and an output device for presenting information to the user.
[0030] Memory 25 is for temporarily storing programs and data processed by programs, etc., and is a volatile memory such as DRAM (Dynamic Random Access Memory).
[0031] Storage 26 is for storing data, and can be, for example, flash memory or an HDD (Hard Disk Drive).
[0032] The processor 29 is hardware for executing the instruction set described in the program, and consists of an arithmetic unit, registers, peripheral circuits, etc.
[0033] Server 91 of the case law search service has a database of case law, allowing users to search for judgments made by courts and other authorities.
[0034] Server 92 of the legal search service has a database of laws and regulations and allows users to search for those laws and regulations.
[0035] The book browsing service server 93 provides a service that makes electronic content such as magazines and books available for viewing. For example, users can access ebooks in fields such as law by paying a fixed fee on a regular basis.
[0036] Server 94 of the Information Media Services provides information services, such as collecting and making available blog posts, Q&A sites, news articles, and IR information. Server 94 of the Information Media Services also stores information on interpretations based on laws and regulations, as well as guidelines from government agencies and other organizations prepared to disseminate operational rules.
[0037] Server 95 of the Large-Scale Language Model Service is a server that executes language processing tasks using language models built through learning processes including artificial intelligence (AI). An LLM (Large Language Model) is a model that has been pre-trained on large amounts of data (such as text data), for example, a large amount of web content on the internet, or a large amount of data stored in a designated database, and can perform various language processing tasks by being given a task.
[0038] The server 95 of the large-scale language model service accepts prompt input in the form of text, images, audio, etc., and generates and responds with answers to those prompts. Examples of LLMs include GPT-3, GPT-4, and GPT-4o developed by OpenAI, and Gemini developed by Google.
[0039] Server 96, operated by the service provider, stores data generated as a result of the service provider's business activities. Access to this data is restricted, with viewing permissions set for users belonging to the service provider and external users not belonging to the service provider, depending on the type of data.
[0040] SNS Server 97 provides services that facilitate interaction between users, such as a service that allows users to mutually view each other's posts. For example, there are services that can be viewed via internet search even without a user account on the SNS, and services that allow users with a user account to view each other's posts.
[0041] <1.2 Functional Configuration of Server 20> Figure 2 shows the configuration of server 20. As shown in Figure 2, server 20 functions as a communication unit 201, a storage unit 202, and a control unit 203.
[0042] The communications unit 201 performs processing to enable the server 20 to communicate with external devices.
[0043] The memory unit 202 stores various databases, including a user database 211, a legal content database 212, a content usage history database 213, a graph-type data structure 214, an abbreviation dictionary database 215, a prompt database 218, and a legal terminology database 219.
[0044] User database 211 is a database that manages information for each user. Further details will be provided later.
[0045] Legal Content Database 212 is a database that holds information on legal content. Further details will be provided later.
[0046] Content Usage History Database 213 is a database of user usage history for legal content. Further details will be provided later.
[0047] The graph-type data structure 214 illustrates a graph-type data structure in which reference relationships are defined for a portion of each content.
[0048] The Abbreviation Dictionary Database 215 is a database that manages the rules for using abbreviations when referring to laws and regulations. Further details will be provided later.
[0049] The prompt database 218 is a database that manages prompt templates to be sent to the server 95 of the large-scale language model service.
[0050] Legal Terminology Database 219 is a database of legal terms. Legal terms include, for example, the names of various legal issues as defined in laws and regulations ("claims," "debts," "damages," "risk allocation," etc.), and terms used in case law.
[0051] The control unit 203 is realized when the processor 29 reads a program stored in the memory unit 202 and executes instructions contained in the program. By operating according to the program, the control unit 203 performs the functions shown as the reception control module 2041, transmission control module 2042, user management module 2043, question processing module 2044, LLM utilization module 2045, data structure definition module 2046, and content presentation module 2047.
[0052] The receive control module 2041 controls the process by which the server 20 receives signals from external devices according to a communication protocol.
[0053] The transmission control module 2042 controls the process by which the server 20 transmits signals to external devices according to a communication protocol.
[0054] The user management module 2043 is a module for managing information for each user using System 1. Specifically, the user management module 2043 accepts registration of each user's information and updates the user database 211.
[0055] The question processing module 2044 is a program module that, in response to a question entered by a user, searches for content managed in the legal content database 212 based on a graph-type data structure 214 and generates an answer to the question.
[0056] The LLM utilization module 2045 is a program module that generates prompts and sends the generated prompts to the server 95 of the large-scale language model service, thereby generating responses to the prompts.
[0057] The data structure definition module 2046 is a program module that defines a graph-type data structure 214 based on the content managed in the legal content database 212.
[0058] The content presentation module 2047 is a program module that provides users with an operation screen for viewing content managed by various devices such as the graph-type data structure 214, the legal content database 212, the case law search service server 91, the legal search service server 92, the book browsing service server 93, and the information media service server 94, and processes the content to be viewed according to the user's operations.
[0059] <1.3 Configuration of Terminal 10> Figure 3 shows the configuration of terminal 10.
[0060] As shown in Figure 3, terminal 10 includes multiple antennas (antenna 111, antenna 112), communication units corresponding to each antenna (first communication unit 120, second communication unit 121), an input device 130 (including a touch-sensitive device 131), a display 132, an audio processing unit 140, a microphone 141, a speaker 142, a position information sensor 150, a camera 160, a motion sensor 170, a storage unit 180, and a control unit 190. Terminal 10 also has functions and configurations not specifically shown in Figure 3 (for example, a battery for maintaining power, a power supply circuit for controlling the supply of power from the battery to each circuit, etc.). As shown in Figure 3, each block included in terminal 10 is electrically connected by a bus or the like.
[0061] Antenna 111 radiates signals emitted by terminal 10 as radio waves. Antenna 111 also receives radio waves from space and provides the received signals to first communication unit 120.
[0062] Antenna 112 radiates signals emitted by terminal 10 as radio waves. Antenna 112 also receives radio waves from space and provides the received signals to the second communication unit 121.
[0063] The first communication unit 120 performs modulation and demodulation processing, etc., for the terminal 10 to transmit and receive signals via the antenna 111 in order to communicate with other wireless devices. The second communication unit 121 also performs modulation and demodulation processing, etc., for the terminal 10 to transmit and receive signals via the antenna 112 in order to communicate with other wireless devices. The first communication unit 120 and the second communication unit 121 are a communication module that includes a tuner, an RSSI (Received Signal Strength Indicator) calculation circuit, a CRC (Cyclic Redundancy Check) calculation circuit, a high-frequency circuit, etc. The first communication unit 120 and the second communication unit 121 perform modulation and demodulation, frequency conversion, etc., of the wireless signals transmitted and received by the terminal 10, and provide the received signal to the control unit 190.
[0064] The input device 130 has a mechanism for receiving user input operations. Specifically, the input device 130 is configured as a touchscreen and includes a touch-sensitive device 131. The touch-sensitive device 131 receives user input operations of the terminal 10. The touch-sensitive device 131 detects the user's contact position with the touch panel, for example, by using a capacitive touch panel. The touch-sensitive device 131 outputs a signal indicating the user's contact position detected by the touch panel to the control unit 190 as an input operation.
[0065] The display 132 displays data such as images, videos, and text in accordance with the control of the control unit 190. The display 132 is implemented by, for example, an LCD or an organic EL display.
[0066] The audio processing unit 140 modulates and demodulates the audio signal. The audio processing unit 140 modulates the signal received from the microphone 141 and provides the modulated signal to the control unit 190. The audio processing unit 140 also provides the audio signal to the speaker 142. The audio processing unit 140 is implemented, for example, by an audio processing processor. The microphone 141 receives an audio input and provides the audio signal corresponding to that audio input to the audio processing unit 140. The speaker 142 converts the audio signal received from the audio processing unit 140 into sound and outputs the sound to the outside of the terminal 10.
[0067] The location information sensor 150 is a sensor that detects the location of the terminal 10, and is, for example, a GPS (Global Positioning System) module. A GPS module is a receiving device used in a satellite positioning system. In a satellite positioning system, signals are received from at least three or four satellites, and the current location of the terminal 10, which is equipped with a GPS module, is detected based on the received signals.
[0068] Camera 160 is a device that receives light using a photodetector and outputs it as an image. Camera 160 is, for example, a depth camera that can detect the distance from camera 160 to the object being photographed.
[0069] The motion sensor 170 includes an acceleration sensor, an angular velocity sensor, etc., and detects the movement of the terminal 10.
[0070] The storage unit 180 is composed of, for example, flash memory and stores data and programs used by the terminal 10. The various types of information stored in the storage unit 180 will be described later.
[0071] The control unit 190 controls the operation of the terminal 10 by reading the program stored in the memory unit 180 and executing the instructions contained in the program. The control unit 190 is, for example, an application processor. By operating according to the program, the control unit 190 performs the functions of an operation reception unit 191, a transmission / reception unit 192, a data processing unit 193, a notification control unit 194, and a memory control unit 195.
[0072] The operation reception unit 191 processes input operations from the user to an input device such as a touch-sensitive device 131. Based on the coordinate information of the touch-sensitive device 131 where the user's finger or the like has made contact, the operation reception unit 191 determines the type of operation, such as whether the user's operation is a flick operation, a tap operation, or a drag (swipe) operation.
[0073] The transmitting / receiving unit 192 performs processing to enable the terminal 10 to send and receive data with an external device such as a server 20 in accordance with a communication protocol.
[0074] The data processing unit 193 performs calculations on the data received as input by the terminal 10 according to the program and outputs the calculation results to memory or other locations.
[0075] The notification control unit 194 performs the following processes: displaying the display image on the display 132, outputting sound to the speaker 142, and generating vibrations.
[0076] The memory control unit 195 controls the storage of data to the memory unit 180.
[0077] The various types of information stored by the memory unit 180 will now be explained. In a given situation, the memory unit 180 stores various types of information, such as user information 181.
[0078] User information 181 is information about a user who uses the services of server 20.
[0079] <2 Data Structure> Figure 4 shows the data structure of the user database 211. The user database 211 includes the following fields: "User ID", "Name", "Email Address", "Business ID", "Department", "Position", "Date of Employment", "Date of Resignation", and "Qualifications Held".
[0080] The "User ID" field is information that identifies each user.
[0081] The "Name" field contains information indicating the user's name.
[0082] The "Email Address" field contains the user's email address information for contact purposes.
[0083] Specifically, the item "email address" contains email address information, which serves as user identification information for accepting user logins to the services provided by server 20.
[0084] The "Business ID" field is information that identifies the organization to which the user belongs.
[0085] Specifically, the "Business ID" field is information that identifies the organization to which the user belongs, and the following types of organizations are possible: • Business company ·Law firm The "Department" field contains information about the department to which the user belongs.
[0086] Specifically, the item "Department" may include the following information as the department to which the user belongs: • Departments such as the legal department that are expected to perform legal duties. • Departments such as business divisions and sales divisions that do not plan to have someone dedicated to legal affairs. The "Job Title" field contains information about the user's job title.
[0087] Specifically, the "Job Title" field may include the following information regarding the user's job title: • Positions with decision-making authority (e.g., management) • A position that does not have decision-making authority but demonstrates expertise and a role in supporting the management department. • Not holding any official position. The "Date of Joining" field contains information about the date the user became a member of the organization.
[0088] Specifically, the item "Date of Joining the Company" may include the following information: • Start date if you are already a member • A date when they are scheduled to join the organization but are not yet affiliated (They may have been hired but not yet have access to the organization's information because their start date has not yet arrived). • While affiliated with the company, the start date is not yet confirmed (the start date may still be being finalized). The "Resignation Date" field contains information about the date on which the user left the organization and resigned.
[0089] Specifically, the item "Retirement Date" may include the following information: • Date of resignation (Upon resignation, access to organizational information may be lost) • Currently employed The "Qualifications Held" item contains information indicating the qualifications held by the user.
[0090] Specifically, the item "Qualifications Held" may include the following qualifications, which serve as information to identify the qualifications held by the user. • National qualifications that demonstrate expertise, such as lawyer, patent attorney, and certified public accountant. • Qualifications recognized by businesses, general incorporated associations, and other organizations as proof of expertise. Figure 5 shows the data structure of the legal content database 212.
[0091] The "Content ID" field is information that identifies each piece of legally relevant content.
[0092] The "Type" field contains information about the type of legal content.
[0093] The "Type" item may include the following as types of legal content: ·Laws Case law • Books on law • Websites containing legal information The item "Source" refers to the information about the source from which the content is provided.
[0094] The item "Sources" may include the following as sources: • Server 92 for the legal search service • Server 91 of the case law search service • Server 93 of the book browsing service that provides access to legal books. • Information media service server 94, which contains legal information. The "Title" field contains information about the name of the content.
[0095] The "Title" field may include the following as the name of the content: In the case of laws and regulations, the name of the law, the article number, the headings attached to the articles, etc. • In the case of precedents, the case number, the name of the case, etc. • In the case of books on law, the title of the book, etc. • In the case of legal guidelines, the name of the guideline, etc. The item "Content Tags" contains information about the names of the tags assigned to the content.
[0096] The item "Content Tags" may include the following as information about tags assigned to the content: • The large-scale language modeling service server 95 extracts tags representing the content by summarizing data indicating the content using keywords, a certain number of characters, etc. • Server 20 maintains a list of tags to be assigned to the entire content in advance, and based on the content, the server 95 of the large-scale language model service identifies tags from the tags included in the list. The item "Data" contains information about the content's data file.
[0097] The item "Tagged Location" refers to information about the location where tags are set, using the names of laws, precedents, legal books, legal guidelines, etc., as tags, and where these tags are set at any point in the data that makes up the content.
[0098] The item "Tagging Location" may include the following as locations where tags are set: In the case of laws and regulations, tags may be set in association with specific clause numbers, etc. In the case of case law, the judgment can be divided into multiple blocks based on the paragraphs, headings, arguments of both parties, and the sections where the court's judgments on each issue are described, and tags can be set at any point in the text within those blocks. • In the case of legal books, the book may be divided into multiple blocks based on chapters, headings, paragraphs, etc., and tags may be set at any point in the text within those blocks. • In the case of legal guidelines, the guideline text may be divided into multiple blocks in the same manner as above, and tags may be set at any point in the text within those blocks. For example, if a book on law contains names such as the names of laws and regulations, the names of articles, or the numbers of precedents, server 20 may set tags on the parts where such names are written.
[0099] The "Name Tag" field contains information about the name of the tag that has been set.
[0100] The item "Name Tag" may include the following as tags attached to the name: • Name of the law, article number, etc. • Name of the case number, etc. of the precedent • Titles of legal books • Name of the legal guidelines The "Name Tag Score" item contains information about the evaluation results for the tags that indicate the name.
[0101] More specifically, the "Name Tag Score" item is an evaluation of tags based on the history of content viewing in various services provided by the legal search service server 92, the case law search service server 91, the book browsing service server 93, the information media service server 94, etc., where tags and content are associated.
[0102] Server 20 may use the results of the evaluation of name tags in this manner to determine the priority of tags corresponding to search terms when a user performs a search (for example, the higher the evaluation of a tag, the higher the priority given to it as a tag corresponding to the user's search), and search for legal content based on the tags with the highest priority.
[0103] The "Name Tag Associated Content ID" field is information that identifies the content whose reference relationship is defined based on where the tag is set, when the same tag is set in multiple content items for the tag set in the "Name Tag" field.
[0104] In the illustrated example, the item "Name Tag Associated Content ID" shows the result of defining reference relationships between content containing the name of a specific law, where each piece of content related to multiple laws contains the name of that specific law.
[0105] The item "Location of Issue Tagging" refers to information about the location where legal issues are tagged and set as tags in arbitrary parts of the data that make up the content.
[0106] The item "Issue Tag" contains information about the names of the tags assigned to the issues.
[0107] The item "Issue Tag Score" contains information on the evaluation results for tags that indicate issues.
[0108] More specifically, the "Issue Tag Score" item is determined by classifying legal books by issue on server 93 of the book browsing service, etc., accepting operations to search for books to be viewed based on issues, and evaluating the issue tags based on the user's browsing history of legal books in response to the search results.
[0109] The item "Issue Tag Associated Content ID" is information that identifies the content that defines the reference relationship based on where the issue tag is set, when the same issue tag is set in multiple pieces of content, as set in the item "Issue Tag".
[0110] Figure 6 shows the data structure of the content usage history database 213.
[0111] The "Search ID" field is information that identifies each search performed to access legal content.
[0112] The "Search User ID" field identifies the user who searched for legal content.
[0113] The item "Search User ID" may be associated with the item "User ID" in user database 211.
[0114] The "Search Terms" field contains information about the search terms used to find the content.
[0115] More specifically, the "search term" field is a word specified by the user, or a sentence entered by the user in natural language as a question.
[0116] The item "Search Issue Tag" contains information about the names of the issue tags assigned to the search terms entered by the user.
[0117] More specifically, the item "Search Issue Tag" is either a tag for an issue extracted by matching the search term entered by the user with a list of issue tags, or information on a legal issue set by the server 95 of the large-scale language model service, etc.
[0118] The "Search Date and Time" field contains information about the timing of the search.
[0119] The "Source" item is database information that is referenced when a user performs a search for content.
[0120] The item "Information Source" may include the following as databases referenced by the search: • Server 92 of the legal search service • Server 91 of the case law search service • Server 93 for the book browsing service • Information Media Services Server 94 The "Viewed Content ID" field is information that identifies the content viewed by the user.
[0121] The item "Viewed Content ID" may be associated with the item "Content ID" in the legal content database 212.
[0122] The item "Content Viewed Location" contains information about the specific locations where the user viewed content.
[0123] More specifically, the item "Content Viewed" refers to information about the sections (pages; it may also be defined as pages viewed for a certain period of time or longer) within a book related to law.
[0124] Figure 7 shows the data structure of graph-type data structure 214.
[0125] In the example shown in Figure 7, it is indicated that information for a set of text 214A, which includes the relevant section of a specific law contained in document A, is stored in the legal content database 212, associated with document A (e.g., a legal book). Similarly, it is indicated that information for a set of text 214B, which includes the relevant section of a specific law contained in document B, is stored in the legal content database 212, associated with document B (e.g., guidelines), as the name of the set of text 214B.
[0126] In this example, a reference relationship (edge) is defined between parts of document A and document B that contain the same legal provision (document 214A, document 214B). In this way, the data structure definition module 2046 generates a graph-type data structure 214 by defining nodes for the parts of the legal content database 212 that contain the names of laws, precedents, legal books, guidelines, etc., and defining nodes and edges based on these names. The example shown illustrates the definition of a reference relationship between two documents, A and B, but it is not limited to just two.
[0127] <3. Operation (First Embodiment)> Figure 8 shows the process flow for creating a graph database using information from multiple legal content sources.
[0128] In step S821, the data structure definition module 2046 of the server 20 obtains information on legal content and information that can be used as tagging candidates by referring to the legal content database 212, the legal terminology database 219, etc.
[0129] In this way, the server 20 stores information on multiple legal content in the legal content database 212 of the storage unit 202.
[0130] In step S823, the data structure definition module 2046 of server 20 identifies a portion of the legal content that includes information identifying laws and regulations, information identifying precedents, information identifying legal books, and information identifying legal guidelines.
[0131] Thus, in each of the multiple legal content items, the data structure definition module 2046 is defined as follows: Information that identifies laws and regulations, Information that identifies a case, Information identifying legal books, Information identifying legal guidelines, Identify a portion that includes at least one of the following.
[0132] More specifically, the data structure definition module 2046 is: As information identifying the law, at least one of the law's name or provision, Information that identifies a case includes the case number, Information that identifies a legal book includes the title of the legal book, Information that identifies legal guidelines, such as the name of the guideline, It may also be necessary to identify a portion that includes at least one of the following.
[0133] More specifically, the data structure definition module 2046 is: As information identifying the law, at least one of the law's name or provision, Information that identifies a case includes the case number, Information that identifies a legal book includes the title of the legal book, Information that identifies legal guidelines, such as the name of the guideline, It is also possible to identify a certain amount of text as a part that contains at least one of the following.
[0134] In step S825, the data structure definition module 2046 of the server 20 associates the identified portion of the content with the content itself and stores it in the legal content database 212.
[0135] In this way, the data structure definition module 2046 stores the portion identified in step S823 in the storage unit 202 in association with the content related to that identification.
[0136] Here, the data structure definition module 2046 may evaluate a portion of the content based on the portion associated with the content, according to the usage of the search results in a service that accepts operations to search for content (for example, the server 91 for the case law search service, the server 92 for the legal search service, the server 93 for the book browsing service, the server 94 for the information media service, etc.). For example, regarding tags associated with content, the tags may be evaluated according to the search history of those tags in the book browsing service mentioned above (e.g., the evaluation value is increased for frequently searched tags and for tags that lead to books being viewed through searches). The data structure definition module 2046 may store the evaluation results of the evaluated portion in the legal content database 212 in association with the content (for example, the item "Name Tag Score"). For example, when the server 20 identifies a tag corresponding to a user's question in a graph-structured database, it may select a tag according to its evaluation value (e.g., prioritizing tags with high evaluation values as tags corresponding to the user's question), and then search the graph-structured database using the selected tag. This improves the response time of tag-based searches using a graph structure compared to searching each piece of content in the legal content database 212 using keywords, while further enhancing the accuracy of search results (presenting legal content suitable for the user).
[0137] By processing steps S823 and S825 described above, it is possible to identify the locations where the names of laws, precedents, books, guidelines, etc., are mentioned, and to assign tags based on these names to the content.
[0138] In step S827, the data structure definition module 2046 of the server 20 extracts identical portions of multiple content items. The data structure definition module 2046 defines reference relationships for the portions of each content item that have been extracted as containing identical items.
[0139] More specifically, the data structure definition module 2046 may define reference relationships by extracting content that has a part of the same type, and by extracting content with the same name relating to the identified part.
[0140] For example, if the same case law number is described in both the first and second ebooks, the section in the first ebook where the case law number is described and the section in the second ebook where the case law number is described are extracted as being of the same type. The data structure definition module 2046 defines reference relationships between the extracted sections in the first and second ebooks where the above-mentioned case law number is described.
[0141] In this way, the data structure definition module 2046 constructs a graph-structured database by extracting content with similar identified parts from multiple content items and defining reference relationships between the content items for each extracted content item. For example, the data structure definition module 2046 may store the reference relationships defined for multiple content items in the legal content database 212.
[0142] Here, the data structure definition module 2046 may define reference relationships between portions of text, which are fixed amounts of text, from multiple contents.
[0143] This allows for the generation of a graph-type data structure that defines reference relationships between multiple pieces of content when those pieces of content contain the same names of laws, precedents, books, and guidelines. For example, if a graph-type data structure is generated based on the name of a law, searching for that law's name can display search results from multiple sources, such as the original text of the law, corresponding precedents, books that mention the law, and guidelines that mention the law. Because the search is performed using a graph-type data structure, the responsiveness of the search can be further improved compared to sequentially searching these sources using the keywords mentioned above.
[0144] In step S829, the data structure definition module 2046 of the server 20 outputs information about the reference relationships defined for multiple contents to the terminal 10. Alternatively, the data structure definition module 2046 may choose not to output the information to the terminal 10 after defining the reference relationships between the contents, or it may present the information about the reference relationships defined between the contents upon request from the terminal 10.
[0145] In step S811, terminal 10 displays information indicating the content reference relationships.
[0146] Figure 9 shows the processing flow using a graph database related to laws.
[0147] In step S921, the question processing module 2044 of the server 20 receives input of a legal question and outputs an operation screen to the terminal 10 that displays the answer.
[0148] As described above, Server 20 holds a data structure for information on legal content (legal content database 212, graph-type data structure 214). Server 20 stores information on multiple legal content items in its storage unit 202. In the storage unit 202, Server 20 stores the results of defining reference relationships between content items for each of the multiple content items, where at least one of the following is included: information that identifies laws and regulations, information that identifies precedents, information that identifies legal books, and information that identifies legal guidelines.
[0149] In step S911, terminal 10 receives input from the user regarding legal questions.
[0150] In step S923, the question processing module 2044 of server 20 assigns a tag representing the content of the question to the question entered by the user of terminal 10, with the help of the server 95 of the large-scale language model service. For example, the question processing module 2044 generates a prompt that includes, along with the content of the question, at least one of the following: an instruction to summarize the content of the question entered by the user, an instruction to refer to legal terms such as the legal terminology database 219 regarding the content of the question and assign a tag to it. The question processing module 2044 sends the prompt to the server 95 of the large-scale language model service via the LLM utilization module 2045, and assigns a tag to the question by receiving an output result corresponding to the prompt from the server 95 of the large-scale language model service.
[0151] In step S925, the question processing module 2044 of the server 20 uses the tags assigned to the question to refer to the graph-type data structure 214 and search for tagged content. This reduces the processing required to search for tagged content compared to searching a database of table-type content, and further improves the response speed of search results.
[0152] In step S927, the question processing module 2044 of server 20 obtains the generated answer by instructing the server 95 of the large-scale language model service to generate an answer to the question based on the search results. Specifically, the question processing module 2044 generates a prompt that includes the search results (legal content searched based on tags assigned to the question) retrieved by referring to the graph-type data structure 214 and the question entered by the user in steps S911 and S923, and sends the generated prompt to the server 95 of the large-scale language model service, thereby receiving the answer generated by the server 95 of the large-scale language model service.
[0153] Furthermore, the question processing module 2044 may, in response to a legal question input from the user, output the search results of the content retrieved in step S925 to the user's terminal 10 without further generating an answer to the question by the server 95 of the large-scale language model service. This allows the system to respond to the user's question with legal content retrieved using the graph-type data structure 214, reducing the processing required for the search and enabling a faster response of search results compared to sequentially searching each data in a table-type content database.
[0154] Thus, the server 20 has a data structure 214 for information on legal content, which is used in the process of responding to a search request to retrieve information on multiple legal content items. This structure identifies a portion (tag) in the database that corresponds to the search request and responds to the search request based on the reference relationships between the content items for the identified portion.
[0155] Here, as described above, the data related to the data structure is used in a process where, in response to a search request that retrieves a portion of multiple contents that corresponds to the question, the server 95 of the large-scale language model service receives a question input from the user and generates an answer to the question. The server 95 retrieves a portion of multiple contents that corresponds to the question and responds with search results for multiple contents based on the reference relationships between the contents for that portion. The large-scale language model generates an answer to the question based on the search results for multiple contents and the question from the user.
[0156] In step S929, the server 20 outputs an answer to the question to the terminal 10 based on the tag corresponding to the question.
[0157] In step S913, terminal 10 displays the answer to the question.
[0158] <4. Screen Example (First Embodiment)> Figure 10 is an example of an operation screen that displays the results of generating answers to user questions while referring to a graph-type data structure 214 related to the law.
[0159] The operation screen 1000 displays the results of a search for laws, precedents, legal books, guidelines, etc., in response to a question, while referring to the graph-type data structure 214.
[0160] The operation screen 1000 corresponds to each process, such as step S921 in Figure 9, step S1121 in Figure 11, and step S1321 in Figure 13.
[0161] The question designation unit 1002 is an operating component that receives the user's designation of the question content.
[0162] More specifically, the question specification unit 1002 accepts questions from the user in natural language.
[0163] The transmission operation unit 1004 is an operation component that receives an operation to generate an answer to a question entered by the user.
[0164] In the illustrated example, the transmission operation unit 1004 transmits the question entered in the question specification unit 1002 to the server 20 in response to the user's operation.
[0165] The account display area 1006 is an area that displays account information for users who use the services provided by server 20.
[0166] In the illustrated example, the account display area 1006 also displays the user's affiliation, but it may also display the user's account pricing plan to make it easier to recognize the differences in available features depending on the pricing plan (for example, limitations on the number of questions that can be entered, and limitations on the scope of laws, precedents, legal books, guidelines, etc. that can be referenced).
[0167] The answer display area 1008 is an area that displays the answer generated by the server 20 in response to the question entered by the user.
[0168] The question content display area 1010 is an area that displays the content of the question entered by the user in the question specification unit 1002.
[0169] In the illustrated example, the question content display area 1010 displays the question for which the server 20 has generated an answer, making it easy to recognize which question the answer is for. For example, if an answer is generated each time a question is entered, even when scrolling back through past questions, it is easy to confirm which question the answer is for.
[0170] The question tag display area 1012 is an area that displays the tags assigned to the entered question.
[0171] The question tag display area 1012 corresponds to the processing shown in step S923 of Figure 9.
[0172] The legal information display area 1014 is an area that displays the legal information retrieved by referring to the graph-type data structure 214 in response to the entered question.
[0173] In the illustrated example, the legal information display area 1014 displays the name of the law and the article number, and also includes a "View Details" button. In response to the user's operation of the "View Details" button, terminal 10 may display the details of the law (for example, the original text of the article) within the operation screen 1000 (for example, by expanding the area of the legal information display area 1014), or it may open a separate window to allow the user to view the details of the law.
[0174] The case law display area 1016 is an area that displays case law retrieved by referring to the graph-type data structure 214 in response to the entered question.
[0175] In the illustrated example, the case law display area 1016 displays the case law number and includes a "View Details" button. Terminal 10 may also display the details of the searched case law (such as an overview of the case, as described later) on the operation screen 1000 or elsewhere, in response to the user's operation of the "View Details" button.
[0176] The legal books display area 1018 is an area that displays legal books found by referring to the graph-type data structure 214 in response to the entered question.
[0177] In the illustrated example, the legal book display area 1018 presents the user with details of the searched legal book (for example, after receiving a user action such as "View Details" shown in the legal information display area 1014, etc.). The legal book display area 1018 displays an overview of the book, such as the title, publisher, and author, while the book details display area 1020, etc., described later, displays an excerpt from the book's text.
[0178] The book details display area 1020 is an area that displays a portion of a legal book extracted according to the searched tags.
[0179] In the illustrated example, the book details display area 1020 displays a portion of the information displayed in the evidence display area 1022, which will be described later, as well as the surrounding text. For example, in a legal book, the text may be divided into multiple blocks by chapters, paragraphs, etc., and the blocks containing the text indicated in the "Tagged Locations" item of the legal content database 212 may be extracted.
[0180] The evidence display area 1022 is an area that displays the location in legal textbooks where the corresponding information is found.
[0181] In the illustrated example, the evidence display area 1022 highlights the text shown in the "Tagged Location" item of the legal content database 212 (for example, by highlighting it with a marker or making it bold), making it easy to identify the location of the text corresponding to the tag.
[0182] The guideline display area 1024 is an area that displays the guidelines found by referring to the graph-type data structure 214 in response to the entered question.
[0183] In the illustrated example, the guideline display area 1024 displays the name of the guideline and includes a "View Details" button. Terminal 10 may also display the details of the case within the operation screen 1000 or elsewhere to allow the user to view the details of the searched guideline in response to the user's operation of the "View Details" button.
[0184] <5 Operation (Second Embodiment)> Next, a second embodiment will be described.
[0185] When a legal consultation is received, it can sometimes be difficult to identify the relevant law from the consultation content alone. This is because legal provisions are often defined using abstract language, and even if search terms are defined based on the consultation content, the results may not correspond to the wording of the legal provision. Therefore, the usual procedure involves searching for cases that correspond to the consultation content and then identifying the relevant law and provision from those cases. For the person providing the consultation, especially in an unfamiliar field, this investigation procedure can be time-consuming.
[0186] Therefore, in the second embodiment, we will describe a technology that accepts natural language input and presents relevant laws and regulations.
[0187] Specifically, in the second embodiment, Accepting questions from users in natural language, The server 95 of the large-scale language model service will be used to identify which law the question pertains to (for example, a specific law such as the Companies Act). Legal books and web articles related to law are pre-tagged (the tagging of such legal books may be performed by the server 95 of the large-scale language model service), and the content of legal books, etc., is searched for those whose content matches the tags (searching by keyword, vectorizing the question text and the content of legal books, etc., and then searching for vectors of law-related content based on the vectors), To obtain laws and regulations cited within the text of searched content, To do so.
[0188] Figure 11 shows the process flow for identifying the relevant laws and regulations in response to the user's question and presenting those laws and regulations to the user.
[0189] In step S1121, the question processing module 2044 of the server 20 receives input of a legal question and outputs an operation screen to the terminal 10 that displays the answer.
[0190] Server 20 stores information on multiple legal content in its storage unit 202, including at least one of the following: legal books, legal guidelines, and articles from legal information media. Storage unit 202 stores information on multiple legal books, and each of these legal books is associated with a tag indicating its content.
[0191] For example, server 20 may set tags indicating the content of legal books by referring to the classification set for each book in server 93 of the book browsing service (classifications entered by a human, such as the service operator or the book publisher), or server 95 of the large-scale language modeling service may assign tags indicating the content of legal books (for example, by sending a prompt to server 95 of the large-scale language modeling service to summarize the content of legal books).
[0192] In step S1111, terminal 10 receives input from the user regarding legal questions.
[0193] In step S1123, the question processing module 2044 of server 20 assigns a tag corresponding to the question entered by the user, using the server 95 of the large-scale language model service.
[0194] The question processing module 2044 assigns tags to questions by providing a prompt to the server 95 of the large-scale language model service that includes instructions to assign legal tags to the questions. The module then obtains information on the tag assignment results output by the server 95 of the large-scale language model service, and assigns tags to the questions based on the obtained information on the tag assignment results.
[0195] Alternatively, the question processing module 2044 may assign tags to questions by inputting the questions entered by the user in step S1111 into a trained model that has been trained to output tags to questions based on the results of tagging questions.
[0196] In addition, the question processing module 2044 also maintains a list of tags to be assigned to questions based on keywords contained in the question text, and may assign tags to the question by comparing the question entered in step S1111 with this list.
[0197] In step S1125, the question processing module 2044 of the server 20 uses the tags assigned to the question to refer to the database and search for the tagged content.
[0198] In this way, the question processing module 2044 searches for information from multiple legal content sources based on the entered question. Here, the question processing module 2044 searches for information from at least one of the following: legal books, legal guidelines, or articles from legal information media. The question processing module 2044 may also search for legal books based on tags assigned to the question and tags associated with legal books.
[0199] Furthermore, in a service that allows users to view multiple legal books (the server 93 for the book viewing service), the scores of tags associated with legal books are determined (legal content database 212, content usage history database 213). The question processing module 2044 may search for legal books based on the scores of the tags associated with them. For example, when searching for tags in the database that correspond to the tags assigned to the question, tags with higher scores may be given priority in the search results.
[0200] As described above, the server 20 stores information on multiple legal content items in the storage unit 202. In the storage unit 202, for each of the multiple content items, the results of defining reference relationships between content items that are of the same type, including at least one of the following: information that identifies laws and regulations, information that identifies precedents, information that identifies legal books, and information that identifies legal guidelines, are stored as a legal data structure (legal content database 212, graph-type data structure 214). The question processing module 2044 retrieves information on multiple legal content items related to a question by referring to the legal data structure.
[0201] In step S1127, the question processing module 2044 of the server 20 identifies information about laws and regulations cited in the search results content. For example, the question processing module 2044 may search for content based on the tags assigned to the question and information associated with legal books held in the legal content database 212 (such as the "Title" and "Content Tag" items in the legal content database 212), and identify information about laws and regulations included in the searched content (for example, the "Tagged Location" and "Name Tag" items in the legal content database 212).
[0202] In step S1129, the question processing module 2044 of the server 20 presents the identified legal information to the user of the terminal 10.
[0203] In step S1113, terminal 10 displays information on laws and regulations corresponding to the question.
[0204] <6. Screen Example (Second Embodiment)> Figure 12 shows an example of an operation screen that presents the user with the relevant laws and regulations identified in response to the user's question.
[0205] The legal details display area 1026 is an area that displays a portion of the legal provisions extracted according to the searched tags.
[0206] In the illustrated example, the legal details display area 1026 displays the legal provisions corresponding to the searched tag. For example, the server 20 may display only the provisions corresponding to the question, rather than all of the provisions contained in the legal text.
[0207] Case details display area 1028 is an area that displays summaries of multiple case precedents extracted according to the searched tags.
[0208] In the illustrated example, the case details display area 1028 displays the results of generating summaries of multiple case precedents, which were retrieved by referring to the graph-type data structure 214 according to the tags corresponding to the question, using the server 95 of the large-scale language model service (for example, the results of generating summaries of case precedents by prompting the generation of an overview of the case and an overview of the conclusion within a predetermined number of characters).
[0209] <7 Operation (Third Embodiment)> Next, a third embodiment will be described.
[0210] Typically, when conducting legal research, searching for legal content using keywords will result in a list of relevant articles. However, especially in areas unfamiliar to the researcher, it can be difficult to determine the importance of the listed articles (e.g., legal books), requiring the researcher to meticulously read through each individual article, which is a significant burden.
[0211] Therefore, in the third embodiment, the system accepts a question in natural language and presents the user with issues related to the entered question. For each issue, the system presents the user with organized content such as legal books.
[0212] Specifically, in the third embodiment, To construct a graph-type data structure for the issues (for example, the issues of each piece of legal content may be identified by the server 95 of the large-scale language model service. For each issue, legal content such as laws, precedents, legal books, and guidelines are associated and constructed as a graph-type data structure), Accepting questions from users in natural language, Based on a database constructed to hold information on legal content based on the issue at hand, the system identifies the issue corresponding to the question (for example, identifying the issue from keywords contained in the question, or vectorizing the question and searching for the issue based on the vector of legal content). Based on the issues corresponding to the questions and a graph-type data structure for those issues, the response will be to provide relevant laws, precedents, legal books, guidelines, etc. To do so.
[0213] Figure 13 illustrates the process flow for generating a database that identifies legal issues from information on multiple legally related content.
[0214] In step S1321, the data structure definition module 2046 of the server 20 obtains information on legal content by referring to the legal content database 212, etc.
[0215] Thus, the server 20 stores information on multiple legal content in its storage unit 202. Specifically, the server 20 stores information on multiple legal content in its storage unit 202, including at least one of the following: information on laws and regulations, information on precedents, information on legal books, and information on legal guidelines.
[0216] In step S1323, the data structure definition module 2046 of the server 20 generates a prompt that includes instructions to identify legal issues corresponding to the content of the legally relevant content. The data structure definition module 2046 may also generate a prompt that includes instructions to extract the portion of the content that formed the basis for identifying the legal issues corresponding to the content.
[0217] In step S1325, the data structure definition module 2046 of server 20 provides the generated prompt to the server 95 of the large-scale language model service, thereby obtaining the results of identifying legal issues regarding the content from the server 95 of the large-scale language model service.
[0218] In this way, the data structure definition module 2046 provides the content for prompt generation and the prompt generated for that content to the server 95 of the large-scale language model service, and by receiving the output results from the server 95 of the large-scale language model service, it obtains the results of identifying the legal issues for each piece of content.
[0219] In step S1327, the data structure definition module 2046 of the server 20 updates the legal content database 212 by associating the content with the identified legal issues.
[0220] Here, the data structure definition module 2046 may associate the extracted portion of the content with the identified legal issues and store them in the legal content database 212.
[0221] Furthermore, based on the legal issues associated with the content, the system may evaluate the legal issues in accordance with the usage of search results in a service that accepts searches for content (for example, a book browsing service server 93), and store the evaluation results of the legal issues in association with the content in a legal content database 212 (item "issue tag score").
[0222] In step S1329, the data structure definition module 2046 of the server 20 identifies a portion of the content relating to the same issue for multiple content items, defines a reference relationship for the identified portion, and updates the database. The data structure definition module 2046 outputs the information of the reference relationships defined for the issues of the multiple content items to the terminal 10.
[0223] Thus, the data structure definition module 2046 may define a graph-type data structure 214 by identifying parts of multiple content related to the same issue and defining reference relationships between the identified parts of the multiple content.
[0224] Furthermore, it is not necessary to update the database defining the reference relationships and then output the update to terminal 10; instead, the system may respond to a request from terminal 10 by indicating the status of the defined reference relationships in the database.
[0225] In step S1311, terminal 10 displays information indicating the reference relationships of content related to the issue.
[0226] Figure 14 shows the process flow for searching for content related to legal issues that address a user's question.
[0227] In step S1421, the question processing module 2044 of the server 20 receives input of a legal question and outputs an operation screen to the terminal 10 that displays the answer.
[0228] Server 20 stores information on multiple legal content items in its legal content database 212 and graph-type data structure 214 in its storage unit 202, and stores information on legal issues associated with each of the multiple content items.
[0229] The graph-type data structure 214 is used to receive legal questions from users, analyze the input questions to identify the legal issues corresponding to the questions, and, based on the identified legal issues, search for legal content related to those issues by referring to the information on legal issues stored in the graph-type data structure 214.
[0230] In the graph-type data structure 214, each of the multiple content items is configured to store information about the legal issue associated with the content and a portion of the content corresponding to that legal issue. In the graph-type data structure 214, reference relationships are defined between multiple content portions relating to the same issue. The graph-type data structure 214 is used in processes that search for multiple content items for which reference relationships are defined regarding a specified issue, based on those reference relationships.
[0231] In step S1411, terminal 10 receives input from the user regarding legal questions.
[0232] In step S1423, the question processing module 2044 of the server 20 identifies the legal issues corresponding to the input question by analyzing the input question. More specifically, the question processing module 2044 generates a prompt for the input question that includes instructions to identify the legal issues related to the question (for example, the prompt may include instructions to refer to a list of issues), and provides the generated prompt to the server 95 of the large-scale language model service. The server 95 then receives the output result of the large-scale language model service and identifies the legal issues related to the question.
[0233] In step S1425, the question processing module 2044 of the server 20 uses the issue identified for the question to refer to the legal content database 212 and the graph-type data structure 214 to search for content tagged with that issue.
[0234] As described above, the question processing module 2044 searches for information on law-related content associated with the identified issue by referring to legal issue information stored in the storage unit 202 based on the issue corresponding to the identified question.
[0235] In step S1427, the question processing module 2044 of the server 20 outputs the content search results using issue tags to the terminal 10, including a part corresponding to the issue for each of these contents (the item "issue tagged portion" in the legal content database 212).
[0236] Here, the question processing module 2044 may provide a prompt including content search results using issue tags and the user's question accepted in steps S1411 and S1423 to the server 95 of the large language model service, thereby causing the server 95 of the large language model service to generate an answer to the user's question while referring to the content search results using issue tags, and output the generated answer to the terminal 10.
[0237] In step S1413, the terminal 10 displays the search results of law-related content along with the issue corresponding to the question.
[0238] <8 Example of Screen (Third Embodiment)> FIG. 15 is an example of an operation screen for responding, to a user's question, with results obtained by searching various databases based on corresponding legal issues.
[0239] The issue display area 1030 is an area for displaying information of the issue corresponding to the question.
[0240] The issue display area 1030 corresponds to processing such as step S1427 in FIG. 14.
[0241] The issue summary display area 1032 is an area for displaying a summary of search results retrieved as issues corresponding to the question.
[0242] In the illustrated example, the issue summary display area 1032 displays information about the issue retrieved by referring to the graph-type data structure 214 in response to the question, and the results of the answer generated by the large-scale language model service server 95 based on the search results and the question.
[0243] The individual issue display area 1034 is an area that displays the details of one of several issues.
[0244] Reference display area 1036 is the area that displays the source material (a legal book in the illustrated example) that supports the issue at hand.
[0245] In the illustrated example, the reference display area 1036 shows multiple legal books that were found by searching for a tag corresponding to the issue.
[0246] The evidence display area 1038 is an area that displays the section in legal textbooks that contains descriptions corresponding to the issue tags.
[0247] The evidence display area 1040 is an area that displays the section in legal textbooks that contains descriptions corresponding to the issue tags.
[0248] <Fourth Embodiment> In the fourth embodiment, we describe a technology that identifies the laws and regulations corresponding to the searched portion in the content and displays the latest laws and regulations and information on legal amendments, thereby making it easy to refer to the laws and regulations that were assumed at the time the content was published, as well as information on subsequent legal amendments.
[0249] As time passes since publication, content such as books may become outdated as it does not reflect legal revisions or the facts underlying those revisions. When referring to such content, there is a risk that one may not be aware of subsequent legal revisions and therefore not realize that the content is no longer relevant to the present time.
[0250] However, checking whether there have been any legal amendments presents a challenge, as it requires considerable effort, such as researching the history of any past legal amendments.
[0251] Therefore, in the fourth embodiment, we will describe a technology that makes it easy to view content even as time passes since its publication and it becomes outdated, taking into account subsequent legal revisions, etc. This will further reduce the effort required for research.
[0252] <9 Operation (Fourth Embodiment)> Figure 16 shows the process flow for identifying the relevant laws and regulations in the content that are identified as the basis for the answers to the questions, and outputting the latest laws and regulations.
[0253] The outline of each process is as follows:
[0254] Terminal 10: The user enters a question, and Terminal 10 sends that question to Server 20. Server 20: Server 20 generates a provisional answer to the question and searches for relevant content based on that provisional answer. Server 20: Identifies relevant laws and regulations from identified portions of content and retrieves the latest legal amendment information. Server 20: Organize and present to users all information that should be presented to them, including legal information prior to the content's publication date. Terminal 10: Terminal 10 displays the received information to the user and highlights important parts. By integrating and displaying related information on the same screen, the user can easily refer to the content and the latest legal amendments. These processing flows allow users to access the latest legal amendments even when referring to older content, such as when time has passed since publication, thereby reducing the effort required for research.
[0255] Next, we will explain each step.
[0256] In step S1621, the question processing module 2044 of the server 20 presents an operation screen that accepts input of a question to the user.
[0257] Here, the server 20 manages a plurality of contents to be browsed in the storage unit 202 (such as the legal content database 212, the graph-type data structure 214, data of the server 91 of the precedent search service, data of the server 92 of the law search service, data of the server 93 of the book browsing service, and data of the server 94 of the information media service).
[0258] The contents managed in the storage unit 202 may include contents that indicate the publication time of the content. More specifically, the contents managed in the legal content database 212 or the like of the storage unit 202 may include at least any one of books, guidelines, and media articles that indicate the publication time of the content. Here, the contents include those mainly composed of text and images such as books and articles, and also include moving images.
[0259] In the storage unit 202, information of a plurality of law-related contents is held as the plurality of contents to be browsed. More specifically, in the storage unit 202, for each of the plurality of law-related contents, information defining cross-content reference relationships for contents of the same type with respect to a part of the content of law specifying information, precedent specifying information, legal book specifying information, and law-related guideline specifying information may be stored.
[0260] In step S1611, the terminal 10 accepts input of a question from a user. The terminal 10 transmits the input question to the server 20. For example, the server 20 may provide an operation screen provided with a text box to the terminal 10, and accept input of a question in natural language, or may accept the input as keywords. Furthermore, a voice guide may be provided on the operation screen, or input of a question by voice input from the user may be accepted.
[0261] In step S1622, the question processing module 2044 of the server 20 obtains a provisional answer to the question entered by the user by sending a prompt to an external information processing system to generate a provisional answer in the form of an answer statement based on the content of the received question.
[0262] More specifically, the question processing module 2044 generates a prompt that includes the content of the received question and an instruction that outputs the content of the question in answer format, and sends the generated prompt to an external information processing system, thereby receiving output corresponding to the prompt from the external information processing system. As described later, the question processing module 2044 searches multiple contents based on the content of the received answer format output. In the case of books, guidelines, and other contents, some are written in the form of answer sentences. For example, in the case of books and other contents, compared to searching by the content of the question, using the question (e.g., "In the case of ~~, is it possible to ~~?", "What should be noted about ~~?") in the form of an answer sentence ("In the case of ~~, it is possible to ~~", "What should be noted about ~~ is ~~") can make it easier to identify a part of the book content that serves as the basis for the answer corresponding to the content of the question. This makes it easier to identify a part of the content that serves as the basis for the answer when generating an answer to a question, and the quality of the answer to the question can be further improved.
[0263] The question processing module 2044 may also analyze the content of the question entered by the user using morphological analysis or the like, as in step S1622, extract words, search for information sources using the extracted words, or identify a part of the content using vector search based on the content of the question, and then generate an answer to the question based on this information.
[0264] In step S1623, the question processing module 2044 of the server 20 searches the database of each content in the graph-type data structure 214 based on the provisional answer text generated in step S1622. The question processing module 2044 identifies a portion (snippet) of the content corresponding to the question and obtains an answer to the question by sending a prompt containing instructions to an external information processing system to generate an answer based on the identified content and the user's question.
[0265] In this way, the question processing module 2044 identifies a portion of the content corresponding to the question by searching multiple contents in the storage unit 202 based on the content of the question received in steps S1611 and S1622. The question processing module 2044 may also identify a portion of the content corresponding to the question based on the reference relationships between the contents.
[0266] In step S1624, the question processing module 2044 of the server 20 outputs the generated answer to the terminal 10.
[0267] In step S1612, terminal 10 displays the answer along with information about a portion of the content that formed the basis for generating the answer. Terminal 10 accepts an operation from the user to specify the content.
[0268] In step S1625, the content presentation module 2047 of the server 20 extracts relevant laws and regulations by search (e.g., vector search) based on the content of a snippet of the specified content that is identified as the basis for generating the answer.
[0269] Thus, the content presentation module 2047 may extract information on laws and regulations related to a portion of the content by performing a vector search of the legal database (such as the data on the legal search service server 92) based on the portion of the content identified in step S1623. The question processing module 2044 may then present the user with information on laws and regulations related to the extracted portion of the content in step S1626, which will be described later.
[0270] In step S1626, the content presentation module 2047 of server 20 refers to a legal database (such as the legal search service server 92) and obtains the latest amendment information for the relevant laws. The content presentation module 2047 may also obtain legal amendment information and the latest legal provisions from after the publication date of the content specified by the user for viewing. The content presentation module 2047 may also obtain legal information prior to the publication date of the content. The content presentation module 2047 may also obtain legal amendment information depending on whether it is before or after the publication date of the content.
[0271] In step S1627, the content presentation module 2047 of the server 20 presents the user with a screen for viewing the content specified by the user. The content presentation module 2047 outputs information on legal amendments, legal provisions, and the publication date of the content to the terminal 10.
[0272] Thus, the content presentation module 2047 presents the user with information on legal amendments to applicable laws when a specified portion of the content contains legal information. For example, the content presentation module 2047 may present the user of terminal 10 with information on legal amendments to applicable laws when a portion of the content displayed on the content viewing screen contains legal information, regardless of when the content was published.
[0273] Content presentation module 2047 may also present users with information on legal amendments, including the most recent legal amendments, that have been in effect since the publication date of the specified content. This allows users to view the content while referring to information on legal amendments made since the publication date of the content, making it easier to make legal judgments and other decisions in consideration of legal amendments.
[0274] The content presentation module 2047 may also display information about legal amendments and information about the relevant articles of law on the user interface that displays the identified portion.
[0275] The content presentation module 2047 may display information on the publication date of the specified content, information on legal amendments, and information on the relevant legal provisions on the user interface. This allows users to view the content while referring to the information on amended legal provisions and provisions compared to the publication date of the content, making it even easier to make legal judgments and other decisions while taking legal amendments into consideration.
[0276] The content presentation module 2047 may also display information on legal amendments made after the publication date of the specified content, along with information on laws and regulations prior to the publication date, on the user interface. This allows users to view the content while confirming the information on the laws and regulations on which the content is based, and also to view the content while considering information on legal amendments made after the publication date, making it even easier to make legal judgments and other decisions in consideration of legal changes.
[0277] The content presentation module 2047 may display the content on the user interface and highlight a portion of the content that corresponds to the question. This allows the user to easily identify the portion of the content that forms the basis of the answer to the question, making it easier to verify the validity of the answer and facilitating legal judgments.
[0278] In step S1613, terminal 10 displays information received from server 20 (such as a screen for viewing content, legal amendments, and legal provisions) to the user.
[0279] In step S1614, terminal 10 presents the user with a screen for viewing content. Terminal 10 highlights a portion of the content (the portion that forms the basis for the answer to the user's question) and displays information about legal amendments related to that portion.
[0280] More specifically, terminal 10 highlights a portion of the content corresponding to the question, and when it accepts the user's selection of content, it displays information on legal amendments related to the selected portion. On the same screen, terminal 10 displays the content's publication date, legal amendment information, and the legal text.
[0281] <10. Screen Example (Fourth Embodiment)> Figure 17 shows an example screen of a user interface for viewing content, which displays both the content itself and a table of contents summarizing it.
[0282] This also includes a description of a screen example relating to the fifth embodiment.
[0283] The target operation screen 1700 is an operation screen that accepts operations for viewing content.
[0284] The browsing target operation screen 1700 corresponds to each process such as steps S1627 and S1613 in Figure 16, and steps S2021 and S2012 in Figure 20, which will be described later. For example, as shown in Figures 10, 12, and 15, the screen that displays the answer to a question displays the content (such as legal books) that formed the basis for generating the answer, and in response to an operation to specify this content, the browsing target operation screen 1700 that allows viewing of the specified content is displayed. For example, in response to an operation in which the user specifies each document (basis display area 1038, 1040) displayed in the reference document display area 1036 of Figure 15, the browsing target operation screen 1700 is displayed with that document as the content to be viewed.
[0285] The account display area 1702 is the area that displays the user's account. The viewing area 1704 is the area that displays the content to be viewed.
[0286] In the illustrated example, the viewing area 1704 displays a portion of a legal book as content.
[0287] Snippet display area 1706 is the area that displays a portion of the content that is highlighted.
[0288] In the illustrated example, the snippet display area 1706, as shown in Figures 10, 12, and 15, displays the content (such as legal books) that formed the basis for generating the answer, and accepts user input to specify content, displaying the target operation screen 1700. In such cases, the section containing the part that formed the basis for generating the answer is displayed in response to the user's specified operation, and that part is highlighted by a highlighter or similar method. This makes it easy for the user to check the answer to a question, as shown in Figure 10, and to easily refer to the source book or other document while also checking the section that formed the basis.
[0289] The sub-window display area 1708 is an area that displays various information related to the content being viewed.
[0290] The bookshelf addition operation unit 1710 is an operation component that accepts the operation to add a bookmark to the displayed content (book).
[0291] The content cover image display area 1712 is the area that displays the cover image of the content (book) being viewed.
[0292] The content bibliographic information display area 1714 is the area that displays the bibliographic information of the content being viewed.
[0293] More specifically, the content bibliographic information display area 1714 displays information such as the name of the content (title in the case of a book), the edition of the content, the date of publication, the entity that published the content, and the author of the content. In the illustrated example, the system accepts the operation of copying this bibliographic information. This makes it easy for users to cite the content information that serves as the basis when reporting their opinions or other opinions based on the content of a book.
[0294] The table of contents specification unit 1716 is an operating component that accepts the specification of a table of contents as various information of the content to be viewed.
[0295] In the illustrated example, the table of contents designation section 1716 displays tabs to switch between different types of information related to the content being viewed. These tabs include "Table of Contents" to show the table of contents of the content, "Other Related Content" to display other content different from the content being viewed (a book in the illustrated example) (laws, precedents, guidelines, etc. in the illustrated example), "Bookmarks" to display a list of locations selected by the user within the content, and "Binder" to manage various types of content together for the content or a part of the content. In the example in Figure 17, the "Table of Contents" tab is designated and displayed to distinguish it from the other tabs.
[0296] The related information designation unit 1718 is an operating component that accepts the designation of other related content as various types of information for the content being viewed.
[0297] In the illustrated example, the related information designation unit 1718 is shown as being designated by the user in the example of Figure 18.
[0298] The bookmark designation unit 1720 is an operating component that receives a designation to display a list of locations selected by the user within the content being viewed, as various information about the content being viewed.
[0299] The binder designation section 1722 is an operating component that accepts designations to register content or a part of content in a "binder" that manages various types of content together.
[0300] In the illustrated example, the binder designation unit 1722 accepts the operation of registering the content itself, or a part of the content (in the case of a book), in association with a "binder" whose name can be arbitrarily set by the user. This allows, for example, if the user sets up a "binder" corresponding to a particular issue, the user can easily check the content or a part of the content registered in the "binder" when organizing their views based on that issue, making it easier to review the content again.
[0301] The detail display area 1724 is an area that displays detailed information related to the content being viewed.
[0302] In the illustrated example, the detail display area 1724 displays the table of contents in Figure 17, highlighting the position in the table of contents that corresponds to the currently viewed section (page number in the book) displayed in the viewing target display area 1704. This makes it easy to understand the relationship between the overall content and the section the user is viewing, and makes it easier for the user to refer to other parts of the content. In Figure 18, information on legal revisions, laws, guidelines, and case law is displayed as other content related to the book.
[0303] The legal reference relationship display area 1726 is an area that displays the legal laws in which the reference relationship is defined.
[0304] In the illustrated example, the legal reference relationship display area 1726 is set to a link to a legal database in which the reference relationship is defined. Depending on the user's operation, the system may transition to the service provided by the legal search service server 92 to display the relevant laws, or the laws may be displayed on a screen such as Figure 17 as an overlay.
[0305] The legal reference relationship display area 1728 is an area that displays the legal laws in which the reference relationship is defined.
[0306] The case reference relationship display area 1730 is an area that displays court cases in which a reference relationship is defined.
[0307] In the illustrated example, the case law reference relationship display area 1730 is set to a case law database in which the reference relationship is defined. Depending on the user's operation, the system may transition to the case law search service server 91 to display the case law provided by that service, or the case law may be displayed on a screen such as Figure 17 as an overlay.
[0308] Figure 18 shows an example of a screen in a content viewing operation screen, which displays the content itself and other related content.
[0309] The related content update operation unit 1732 is an operation component that receives an operation to update the display content of other related content for the content being viewed, which is displayed in the detailed display area 1724.
[0310] In the illustrated example, the related content update operation unit 1732 updates the display of information such as legal revisions, laws, guidelines, and precedents related to a portion of the content displayed in the viewable display area 1704, in response to user operations. This corresponds to the process shown in step S2023 of Figure 20.
[0311] The legal amendment information display area 1734 is an area for displaying information about legal amendments.
[0312] The legal amendment information display area 1734 corresponds to the processing shown in step S1627 of Figure 16. In the illustrated example, it accepts an operation to specify information about legal amendments, and in response to this operation, it displays the details of the legal amendment information as shown in Figure 19.
[0313] The legal information display area 1736 is an area for displaying legal information.
[0314] The legal information display area 1736 corresponds to the process shown in step S1627 of Figure 16. In the illustrated example, links are set to the legal information, and information such as the legal database is displayed as described above in response to the operation of specifying the link.
[0315] Guideline information display area 1738 is the area for displaying guideline information.
[0316] The guideline information display area 1738 corresponds to the process shown in step S2022 of Figure 20.
[0317] The case law information display area 1740 is an area for displaying case law information.
[0318] The case law information display area 1740 corresponds to the processing shown in step S2022 of Figure 20.
[0319] Figure 19 shows an example screen of a user interface for viewing content, where information on laws and legal amendments is displayed as other content related to the content being viewed.
[0320] The legal amendment details display area 1742 is an area that displays detailed information about legal amendments.
[0321] In the illustrated example, the legal amendment details display area 1742 is displayed in place of the sub-window display area 1708 shown in Figures 17, 18, etc. Alternatively, the legal amendment information may be displayed in a separate window, or the legal amendment information may be displayed in the sub-window display area 1708 in response to user operations.
[0322] The legal compliance confirmation operation unit 1744 is an operating component that accepts operations to confirm the legal compliance that is being displayed in detail.
[0323] In the illustrated example, the legal information verification unit 1744 displays legal information, such as that found in the legal database, in response to user operations. This makes it easy to verify each article of the law.
[0324] The window close operation unit 1746 is an operating member that receives an operation to close the legal amendment details display area 1742.
[0325] In the illustrated example, the window close operation unit 1746, in response to user operation, erases the display of the legal amendment details display area 1742 and displays the sub-window display area 1708 shown in Figures 17 and 18, etc.
[0326] The legal amendment information display area 1748 is an area that displays details of legal amendment information.
[0327] The legal amendment information display area 1748 corresponds to the processing shown in step S1627 of Figure 16. In the illustrated example, the latest legal amendment information and the information before the amendment are displayed. Alternatively, information on the publication date of the content being viewed may be displayed, along with information on legal amendments before and after that publication date.
[0328] <Fifth Embodiment> The fifth embodiment describes a technology that displays other content (such as laws, precedents, guidelines, etc.) related to the currently open section (a portion of a page in a book) while viewing content such as books.
[0329] When conducting research on content, asking questions to a large-scale language model and examining the responses it generates can provide an overview. However, the responses may be too abstract, leading to misunderstandings or even errors. Therefore, it is necessary to verify the accuracy of the answers. For example, when organizing legal opinions, it is necessary to thoroughly understand the basis for those opinions and then summarize them while referring to evidence.
[0330] While there is a book browsing service that allows you to view books, it's not easy to find the specific book you need for what you want to know, or to locate the relevant section within that book.
[0331] Therefore, this document describes a technology that allows users to view various pieces of information related to the section they are viewing simultaneously, thereby further reducing the effort required to research original sources that serve as evidence.
[0332] <11 Operation (Fifth Embodiment)> Figure 20 shows the process flow for displaying a portion of other content related to the content being displayed, while simultaneously presenting an operation screen for viewing content.
[0333] The outline of each process is as follows:
[0334] Content Viewing and Presentation of Related Information: Users can view legal books and other content while simultaneously referencing relevant laws, precedents, and guidelines for the page they are viewing. This facilitates the collection of detailed information and accurate understanding. User Interface Configuration: The screen is broadly divided into a first area (main content display) and a second area (overview and related information display). The content displayed in the second area can be flexibly switched by the user's operation. Identifying related content: Server 20 identifies other related content based on the content being viewed or selected. Information about the relationships between content is used in this process. Related Content Information Update: When a user changes the content they are viewing, related information is dynamically updated accordingly (or in response to user actions). This makes it easier to obtain the most relevant information for the content currently being viewed. Integration with the question answering function: In response to user questions, the server 20 performs searches and integrates with external information processing systems (the server 95 of the large-scale language model service) to present answers and their supporting evidence, making it easy to check the details of the original content. Research Support: This system allows users to easily access primary sources of other relevant content, such as laws and precedents, when reading specialized books and other materials, thereby facilitating information gathering and understanding.
[0335] These processes allow users, for example in legal research, to efficiently acquire legal knowledge and reduce the effort required for research and learning. Furthermore, the centralized provision of relevant information helps prevent overlooking or misunderstanding information, supporting more reliable judgment and analysis. Next, we will explain each step.
[0336] In step S2011, terminal 10 presents the user with a screen for viewing one of the following contents: laws, precedents, legal books, guidelines, etc. Terminal 10 presents the user with a screen for viewing the content specified by the user (one of the following: laws, precedents, legal books, guidelines, etc.). Terminal 10 accepts an operation from the user to specify the content they wish to view.
[0337] Here, the storage unit 202 of the server 20 stores at least multiple types of content, including laws, precedents, legal books, and legal guidelines, by associating parts of the content with each other (legal content database 212, graph-type data structure 214). The content presentation module 2047 presents the user with a screen that allows them to view any of the content stored in the storage unit 202, including laws, precedents, legal books, and legal guidelines.
[0338] The content presentation module 2047 identifies parts of other content that are associated with the part of the content being viewed on the screen being viewed.
[0339] Furthermore, the server 20 may generate answers to user questions. The question processing module 2044 receives a user question, searches various contents in the memory based on the question, creates a prompt that includes the search results in the search step and the user's question, and includes instructions to answer the question by referring to the search results, sends the created prompt to an external information processing system (the server 95 of the large-scale language model service) to obtain the answer generated for the question from the external information processing system (the server 95 of the large-scale language model service), and displays the content of the answer and a portion of the content that formed the basis for generating the answer to the user as the obtained answer. The content presentation module 2047 may accept an operation from the user to specify a portion of the content that is displayed as the basis for generating the answer, and then present a screen to the user displaying the portion of the content that was specified.
[0340] In step S2021, the content presentation module 2047 of the server 20 sends the content specified by the user to the terminal 10.
[0341] In step S2012, terminal 10 displays the acquired content in the first area of the operation screen. Terminal 10 displays a second area on the operation screen that is different from the first area, and initially displays an overview of the content (such as a table of contents). Terminal 10 may also display the first area and the second area side by side on the operation screen. For example, tabs for changing the displayed content may be placed in the second area, and the displayed content may be switched according to the tab specified by the user.
[0342] In step S2013, terminal 10 accepts an operation to display "other related content" in the second area. For example, terminal 10 may display "other related content" in the second area in response to an operation on a tab to change the content to be displayed.
[0343] Thus, the content presentation module 2047 of server 20 provides a function to switch the content displayed in the second area between a "content overview" (such as a table of contents) and "other related content" in response to user interaction. Terminal 10 accepts the user's operation to select the display of "other related content" in the second area. In this way, the content presentation module 2047 displays a portion of the identified other content in response to user interaction.
[0344] As described above, the content presentation module 2047 displays a first area on the content viewing screen that displays the content to be viewed, and a second area different from the first area. Depending on the user's operation, the module switches whether the content displayed in the second area is an overview of the content to be viewed or a part of other specified content.
[0345] More specifically, the content presentation module 2047 displays the contents of a legal book as the content to be viewed in the first area of the screen where the content is viewed, and in the second area, depending on the user's operation, switches between displaying an overview including the table of contents of the legal book, or displaying at least one part of laws, precedents, or legal guidelines as other content associated with the part of the legal book that is to be viewed.
[0346] In step S2022, the content presentation module 2047 of server 20 outputs portions of other related content, such as laws, precedents, and guidelines, associated with the portion of the content currently being viewed (e.g., a book). More specifically, the content presentation module 2047 identifies other content, such as laws, precedents, and guidelines, associated with the portion of the content currently being viewed. The content presentation module 2047 retrieves portions of the identified other content from various databases, etc.
[0347] In step S2014, terminal 10 displays a portion of other related content in the second area.
[0348] In step S2015, terminal 10 accepts operations to change the content viewing range (such as scrolling or page turning) and operations to select a portion of the content displayed in the first area (such as text selection).
[0349] In step S2023, the content presentation module 2047 of the server 20 identifies other content related to the selected area in response to a change in the content viewing range (in the case of a book, changing the page being viewed) or an operation to select a portion of the content (in the case of a book, the user selects a portion of the text), and outputs it to the terminal 10. For example, the content presentation module 2047 may be equipped with an operation component that accepts an operation to identify other related content (for example, an "update" button), and in response to an operation on the operation component, it may identify other content related to the range being viewed at the time the operation was performed, or the range in which a portion of the content has been selected, and output it to the terminal 10.
[0350] Thus, the content presentation module 2047 may re-identify other related content based on the scope of newly viewed content and update its display. Alternatively, the content presentation module 2047 may identify portions of other content (laws, precedents, guidelines) related to portions of content selected by the user and update their display.
[0351] In step S2016, terminal 10 displays the updated related content in the second area.
[0352] <12. Screen Example (Fifth Embodiment)> The example screen in the fifth embodiment is the same as that shown in Figure 18, etc., described in the fourth embodiment above.
[0353] <Sixth Embodiment> In the sixth embodiment, when generating answers to questions using an external information processing system such as the server 95 of the large-scale language model service, a function is provided that allows switching between simple, high-speed answers and slower but more in-depth, point-based answers, according to the user's desired answer style and search depth, thereby meeting a wide range of search needs.
[0354] Furthermore, we will describe the technology for determining the mode of response generation, more specifically, the technology for analyzing the user-entered question query and the literature searched based on that query, and determining the user's desired response style and search depth based on the nature of the question query, the search target field, and the variability of the search results. Here, we will describe the technology for automatically switching between simple, fast responses and slower but more in-depth, point-based responses depending on the determination result, thereby meeting a wide range of search needs.
[0355] For users, when generating answers to questions using external information processing systems such as large-scale language models, it is not easy to determine which of the various models to use. Furthermore, if a user always relies on a model that generates detailed answers, and there are restrictions on the number of times such models can be used (for example, a limit on the number of questions that can be asked within a certain period), there is a risk of unnecessarily reducing the number of times the model can be used. Also, if a user always tries to use a model that generates detailed answers, the time it takes to receive answers will increase. The more questions a user asks, the longer it will take to receive answers, which may hinder the smooth progress of the user's work.
[0356] Therefore, in the sixth embodiment, in order to generate answers that correspond to user questions while also shortening and optimizing the time required for research work and supporting the smooth progress of the user's work, the present invention describes a technology that determines whether to narrow the scope of the research to generate an answer or broaden the scope of the research to generate an answer, and if it is possible to generate an answer quickly by narrowing the scope of the research, the present invention describes a technology that responds in a mode that generates an answer quickly.
[0357] <13 Operation (Sixth Embodiment)> Figure 21 shows the process flow for generating an answer, which determines whether to use a first mode that generates answers quickly or a second mode that generates answers slowly and in detail, depending on the content of the question.
[0358] The outline of each process is as follows:
[0359] Mode Determination: Server 20 analyzes the content and difficulty of the question, the variability of the search results, the number of points of discussion in each field, etc., to determine whether fast mode or detailed answer mode is appropriate. User behavior feedback: Record the user's actions after confirming their answers, such as whether they ask additional questions or leave the system quickly, and use this information to determine the next mode. By providing such a mode determination for each user, it is possible to offer the most suitable response style for each user. Optimizing waiting times: By displaying the current processing mode while generating answers, users can more easily predict waiting times. Validity of the answer: Along with the final answer, the mode in which the answer was generated is displayed. This allows users to easily understand the depth and detail of the answer. According to the sixth embodiment, users can obtain answers of a depth that suits their needs at an appropriate time, which is expected to improve work efficiency and reduce the time required for research.
[0360] Next, we will explain each step.
[0361] In step S2121, the question processing module 2044 of the server 20 presents the user with an operation screen that accepts the input of a question.
[0362] In step S2111, terminal 10 receives a question input from the user. Terminal 10 sends the question to server 20.
[0363] In step S2122, the question processing module 2044 of the server 20 analyzes the content of the question received from the terminal 10 and determines whether it is in high-speed mode or detailed mode.
[0364] In this way, the question processing module 2044 analyzes the content of the user's question to determine whether to generate the answer in a first mode that generates an answer quickly, or in a second mode that generates an answer more slowly and in more detail than the first mode.
[0365] (Method for determining the mode to operate the external information processing system for generating the answer)
[0366] (1) Variation in the sources of information that answer the questions The question processing module 2044 analyzes the content of the user's question, referring to a graph-type data structure 214, a legal content database 212, etc., and searching for information sources, including books. The question processing module 2044 determines whether to generate an answer in the first mode or the second mode, depending on the variability in the fields of the books. More specifically, the question processing module 2044 may determine to generate a detailed answer in the second mode if the variability in the fields of the books is large, and to generate an answer quickly in the first mode if the variability in fields is small.
[0367] For example, the question processing module 2044 may determine whether the genres of the books are scattered, based on whether the proportion of books in the same field (for example, books that have been pre-classified according to the field of law, etc.) is above a certain level, and the degree of dispersion across genres.
[0368] The question processing module 2044 may determine, based on the variability of the fields, whether to generate the answer in the first mode or the second mode, depending on the number of issues in the field. If the number of issues in the question is above a certain level, the answer may be generated in the second mode, and if the number of issues is below a certain level, the answer may be generated in the first mode.
[0369] (2) The server 95 of the large-scale language model service determines whether or not to generate a detailed response. The question processing module 2044 analyzes the content of the user's question and generates an evaluation prompt that includes the content of the question and instructions to evaluate whether further consideration of the question is necessary. The question processing module 2044 sends the generated evaluation prompt to an external information processing system (the server 95 of the large-scale language model service). The question processing module 2044 obtains the evaluation results from the external information processing system. Depending on the obtained evaluation results, the question processing module 2044 may determine whether to generate an answer in the first mode or the second mode.
[0370] Here, the question processing module 2044 may generate an evaluation prompt that includes the content of the question and instructions to evaluate whether the question requires detailed consideration in terms of whether or not there are many points to discuss.
[0371] (3) Determination based on the user's behavior history when presented with the generated answers to the questions. Server 20 may store in its memory unit 202 behavioral history information, which is a history of the user's actions regarding a question after the answer has been presented to the user. For example, Server 20 stores information relating the content of a question, the answer generated by the server 95 of the large-scale language model service in response to that question, and the user's actions after the answer has been presented to the user, such as the time spent viewing the answer and the number of additional questions asked.
[0372] The question processing module 2044 analyzes the content of the user's question and may determine whether to generate an answer in the first mode or the second mode based on the user's behavior history shown in the behavior history information regarding past questions related to the content of the question.
[0373] More specifically, the question processing module 2044 may make a determination based on the user's behavior history information regarding past questions, specifically the extent of additional questions the user asked in past questions, or the viewing time the user spent reviewing the answers, which is a history of the user's behavior in past questions. For example, if an answer was provided to the user for a similar past question and there were few additional questions (the number of additional questions, the number of characters in the additional questions, etc., were below a certain level), the module may determine that the provided answer to the question was sufficient and decide to respond in the first mode. Alternatively, if the viewing time the user spent reviewing the answers to similar past questions was below a certain level, the module may determine that the provided answer to the question was sufficient and decide to respond in the first mode.
[0374] Server 20 may be configured in its memory unit 202 to store a trained model that uses the content of the question and the history of the user's actions as training data based on the behavior history information, and is trained to output the content of the user's actions in response to the content of the question. Question processing module 2044 may make a decision based on the content of the question received in steps S2111 and S2122 and the output result of the content of the user's actions output based on the trained model.
[0375] (The timing for displaying the determination result of the mode for generating answers to questions) The question processing module 2044 displays an input field for receiving questions and an operating component for receiving the questions entered in the input field on the screen. When the module accepts the input of questions entered in the input field in response to user operation on the operating component, it determines the mode for generating an answer to the question entered in the input field. As a result of this determination, it may also present information to the user indicating whether the answer will be generated in the first mode or the second mode, even without accepting user operation on the operating component (i.e., regardless of whether an operation to generate an answer to the entered question has been performed). This makes it even easier for the user to understand whether detailed consideration is necessary before sending a question, and how long the waiting time until an answer is received will be.
[0376] In step S2112, terminal 10 displays that it is processing until the answer is generated, and whether it is in fast mode or detailed mode. Terminal 10 displays to the user that it is processing until the answer is generated by the question processing module 2044 of server 20, and the mode being used to generate the answer.
[0377] Thus, the question processing module 2044 may provide the user with information indicating whether it is generating the answer in the first mode or the second mode during the period in which it generates the answer before presenting it to the user. This makes it easier for the user to anticipate the waiting time until the answer is generated, eliminating the difficulty of operation caused by not knowing the waiting time, and potentially making the waiting time feel shorter.
[0378] In step S2123, the question processing module 2044 of the server 20, based on the determination result in step S2122, reduces the amount of information sources referenced based on the question to a certain level or less in the case of high-speed mode (first mode), while referencing more information sources in the case of detailed answer mode (second mode).
[0379] The question processing module 2044 prepares to generate an answer in the selected mode based on the determination result in step S2122. In fast mode, the amount of information sources referenced is reduced. In detailed mode, more information sources are referenced. For example, in fast mode, the number of books referenced may be limited to a certain number (e.g., 10 books), and the top a certain number (e.g., the top 3) may be presented as the basis for generating the answer. In detailed answer mode, the number of books referenced may be increased to a certain multiple compared to fast mode to generate the answer.
[0380] In step S2124, the question processing module 2044 of the server 20 generates a prompt that includes the question and the information sources referenced in response to the question, and sends it to an external information processing system (the server 95 of the large-scale language model service).
[0381] Here, the question processing module 2044 may refer to a smaller amount of information sources when generating an answer in the first mode than in the second mode. The question processing module 2044 generates a prompt that includes the content of the user's question and the information sources it referenced, and includes instructions to generate an answer to the question based on the information sources, and obtains an answer from an external information processing system by sending the generated prompt to the external information processing system.
[0382] In step S2125, the question processing module 2044 of server 20 obtains an answer from an external information processing system (server 95 of the large-scale language model service). The question processing module 2044 sends the obtained answer to the terminal.
[0383] In step S2126, the question processing module 2044 of the server 20 may record the user's behavior history when viewing the answers to the questions and store it in the storage unit 202. In this way, the question processing module 2044 may record the user's behavior history when viewing the answers to the questions and use it for future mode determination.
[0384] The question processing module 2044 may also present the answer along with information indicating whether the answer was generated in the first or second mode. This allows the user to easily recognize which mode was used to generate the answer.
[0385] In step S2113, terminal 10 displays the generated response. Terminal 10 also displays whether the response was generated in high-speed mode or detailed viewing mode.
[0386] <14. Screen Example (Sixth Embodiment)> Figure 22 is an example screen that displays the result of determining whether to answer in the first mode (high-speed mode) or the second mode (detailed answer mode) before accepting the operation to submit the question, based on the content of the input question.
[0387] The mode determination result display area 2202 is an area that displays the determination result of whether to use high-speed mode or detailed answer mode as the mode for generating the answer, at the stage before the question is sent (when the question has been entered into the question specification unit 1002 but before the operation is performed on the transmission operation unit 1004).
[0388] The mode determination result display area 2202 corresponds to each process shown in Figure 21. In the illustrated example, it is displayed that it has been determined to generate an answer in detailed response mode.
[0389] The high-speed mode specification unit 2204 is an operating member that receives a specification to generate an answer in high-speed mode.
[0390] In the illustrated example, the high-speed mode selection unit 2204 is shown to be located inside the mode determination result display area 2202, but it may also be located near the transmission operation unit 1004, etc., and the user may be able to manually switch between generating an answer in high-speed mode and generating an answer in detailed answer mode.
[0391] Figure 23 is an example of a screen that displays the determination result of the mode for generating the answer on the user interface during the period from when a question is submitted until an answer is generated using an external information processing system.
[0392] The processing progress display area 2206 is an area that displays an overview of each process involved in generating an answer to a question, as well as the status of the processing.
[0393] In the illustrated example, the processing progress display area 2206 displays steps that have been completed, steps that are in progress, and steps that have not yet been processed, distinguishing between them.
[0394] The mode display area 2208 is an area that displays the mode used to generate the answer.
[0395] The mode display area 2208 corresponds to the process shown in step S2112 of Figure 21.
[0396] Figure 24 is an example of a screen that displays the result of determining the mode in which the answer was generated, on top of the screen that displays the answer generated for the question.
[0397] The mode display area 2210 is an area that displays the mode used to generate the answer, if an answer has been generated.
[0398] The mode display area 2210 corresponds to the processes in steps S2125, S2113, etc., in Figure 21.
[0399] <Seventh Embodiment> In the seventh embodiment, a technology is described that assists in refining questions in a system that generates answers to questions, in order to make it easier to obtain suitable answers.
[0400] In interactive systems that generate answers to questions, users may not know how to formulate their questions effectively. Even if they input a vague, abstract question, the resulting answer may be broad and superficial, lacking specificity and therefore unhelpful.
[0401] Therefore, this section describes technologies that enable users to quickly recognize the questions they want to ask, thereby making it easier for them to obtain answers to their questions. More specifically, it describes technologies that make it easier for users to input questions, such as by presenting suggested questions in response to their initial questions.
[0402] <15 Operation (Seventh Embodiment)> Figure 25 shows the processing flow for receiving follow-up questions that delve deeper into the questions entered by the user.
[0403] The outline of each process is as follows:
[0404] Question Input and Analysis: Analyzes user-entered questions and evaluates their specificity. Facilitating the Specificity of Questions: When the question needs to be more specific, provide relevant information to encourage the input of a more detailed question. Answer generation: Based on the specified question, information sources are searched and the answer is generated by the server 95 of the large-scale language model service. Providing answers: Present the obtained answers to the user. These processing flows allow users to clarify what they want to know and receive appropriate answers. Even vague questions can be helped to become more specific through the processing of server 20, improving the user experience.
[0405] Next, we will explain each step.
[0406] In step S2521, the question processing module 2044 of the server 20 presents the user with an operation screen that accepts the input of a question.
[0407] Server 20 maintains a graph database (graph-type data structure 214) in its storage unit 202 as an information source. This graph database is a graph-type data structure in which reference relationships are defined between parts of the content of multiple contents. The graph database also maintains information about attributes set on at least one of the parts of the content for which a reference relationship is defined, or on the reference relationship itself.
[0408] In the memory unit 202, as a graph database, at least multiple types of content, including laws, precedents, legal books, and legal guidelines, are stored, with reference relationships defined by linking parts of the content together, and information on attributes, including at least information on the issues at stake, is also stored.
[0409] In step S2511, terminal 10 receives a question (first input) from the user. Terminal 10 sends the entered question to the server device.
[0410] In step S2522, the question processing module 2044 of the server 20 receives the user's question. The question processing module 2044 determines whether the received question is specific.
[0411] (Assess whether the question is specific and whether additional questions or further clarification are needed.) The question processing module 2044 determines whether further specific questions are needed to generate an answer to the question related to the first input. If it is determined that further specific questions are needed, the question processing module 2044 performs the following step S2544 (prompting the specificization of the question). If it is determined that further specific questions are not needed, the module may generate a prompt that includes the content of the question related to the first input and the search results obtained by searching for information sources based on the content of the question, and includes an instruction to generate an answer by referring to the search results obtained by searching for information sources for the question. The generated prompt may also be sent to an external information processing system (large-scale language model service server 95) to obtain the answer created by the external information processing system (large-scale language model service server 95) in response to the prompt, and to output the obtained answer to the user.
[0412] (The server 95 of the large-scale language model service determines whether further questioning is necessary.) The question processing module 2044 may also perform a determination by generating a prompt that includes the content of a question relating to a first input, as shown in step S2511, and includes an instruction to determine whether a further specific question is required to generate an answer to the said question, and by sending the generated prompt to an external information processing system (large-scale language model service server 95), and obtaining the result of the determination created by the external information processing system (large-scale language model service server 95) in response to the prompt.
[0413] (Determined based on the scope of the graph database reference) The question processing module 2044 may obtain information within the range referenced from the graph database (graph-type data structure 214) by referring to the graph database as an information source based on the content of the question relating to the first input as shown in step S2511, and may make a determination according to the extent of the referenced range obtained.
[0414] The question processing module 2044 may make a determination based on at least one of the following as information about the range referenced in the graph database (graph-type data structure 214): the number of content items included in the referenced range, or the amount of data including the number of characters in the content. For example, the module may determine that the more content items included in the referenced range (for example, the more books extracted), the more the module needs to delve deeper into the question and make it more specific. Alternatively, the more characters in the referenced content, the more the module may determine that the more the module needs to delve deeper into the question and make it more specific.
[0415] If the question processing module 2044 determines that a more specific question is required, it may present the user with candidate questions in the following step S2524.
[0416] (End of questioning) The server 20 may store in its memory unit 202 information that is sufficient to prompt the user to input specific questions. For example, it may store information such as the upper limit on the number of questions asked and the range within which the graph-type data structure 214 is referenced as the above-mentioned upper limit information.
[0417] The question processing module 2044 may terminate the process of prompting the user to input a more specific question when it has repeatedly received question input from the user through processes such as step S2511 in this embodiment and step S2513 described later (for example, when the user asks a question after being prompted to ask a more specific question), and the extent to which the information sources have been consulted has reached the upper limit of information. Here, the question processing module 2044 may terminate the process of prompting the user to input a more specific question and present the user with a response indicating that it consulted the information sources for the question but could not find information corresponding to the question. This allows the user to confirm that they have investigated but could not find the relevant information.
[0418] In step S2523, if the question processing module 2044 of server 20 is insufficient (abstract), it creates candidate specific questions to present to the user by referencing information sources within a certain scope and generating a list of relevant keywords and points of discussion. For example, it generates a list prioritizing keywords and points of discussion that have a high score related to the question among the results of referencing the above information sources within a certain scope.
[0419] In step S2524, the question processing module 2044 of the server 20 sends relevant information to the terminal 10 along with a message prompting the terminal to elaborate on the question.
[0420] In this way, the question processing module 2044 refers to information sources within a certain range in response to the content of the user's question in the first input, and presents the information within that range to the user by responding with the information within that range, while also prompting the user to enter a more specific question than the one in the first input.
[0421] The question processing module 2044, in responding to the user's question related to the first input with referenced information within its scope, may present the referenced information to the user without using a process that causes an external information processing system, which generates an answer in response to prompt input, to identify the information source corresponding to the question.
[0422] The question processing module 2044 refers to the graph database (graph-type data structure 214) as an information source based on the content of the question related to the first input. The question processing module 2044 presents the user with information extracted based on multiple attributes (such as issues) included in the referenced range. At this time, the question processing module 2044 may also prompt the user to input a more specific question using the extracted information as an example.
[0423] The question processing module 2044 may present the user with a list containing the issues and keywords included in the range referenced from the graph database based on the content of the first input question, and prompt the user to input a more specific question.
[0424] (Determine if the question includes multiple themes) The question processing module 2044 may also determine whether a question received from the user contains multiple themes.
[0425] The question processing module 2044 may, if it determines that a question contains multiple themes, classify the questions according to the themes and present the classification results to the user.
[0426] More specifically, the question processing module 2044 generates a prompt that includes the question it has received as input, and includes an instruction that determines whether the question contains multiple themes. The question processing module 2044 sends the generated prompt to an external information processing system (the server 95 of the large-scale language model service). The question processing module 2044 obtains information from the external information processing system (the server 95 of the large-scale language model service) regarding the determination made according to the instruction. The question processing module 2044 may also make a determination based on the determination result information obtained from the server 95 of the large-scale language model service.
[0427] Furthermore, the question processing module 2044 may also obtain information about the range referenced in response to the question by referring to the graph database (graph-type data structure 214) based on the question it has received as input, and then make a determination based on the information about the range it has obtained. For example, when the graph-type data structure 214 is referred to in response to a question, the module may determine whether multiple themes (e.g., multiple issues) are included based on the information about the nodes and edges set in the range it has referenced.
[0428] In step S2512, terminal 10 displays candidate questions and related information to the user, and displays a message prompting the user to enter further questions.
[0429] In step S2513, terminal 10 receives a specific question (second input) from the user. Terminal 10 sends the question to server 20.
[0430] In step S2525, the question processing module 2044 of the server 20 searches for information sources based on the question and retrieves the search results. In this way, the question processing module 2044 receives a specific question entered by the user, searches for information sources based on the specific question, and retrieves the search results.
[0431] In step S2526, the question processing module 2044 of server 20 generates a prompt containing the question and search results and sends it to an external information processing system (server 95 of the large-scale language model service). The question processing module 2044 retrieves the answer from the external information processing system (server 95 of the large-scale language model service).
[0432] Thus, the question processing module 2044 generates a prompt that includes the content of the user's question related to the second input and the search results of the search step, and that includes instructions to generate an answer to the question while referring to the search results. The question processing module 2044 sends the generated prompt to an external information processing system (the server 95 of the large-scale language model service), thereby obtaining the answer created by the external information processing system in response to the prompt.
[0433] In step S2527, the question processing module 2044 of server 20 sends the answer obtained from the server 95 of the large-scale language model service to terminal 10.
[0434] In step S2514, terminal 10 displays the answer to the user.
[0435] <16. Screen Example (Seventh Embodiment)> Figure 26 shows an example screen for receiving follow-up questions that delve deeper into the questions entered by the user.
[0436] The question content display area 2602 is the area that displays the question entered by the user (the first question).
[0437] In the illustrated example, the question content display area 2602 displays the content of the question received in step S2511 of Figure 25.
[0438] The question candidate display area 2604 is an area that displays candidate questions when it is determined that further investigation of the question is necessary.
[0439] The question candidate display area 2604 corresponds to the processes in steps S2524, S2512, etc., in Figure 25.
[0440] The question content display area 2606 is an area that displays further questions (second inputs) entered by the user.
[0441] In the illustrated example, the question content display area 2606 displays the content of the question received in step S2513 of Figure 25.
[0442] <Eighth Embodiment> In the eighth embodiment, for example, a system that generates answers to questions will display a list of content that forms the basis of the answer in multiple stages, and will describe a technique for switching between a mode that makes it easier to grasp the overall picture of many contents and to browse through the quoted sections of each content, and a mode that makes it easier to compare and consider which content to view in detail.
[0443] If the content is a book, in the first stage, the book title and a certain number of characters are displayed for each piece of content. In the second stage, a larger number of characters from the quoted passages of the book are displayed. The window is enlarged compared to the first stage to allow users to view more of the quoted text and make it easier to select a book.
[0444] Since responses from large-scale language models are not always accurate, it is difficult to rely solely on their responses when conducting research, and it is necessary to consult more evidence-based materials. On the other hand, content such as books is often created by authors according to their own purposes, and may contain a variety of information.
[0445] Therefore, it is not easy for users to determine which materials they should consult to find the information they need.
[0446] Therefore, this document describes a technology that further reduces the workload required for research by providing users with an overall understanding of what they want to know and making it easier for them to identify evidence-based materials.
[0447] <17 Operation (Eighth Embodiment)> Figure 27 shows the process flow for displaying a list of content related to a question while presenting the answer, and changing the display of the list according to the user's actions.
[0448] The outline of each process is as follows:
[0449] Two-stage display: The user first receives the answer and a list of concise content (first form). If the user wants to see more details, the number of quoted sections for each content increases through user interaction, resulting in a more information-rich display (second form). Display Mode Switching: When switching from the first mode to the second mode, an animation is performed in which the display area expands sequentially both vertically and horizontally, visually notifying the user that the amount of quoted text displayed is increasing. Conversely, when returning from the second mode to the first mode, an animation is performed in which the display area shrinks sequentially. Content List by Issue: The content list is organized and presented to the user by legal issue. When the user switches issues, the name of the displayed issue and the list of related content are updated. Research Scope Information Provision: Users can see the number of information sources consulted, the number of words, and the estimated time taken based on their average reading speed. This makes it easy to understand the extent to which the time and amount of information required for research have been reduced compared to individually viewing each book or other content. Utilizing a graph database: Server 20 refers to the graph database to identify appropriate information sources for a question. The graph database includes laws, precedents, legal books, guidelines, etc. Answer generation by server 95 of the large-scale language model service: Based on the identified information sources, answers are generated using server 95 of the large-scale language model service. The answers include evidence of the information sources, providing reliable information to the user. Improved user experience: Switching display modes and selecting discussion points is easier, allowing users to easily access the information they need. This processing flow allows users to efficiently review the documents that formed the basis of their answers and easily access detailed information as needed. As a result, the workload associated with the survey is reduced, and higher quality information can be collected.
[0450] Next, we will explain each step.
[0451] In step S2721, the question processing module 2044 of the server 20 presents the user with an operation screen that accepts the input of a question.
[0452] The memory unit 202 holds a graph database (graph-type data structure 214), which is a graph-type data structure that defines reference relationships between parts of the content of multiple contents. The memory unit 202 holds at least the information of the legal book content in the graph database (graph-type data structure 214) as an information source.
[0453] In step S2711, terminal 10 receives a question from the user. Terminal 10 sends the entered question to server 20.
[0454] In step S2722, the question processing module 2044 of the server 20 refers to the graph database (graph-type data structure 214) to identify the content that will serve as the information source to be referenced for the question.
[0455] In step S2723, the question processing module 2044 of the server 20 creates a prompt to generate an answer based on the identified information source (including books) and sends it to an external information processing system (the server 95 of the large-scale language model service), thereby obtaining the answer from the external information processing system (the server 95 of the large-scale language model service).
[0456] The question processing module 2044 generates answer information that includes the content of the answer to the question and information about the source on which the answer is based, based on the identified information sources.
[0457] More specifically, the question processing module 2044 creates a prompt that includes the content of the question and identified information sources, and that includes instructions to generate an answer by referring to the information sources regarding the question. By sending the created prompt to an external information processing system (the server 95 of the large-scale language model service), the module generates answer information by obtaining the answer generated by the external information processing system (the server 95 of the large-scale language model service). In this way, the module causes the server 95 of the large-scale language model service to generate an answer to the question, along with the basis for the answer.
[0458] In step S2724, the question processing module 2044 of the server 20 generates the content of the answer and a list of the content (information sources) on which the answer is based. The question processing module 2044 sends the answer and the list of content to the terminal 10.
[0459] Here, the question processing module 2044 may calculate the number and character count of the referenced sources when generating the answer and a list of the content (sources) on which the answer is based, and calculate the time taken from the average reading speed. It may also send this calculated information (such as the number of referenced sources) along with the answer and the list of content to terminal 10.
[0460] More specifically, the question processing module 2044 presents the user with a list of multiple source contents that formed the basis of the answer, either in a first manner or in a second manner that presents a larger portion of the relevant sections of each source content that formed the basis of the answer than in the first manner.
[0461] (Display mode and switching between display modes) The question processing module 2044 displays both the answer and a list of the supporting content on the screen. The question processing module 2044 accepts user input to switch between displaying the content list in either the first or second format. The question processing module 2044 switches between the first and second formats to display the list according to the user's input.
[0462] More specifically, in the second embodiment, the question processing module 2044 may display the list in a manner that increases the display area and amount of content for each item in the list compared to the first embodiment. This makes it easy to grasp the overall picture of the content included in the list in the first embodiment, while in the second embodiment, it makes it easy to consider and select from among the candidates for content to which the user wishes to check the details.
[0463] The question processing module 2044 may display the answer content and a list of supporting content on the screen, and may also display an operation member on the screen that accepts an operation to switch whether to display the content list in the first or second form, in association with the list. When switching from the first form to the second form in response to the user's operation on the operation member, the question processing module 2044 may transition to the second form by drawing the display area of each content in the first form to expand sequentially in at least one of the vertical or horizontal directions. This makes it possible to visually recognize to the user that the amount of content that can be viewed in each content is greater in the second form than in the first form, in response to the user's operation, making the user operation even easier.
[0464] The question processing module 2044 may transition to the first mode by sequentially shrinking the display area of each content in the second mode in at least one of the vertical or horizontal directions when switching from the second mode to the first mode in response to user operation on the control element displayed in association with the list. This makes it easier for the user to recognize that the amount of content that can be viewed is reduced when switching from the second mode to the first mode, making it easier to grasp the overall picture of each content included in the list, and thus makes user operation even easier.
[0465] More specifically, the question processing module 2044 may display a list of legal books related to each legal issue included in the answer to the question, as a list of multiple contents. The question processing module 2044 displays an operation element that accepts the operation of switching between legal issues in the list of legal books. In response to the user's operation on the operation element for switching legal issues, the question processing module 2044 may update the name of the legal issue as the title of the legal book list and display a list of legal books corresponding to the legal issue to be displayed. This makes it easy to check a list of legal books for each legal issue, and makes it easy to select a book to check in detail from the list displayed by narrowing down the legal issues.
[0466] (Checking the amount of data in the referenced information sources) When the question processing module 2044 presents the user with a list of the legal books that served as the basis or source of information, it may also present the user with information about the amount of information contained in each source.
[0467] More specifically, the question processing module 2044 may present the user with information about the amount of information in the information source, by presenting the user with at least one of the following: the number of contents in the referenced information source or the amount of data contained in the contents of the referenced information source. The question processing module 2044 may also present the user with information about the amount of information in the information source, by calculating the time it would take a human to read the referenced information source and presenting the calculated time to the user.
[0468] In step S2712, terminal 10 presents the user with the response received from the server and a list of content (first mode). Terminal 10 also presents the user with the number of information sources referenced, the number of characters, and the time taken, as information within the scope of the research.
[0469] In step S2713, terminal 10 switches the display mode to change the content, amount of information, and display area of the list in response to user operations on the list of content (switching from the first mode to the second mode, switching the topic, switching from the second mode to the first mode).
[0470] Terminal 10 switches to the second mode when it receives an operation from the user to display a detailed view of the list of content displayed in the first mode. At this time, in response to the user's operation, terminal 10 displays an animation that sequentially expands the display area and amount of information for each content in the list of content as a switch to the second mode. In the second mode, the quoted portion of each content is increased and detailed information is displayed. If the user performs an operation to switch the content display mode again, an animation that sequentially shrinks the display area is displayed and the system transitions from the second mode to the first mode.
[0471] Terminal 10 accepts user requests to switch topics. In response to the topic switch, Terminal 10 updates the list of displayed topic names and related content.
[0472] In step S2714, terminal 10 receives a specification from the user regarding the content they wish to view in detail.
[0473] In step S2725, the content presentation module 2047 of the server 20 retrieves detailed information of the specified content and responds to the terminal 10. For example, if a legal book is specified, the contents of that legal book are output to the terminal 10.
[0474] In step S2715, terminal 10 presents the user with details of the acquired content.
[0475] <18 Screen example (8th embodiment)> Figure 28 is an example of a screen that displays a list of content related to a question while presenting the answer, in a first manner that facilitates understanding the overall picture of the content.
[0476] Figure 29 is an example screen of a second mode in which a list of content related to a question is displayed, with the amount of information in a portion of each content increased to assist in selecting content to view in more detail.
[0477] Operation screen 2800 is an operation screen that displays the generated answers to questions.
[0478] The operation screen 2800 corresponds to the process shown in step S2724 and other steps in Figure 27.
[0479] The answer display area 2802 is the area that displays the answers generated in response to the question.
[0480] The input question display area 2804 is an area that displays the content of the question entered by the user.
[0481] The answer summary display area 2806 is an area that displays a summary of the answers included in the generated responses.
[0482] The issue display area 2808 is an area that displays a list of issues included in the generated response.
[0483] The research scope display area 2810 is the area that displays the range of information sources referenced to generate the answer.
[0484] The first issue display area 2812 is the area that displays the answer to the first of the multiple issues included in the generated answer.
[0485] The bibliography selection unit 2814 is an operating member that accepts a specification to display a list of documents that served as the basis for generating the answer to the first point of contention.
[0486] In the illustrated example, the bibliography selection unit 2814 switches the list of bibliographies displayed in the bibliography display area 2902 to one that corresponds to the first issue, in response to user operation (switching the group of bibliographies to be displayed in the list group display area 2904).
[0487] The reference list display area 2902 is an area that displays a list of content (books, guidelines, etc.) that was referenced when generating the answer.
[0488] The list group display area 2904 is the area that displays the group of content to be displayed.
[0489] In the illustrated example, the list group display area 2904 is divided into groups corresponding to the issues shown in the issue display area 2808. The displayed groups are switched according to the user's actions on the list group display area 2904.
[0490] The list expansion / contraction operation unit 2906 is an operation member that, in a manner that displays a list of content displayed in the bibliography list display area 2902, accepts operations to expand or contract the display area for each content in order to increase the amount of information cited from each content.
[0491] In the illustrated example, the list expansion / contraction unit 2906 transitions from the state shown in Figure 28 to the state shown in Figure 29 in response to user operation. Specifically, in response to user operation on the list expansion / contraction unit 2906, the horizontal width of the bibliography display area 2902 is expanded while increasing the amount of information displayed for citations in each content item in the content display area 2910. In the illustrated example, the width of the answer display area 2802 is reduced in response to user operation on the list expansion / contraction unit 2906, and the line breaks of the answer text are changed accordingly. However, the width of the answer display area 2802 itself may not be changed, and the display position of the answer display area 2802 may be moved horizontally in accordance with the expansion of the display width of the bibliography display area 2902. In the example in Figure 28, the display mode of the list expansion / contraction unit 2906 is shown as expanding the bibliography display area 2902. In the example shown in Figure 29, the display mode of the list expansion / contraction operation unit 2906 is shown as reducing the size of the bibliography list display area 2902.
[0492] The window close operation unit 2908 is an operating member that receives an operation to close the bibliography list display area 2902.
[0493] Content display area 2910 is an area that displays an overview of each reference content (such as books).
[0494] In the illustrated example, the content display area 2910 displays the content title, content classification (book, guideline, etc.), bibliographic information such as the publication date of the content, and a citation of a portion of the content that served as the basis for generating the answer to the question. Compared to the example shown in Figure 29, the example in Figure 28 reduces the amount of citation information in the content display area 2910. As a result, in the reference list display area 2902, compared to the example in Figure 29, the example in Figure 28 displays the content in a way that makes it easier to grasp the overall picture of the references. For example, when a user refers to the answers displayed in the answer display area 2802 and considers which answers to refer to the original source such as a book, it becomes easier to narrow down the points by viewing the answers displayed in the answer display area 2802 while also checking the list of references in the reference list display area 2902 for each point. On the other hand, in the example in Figure 29, the amount of information on the citations for each piece of content is increased, making it easier to decide which pieces of content to examine in detail. Depending on the user's specified operation, each piece of content transitions to a content viewing screen as shown in Figure 17, etc.
[0495] <Ninth Embodiment> The ninth embodiment will be described as follows.
[0496] (1) Flowchart for creating a graph-type data structure that corresponds to abbreviated notations of laws and regulations. Depending on the various forms of notation of laws and regulations, the section in the content that refers to the law is identified, and the reference relationship with the information of the law that has a graph-type data structure is defined to generate a graph-type data structure.
[0497] (2) A flow in which a vector search is performed based on the question, information is obtained from the citation graph using the retrieved documents, and the information obtained from the question, retrieved documents, and graph database is provided to the LLM to generate an answer. Various content is searched according to the question, and relevant content is extracted by referring to a graph-type data structure for the content of the search results, which is used as the basis for generating the answer to the question.
[0498] In order to make appropriate judgments in legal practice, it is necessary to comprehensively consider a variety of information sources, including relevant laws and regulations, similar precedents, and explanations in specialized books and guidelines.
[0499] Therefore, this paper explains a technique for creating citation graphs that comprehensively link laws, precedents, books, and guidelines, thereby serving as a source of information to be referenced in legal practice.
[0500] Specifically, the following will be explained. • How to construct citation graphs that include diverse information sources such as laws and precedents, as well as books and guidelines. • How to handle question-answering tasks that require presenting the basis for the answer, simulating legal practice. • A method for presenting appropriate evidence in response to questions by combining large-scale language models with constructed citation graphs. A method for generating more favorable answers for legal professionals. <19 Data structure of the ninth embodiment> Figure 30 shows a graph-type data structure for each type of content: laws, precedents, books, and guidelines.
[0501] The example shown illustrates a graph-type data structure where each vertex represents a part (chunk) of the book's text, and edges are defined in a parent-child relationship between the vertices representing the book-identifying information and the vertices representing the chunks.
[0502] The illustrated example shows a graph-type data structure for laws and regulations, where information identifying the law (such as the name of the law, the date of amendment, and the date of enforcement) and the articles, paragraphs, and subparagraphs are each used as vertices, and edges are defined by parent-child relationships.
[0503] Although not shown in the diagram, the guidelines are similarly structured as a graph-type data structure, with each vertex representing information that identifies the guideline (guideline name, issuing body, publication date, etc.) and each part of the guideline text divided into sections, and edges defined in a parent-child relationship between the vertex representing the information that identifies the guideline and the vertex representing the chunk.
[0504] Although not shown in the diagram, the data structure for precedents is similar, with each vertex representing either information identifying the precedent (date of trial, type of trial, court information, trial number, etc.) or at least one of the summaries of the precedents or chunks of the judgment text, and edges defined in a parent-child relationship as described above.
[0505] Figure 31 shows the data structure of an abbreviation dictionary database 215 that defines rules for abbreviating laws and regulations. The abbreviation dictionary database 215 includes the following items: "Abbreviation Rule ID", "Official Name of Law or Regulation", "Abbreviation of Law or Regulation", "Abbreviation Rule", "Law or Regulation ID", and "Content ID".
[0506] The item "Abbreviation Rule ID" is an ID used to uniquely identify an abbreviation rule.
[0507] The item "Official Name of Law" contains information about the official name of the law.
[0508] The item "Abbreviated Name of Law" contains information on the abbreviated name of the law.
[0509] The "Abbreviation Rules" section provides detailed information on the rules and patterns for applying abbreviations.
[0510] The item "Legal ID" is an ID used to uniquely identify a legal or statutory law.
[0511] The "Content ID" field is an ID used to uniquely identify content to which abbreviation rules apply.
[0512] <20 Operation (Processing flow of the 9th embodiment)> Figure 32 shows the process flow for identifying sections in the content that refer to laws and regulations, depending on the various forms of legal notation, defining reference relationships with the information of the laws and regulations which have a graph-type data structure, and generating a graph-type data structure.
[0513] In step S3221, the data structure definition module 2046 of server 20 acquires information on various content to be viewed, such as book information from server 93 of the book viewing service, guideline information from server 94 of the information media service, and case law information from server 91 of the case law search service.
[0514] The data structure definition module 2046 integrates the book title and its chunks (each divided part of the content) into a graph structure with parent-child relationships. The data structure definition module 2046 extracts chunks from the book's text, detecting periods, commas, and line breaks, and each section reaching a certain number of characters. For example, in the book's text, it may extract a single chunk up to the point where a period, comma, or line break is detected after the sentence exceeds a certain number of characters.
[0515] The data structure definition module 2046 integrates the guideline titles and their chunks (each part of the content divided into sections) into a graph structure with parent-child relationships.
[0516] The data structure definition module 2046 integrates information identifying a case (such as the case number and the date of the judgment) and information about the content of the judgment (for example, a summary of the case and chunks of the judgment text) into a parent-child relationship graph structure.
[0517] Server 20 manages, in its storage unit 202, information on content that is subject to viewing and is at least one of books or guidelines (such as the legal content database 212, data from the book viewing service server 93, and data from the information media service server 94), and information on laws and regulations that include provisions of laws and regulations (such as the data from the law search service server 92).
[0518] Server 20 manages content information and legal information as graph-type data structures in its storage unit 202. More specifically, Server 20 stores legal information in its storage unit 202 using a graph-type data structure in which the legal name, article, paragraph, and item are each vertices, and reference relationships are defined such that the vertices of the legal name are in a parent-child relationship with the vertices of the article, paragraph, and item in order.
[0519] Here, the server 20 stores information on laws and regulations in the storage unit 202, distinguishing them based on at least one of the amendment date or the enforcement date, and storing them in a graph-type data structure according to at least one of the respective amendment date or enforcement date.
[0520] Server 20 manages content information in its storage unit 202 as data in a graph-type data structure, where the title of the book or guideline is at one end, and each part of the book or guideline's content is at another end, with each part's vertex representing a parent-child relationship defined between the title of the book or guideline and the vertex representing the part.
[0521] Server 20 manages case law information in its memory unit 202, and manages it as data in a graph-type data structure in which the information identifying the case is at the top, and at least one of the case summary or each part obtained by dividing the judgment text of the case is at the top, with reference relationships defined as parent-child relationships between the vertices of the case summary or each part of the case and the vertices of the information identifying the case. Here, the information identifying the case includes the following elements. ・Court name: The name of the court where the trial was held. Abbreviations of the court name may also be used. Examples include "Supreme Court (Saiko)" and "Tokyo High Court (Tokyo Kousai)" ・Trial type: Judgment (Han), Order (Mei), Ruling (Ketsu) ・Trial date: The year, month and day of the trial. Examples include "Hei 5.9.9" and "December 10, 2020 (Reiwa 2)" ・Source: The published case reporter or journal (abbreviated notations allowed) and page number. Examples include "Hanrei Jiho (Hanji)" and "Hanrei Times (Hanta)" As described above, the information for identifying a precedent is constituted by a regular expression that includes character strings of the above elements. The data structure definition module 2046 extracts character strings of the above abbreviations using a regular expression, and performs processing to restore abbreviations such as court name, trial type, and source to their official names. This enables unique identification of precedents, and allows defining reference relationships from portions of books and guidelines to precedents.
[0522] In step S3222, the data structure definition module 2046 of the server 20 identifies portions referring to laws and regulations in the content of the content, based on rules for referring to laws and regulations in the content, even if the notation is different from the official name of the law. The data structure definition module 2046 extracts portions referring to laws and regulations in the content through text analysis.
[0523] (1) Extraction of laws and regulations by referring to the abbreviated notation dictionary database 215 that manages abbreviated notations The server 20 refers to dictionary information (the abbreviated notation dictionary database 215) that has a correspondence relationship between abbreviated notations of laws and regulations and their official names, and identifies portions where abbreviated notations of laws and regulations are used in the content. The data structure definition module 2046 identifies the official name of the law in the portion referring to the law in the content by referring to the abbreviated notation dictionary database 215. As described above, the data structure definition module 2046 may identify the official name of the law based on dictionary information corresponding to the type of abbreviated notation associated with the content.
[0524] The abbreviation dictionary database 215 manages abbreviated expressions such as the following examples. • Unique abbreviations used in books, etc. For example, the phrase "Article 124, Item 3 of the Company Regulations" may refer to "Article 124, Paragraph 1, Item 3 of the Companies Act Enforcement Regulations" (in books and other publications, the names of laws and regulations may be abbreviated in unique ways). For example, the phrase "Article 27-2, Paragraph 1, Item 3 of the Financial Instruments and Exchange Act" may refer to "Article 27-2, Paragraph 1, Item 3 of the Financial Instruments and Exchange Act." For example, the notation "Commercial Registration 54IV" may refer to "Article 54, Paragraph 4 of the Commercial Registration Act" (the notation of numbers differs between the article and the paragraph; for example, the article is written in Arabic numerals and the paragraph in Roman numerals). Here, the official name of the law may be obtained by acquiring information on the official name provided by the legal search service server 92 (for example, the official name of the law listed in the file name obtained from e-GOV legal search).
[0525] (2) Extraction of laws and regulations when the name of the law or regulation is abbreviated. The data structure definition module 2046 may extract portions of the content that do not contain the name of a law but do contain descriptions of articles, paragraphs, or subparagraphs, and then identify the official name of the law by referring to the title of the content or other parts of the content for the extracted portions.
[0526] In some books and guidelines, the full name of a law may not be stated in the text, and instead, it may simply be referred to as "the Law" or as "the Law," suggesting that it is following the name of a preceding law. In response to these omissions (e.g., "the Law," "the Law"), Data Structure Definition Module 2046 may infer the name of the law in the abbreviated notation by referring to the title of the book or guideline, the table of contents, and the text surrounding the abbreviated section to extract the name of the law. For example, even if the text is simply abbreviated as "the Law," the name of the law may be inferred by extracting the name of the law as shown in the book's title, the table of contents, and the name of the law written around the abbreviated section. In this way, when the name of a law is abbreviated, Data Structure Definition Module 2046 infers the correct law by referring to the context of the description and other parts of the content, and defines the article number and reference relationships of the law.
[0527] (3) Definition of reference relationships between multiple articles, paragraphs, and clauses The data structure definition module 2046 may extract portions of the content that contain specific phrases used with articles, paragraphs, or subparagraphs, and define reference relationships between those portions and multiple articles, paragraphs, or subparagraphs, depending on the specific phrases in those portions.
[0528] Data structure definition module 2046 detects when phrases such as "and" and "·" are included in the text of books and guidelines, along with abbreviations of laws and regulations. When these specific phrases are used, data structure definition module 2046 analyzes that multiple articles, paragraphs, or items are being referred to and defines the reference relationships between these multiple articles, paragraphs, or items in the graph.
[0529] For example, if the book contains the phrase "Articles 364 and 367 of the Civil Code," the data structure definition module 2046 will define the reference relationships for each of these articles, treating the portion (chunk) of the book containing that phrase as legal information, signifying a reference to "Article 364 of the Civil Code" and "Article 367 of the Civil Code."
[0530] Similarly, when the phrase "·" is used in a book or guideline (e.g., when it is written as "Article 210, Paragraph 1 · Article 214 of the Corporate Reorganization Act"), the data structure definition module 2046 defines a reference relationship to multiple articles, paragraphs, and items, just as when the phrase "and" is used in a book or guideline along with information about the law (e.g., it defines a reference relationship between "Article 210, Paragraph 1 of the Corporate Reorganization Act" and "Article 214 of the Corporate Reorganization Act")).
[0531] Furthermore, if a book or guideline uses expressions that specify the scope of an article (for example, "from," "~") along with legal information such as abbreviations of laws and regulations, the data structure definition module 2046 extracts the expressions that specify the scope of the article and defines the reference relationships between the main text of the book or guideline and each article within that scope (e.g., if the main text of the book states "Articles 258 to 260 of the Civil Rehabilitation Act," the module defines the reference relationships between the part of the book in which that statement is made and "Article 258 of the Civil Rehabilitation Act," "Article 259 of the Civil Rehabilitation Act," and "Article 260 of the Civil Rehabilitation Act"). (e.g., if the main text of the book states "Articles 138 to 141 of the Civil Code," the module defines the reference relationships between the part of the book in which that statement is made and "Article 138 of the Civil Code," "Article 139 of the Civil Code," "Article 140 of the Civil Code," and "Article 141 of the Civil Code").
[0532] As described above, the abbreviation dictionary database 215 of the memory unit 202 is configured to store dictionary information showing the correspondence between the abbreviated form of a law and its formal name. There are multiple types of abbreviated forms of laws. The server 20 is configured to store dictionary information for each type of abbreviated form of a law in the abbreviated dictionary database 215 of the memory unit 202. For each piece of content, it is configured to store information associated with identifying the type of abbreviated form of the law.
[0533] In step S3223, the data structure definition module 2046 of the server 20 parses the law names, articles, paragraphs, and subparagraphs identified within the content and identifies the corresponding vertices in the graph-type data structure of the law.
[0534] The data structure definition module 2046 refers to the publication date of the content (e.g., the publication date of a book, the publication date of a guideline) and identifies the vertex of the latest legal version (revision history) for that period in the graph-type data structure for the law.
[0535] In step S3224, the data structure definition module 2046 of server 20 defines the reference relationship between the identified legal information and the references in the content, adding a new edge to the graph-type data structure. In this way, the data structure definition module 2046 defines a reference relationship between the parts of the content that refer to the identified legal information and the information of the referred legal information. More specifically, the data structure definition module 2046 defines a reference relationship between the parts of the content that refer to the identified legal information and the vertices corresponding to the legal name, article, paragraph, or subparagraph of the referred legal information.
[0536] The data structure definition module 2046 may refer to information about the publication date of the content and define reference relationships between the portion of the content that refers to a specific law and the corresponding vertex of the law, article, paragraph, or subparagraph of the law, with respect to the most recent legal information available at the time of publication of the content.
[0537] Thus, the data structure definition module 2046 defines a reference relationship between the portion of the content that is identified as mentioning a law, even if the full name of the law is not explicitly stated, and the information of the law being mentioned.
[0538] The data structure definition module 2046 may define reference relationships between each vertex of the content that refers to a case and at least one of the vertices of each part of the case summary or judgment text.
[0539] In step S3225, the data structure definition module 2046 of the server 20 calculates a parameter indicating similarity based on the reference relationships defined between the contents and assigns it to the parent vertex of each content. The final graph-type data structure may be stored in a database in preparation for use in processes that respond to content searches, such as a question answering system.
[0540] The data structure definition module 2046 calculates a parameter indicating the similarity of each content based on the reference relationship information defined for each content. For example, by vectorizing other content (including parts of other content) that has a reference relationship defined for each content, a vector for the reference relationship of each content can be identified. For example, the text of parts of other content that have a reference relationship defined for the content may be aggregated and vectorized, or a vector based on the reference relationship of the content may be obtained by performing calculations on the vectors of parts of other content that have a reference relationship defined for the content. By comparing these vectors, the similarity of the reference relationships of each content (the degree to which they refer to similar other content) can be calculated. The data structure definition module 2046 associates the calculated similarity parameter with the content and stores it in the storage unit 202.
[0541] Here, server 20 may be a system that generates answers to questions. The question processing module 2044 may search for content based on the content of the question entered by the user, further extract content similar to the searched content based on a parameter indicating the similarity of the searched content, and generate an answer to the question using the extracted content.
[0542] For example, the question processing module 2044 may perform the following functions: extract content by referring to a graph-type data structure based on the content of the question entered by the user; generate a prompt that includes the content of the question and the extracted content, and includes instructions that refer to the extracted content in response to the question to generate an answer; and provide the generated prompt to an information processing system (large-scale language model service server 95) that performs language processing, thereby obtaining the answer generated in response to the prompt from the information processing system that performs language processing, and presenting the obtained answer to the user.
[0543] Figure 33 shows the process flow for generating answers to questions by searching for various types of content in response to a question, extracting relevant content by referring to a graph-type data structure for the search results, and using this as the basis for generating an answer to the question.
[0544] In step S3311, terminal 10 receives a question from the user. Terminal 10 sends the content of the received question to server 20. In step S3321, the question processing module 2044 of the server 20 vectorizes the content of the user's question that it has received.
[0545] Here, we will explain the data structure of the various data held by server 20 that are used in the processing.
[0546] The server 20 is configured to store data in a graph-type data structure 214 of the storage unit 202, in which reference relationships are defined between the parts that represent the content of each piece of content for multiple pieces of content that are to be viewed.
[0547] More specifically, the graph-type data structure 214 of the memory unit 202 is configured to store data in a graph-type data structure that holds multiple types of content as multiple contents, and defines reference relationships between different types of content. The multiple types of content are configured to store at least one of either book information including books on law or guideline information including legal guidelines, as well as information on laws and regulations and information on precedents.
[0548] In the graph-type data structure 214 of the memory unit 202, information is stored for each part (chunk) obtained by dividing the content of the book or guideline into multiple parts, for at least one of the book information or guideline information that is held as multiple types of content. In the graph-type data structure 214 of the memory unit 202, data is stored as graph-type data that defines the reference relationships between parts of the content of at least one of the book or guideline and parts of other content.
[0549] More specifically, the memory unit 202 holds data in a graph-type data structure, where the title of at least one of the books or guidelines is a vertex, and the content portion of at least one of the books or guidelines is a vertex, with a parent-child relationship defined between the title vertex and the content portion vertex.
[0550] More specifically, the memory unit 202 stores data in a graph-type data structure, with the name of the law as one vertex, and each of the articles, paragraphs, and clauses that make up the law as another vertex, and a parent-child relationship defined between the name of the law and each vertex in the order of articles, paragraphs, and clauses.
[0551] More specifically, the memory unit 202 holds data in a graph-type data structure, with information identifying the precedent as the vertex, and the content of the precedent as the vertex, with at least one of the summaries of the precedent or the content of the judgment divided into multiple parts as the vertex, defining a parent-child relationship between the vertex identifying the precedent and at least one of the vertices of the summaries of the precedent or the content of the judgment.
[0552] Furthermore, the data structure definition module 2046 defines reference relationships between parent vertices by aggregating the reference relationships of child vertices to parent vertices based on the reference relationships defined from one part of the content to other parts of the content, and stores them as a graph-type data structure 214.
[0553] In step S3322, the question processing module 2044 of the server 20 vectorizes the content of the question and searches for the part corresponding to the content of the question by comparing the vector of the question content with the vectors of each part of the content. The question processing module 2044 compares the vectorized question with the vectors of each part (chunk) of the content such as books and guidelines in the database to search for chunks related to the question (searches for chunks with similar vectors).
[0554] Thus, the question processing module 2044 searches multiple types of content based on the content of the question. More specifically, the question processing module 2044 searches multiple types of content. Based on the content of the question, the question processing module 2044 searches at least one of the multiple types of content, namely information from books or information from guidelines.
[0555] In step S3323, the question processing module 2044 of the server 20 extracts relevant content (first content) based on the ranking of the search result chunks.
[0556] More specifically, the question processing module 2044 searches for the relevant section of a book or guideline based on the content of the question, and determines the ranking of the content based on the search ranking of each section of the content. For example, the search ranking of the content can be determined by aggregating the search rankings of each chunk of content and weighting those with higher rankings, calculating the average of the search rankings of the chunks for each piece of content, or extracting a certain number of top-ranking chunks from the search rankings and aggregating the number of chunk search results included in each piece of content. For example, content with a large number of high-ranking chunks based on vector comparisons may be given a higher search ranking.
[0557] In step S3324, the query processing module 2044 of the server 20 selects a second content that has a defined reference relationship with the first content, based on the data in a graph-type data structure, for one or more first content items that are the search results. The query processing module 2044 then references the graph-type data structure 214 for the first content and extracts other related content (second content).
[0558] As described above, the data structure definition module 2046 extracts multiple content items as a group if the reference relationships defined from one content item to another are similar. The question processing module 2044 may select the other content items extracted as a group due to similar reference relationships with respect to the first content item as the second content item. In this way, the question processing module 2044 may extract the second content item related to the first content item by finding that the reference relationships of the graph-type data structure are similar for the first content item (by comparing the vectorized reference relationships defined for the content items with each other, thereby extracting content items with similar reference relationships).
[0559] In step S3325, the question processing module 2044 of the server 20 generates a prompt that includes a question statement, a first content, and a second content, and formats it as input data to the LLM so that the prompt generates an answer to the question.
[0560] Thus, the question processing module 2044 generates a prompt that includes the content of the question, the retrieved first content, and the selected second content, and includes instructions that refer to the first content and the second content to generate an answer to the content of the question.
[0561] In step S3326, the question processing module 2044 of server 20 provides the generated prompt to the LLM (Large-Scale Language Model Service server 95) to generate an answer. In this way, the question processing module 2044 provides the generated prompt to the information processing system that performs language processing, thereby obtaining the answer output by the information processing system that performs language processing.
[0562] In step S3327, the question processing module 2044 of server 20 retrieves the answer output from LLM (Large-Scale Language Model Service server 95) and presents it to the user.
[0563] In step S3312, terminal 10 presents the acquired response to the user.
[0564] <Variation> The embodiments described above may also be combined in various ways.
[0565] Furthermore, each embodiment described above may be modified as follows.
[0566] (1) Sources of information on which answers to questions are based In the above description of the embodiment, a graph-type data structure is defined for laws, precedents, legal books, and legal guidelines, and an example is described in which a user's question is answered by searching using this graph-type data structure and outputting an answer that includes the search results.
[0567] In addition, a graph-type data structure may be defined in the same manner as above, based on information that is not necessarily made public to an unspecified number of people, such as information held on the operator's server 96. Then, in response to a question, the same type of information as that held on the operator's server 96 may be searched based on the graph-type data structure, and the answer to the question may be output. For example, if the operator has accumulated results of judgments made within the company regarding specific cases based on legal issues (for example, issues related to the Premiums and Representations Act, etc.), a summary of the internal cases of the operator may be presented, similar to presenting court cases such as those shown in Figure 12 above.
[0568] For example, the information held by a business operator may include internal regulations, manuals, incident reports, documents such as Q&A and meeting minutes stored in document tools, and internal user inquiries in internal communication tools (for example, a system where users are assigned to groups such as channels, enabling the sending and receiving of messages between users) (for example, user responses to inquiries to accounting, legal, etc.). A graph-type data structure may be defined for this data using information from the company's organizational chart, internal terminology, general terminology, etc. As described in the above embodiment, answers to user questions can be provided by having the server 95 of the large-scale language model service generate answers by referring to the graph-type data structure. For example, these data files may be stored in a specific folder, and the server 95 of the large-scale language model service may generate answers by referring to the information in that folder.
[0569] In addition to the above, a graph-type data structure similar to the above can be defined based on information that may be made public to an unspecified number of people, such as on SNS server 97, and this structure can be referenced when answering user questions. For example, the content of posts made by a specific person's account may be cited as the basis for answering a question.
[0570] (2) How to generate a prompt As described in each of the embodiments above, it has not been easy to create prompts to provide to a large-scale language model in a system that generates answers to questions.
[0571] Therefore, prompts may be created using the methods described in each of the above embodiments.
[0572] A network consists of various mobile communication systems, such as the internet, LANs, and wireless base stations. For example, a network includes 3G, 4G, and 5G mobile communication systems, LTE (Long Term Evolution), and wireless networks that can connect to the internet via designated access points (e.g., Wi-Fi®). When connecting wirelessly, communication protocols include, for example, Z-Wave®, ZigBee®, and Bluetooth®. When connecting via a wired connection, the network also includes connections made directly via USB (Universal Serial Bus) cables, etc.
[0573] Furthermore, by distributing all or part of each hardware configuration across multiple computers and connecting them to each other via a network, a computer can be virtually realized. Thus, the concept of a computer includes not only computers housed in a single enclosure or case, but also virtualized computer systems.
[0574] A database, specifically a relational database, is used to manage and link together tabular data sets called masters, which are structurally defined by rows and columns. In a database, tables are called tables, masters are called masters, the columns of tables are called columns, and the rows of tables are called records. In a relational database, relationships can be established and linked between tables and masters.
[0575] Typically, each table and master has a primary key column to uniquely identify records, but setting a primary key column is not mandatory. The control unit can instruct the processor 901 to add, delete, or update records in specific tables and masters stored in the memory unit, according to various programs.
[0576] Furthermore, by storing data, various programs, and various databases in the memory unit, the information processing device and information processing system related to this disclosure can be considered to have been manufactured.
[0577] Furthermore, the databases and masters in this disclosure may include any data structures (lists, dictionaries, associative arrays, objects, etc.) in which information is structurally defined. Data structures also include data that can be considered as data structures by combining data with functions, classes, methods, etc., written in any programming language.
[0578] Furthermore, each of the above-mentioned configurations, functions, processing units, processing means, etc., may be implemented in hardware, either partially or entirely, by designing them as integrated circuits, for example. The present invention can also be implemented by software program code that realizes the functions of the embodiment. In this case, a storage medium on which the program code is recorded is provided to a computer, and the processor of that computer reads the program code stored in the storage medium. In this case, the program code read from the storage medium itself realizes the functions of the embodiment described above, and the program code itself and the storage medium on which it is stored constitute the present invention. Examples of storage media used to supply such program code include flexible disks, CD-ROMs, DVD-ROMs, hard disks, SSDs, optical disks, magneto-optical disks, CD-Rs, magnetic tapes, non-volatile memory cards, ROMs, and the like.
[0579] Furthermore, the program code that implements the functions described in this embodiment can be implemented in a wide range of programming or scripting languages, such as assembler, C / C++, Perl, Shell, PHP, and Java (registered trademark).
[0580] Furthermore, the program code for the software that implements the functions of the embodiment may be distributed via a network and stored in a storage means such as a computer's hard disk or memory, or in a storage medium such as a CD-RW or CD-R, and the computer's processor may read and execute the program code stored in the storage means or storage medium.
[0581] The functions realized by the components described herein may be implemented in a circuit or processing circuitry, including general-purpose processors, application-specific processors, integrated circuits, ASICs (Application Specific Integrated Circuits), CPUs (a Central Processing Unit), conventional circuits, and / or combinations thereof, programmed to realize the functions described herein. A processor is considered to be a circuit or processing circuitry, including transistors and other circuits. A processor may be a programmed processor that executes a program stored in memory.
[0582] In this specification, circuitry, unit, and means are hardware programmed to perform or execute the functions described herein. Such hardware may be any hardware disclosed herein, or any hardware known to be programmed to perform or execute the functions described herein.
[0583] If the hardware is a processor that is considered to be a type of circuitry, then the circuitry, means, or unit is a combination of hardware and software used to constitute the hardware and / or processor.
[0584] While several embodiments of this disclosure have been described above, these embodiments can be implemented in a variety of other forms, and various omissions, substitutions, and modifications are permitted without departing from the spirit of the invention. These embodiments and their variations are included in the scope and spirit of the invention, as well as in the claims and their equivalents.
[0585] (Note) The details described in each of the above embodiments are noted below.
[0586] (First Addendum)
[0587] (Note 1) A program for operating a computer having one or more computer processors, wherein in a memory unit, it manages information about content that is the subject of viewing and is at least one of books or guidelines, and information about laws and regulations that include provisions of laws and regulations, and in the memory unit, it manages the information about the content and the information about laws and regulations as data in a graph-type data structure, and the program causes one or more computer processors to perform the steps of identifying the portion of the content in the memory unit that refers to laws and regulations, even if it is not written in a way that differs from the official name of the law, based on rules for referring to laws and regulations, and defining a reference relationship between the portion of the content that refers to the identified law and the information about the referred law.
[0588] (Note 2) The program described in Appendix 1 stores information about laws and regulations in a graph-type data structure in the memory unit, where the law name, article, paragraph, and item are each vertices, and reference relationships are defined such that each vertex of the article, paragraph, and item are sequentially parent-child relationships to the law name vertex. In the definition step, a reference relationship is defined between the part of the content that refers to a specific law and the vertex corresponding to the law name, article, paragraph, or item of the referred law.
[0589] (Note 3) The program described in Appendix 2, which, in its memory unit, distinguishes information on laws and regulations based on at least one of the amendment date or the effective date, and stores each in a graph-type data structure according to at least one of the respective amendment date or effective date, and in the defining step, refers to information on the publication date of the content, and defines a reference relationship between the portion of the content that refers to the law specified, and the vertex corresponding to the name of the law, article, paragraph, or item of the law mentioned, with respect to the most recent information on laws and regulations at the time of publication of the content.
[0590] (Note 4) The program described in Appendix 2 manages the content information in the memory unit as data in a graph-type data structure, where the information indicating the title of the book or guideline is at one vertex, and each part of the content divided into sections is at another vertex, with reference relationships defined between the vertex indicating the title of the book or guideline and the vertex of each section as parent-child relationships. In the definition step, reference relationships are defined between the vertex of each section in the content that refers to the law and the vertex of each article, paragraph, and item of the law.
[0591] (Note 5) The program described in Appendix 4 manages case law information in its memory unit, using the information identifying the case as the vertex, and at least one of the case summary or each part of the case judgment as the vertex, defining a parent-child relationship between the vertex identifying the case and the vertex of the case summary or each part of the case judgment. In the defining step, a reference relationship is defined between the vertex of each part that refers to the case in the content and at least one of the vertex of the case summary or each part of the case judgment, as described in Appendix 4.
[0592] (Note 6) The program, as described in any of Appendix 1 to 5, is configured in its memory unit to store dictionary information indicating the correspondence between abbreviated forms of laws and regulations and their formal names, and in the identification step, by referring to the dictionary information, identifies the formal name of the law in the portion of the content that refers to the law, and in the defining step, defines a reference relationship between the portion identified as referring to the law and the information of the law that was referred to, even if the formal name of the law is not stated in the content.
[0593] (Note 7) The program described in Appendix 6 has multiple forms of abbreviated notation for laws and regulations, and is configured to store dictionary information for each type of abbreviated notation for laws and regulations in its memory unit, and is configured to store information that identifies the type of abbreviated notation for laws and regulations for each piece of content, and in the identification step, identifies the full name of the law or regulation using the dictionary information corresponding to the type of abbreviated notation associated with the content.
[0594] (Note 8) A program described in any of the appendices 1 to 7, which, in the identification step, extracts portions of the content that do not contain the name of a law but do contain descriptions of articles, paragraphs, or clauses, and then identifies the official name of the law by referring to the title of the content or other parts of the content for the extracted portions.
[0595] (Note 9) A program as described in any of Appendix 2 to 8, which, in the identification step, extracts portions of the content that contain specific phrases used in conjunction with articles, paragraphs, or subparagraphs, and defines reference relationships between those portions and multiple articles, paragraphs, or subparagraphs, depending on the specific phrases in the extracted portions.
[0596] (Note 10) The program is one of the programs described in any of Appendix 5 to 9, which causes one or more computer processors to perform the steps of: calculating a parameter indicating the similarity of each piece of content based on reference relationship information defined for the content; and storing the calculated similarity parameter in a memory unit in association with the content.
[0597] (Note 11) The computer is a system that generates answers to questions, and the program is the program described in Appendix 10, which causes one or more computer processors to further search for content based on the content of a question entered by the user, to further extract content similar to the searched content based on a parameter indicating the similarity of the searched content, and to generate answers to questions using the extracted content.
[0598] (Note 12) The computer is a system that generates answers to questions, and the program is a program as described in any of Appendix 1 to 11, which further causes one or more computer processors to: extract content by referring to a graph-type data structure based on the content of a question entered by the user; generate a prompt that includes the content of the question and the extracted content, and a prompt that includes instructions to generate an answer by referring to the extracted content in response to the question; provide the generated prompt to an information processing system that performs language processing, thereby obtaining the answer generated in response to the prompt from the information processing system that performs language processing; and present the obtained answer to the user.
[0599] (Note 13) A method performed by a computer having one or more computer processors, wherein a storage unit manages information about content to be viewed, which is at least one of a book or a guideline, and information about laws and regulations, which includes provisions of laws and regulations, and the storage unit manages the information about the content and the information about laws and regulations, respectively, as data in a graph-type data structure, the method comprising: one or more computer processors performing the steps of: identifying a portion of the content in the storage unit that refers to a law or regulation, based on rules for referring to laws and regulations, even if the notation is different from the official name of the law or regulation; and defining a reference relationship between the portion of the content that refers to the identified law or regulation and the information about the referred law or regulation.
[0600] (Note 14) Information processing device, wherein in a storage unit, information of content that is subject to viewing and is at least one of books or guidelines, and information of laws and regulations including provisions of laws and regulations, the storage unit manages the information of content and the information of laws and regulations as data in a graph-type data structure, and the control unit of the information processing device performs the steps of identifying the portion of the content in the storage unit that refers to laws and regulations, even if it is not written in a way that differs from the official name of the law, based on rules for referring to laws and regulations, and defining a reference relationship between the portion of the content that refers to laws and regulations identified and the information of the referred law.
[0601] (Note 15) A graph-type data structure having a graph-type data structure for content information that is the subject of viewing and is at least one of books or guidelines, a graph-type data structure for information on laws and regulations that include provisions of laws and regulations, a graph-type data structure in which a reference relationship is defined between the part of the content that refers to laws and regulations and includes a notation different from the official name of the law, and the information of the referred law, and a graph-type data structure used in a process to identify the information of laws and regulations for which a reference relationship is defined in the content of the searched content, and to respond with the information of the identified law.
[0602] (Second addendum)
[0603] (Note 1) A program for operating a computer having one or more computer processors, wherein the memory unit is configured to store data in the form of a graph-type data structure in which reference relationships are defined between parts that represent the content of each piece of content for a set of pieces of content to be viewed, and the program causes one or more computer processors to perform the following steps: receiving input of a question from a user; searching for a set of pieces of content based on the content of the question; selecting a second piece of content that has a reference relationship defined with the first piece of content based on the data in the graph-type data structure for one or more first pieces of content that are the search results; generating a prompt that includes the content of the question, the searched first piece of content, and the selected second piece of content, and includes instructions that generate an answer to the content of the question by referencing the first piece of content and the second piece of content; providing the generated prompt to an information processing system that performs language processing, thereby obtaining an answer output by the information processing system that performs language processing; and presenting the obtained answer to the user.
[0604] (Note 2) The memory unit is configured to store data in a graph-type data structure that holds multiple types of content as multiple content items, and defines reference relationships between different types of content items, and is configured to store at least one of the following as multiple types of content items: information on books including books on law or information on guidelines including guidelines on law, information on laws and regulations, and information on precedents, and the program described in Appendix 1 searches for multiple types of content items in the search step.
[0605] (Note 3) The program described in Appendix 2 searches for at least one of several types of content, either book information or guideline information, based on the content of the question during the search step.
[0606] (Note 4) The program described in Appendix 3 has a memory unit that stores information on each of the parts into which the content of a book or guideline is divided, for at least one of the book information or guideline information that is held as multiple types of content, and in the search step, it searches for the book or guideline part corresponding to the content of the question and determines the ranking of the content corresponding to the content of the question according to the ranking of the search for each part of the content.
[0607] (Note 5) The program described in Appendix 4 performs the following steps in the search: vectorizing the content of the question and comparing the vector of the question content with the vectors of each part of the content to search for the part corresponding to the content of the question.
[0608] (Note 6) The program described in any of Appendix 4 to 5, wherein the memory unit holds data in a graph-type data structure that defines reference relationships between parts of the content of at least one of the contents of a book or a guideline and parts of other content, and the program causes one or more computer processors to perform the step of extracting multiple contents as groups in which the reference relationships defined from parts of content to other contents are similar.
[0609] (Note 7) The program described in Appendix 6, wherein in the selection step, the first content is selected, and in the extraction step, other content extracted as a group is selected as the second content.
[0610] (Note 8) The memory unit holds data in a graph-type data structure, where the title of at least one of the books or guidelines is at each vertex, and the content portion of at least one of the books or guidelines is at each vertex, defining a reference relationship between the title vertex and the content portion vertex in a parent-child relationship; the memory unit also holds data in a graph-type data structure, where the name of a law is at each vertex, and each of the articles, paragraphs, and clauses constituting the law is at each vertex, defining a reference relationship between the name of the law and each vertex in the order of articles, paragraphs, and clauses in a parent-child relationship; and the program causes one or more computer processors to execute a step of defining a reference relationship between parent vertices by aggregating the reference relationships of child vertices to parent vertices based on the reference relationships defined from content portion to other content.
[0611] (Note 9) The program described in Appendix 8 stores data in its memory unit as graph-type data structure data, with information identifying a precedent as the vertex, and the content of the precedent as at least one of the summaries of the precedents and the content of the judgment text divided into multiple parts as vertices, and defines a parent-child relationship between the vertex of information identifying the precedent and at least one of the vertices of the summaries of the precedents and the content of the judgment text.
[0612] (Note 10) A method performed by a computer having one or more computer processors, The memory unit is configured to store data in a graph-type data structure in which reference relationships are defined between parts that represent the content of each piece of content for multiple pieces of content to be viewed. The method is a method in which one or more computer processors perform the following steps: receiving input of a question from a user; searching for multiple pieces of content based on the content of the question; selecting a second piece of content that has a reference relationship defined with the first piece of content based on the graph-type data structure data for one or more first pieces of content that are the search results; generating a prompt that includes the content of the question, the searched first piece of content, and the selected second piece of content, and includes instructions that generate an answer to the content of the question by referencing the first piece of content and the second piece of content; providing the generated prompt to an information processing system that performs language processing, thereby obtaining an answer output by the information processing system that performs language processing; and presenting the obtained answer to the user.
[0613] (Note 11) Information processing device, wherein the storage unit is configured to store data in a graph-type data structure in which reference relationships are defined between parts that represent the content of each piece of content for a plurality of pieces of content to be viewed, and the control unit of the information processing device performs the following steps: receiving input of a question from a user; searching for a plurality of pieces of content based on the content of the question; selecting a second piece of content that has a reference relationship defined with the first piece of content based on the graph-type data structure data for one or more first pieces of content that are the search results; generating a prompt that includes the content of the question, the searched first piece of content, and the selected second piece of content, and includes an instruction that generates an answer to the content of the question by referencing the first piece of content and the second piece of content; obtaining an answer output by an information processing system that performs language processing by providing the generated prompt to an information processing system that performs language processing; and presenting the obtained answer to the user.
[0614] (Note 12) A graph-type data structure that has data in which, for multiple types of content to be viewed, reference relationships are defined between parts that represent the content of each content, even between different types of content, and the multiple types of content include at least one of either information on books including books on law or information on guidelines including guidelines on law, information on laws and regulations, and information on precedents, and is used in a system that generates answers to questions, to respond to the content of a question entered by a user with information on content corresponding to the content of the question, based on the reference relationships of each vertex defined in the data of the graph-type data structure.
Claims
1. A program for operating a computer having one or more computer processors, The memory unit is configured to store data in a graph-type data structure that defines reference relationships between the parts representing the content of each piece of content for multiple pieces of content to be viewed. The program is configured on one or more computer processors. Steps to receive user input for questions, Based on the content of the above question, the steps include: searching for the above multiple contents, A step of selecting a second content that has a defined reference relationship with the first content, based on the data of the graph-type data structure, for one or more first content items that are the search results, A step of generating a prompt that includes the content of the question, the retrieved first content, and the selected second content, and which includes an instruction that generates an answer to the content of the question by referring to the first content and the second content; The steps include: providing the generated prompt to an information processing system that performs language processing, thereby obtaining a response output by the information processing system that performs language processing; A program that performs the steps of presenting the obtained answer to the user.
2. The memory unit is configured to store data in a graph-type data structure that holds multiple types of content as the aforementioned multiple contents, and defines reference relationships between different types of content. As for the aforementioned multiple types of content, Information from books, including legal textbooks, or information from guidelines, including legal guidelines, at least one of the following: Information on laws and regulations, It is configured to store information on case precedents, The program according to claim 1, wherein the search step involves searching for the multiple types of content.
3. The program according to claim 2, wherein in the search step, based on the content of the question, searches for at least one of the multiple types of content, namely the information of the book or the information of the guideline.
4. In the aforementioned storage unit, information is stored for each of the parts obtained by dividing the content of the book or the content of the guidelines into multiple parts, for at least one of the information of the book or the information of the guidelines that is held as multiple types of content. In the aforementioned search step, Search the relevant section of the aforementioned book or guideline that corresponds to the content of the aforementioned question, The program according to claim 3, which determines the ranking of content corresponding to the content of the aforementioned question according to the ranking of each part of the content in the search results.
5. In the aforementioned search step, To vectorize the content of the above question, The program according to claim 4, which searches for the part corresponding to the content of the question by comparing the vector of the content of the question with the vectors of each part of the content.
6. In the aforementioned storage unit, data is held as data in the graph-type data structure that defines reference relationships between parts of the content of at least one of the book or guidelines and parts of other content. The program further provides the one or more computer processors with: The program according to claim 4, which performs the step of extracting multiple pieces of content as a group that have similar reference relationships defined from one part of the aforementioned content to other pieces of content.
7. The program according to claim 6, wherein in the selection step, other content extracted as a group in the extraction step is selected as the second content for the first content.
8. In the storage unit, as data for the graph-type data structure, data is held in which the title of at least one of the books or guidelines is a vertex, the content portion of at least one of the books or guidelines is a vertex, and a reference relationship is defined between the title vertex and the content portion vertex in a parent-child relationship. In the aforementioned storage unit, the data of the graph-type data structure is stored with the name of the law as the vertex, and each of the articles, paragraphs, and items constituting the law as the vertex, and a parent-child relationship is defined between the vertex of the name of the law and each vertex in the order of articles, paragraphs, and items, and reference relationships are defined between each vertex. The program further provides the one or more computer processors with: The program according to claim 2, which performs the step of defining a reference relationship between parent vertices by aggregating the reference relationships of child vertices to parent vertices based on the reference relationships defined from the aforementioned content to other content.
9. The program according to claim 8, wherein the memory unit holds data in which, as data of the graph-type data structure, information identifying a precedent is used as a vertex, and the content of a precedent is used as a vertex, with at least one of the summaries of precedents and the content of judgments divided into multiple parts as vertices, and a reference relationship is defined between the vertex of information identifying a precedent and at least one of the vertices of the summaries of precedents and the content of judgments in a parent-child relationship.
10. A method performed by a computer having one or more computer processors, The memory unit is configured to store data in a graph-type data structure that defines reference relationships between the parts representing the content of each piece of content for multiple pieces of content to be viewed. The above method involves one or more computer processors, Steps to receive user input for questions, Based on the content of the above question, the steps include: searching for the above multiple contents, A step of selecting a second content that has a defined reference relationship with the first content, based on the data of the graph-type data structure, for one or more first content items that are the search results, A step of generating a prompt that includes the content of the question, the retrieved first content, and the selected second content, and which includes an instruction that generates an answer to the content of the question by referring to the first content and the second content; The steps include: providing the generated prompt to an information processing system that performs language processing, thereby obtaining a response output by the information processing system that performs language processing; A method for performing the steps of presenting the obtained response to the user.
11. An information processing device, The memory unit is configured to store data in a graph-type data structure that defines reference relationships between the parts representing the content of each piece of content for multiple pieces of content to be viewed. The control unit of the information processing device, Steps to receive user input for questions, Based on the content of the above question, the steps include: searching for the above multiple contents, A step of selecting a second content that has a defined reference relationship with the first content, based on the data of the graph-type data structure, for one or more first content items that are the search results, A step of generating a prompt that includes the content of the question, the retrieved first content, and the selected second content, and which includes an instruction that generates an answer to the content of the question by referring to the first content and the second content; The steps include: providing the generated prompt to an information processing system that performs language processing, thereby obtaining a response output by the information processing system that performs language processing; An information processing device that performs the steps of presenting the acquired response to the user.
12. A graph-type data structure, For multiple types of content that are viewed, the system has a graph-type data structure that defines reference relationships between the parts representing the content of each piece of content, even between different types of content. As for the aforementioned multiple types of content, Information from books, including legal textbooks, or information from guidelines, including legal guidelines, at least one of the following: Information on laws and regulations, Including case law information, A graph-type data structure used in a system that generates answers to questions, for which the system responds to the content of a question entered by a user by providing information about content corresponding to the content of the question, based on the reference relationships of each vertex defined in the data of the graph-type data structure.
Citation Information
Patent Citations
Methods for electronic document searching and graphically representing electronic document searches
JP2017010580A