Information processing system, information processing apparatus, information processing method, and program

The information processing system enhances the efficiency of identifying relevant content in electronic documents by using vector indexes to manage document portions and event attributes, improving the accuracy of information retrieval.

JP2026005878APending Publication Date: 2026-01-16HIGASHI NIHON MEDICOM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024104491
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-27
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Conventional techniques for searching electronic document data are inefficient, as users often struggle to identify relevant information without keyword-based searches, leading to reduced efficiency in finding appropriate content related to a target event.

Method used

An information processing system that utilizes a terminal device and a database server configured to communicate, where the database server manages electronic document data using vector indexes based on document portions and event attributes, enabling precise identification of relevant content through a vector database.

Benefits of technology

Facilitates easier and more accurate identification of appropriate content related to a target event by converting document content into vector indexes, allowing for efficient retrieval of relevant information from electronic data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026005878000001_ABST
    Figure 2026005878000001_ABST
Patent Text Reader

Abstract

To more easily specify appropriate contents related to a target event from electronic data of a document.SOLUTION: In the information processing system 1, the search request unit 154 transmits a search request for acquiring electronic data of a document related to a target event to the database server 20. A document data storage part 271 manages the electronic data of the document on the basis of a vector index constituted with the part of the electronic data of the document as a unit and constituted corresponding to the attribute of an event to be an object. The retrieval request reception unit 253 receives a retrieval request from the terminal 10. A retrieval execution part 255 specifies a part of the electronic data of the document managed by the vector database as a retrieval result based on a vector index showing the attribute of an event to be a target in the retrieval request. A retrieval result transmission part 256 transmits the retrieval result specified by the retrieval execution part 255 to the terminal 10.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing system, an information processing device, an information processing method, and a program. [Background technology]

[0002] In recent years, opportunities to refer to electronic data of documents such as digitized books and articles on the Internet have increased. For example, medical professionals diagnose or provide guidance to patients by appropriately referring to electronic books or articles on the Internet that contain medical or pharmaceutical information. Patent Document 1 describes a technique for referencing electronic data of a document. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2023-066551 Summary of the Invention [Problem to be solved by the invention]

[0004] However, in conventional techniques, when searching for desired information, it is necessary to search for electronic data of documents using keywords or to refer to the table of contents to find pages that are thought to be relevant. In this case, the user may not necessarily be able to reach the desired information, which reduces the efficiency of searching for information from electronic document data. That is, in the conventional technology, it is difficult to identify appropriate content related to a target event from electronic data of a document.

[0005] An object of the present invention is to more easily identify appropriate content related to a subject matter from electronic data of a document. [Means for solving the problem]

[0006] In order to solve the above problem, an information processing system according to one aspect of the present invention comprises: An information processing system in which a terminal device used by a user and a database server that stores electronic data of documents are configured to be able to communicate with each other, The terminal device a search request means for transmitting a search request to the database server to acquire electronic data of the document related to the event of interest; a search result display means for displaying the search results transmitted from the database server in response to the search request; Equipped with The database server a vector database for managing the electronic data of the document based on a vector index configured in units of portions of the electronic data of the document, the vector index having a configuration corresponding to the attribute of the target event; a search request receiving means for receiving the search request from the terminal device; a document specifying means for specifying, as a search result, a portion of the electronic data of the document managed by the vector database based on the vector index that indicates an attribute of the target event in the search request; a search result transmission means for transmitting the search results identified by the document identification means to the terminal device; The present invention is characterized by comprising: [Effects of the Invention]

[0007] According to the present invention, it is possible to more easily identify appropriate content related to a target event from electronic data of a document. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a schematic diagram showing the system configuration of an information processing system 1 according to one embodiment of the present invention. [Figure 2] FIG. 8 is a diagram showing the hardware configuration of an information processing device 800 that constitutes each device. [Figure 3] FIG. 2 is a block diagram showing the functional configuration of a user terminal 10. [Figure 4] FIG. 2 is a block diagram showing the functional configuration of a database server 20. [Figure 5] 10 is a flowchart showing the flow of a vector database generation process executed by the information processing system 1. [Figure 6] FIG. 10 is a schematic diagram illustrating the concept of a vector database generation process. [Figure 7] 10 is a flowchart showing the flow of document information presentation processing executed by the information processing system 1. [Figure 8] FIG. 1 is a schematic diagram illustrating the concept of document information presentation processing. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0010] [First embodiment] The information processing system according to this embodiment converts the contents of electronic data of documents such as electronic books into vector indexes for each part, constructs a vector database, converts the phenomenon to be searched into a vector index, refers to the constructed vector database, and identifies the part of the electronic data of the document stored in the vector database that is similar to the vector index of the phenomenon to be searched. For example, the content of an e-book related to medicine or pharmacy can be converted into a vector index for each part and stored in a vector database, and the e-book's vector database can be searched based on the vector index of the patient to be searched (e.g., a patient who visited a medical institution). This makes it possible to appropriately identify the parts of the e-book's content that are relevant to the patient to be searched. Therefore, according to the present invention, it is possible to more easily identify appropriate content related to a target event from electronic data of a document such as an electronic book. Specific embodiments will be described below.

[0011] [System configuration of Information Processing System 1] 1 is a schematic diagram showing the system configuration of an information processing system 1 according to one embodiment of the present invention. In the following embodiment, an example will be described in which the information processing system 1 is installed in a dispensing pharmacy and is applied to searching for electronic books related to medicines. 1, the information processing system 1 includes a plurality of user terminals 10 and a database server 20, and the user terminals 10 and the database server 20 are configured to be able to communicate with each other via a network 30. In the information processing system 1 according to this embodiment, the user terminals 10 and the database server 20 can also be configured to be able to communicate with other information processing devices (such as other terminal devices or servers that provide electronic book data) not shown. In addition, the server included in the information processing system 1 can be configured as a single server computer, or as a virtual server in a server system configured by a plurality of server computers.

[0012] The user terminal 10 is a terminal device used by a user to search for electronic document data (here, e-book data) using the information processing system 1, and is configured, for example, by an information processing device such as a smartphone, mobile phone, tablet terminal, or PC (Personal Computer). The user terminal 10 also functions as a terminal device for using a predetermined business application. For example, the user terminal 10 functions as a terminal device for using an application of a medication history management system that manages the medication history of patients in a pharmacy. When the user terminal 10 needs to search for information stored in the database server 20 while using a business application, for example, the user terminal 10 searches for the desired information by sending a search request together with information about the target event (here, the target patient) to the database server 20. Note that the user terminal 10 may also send to the database server 20 information about the target patient that includes changes from information about the past patient, such as information extracted by comparing the target patient's past prescription with the target patient's current prescription.

[0013] Furthermore, the user terminal 10 displays a list of search results sent from the database server 20 in response to the search request on the display screen. At this time, the search results sent from the database server 20 display a list of summaries of the information searched for in the electronic data of the document (summary search results). When the user selects any information from the displayed list of summary search results, the user terminal 10 requests the database server 20 to send detailed information corresponding to the selected summary information. Upon receiving the detailed information from the database server 20, the user terminal 10 displays the received detailed information on the display screen. Note that the database server 20 may transmit the summary of the searched information and the detailed information corresponding to the summary to the user terminal 10 as search results, and when any information is selected from the displayed list of summary information, the user terminal 10 may display the already received detailed information on the display screen without communicating with the database server 20.

[0014] The database server 20 is configured with an information processing device such as a PC or a server computer, and has the functionality of a database for storing electronic document data. Furthermore, when the database server 20 acquires electronic document data to be stored, it breaks the document down into parts with specific semantic content and extracts patient condition chunks representing patient attributes from the content represented by the parts. Hereinafter, patient condition chunks generated from the electronic document data are referred to as "document-based patient condition chunks." Document-based patient condition chunks are data (vector data) representing a patient model described in the electronic document data as patient attributes highly related to diseases and medications, such as subjects prone to diseases and subjects taking medications. They are identified by elements representing medical and pharmaceutical information related to the patient in the electronic document data, such as the patient's gender, age, chief complaint, symptoms, prescribed medication, test data, changes in the patient's condition, medication history (medication history), medical history, and the patient's occupation. Changes in patient condition may include information obtained by comparing with past patient conditions, such as an increase, decrease, or change in prescribed medication, a change in dosage or administration, improvement or worsening of symptoms, initiation or discontinuation of over-the-counter medication, etc. Changes in patient condition may also include newly discovered patient preferences (e.g., frequent consumption of certain foods such as grapefruit) or lifestyle habits (e.g., starting or stopping exercise).

[0015] The database server 20 then generates a vector index based on the document-based patient condition chunk, associates the generated vector index with the electronic data of the document (electronic data of the document parts), and stores the vector index in the database. In this embodiment, the types of patient attributes extracted from the electronic data parts of the document are counted for all documents, and a vector data format (document-based patient condition chunk) is generated in which the counted number of types is the number of components (dimensions). When generating document-based patient condition chunks for each part of the electronic data of the document, the document-based patient condition chunks are generated using the elements included in each part of the electronic data of the document, and for elements not included in each part of the electronic data of the document, it is assumed that there is no corresponding component data in the document-based patient condition chunk.

[0016] Furthermore, when the database server 20 receives a search request from the user terminal 10, it generates a vector index based on information about the target event (here, the target patient) that was also sent. That is, the database server 20 generates a patient status chunk for the target patient from the information about the patient sent along with the search request, using data corresponding to the elements that make up the document-based patient status chunk (the target patient's gender, age, chief complaint, symptoms, prescribed medications, test-related data, changes in the patient's condition, medication history, medical history, the patient's occupation, etc.). The patient status chunk for the target patient is hereinafter referred to as the "target patient status chunk." The patient information sent along with the search request can include the contents of the prescription issued to the target patient (prescribed medications, etc.), the patient's medication history (medication history), the patient's interview results (information including the results of an interview about changes in the patient's condition, etc.), the patient's test value data, etc., and from this information, elements for generating the target patient status chunk (the patient's gender, age, chief complaint, symptoms, prescribed medications, changes in the patient's condition, medication history (medication history), medical history, the patient's occupation, etc.) are selected and used. Then, the database server 20 generates a vector index based on the target patient condition chunk.

[0017] Hereinafter, the vector index of the document-based patient condition chunk will be referred to as the "document vector index", and the vector index generated from the target patient condition chunk will be referred to as the "target vector index".

[0018] In this embodiment, the database server 20 is constructed as a vector database that implements large-scale language models (LLMs), and can utilize functions such as chunking information provided by the large-scale language models and generating vector indexes by embedding. In this embodiment, when the database server 20 acquires electronic data of a document, it uses the information chunking function of the large-scale language model to decompose the document into parts based on certain semantic content (for example, content separated by chapters, sections, headings, or pages shown in the table of contents). Hereinafter, the data of the document decomposed into parts based on certain semantic content will be referred to as "table of contents chunks." The database server 20 also extracts document-based patient condition chunks from the table of contents chunks. The database server 20 also generates vector indexes for the document-based patient condition chunks using the vector index generation function of the large-scale language model. Furthermore, when the database server 20 receives the information about the patient transmitted together with the search request, it generates a target patient condition chunk using the vector index generation function provided in the large-scale language model.

[0019] [Hardware configuration] Next, the hardware configuration of each device in the information processing system 1 will be described. In the information processing system 1, each device is configured by an information processing device such as a PC, a server computer, or a tablet terminal, and the basic configurations thereof are the same.

[0020] FIG. 2 is a diagram showing the hardware configuration of an information processing device 800 that constitutes each device. As shown in Figure 2, the information processing device 800 that constitutes each device includes a CPU (Central Processing Unit) 811, a ROM (Read Only Memory) 812, a RAM (Random Access Memory) 813, a bus 814, an input unit 815, an output unit 816, a memory unit 817, a communication unit 818, a drive 819, and an imaging unit 820.

[0021] The CPU 811 executes various processes according to a program recorded in the ROM 812 or a program loaded from the storage unit 817 into the RAM 813 . The RAM 813 also stores data and the like necessary for the CPU 811 to execute various processes.

[0022] The CPU 811, ROM 812, and RAM 813 are connected to one another via a bus 814. To the bus 814, an input unit 815, an output unit 816, a storage unit 817, a communication unit 818, a drive 819, and an imaging unit 820 are connected.

[0023] The input unit 815 is composed of various buttons, a microphone, etc., and inputs various information in response to instruction operations. The output unit 816 is composed of a display, a speaker, etc., and outputs images and sounds. The storage unit 817 is configured with a hard disk or a DRAM (Dynamic Random Access Memory), etc., and stores various data managed by each server. The communication unit 818 controls communication with other devices via the network 50 .

[0024] Removable media 831, such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory, is appropriately loaded into the drive 819. A program read from the removable media 831 by the drive 819 is installed in the storage unit 817 as needed. The imaging unit 820 is configured by an imaging device equipped with a lens, an imaging element, etc., and captures a digital image of a subject. When the information processing device 800 is configured as the database server 20, it is also possible to omit the imaging unit 820. When the information processing device 800 is configured as a tablet terminal, it is also possible to configure the input unit 815 using a touch sensor and place it over the display of the output unit 816, thereby providing a touch panel.

[0025] [Functional configuration] Next, the functional configuration of the information processing system 1 will be described.

[0026] [Functional configuration of user terminal 10] FIG. 3 is a block diagram showing the functional configuration of the user terminal 10. As shown in FIG. 3, a user interface display control unit (UI display control unit) 151, an application execution unit 152, a patient data acquisition unit 153, a search request unit 154, and a search result acquisition unit 155 function in the CPU 811 of the user terminal 10. In addition, a patient data storage unit 171, an application data storage unit 172, and a search result storage unit 173 are formed in the storage unit 817 of the user terminal 10.

[0027] The patient data storage unit 171 stores various data of patients who are the subject of conversation at the dispensing pharmacy, and stores, for example, the patient's gender, age, chief complaint, symptoms, prescribed medication, test data, changes in the patient's condition, medication history (medication history), medical history, and information on the patient's occupation type in association with the patient's name, health insurance card number, and other identification information. The patient data storage unit 171 also stores patient data acquired from a database that integrates and stores data on patients who visit the dispensing pharmacy, as well as newly entered patient data (new patient data). The application data storage unit 172 stores data related to various applications that can be used on the user terminal 10, for example, data related to applications for medication history management that can be used on the user terminal 10 (such as the address of the server that performs medication history management and the client program for medication history management). The search result storage unit 173 stores the search results (search results for electronic data of documents) sent from the database server 20 in response to the search request sent by the user terminal 10 in association with the search request.

[0028] The UI display control unit 151 displays a user interface screen (UI screen) for the user to perform various processes on the user terminal 10. For example, the UI display control unit 151 displays a screen for making a search request to the database server 20 and a screen for presenting search results sent from the database server 20. The UI display control unit 151 also displays a screen for applications available on the user terminal 10.

[0029] When an available application is started on the user terminal 10, the application execution unit 152 executes a program for that application. For example, when an application for medication history management is started on the user terminal 10, the application execution unit 152 executes a client program for communicating with the medication history management server, refers to the patient's medication history data stored in the medication history management server, and adds and stores the patient's medication history data.

[0030] The patient data acquisition unit 153 acquires data of a patient with whom a user (e.g., a pharmacist) who uses the user terminal 10 is having a conversation. For example, the patient data acquisition unit 153 acquires patient data from a database that integrates and stores data of patients who visit a dispensing pharmacy, or acquires data of a newly entered patient (data of a new patient).

[0031] The search request unit 154 transmits a search request to the database server 20 to search for electronic data of documents related to the target patient. At this time, the search request unit 154 transmits information about the target patient along with the search request to the database server 20. For example, information about the target patient, such as "gender: male, age: 60 years old, chief complaint: back pain for the past 2 days, symptoms: fever 38°C, medication history: Lipitor...", is transmitted to the database server 20. Furthermore, as described above, the search request unit 154 can transmit to the database server 20 information about the target patient that includes changes from previous information about the patient, such as information extracted by comparing the target patient's previous prescription with the target patient's current prescription. Note that when an application is launched by the application execution unit 152, if a user issues a request to search for electronic data of documents related to the target patient in response to a user instruction (such as inputting a search request command or clicking a search request button), the search request unit 154 transmits a search request to the database server 20 to search for electronic data of documents related to the target patient. In other words, when a user uses an application and needs to search for information such as literature related to a specific patient, the user can easily search the electronic data of the document stored in the database server 20 and obtain the appropriate information.

[0032] The search result acquisition unit 155 acquires the search results transmitted from the database server 20 in response to the search request transmitted by the search request unit 154. For example, the search result acquisition unit 155 acquires information representing the summary search results transmitted from the database server 20. Furthermore, when the user selects any information from the summary search results, the search result acquisition unit 154 requests the database server 20 to transmit detailed information corresponding to the selected summary information, and acquires the detailed search result information transmitted from the database server 20 in response to this request.

[0033] [Functional configuration of database server 20] FIG. 4 is a block diagram showing the functional configuration of the database server 20. As shown in FIG. 4, the CPU 811 of the database server 20 includes a document data acquisition unit 251, a chunking unit 252, a search request receiving unit 253, a vector index generating unit 254, a search execution unit 255, and a search result transmitting unit 256. The storage unit 817 of the database server 20 includes a document data storage unit 271 and a search result storage unit 272.

[0034] The document data storage unit 271 stores electronic data of various documents input to the database server 20, as well as document-based patient condition chunks and document vector indexes associated with this electronic data. In other words, the document data storage unit 271 constitutes a vector database. The electronic data of various documents stored in the document data storage unit 271 is broken down into parts based on certain semantic content (for example, content separated by chapters, sections, headings, or pages shown in the table of contents), document-based patient condition chunks are generated from the content represented by the parts, and document vector indexes are generated based on the document-based patient condition chunks. Therefore, when a vector index (target vector index) of an event to be searched (target patient condition chunk) is provided, it is possible to identify parts of a document that are similar in semantic content by calculating the similarity between them. The search result storage unit 272 stores the search result data acquired by the search execution unit 255 when a search request for searching electronic data of documents related to a target patient is received from the user terminal 10. In this embodiment, the search result data is output as a summary of the information searched for in the electronic data of the documents (summary search result).

[0035] The document data acquisition unit 251 acquires electronic data of documents stored in the database server 20. The document data acquisition unit 251 acquires, as electronic data of documents, documents published as electronic data (such as electronic books, data of academic papers, or articles on the Web), data of documents converted manually or automatically from paper documents, and the like.

[0036] The chunking processor 252 decomposes the electronic data of the document acquired by the document data acquisition unit 251 into parts for each specific semantic content of the document (for example, each content separated by chapters, sections, headings, or pages shown in the table of contents), and extracts document-based patient condition chunks representing patient attributes from the content represented by the parts. For example, the chunking processor 252 can use an information chunking function provided in the large-scale language model to decompose the electronic data of the document into parts and generate document-based patient condition chunks. The document-based patient condition chunks generated by the chunking processor 252 are stored in the document data storage unit 271. The chunking processor 252 also generates target patient condition chunks using data corresponding to elements constituting the document-based patient condition chunks from the information about the patient sent along with the search request. The chunking processor 252 can also generate target patient condition chunks using the information chunking function provided in the large-scale language model.

[0037] The search request receiving unit 253 receives a search request for searching electronic data of documents related to the target patient, which is transmitted from the user terminal 10. Since the search request received at this time includes information about the target patient, the search request receiving unit 253 also receives the information about the target patient.

[0038] The vector index generation unit 254 generates a document vector index based on the document-based patient condition chunk generated by the chunking processor 252. The document vector index generated by the vector index generation unit 254 is stored in the document data storage unit 271. A vector database is constructed by storing the document-based patient condition chunk and the document vector index in the document data storage unit 271 in association with the electronic data of the document. Furthermore, the vector index generation unit 254 generates a target vector index based on the target patient condition chunk generated by the chunking processing unit 252. Note that the vector index generation unit 254 can generate a document vector index and a target vector index using a vector index generation function provided in the large-scale language model.

[0039] The search execution unit 255 searches the electronic data of the document stored in the document data storage unit 271 based on the target vector index and the document vector index, and identifies the electronic data of the document (electronic data of a part of the document) corresponding to the content represented by the target vector index. At this time, the search execution unit 255 identifies the document vector index that is close in distance to the target vector index using an algorithm for determining similarity, such as cosine similarity calculation. In other words, a document-based patient condition chunk with content similar to the patient condition chunk of the target patient is identified, and search results according to the target vector index are identified. The search result sending unit 256 sends the search results corresponding to the target vector index identified by the search execution unit 255 to the user terminal 10 as search results (summary search results) in response to a search request from the user terminal 10. Furthermore, when the user terminal 10 requests transmission of detailed search result information by selecting any information from the summary search results, the search result sending unit 256 sends detailed information corresponding to the selected summary information to the user terminal 10. Note that the summary search results can be generated by summarizing electronic data of the document portion that is the search result identified according to the target vector index using a large-scale language model.

[0040] [Operation] Next, the operation of the information processing system 1 will be described.

[0041] [Vector database generation process] Fig. 5 is a flowchart showing the flow of the vector database generation process executed by the information processing system 1. Fig. 6 is a schematic diagram showing the concept of the vector database generation process. The flow of the vector database generation process shown in FIG. 5 will be described below with reference to FIG. 6 as needed. The vector database generation process is started in response to an instruction to execute the vector database generation process in the database server 20.

[0042] When the vector database generation process is started, in step S1, the document data acquisition unit 251 acquires electronic data of documents stored in the database server 20 (see FIG. 6(1)). In step S2, the chunking processing unit 252 breaks down the electronic data of the document acquired by the document data acquisition unit 251 into parts (table of contents chunks) for each certain semantic content in the document (for example, each content separated by chapters, sections, headings, pages, etc. shown in the table of contents) (see Figure 6(2)).

[0043] In step S3, the chunking processor 252 extracts a document-based patient condition chunk from one table of contents chunk (see FIG. 6(3)). In step S4, the vector index generating unit 254 generates a document vector index based on the document-based patient condition chunk (see FIG. 6(4)). In step S5, the vector index generating unit 254 stores the document-based patient condition chunk and the document vector index together with the electronic data of the document (electronic data of the document portion) in the document data storage unit 271 (see FIG. 6(5)).

[0044] In step S6, the chunking processor 252 determines whether or not processing has been completed for all table of contents chunks. If processing has not been completed for all table of contents chunks (that is, if there are unprocessed table of contents chunks), the determination in step S6 is NO, and the processing proceeds to step S3. On the other hand, if the processing has been completed for all table of contents chunks (that is, if there are no unprocessed table of contents chunks), the determination in step S6 is YES, and the processing proceeds to step S7.

[0045] In step S7, document data acquisition unit 251 determines whether or not processing of electronic data of all documents has been completed. If the processing of the electronic data of all documents has not been completed (that is, if there is electronic data of unprocessed documents), the determination in step S7 is NO, and the processing proceeds to step S1. On the other hand, if the processing has been completed for the electronic data of all documents (that is, if there is no electronic data of unprocessed documents), the determination in step S7 is YES, and the vector database generation processing ends. As a result of this processing, a vector database is constructed that aggregates electronic data of documents.

[0046] [Document information presentation processing] Fig. 7 is a flowchart showing the flow of the document information presentation process executed by the information processing system 1. Fig. 8 is a schematic diagram showing the concept of the document information presentation process. The flow of the document information presentation process shown in FIG. 7 will be described below with reference to FIG. 8 as needed. The document information presentation process is started in response to an instruction to execute the document information presentation process in the database server 20. When the document information presentation process is started, in step S11, the search request receiving unit 253 receives a search request for searching electronic data of documents related to the target patient, which is transmitted from the user terminal 10 (see FIG. 8(1)). The search request received at this time includes information about the target patient.

[0047] In step S12, the chunking processor 252 generates a target patient condition chunk from data corresponding to elements constituting the document-based patient condition chunk, out of the information about the patient sent together with the search request (see FIG. 8(2)). In step S13, the vector index generating unit 254 generates a target vector index based on the target patient chunk (see FIG. 8(3)). In step S14, the search execution unit 255 uses an algorithm for determining similarity, such as cosine similarity calculation, to identify document vector indexes that are close to the target vector index (see FIG. 8(4)). That is, a search of the vector database is executed.

[0048] In step S15, the search execution unit 255 identifies electronic data of a document (electronic data of a portion of a document) corresponding to the content represented by the target vector index (see FIG. 8(5)). In step S16, the search result transmission unit 256 transmits the search results corresponding to the target vector index identified by the search execution unit 255 to the user terminal 10 as search results (summary search results) in response to the search request from the user terminal 10 (see FIG. 8(6)). The search results transmitted at this time can be, for example, data in a list format in which summary search results are sorted in order of similarity. The user terminal 10 that has received the search results (summary search results) displays the search results in a list using the UI display control unit 151. After step S16, the document information presentation process ends.

[0049] Through this process, the contents of the electronic data of the document are converted into vector indexes for each part (each document-based patient condition chunk), and a vector database is constructed.The event to be searched (target patient condition chunk) is then converted into a vector index, and the constructed vector database is referenced to identify, among the electronic data of the document stored in the vector database, the part of the electronic data of the document that contains content similar to the vector index of the event to be searched. Therefore, according to the present invention, it is possible to more easily identify appropriate content related to a target event from electronic data of a document such as an electronic book.

[0050] [Variation 1] In the above-described embodiment, if the electronic data of the document includes extended data other than simple text data such as diagrams and tables, this extended data can also be stored in the vector database. For example, for extended data such as charts and tables, semantic content can be set using alternative text, and the alternative text can be stored in a vector database in association with the charts and tables. In this case, by generating a document vector index for the alternative text, the alternative text can be used as a search target based on the target vector index. In addition, since a figure, table, etc. may not constitute a semantic division equivalent to a section or chapter by itself, extended data can be stored in the vector database as data accompanying the text that refers to the figure, table, etc., and if the text has a high degree of similarity, the figure, table, etc. can be presented as a search result. This allows a variety of data other than text data to be included in searches based on vector indexes, making it easier to identify appropriate content related to the target event from the electronic data of documents.

[0051] As described above, the information processing system 1 according to this embodiment includes the user terminal 10 and the database server 20. The user terminal 10 includes the UI display control unit 151 and the search request unit 154, and the database server 20 includes the document data storage unit 271 (vector database), the search request receiving unit 253, the search execution unit 255, and the search result sending unit 256. The search request unit 154 sends a search request to the database server 20 to acquire electronic data of documents related to the event of interest. The UI display control unit 151 displays the search results sent from the database server 20 in response to the search request. The document data storage unit 271 manages electronic data of a document based on a vector index that is configured in units of portions of electronic data of the document and that corresponds to the attributes of the target event. The search request receiving unit 253 receives a search request from the terminal device 10 . The search execution unit 255 identifies, as search results, portions of electronic data of documents managed by the vector database based on vector indexes that indicate attributes of events that are targets of a search request. The search result transmission unit 256 transmits the search results identified by the search execution unit 255 to the terminal device 10 . As a result, the contents of the electronic data of the document are converted into vector indexes for each part, and a vector database is constructed. Then, by referring to the constructed vector database, parts of the electronic data of the document stored in the vector database that contain content similar to the vector index of the event to be searched are identified. Therefore, according to the present invention, it is possible to more easily identify appropriate content related to a target event from electronic data of a document such as an electronic book.

[0052] The database server 20 includes a vector index generation unit 254 . The vector index generating unit 254 chunks the electronic data of the document to be managed for each predetermined portion, and generates a vector index for each chunked portion, thereby generating a vector database. This makes it possible to generate a vector database that allows searches based on semantic content for each part of the electronic data of a document.

[0053] The search execution unit 255 identifies parts of the electronic data of the document that are highly relevant in content to the target event based on the similarity between the vector index representing the attributes of the target event and the vector index of the part of the electronic data of the document. This makes it possible to identify the portion of the electronic data of the document that contains content that is thought to match the attributes of the target event.

[0054] The electronic data portion of the document is divided into units of at least one of chapters, sections, headings, and pages shown in the table of contents. This makes it possible to perform a search using a vector index, with the semantic content division in the electronic data of the document as a unit.

[0055] The vector index of the portion of the electronic data of the document includes the contents of the diagrams included in the electronic data of the document as attributes. This allows searching the electronic data portion of a document, including not only the text data but also the content represented by diagrams and tables.

[0056] The document data storage unit 271 (vector database) manages electronic data of documents related to medicine or pharmacy. The event of interest is to search for electronic data of documents related to patients who visited a medical institution. This allows easy identification of portions of electronic medical or pharmaceutical documents that may be relevant to an individual patient.

[0057] The present invention is not limited to the above-described embodiment, and any modifications and improvements that can achieve the object of the present invention are included in the present invention. For example, in the above-described embodiment, the types of patient attributes extracted from the electronic data portions of a document are counted for all documents, and a vector data (document-based patient condition chunk) format is generated in which the counted number of types represents the number of components (dimensions). However, this is not limited to this. That is, the document-based patient condition chunk may be generated as a format of vector data containing only major specific attributes as components, and if content related to the major specific attributes is included in each portion of the electronic data of the document, it may be extracted as an element of the document-based patient condition chunk. In this case, the size of the document-based patient condition chunk and the target patient condition chunk can be reduced and configured as vector data with a fixed number of components, thereby improving search speed in the vector database.

[0058] Furthermore, for example, the database server 20 in the above-described embodiment can be configured as an on-premise server or a cloud server. Furthermore, the information processing system 1 according to the present invention can be configured by each device in the above-described embodiments, or the functions of multiple devices in the above-described embodiments can be implemented together in one device, or the functions of one device can be distributed and implemented across multiple devices.

[0059] The above-described series of processes can be executed by hardware or software. In other words, the functional configurations in the above-described embodiments are merely examples and are not particularly limited. That is, it is sufficient that any of the computers constituting the information processing system 1 has a function capable of executing the above-described series of processes as a whole, and the functional blocks used to realize this function are not particularly limited to the examples shown. Furthermore, one functional block may be configured as a single piece of hardware, a single piece of software, or a combination thereof.

[0060] Furthermore, the recording medium containing the program for executing the above-mentioned series of processes may be configured not only as a removable medium distributed separately from the device main body in order to provide the program to the user, but also as a recording medium provided to the user in a state where it is pre-installed in the device main body.

[0061] Although the embodiments of the present invention have been described above, the present invention is not limited to the above-described embodiments. Furthermore, the effects described in the present embodiments are merely a list of the most preferable effects resulting from the present invention, and the effects of the present invention are not limited to those described in the present embodiments. [Explanation of symbols]

[0062] 1 Information processing system, 10 User terminal, 20 Database server, 30 Network, 151 User interface display control unit (UI display control unit), 152 Application execution unit, 153 Patient data acquisition unit, 154 Search request unit, 155 Search result acquisition unit, 171 Patient data storage unit, 172 Application data storage unit, 173 Search result storage unit, 800 Information processing device, 811 CPU, 812 ROM, 813 RAM, 814 Bus, 815 Input unit, 816 Output unit, 817 Storage unit, 818 Communication unit, 819 Drive, 820 Imaging unit, 831 Removable media

Claims

1. An information processing system in which a terminal device used by a user and a database server that stores electronic data of documents are configured to be able to communicate with each other, The terminal device a search request means for transmitting a search request to the database server to acquire electronic data of the document related to the event of interest; a search result display means for displaying the search results transmitted from the database server in response to the search request; Equipped with The database server a vector database for managing the electronic data of the document based on a vector index configured in units of portions of the electronic data of the document, the vector index having a configuration corresponding to the attribute of the target event; a search request receiving means for receiving the search request from the terminal device; a document specifying means for specifying, as a search result, a portion of the electronic data of the document managed by the vector database based on the vector index that indicates an attribute of the target event in the search request; a search result transmission means for transmitting the search results identified by the document identification means to the terminal device; An information processing system comprising:

2. The database server The information processing system according to claim 1, further comprising a vector database generation means for generating the vector database by chunking the electronic data of the document to be managed into predetermined parts and generating the vector index for each chunked part.

3. The information processing system according to claim 1 or 2, characterized in that the document identification means identifies parts of the electronic data of the document that are highly relevant in content to the target event based on the similarity between the vector index representing the attributes of the target event and the vector index of a part of the electronic data of the document.

4. 3. The information processing system according to claim 1, wherein the electronic data portion of the document is in units of at least one of chapters, sections, headings, and pages shown in a table of contents.

5. 3. The information processing system according to claim 1, wherein the vector index of the portion of the electronic data of the document includes, as an attribute, the contents of a diagram or table included in the electronic data of the document.

6. The vector database manages electronic data of the medical or pharmaceutical documents; 3. The information processing system according to claim 1, wherein the electronic data of the document related to a patient who visited a medical institution is searched for as the event of interest.

7. An information processing device configured to be able to communicate with a terminal device used by a user and storing electronic data of a document, a vector database for managing the electronic data of the document based on a vector index configured in units of portions of the electronic data of the document, the vector index having a configuration corresponding to an attribute of a target event; a search request receiving means for receiving a search request for acquiring electronic data of the document related to the event of interest from the terminal device; a document specifying means for specifying, as a search result, a portion of the electronic data of the document managed by the vector database based on the vector index that indicates an attribute of the target event in the search request; a search result transmission means for transmitting the search results identified by the document identification means to the terminal device; An information processing device comprising:

8. An information processing method executed by an information processing system configured to enable communication between a terminal device used by a user and a database server storing electronic data of documents, comprising: The terminal device, a search request step of sending a search request to the database server to acquire electronic data of the document related to the event of interest; a search result display step of displaying the search results transmitted from the database server in response to the search request; Including, The database server, a vector database providing step of providing a function of a vector database that manages the electronic data of the document based on a vector index configured in units of portions of the electronic data of the document and configured to correspond to attributes of the target event; a search request receiving step of receiving the search request from the terminal device; a document specifying step of specifying, as a search result, a portion of electronic data of the document managed by the vector database based on the vector index representing an attribute of the target event in the search request; a search result transmission step of transmitting the search results identified in the document identification step to the terminal device; An information processing method comprising:

9. A computer constituting an information processing device configured to be able to communicate with a terminal device used by a user and storing electronic data of a document, a vector database providing function that provides a vector database function for managing the electronic data of the document based on a vector index configured in units of portions of the electronic data of the document and configured to correspond to attributes of a target event; a search request receiving function for receiving a search request for acquiring electronic data of the document related to the event of interest from the terminal device; a document identification function that identifies, as a search result, a portion of the electronic data of the document managed by the vector database based on the vector index that represents the attribute of the target event in the search request; a search result transmission function that transmits the search results identified by the document identification function to the terminal device; A program characterized by realizing the above.

Citation Information

Patent Citations

  • Browsing support system, browsing support program, information terminal, and browsing support method

    JP2023066551A