Information processing device and information processing system

By adopting large-scale language models and retrieval enhancement generation technology in the health management system, combining personal health examination data and impersonal data to generate personalized health management and disease prevention suggestions, it solves the shortcomings of traditional systems that cannot effectively deal with personalized health problems and data privacy and security issues, and achieves high-precision and safe health management and prevention suggestions.

JP7678633B1Active Publication Date: 2025-05-16PHENOGEN MEDICAL INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024184131
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-10-18
Publication Date
2025-05-16
Estimated Expiration
2044-10-18

AI Technical Summary

Technical Problem

Traditional health management systems have failed to make full use of personal health data, provided health advice too widespread, unable to effectively deal with individual health status and disease risks, and have data privacy and information security issues.

Method used

Using information processing equipment and systems, using large-scale language model (LLM) and search augmented generation (RAG) technology, personalized health management and disease prevention recommendations are generated through personal health examination data, lifestyle data and genetic information. The system ensures data privacy and security by searching for personal-related databases and impersonal databases.

Benefits of technology

Real-time health management and disease prevention suggestions based on personalized health data are realized, improving the accuracy and security of health predictions and suggestions, and ensuring the privacy and security of user data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007678633000001_ABST
    Figure 0007678633000001_ABST
Patent Text Reader

Abstract

Provided is an information processing device and an information processing system capable of presenting health advice that reflects personal health data while appropriately dealing with issues such as data privacy protection and security. [Solution] An information processing device equipped with a control unit, the control unit comprising: a search expansion generation unit that generates an inference request to be input to a machine-learned large-scale language model based on input information including personal identification information that can identify an individual; and a transmission / reception unit that sends an inference request to the large-scale language model and receives an inference result from the large-scale language model, wherein the search expansion generation unit generates an inference request influenced by the personal identification information based on a first search result obtained by searching a first database in which first information linked to an individual identified by the personal identification information is stored, and a second search result obtained by searching a second database in which second information consisting of non-personal information is stored.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to the fields of medical information technology and artificial intelligence technology, and in particular to a system that utilizes an individual's health checkup data, lifestyle data, genetic information, etc., and provides personalized health management and disease prevention advice using a large-scale language model (LLM) and Retrieval-Augmented Generation (RAG). [Background technology]

[0002] As an information processing technique, Patent Document 1 discloses a similar document search device having a search target document storage means for storing a plurality of search target documents, a search key document storage means for storing a plurality of search key documents that serve as keys for deriving the search target documents, a search means for searching for a predetermined document based on an input search string, an output means for outputting the predetermined document searched by the search means, a search example storage means for storing a plurality of search examples, and a registration means for registering a new search example in the search example storage means, in which the search means calculates a similarity for a combination of the search key document and the search target document based on a similarity between the input search string and each search key document stored in the search key document storage means and a similarity between the input search string and each search target document stored in the search target document storage means, and the output means outputs the searched specified search target document; the registration means, if the outputted specified search target document is one desired by a user, registers the specified search key document as a new search example in the search example storage means based on a user instruction; the search means, if the outputted specified search target document is not one desired by the user, searches for search examples in order of decreasing similarity based on a similarity between an input search string and each search example stored in the search example storage means, based on a user instruction; the output means outputs the search examples in order of decreasing similarity; and the registration means registers a search sentence edited based on the outputted search example in the search example storage means as a new search example.

[0003] As an artificial intelligence technology, Non-Patent Document 1 shows the effect of a large-scale language model in learning by several examples. In a large-scale language model, it is preferable to include appropriate information in the inference request (prompt) from the viewpoint of obtaining a desirable inference result, and for this purpose, Retrieval-Augmented Generation, a method of searching for related information from an external database and generating an answer using that information, may be introduced. Non-Patent Document 2 is known as technical information related to Retrieval-Augmented Generation.

[0004] The fusion of medical information technology and artificial intelligence technology has progressed rapidly in recent years. For example, Patent Document 1 proposes a medical support device that includes a machine learning unit that performs machine learning related to the generation of advice, using a combination of text indicating a patient's condition and advice to be used in the treatment of the patient as training data for a large-scale language model that has been pre-trained with more than 500 GB of text, an acquisition unit that acquires text indicating the patient's condition (first text), and a generation unit that inputs the first text into the large-scale language model and causes the large-scale language model to generate advice to be used in the treatment of the patient, wherein the machine learning unit performs machine learning using a combination of a medical record of a subject of a clinical trial and the number of subjects required for the clinical trial as training data, the acquisition unit acquires text of a plurality of the medical records (sixth text), and the generation unit inputs the sixth text and generates advice indicating the number of subjects, and generates advice indicating the number of subjects required for the clinical trial related to the medical record.

[0005] In addition, Non-Patent Document 3 discloses a technology that uses an AI system that has learned from university medical record information to provide health risk prediction information as a tool for health management. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] Patent No. 7300237 [Patent Document 2] Patent No. 7454090 [Non-patent literature]

[0007] [Non-Patent Document 1] Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, Dario Amodei, Language Models are Few-Shot Learners, https: / / arxiv.org / pdf / 2005.14165, arXiv:2005.14165, 22 Jul 2020. [Non-Patent Document 2] Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Kuttler, Mike Lewis, Wen-tau Yih, Tim Rocktaschel, Sebastian Riedel, Douwe Kiela, Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, https: / / arxiv.org / pdf / 2005.11401, arXiv:2005.11401, 12 Apr 2021. [Non-Patent Document 3] Shinji Ueno, "Concrete Results of AI Lifestyle Disease Risk Prediction Expanded to Six Diseases for Health Management and Healthy Business Management," New Medical, October 2023 issue, page 90 Summary of the Invention [Problem to be solved by the invention]

[0008] Conventional health management systems are unable to fully utilize personal health data and are limited to providing general health advice. As a result, they are unable to respond appropriately to individual health conditions and risks, and are limited in their ability to predict disease early or provide specific preventive measures. In addition, issues regarding data privacy protection and information security are also an issue.

[0009] Therefore, the present invention aims to provide an information processing device and information processing system that provide health advice using a large-scale language model, and that is capable of providing personalized health advice that appropriately reflects personal health data while appropriately responding to issues such as data privacy protection and information security. [Means for solving the problem]

[0010] The present invention, provided to solve the above problems, includes the following aspects. (1) An information processing device comprising: a control unit having a search expansion generation unit that generates an inference request to be input to a large-scale language model that has been machine-learned based on input information including personal identification information that can identify an individual; and a transceiver unit that sends the inference request to the large-scale language model and receives an inference result from the large-scale language model, wherein the search expansion generation unit generates the inference request influenced by the personal identification information based on a first search result obtained by searching a first database in which first information linked to the individual identified by the personal identification information is stored, and a second search result obtained by searching a second database in which second information consisting of non-personal information is stored.

[0011] (2) The information processing device according to (1) above, wherein the inference request created by the search extension generation unit does not include the personally identifiable information.

[0012] (3) The information processing device described in (1) above, wherein the search extension generation unit searches the first database before searching the second database, and sets search conditions for the second database based on the first search results.

[0013] (4) The information processing device according to (3) above, wherein the search conditions of the second database do not include the personally identifiable information.

[0014] (5) The information processing device described in (1) above, further comprising a personal-related information database in which personal-related information useful for searching the first information linked to the personally identifying information is stored, and the search extension generation unit sets first search conditions for obtaining the first search results based on the personal-related information obtained by searching the personal-related information database.

[0015] (6) The information processing device according to (5) above, wherein the first search result does not include the personally identifying information and does not include the personally related information.

[0016] (7) The information processing device according to (5) above, wherein the personal information includes at least one of health check information and social insurance medical fee statement information.

[0017] (8) The information processing device according to (1) above, wherein the first information includes one or more types selected from the group consisting of health risk data, stress visualization data, and health risk fluctuation data of the individual.

[0018] (9) The information processing device according to (1) above, wherein the second information includes one or more types selected from the group consisting of medical papers, health management related information, nutrition and diet related information, motor function related information, and doctor's findings related information.

[0019] (10) The information processing device according to (1) above, wherein the large-scale language model is managed in a closed environment with restricted access.

[0020] (11) The information processing device according to (1) above, wherein the first database is managed in a closed environment with restricted access.

[0021] (12) The information processing device according to (1) above, wherein the control unit further includes a display signal generation unit that receives the inference result from the large-scale language model and generates a display signal including at least a portion of the inference result.

[0022] (13) The information processing device according to (1) above, wherein the control unit further includes a teacher data creation unit that creates teacher data based on the inference request, the inference result, and corrections to the inference result.

[0023] (14) An information processing system comprising: a first device having an input / output function; a second device having a data storage function; and a third device having a control unit, wherein the first device is capable of communicating with the second device and the third device, wherein the second device has a model storage device in which a large-scale language model that has been machine-learned is stored, a first database in which first information linked to an individual identified by personal identification information input to the first device is stored, and a second database in which second information consisting of non-personal information is stored, wherein the control unit of the third device has a search expansion generation unit that generates an inference request to be input to the large-scale language model based on information input from the first device, wherein the first device has a transceiver unit that transmits the inference request input from the third device to the second device having the large-scale language model and receives an inference result from the second device, wherein the search expansion generation unit generates the inference request influenced by the first information included in the first search result based on a first search result including the first information obtained by searching the first database and a second search result including the second information obtained by searching the second database.

[0024] (15) The information processing system according to (14) above, wherein at least two of the first device, the second device and the third device are integrated.

[0025] (16) An information processing system having a plurality of computer devices constituting a peer-to-peer network, wherein at least one of the plurality of computer devices has a search expansion generation unit that generates an inference request to be input to a large-scale language model that has been machine-learned based on input information, and at least one of the plurality of computer devices has a transceiver unit that transmits the inference request to the large-scale language model and receives an inference result from the large-scale language model, and at least one of the plurality of computer devices has an input information generation unit that generates information to be input to the search expansion generation unit, and at least one of the plurality of computer devices has a display signal generation unit that generates a display signal including at least a portion of the inference result, wherein the search expansion generation unit generates the inference request influenced by the first information included in the first search result based on a first search result including the first information obtained by searching a first database in which first information linked to an individual identified by personal identification information is stored, and a second search result including the second information obtained by searching a second database in which second information consisting of non-personal information is stored. Effect of the Invention

[0026] According to the present invention, the following effects can be obtained. (1) By utilizing diverse health data of individuals, personalized health management and disease prevention advice can be provided in real time. (2) By leveraging large-scale language models and search expansion generation, users are provided with advanced and accurate health predictions and advice. (3) It enables safe management and analysis of data while protecting user privacy. [Brief description of the drawings]

[0027] [Figure 1] 1 is a block diagram illustrating an information processing system having an information processing device according to an embodiment of the present invention. [Diagram 2] FIG. 1 is a flow diagram (part 1) illustrating information processing performed by an information processing device according to an embodiment of the present invention. [Diagram 3] FIG. 2 is a flow diagram (part 2) illustrating information processing performed by the information processing device according to one embodiment of the present invention. [Figure 4] FIG. 11 is a block diagram illustrating a modified example of an information processing system having an information processing device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0028] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0029] FIG. 1 is a block diagram illustrating the functions of an information processing system having information processing according to an embodiment of the present invention.

[0030] As shown in FIG. 1, an information processing system 1000 according to one embodiment of the present invention includes an information processing device 100, a large-scale language model 200 that has undergone machine learning, a personal-related information database (in FIG. 1, "database" is abbreviated to "DB") 300, a first database 400 that stores personal information, and a second database 500 that stores non-personal information.

[0031] The information processing device 100 includes an input / output unit 110 that exchanges data with a user interface, a control unit 100C that performs information processing, and a transmission / reception unit 130 that exchanges data with a large-scale language model 200. The control unit 100C has a search extension generation unit 120, a display signal generation unit 140, and a teacher data creation unit 150.

[0032] The input / output unit 110 receives information from a user of the information processing system 1000, which is input to a user interface such as a wearable device typified by a smartphone or a personal computer installed in a medical institution, and outputs the results of information processing performed by the information processing system 1000 to the user interface as data to be displayed to the user of the system. In this specification, the term "display" includes showing an image or video on an image display device, emitting sound from an audio output device such as a speaker or earphone, and printing text information or image information using a printing device. The input / output unit 110 may have a user interface function, in which case it may include input devices such as a keyboard or a microphone, and may include output devices such as an image display device, a speaker, or a printing device.

[0033] The search extension generation unit 120 receives input information data including individual identifying information capable of identifying an individual from the input / output unit 110, and generates an inference request based on this input to be input to the machine-learned large-scale language model 200. The specific processing performed by the search extension generation unit 120 will be described later.

[0034] The transmission / reception unit 130 transmits an inference request to the large-scale language model 200 and receives an inference result from the large-scale language model 200. The transmission / reception unit 130 can also transmit the teacher data created in the teacher data creation unit 150 to the large-scale language model 200. The communication environment between the transmission / reception unit 130 and the large-scale language model 200 is not limited. They may be arranged in a common device and connected by wiring within the device, or the transmission / reception unit 130 and the large-scale language model 200 may be arranged in different devices but connected by a dedicated line. The Internet may exist between the transmission / reception unit 130 and the large-scale language model 200. In that case, it is preferable that a process for increasing the level of information security is performed, such as encrypting the communication data.

[0035] The display signal generation unit 140 receives as input the inference result received by the transmission / reception unit 130, and creates data for displaying the answer / report to be presented to the user of the system based on the information described in the inference result, and outputs the data to the input / output unit 110.

[0036] When a user of the system checks the answer / report generated by the display signal generating unit 140 and presented via the input / output unit 110, corrects the contents of the answer / report, and the corrections are input via the input / output unit 110, the teacher data creating unit 150 creates teacher data including the corrections and sends the teacher data to the large-scale language model 200 via the transmitting / receiving unit 130.

[0037] The large-scale language model 200 has already undergone machine learning, and when an inference request is input from the transmission / reception unit 130, the large-scale language model 200 performs an inference process, generates an inference result, and transmits it to the transmission / reception unit 130. The large-scale language model 200 may be provided as an open source or a closed source. When provided as an open source, the large-scale language model 200 can be managed in a closed environment with limited access, and by managing it in this way, the learning contents of the large-scale language model 200 can be exclusively used by users of the present system. This makes it impossible for third parties to enjoy the results of fine-tuning of the large-scale language model 200, and such operation may be preferable from the viewpoint of increasing the level of information security.

[0038] The individual-related information database 300 is a database that accumulates related information linked to an individual indicated by individual identification information input by a user of the present system to the search extension generation unit 120 via the input / output unit 110. The type of related information is not particularly limited, but examples include individual health checkup information (individual health checkup information 310) and social insurance medical fee statement information (receipt information 320) issued by a medical institution, and may include social insurance medical fee information and information on medical history. Since the individual-related information database 300 includes personal information, in a specific example, personal health information, in one example, it is managed in a closed environment with restricted access.

[0039] The first database 400 is a database that accumulates first information linked to an individual identified by the individual identifying information. The first database 400 is searched by the first search criteria created by the search extension generating unit 120, and outputs data including information on the health of the individual identified by the individual identifying information as the first search result. Examples of data included in the first search result include data including health risk prediction information (health risk data 410), data including information that visualizes stress (stress visualization data 420), and data indicating the degree of fluctuation in health risk (health risk fluctuation data 430).

[0040] In one example, the data included in these first search results are accumulated data generated in external systems. Specifically, health risk data 410 is a calculation result of an externally provided health risk prediction system, stress visualization data 420 is a calculation result of an externally provided stress (fatigue level) calculation system, and health risk fluctuation data 430 is a calculation result of an externally provided health risk fluctuation prediction system. First database 400 may acquire and accumulate data that is a calculation result from each external system via, for example, a WebAPI.

[0041] A specific example of the health risk data 410 is a three-stage prediction result of the possibility of developing diabetes in three years, five years, and ten years. A specific example of the stress visualization data 420 is data showing which of a plurality of stages the stress level or mood falls into, data showing the cortisol level that increases when one feels stress, and quantitative data of fatigue level evaluated by integrating these data. A specific example of the health risk fluctuation data 430 is data showing the fluctuation trend of the current health risk based on data on health risks at different times in the past.

[0042] Since the first database 400 includes the first information, which is personal information, in one example, it is managed in a closed environment with limited access. The first information stored in the first database 400 may be personal-related information. In that case, the first database 400 also has the function of the personal-related information database 300.

[0043] The second database 500 accumulates the second information, which is non-personal information. The second database 500 is searched by the second search conditions created by the search extension generating unit 120, and outputs data including non-personal information as the second search result. Examples of data included in the second search result include medical papers 510 of each medical department, information related to health management (health management related information 520), information related to nutritional diet (nutritional diet related information 530), information related to motor function (motor function related information 540), and information provided by a doctor in response to the health risk data 410 (specific examples include the doctor's opinion, treatment plan, and measures to prevent exacerbation, and are referred to as "doctor's opinion related information 550" in this specification).

[0044] This secondary information is not information that identifies an individual, but general information related to medical care and health, and much of it can be collected from the Internet. In addition, for academic papers such as medical papers 510 on each medical department, if a contract with the organization that publishes the academic paper is properly made, the desired paper can be easily collected.

[0045] For example, the doctor's findings related information 550 includes comments that a doctor, who is a user of a health risk prediction engine 610 (see FIG. 4) that outputs health risk data 410 as shown in Non-Patent Document 3, gives advice to improve the health condition of a specific individual based on information indicated by the health risk data 410 (e.g., the current risk of diabetes is the second of a three-level evaluation) obtained by inputting the medical and health data of the specific individual. Such comments included in the doctor's findings related information 550 are linked to specific first information (e.g., the current risk of diabetes is the second of a three-level evaluation) stored in the first database 400, such as the health risk data 410, and are stored in the second database 500. However, the doctor's findings related information 550 is not linked to information that identifies an individual, and therefore belongs to non-personal information.

[0046] FIG. 2 is a flow diagram (part 1) illustrating information processing performed by an information processing device according to one embodiment of the present invention, and FIG. 3 is the same flow diagram (part 2).

[0047] First, the user interface of the user outputs input information data indicating a question including personal identification information to the information processing device 100, and the input information data including the personal identification information is input to the input / output unit 110 of the information processing device 100 (step S101). The input / output unit 110 outputs the input information data to the search extension generation unit 120 of the control unit 100C, and upon receiving the input of the input information data, the search extension generation unit 120 creates search conditions including the personal identification information and searches the personal-related information database 300 with the search conditions. As a result, the personal-related information database 300 outputs personal-related information such as individual medical checkup information 310 and medical receipt information 320 to the search extension generation unit 120 (step S102).

[0048] The search extension generation unit 120 uses the input personal related information to create a first search condition for searching the first database 400, and searches the first database 400 (step S103). If the data stored in the first database 400 is managed by facility where the medical checkup or treatment was received, the search extension generation unit 120 identifies the facility where the target individual received the medical checkup or treatment from the input personal related information, and adds information about the authority to access information in the facility to the first search condition. The first database 400 outputs the health information of the individual that satisfies the first search condition as the first search result (step S104).

[0049] In the case where the first database 400 not only stores data but also has an information processing unit 600 (see FIG. 4) that executes the health risk prediction engine 610 and the like, the search extension generating unit 120 sets the first search condition in step S103 so as to include the health-related information (height, weight, blood sugar level, etc.) of a specific individual that reflects the individual-related information obtained from the individual-related information database 300. The first database 400 executes the health risk prediction engine 610 using the health-related information of the specific individual included in the first search condition as an input to generate first information such as the health risk data 410, and outputs the first search result including the first information of the specific individual thus generated (step S104). The stress visualization data 420 and the health risk fluctuation data 430 included in the first search result may also be the result generated by the information processing executed by the first database 400 (the result of executing the stress visualization engine 620, the result of executing the health risk fluctuation prediction engine 630).

[0050] An information processing unit 600 having a health risk prediction engine 610 for generating health risk data 410, a stress visualization engine 620 for generating stress visualization data 420, and a health risk fluctuation prediction engine 630 for generating health risk fluctuation data 430 may be provided separately from the first database 400, and at least a part of the first search condition may be input to this information processing unit 600. In this case, the information processing unit 600 performs a predetermined information processing, and the data resulting from the processing (health risk data 410, stress visualization data 420, health risk fluctuation data 430) may become a part of the first information, and the first search result including the first information may be output from the information processing unit 600. Such a configuration is shown in FIG. 4 as a modified example of the information processing system 1000 according to this embodiment. As described above, the part surrounded by the dashed line in FIG. 4 may become the first database 400.

[0051] The search extension generating unit 120 uses the first search result from the first database 400 as an input, generates second search criteria, and searches the second database 500 (step S105). The first search result includes information that identifies an individual's health condition, such as the health risk data 410, and the second search criteria are set to include such a specific health condition. Therefore, the second search result searched for using the second search criteria and output from the second database 500 (step S106) is more likely to include information that is useful for the specific health condition.

[0052] For example, if the second search condition includes information that the current risk of diabetes is the second of a three-level evaluation, the second database 500 outputs, as the second search result, doctor's finding-related information 550 that is valid when the risk of diabetes is the second of a three-level evaluation. As a result, information that is valid when the risk of diabetes is the third of a three-level evaluation, or conversely, information that is valid when the risk is the first of a three-level evaluation, is not included in the second search result. If the second search result includes information that is not appropriate for the level of a specific individual's risk of diabetes, the content of the inference request generated by the search extension generation unit 120 thereafter will differ from the content that should be. In such a case, the accuracy (degree of validity) of the information included in the inference result will naturally decrease.

[0053] Moreover, when the second search condition includes information that the risk of diabetes is moderate, the doctor's findings related information 550 included in the second search result will not be only general information for not contracting or aggravating diabetes, such as "Get moderate exercise and sleep well." Such general information generalizes the content of the inference request subsequently generated by the search extension generation unit 120, and is therefore rather intrusive from the viewpoint of improving the accuracy of the information included in the inference result.

[0054] In this way, the second information stored in the second database 500 is non-personal information, but since the second search conditions are based on the first search results reflecting the individual-related information, the second search results obtained by searching the second database 500 include information that is meaningful to a specific individual. In other words, the non-personal information included in the second search results output in step S106 has a bias in content corresponding to a specific individual.

[0055] The search extension generation unit 120 generates an inference request using the second search result including non-personal information with biased content and the first search result including personal health information (step S107). The inference request generated by the search extension generation unit 120 is output to the transmission / reception unit 130, and the data including the inference request is transmitted to the large-scale language model 200 by the transmission / reception unit 130 (step S108).

[0056] The large-scale language model 200 to which the transmission / reception unit 130 transmits is not limited to one type, and may be transmitted to a plurality of language models. Furthermore, if the inference request includes information for determining which language model to transmit to, the transmission / reception unit 130 transmits the inference request to a specific language model identified based on the information.

[0057] The large-scale language model 200 receives an inference request from the transmission / reception unit 130, performs inference processing, and generates an inference result (step S109). The inference result generated by the large-scale language model 200 is transmitted to the transmission / reception unit 130 of the information processing device 100, and the transmission / reception unit 130 receives the inference result (step S110).

[0058] The inference result received by the transmitting / receiving unit 130 may be displayed as it is on the user interface, or may be converted into information that is easier for the user to understand as a primary answer. Specifically, the inference result received by the transmitting / receiving unit 130 is input to the display signal generating unit 140, which creates a display signal such as text or a chart based on the information included in the inference result, and creates a secondary answer such as an answer or a report (step S111). Data including this secondary answer is output to the input / output unit 110, and the user interface receives this data, and based on the data, displays the secondary answer such as an answer or a report on an image display device or prints it on a printer (step S112).

[0059] The user who recognizes the inference result output by the information processing device 100 through the above information processing can input corrections, such as adding further comments to the result, to the user interface (step S113). The input corrections are input from the user interface to the input / output unit 110 of the information processing device 100 (step S114). Then, they are input to the teacher data creation unit 150 together with the inference request generated by the search extension generation unit 120 and the secondary answer from the display signal generation unit 140. The teacher data creation unit 150 generates teacher data including the inference request and the inference result reflecting the corrections based on the input data (step S115) and outputs the generated teacher data to the transmission / reception unit 130. Note that the primary answer may be input to the teacher data creation unit 150 in addition to the secondary answer, or the primary answer may be input instead of the secondary answer.

[0060] The transmitting / receiving unit 130 outputs the training data to the large-scale language model 200 (step S116), and the large-scale language model 200 performs fine tuning using the input training data (step S117).

[0061] When the large-scale language model 200 is located at a location distant from the information processing device 100, the inference request, the inference result, and the teacher data are transmitted from the information processing device 100 to the large-scale language model 200 via an open communication environment such as the Internet, but the large-scale language model 200 may be located at the same physical location as the information processing device 100, for example, in the same device. When the large-scale language model 200 is localized in this way, the resources (such as power) required for communication are reduced, and the operation efficiency of the information processing system 1000 may be improved. In addition, the large-scale language model 200 can be managed in a closed environment with limited access. In this case, it is difficult for a third party to access the large-scale language model 200, so the possibility of information leakage may be lower than when an open communication environment is used. This is advantageous from the viewpoint of increasing the level of information security.

[0062] On the other hand, since the localized large-scale language model 200 can only obtain the teacher data from the teacher data creation unit 150, the level of learning can only be improved by operating the information processing system 1000. From the viewpoint of efficiently improving the level of learning, the large-scale language model 200 may have a federated learning function. Specifically, the improvements and local language models generated by the localized large-scale language model 200 included in the information processing system 1000 are transmitted to a large-scale language model (central language model) that has a common basic configuration and belongs to a central integration environment.

[0063] The central language model generates a comprehensive improvement or global language model by integrating the improvements or local language models received from the information processing system 1000, and improvements or local language models from other large-scale language models (local language models) that are not included in the information processing system 1000 and have a common basic configuration with the central language model. The generated comprehensive improvement or global language model is transmitted from the central integration environment to localized large-scale language models including the large-scale language model 200, and is shared (updated).

[0064] According to such federated learning, specific information of the inference requests of each localized large-scale language model is ignored, and only abstracted information is sent to the central integration environment, so that the independence of each large-scale language model (non-sharing of inference requests and inference results) can be ensured. Note that in the integrated learning, there may be no central integration environment, and only multiple localized large-scale language models exist, which share information with each other to improve the learning level.

[0065] The above-described embodiments are described for the purpose of facilitating understanding of the present invention, and are not described for the purpose of limiting the present invention. Therefore, each element disclosed in the above embodiment is intended to include all design modifications and equivalents that fall within the technical scope of the present invention.

[0066] In the above embodiment, the information processing device 100 includes the input / output unit 110, the control unit 100C, and the transmission / reception unit 130, but is not limited thereto. An information processing system 1000 according to another embodiment of the present invention includes a first device having an input / output function, i.e., at least the transmission / reception unit 130, a second device having a data storage function, and a third device having the control unit 100C, and the first device can communicate with the second device and the third device.

[0067] The second device has a model storage device in which a large-scale language model 200 that has been machine-learned is stored, a first database 400 in which first information linked to an individual identified by personal identification information input to the first device is stored, and a second database 500 in which second information consisting of non-personal information is stored.

[0068] The control unit 100C of the third device has a search expansion generation unit 120 that generates an inference request to be input to the machine-learned large-scale language model 200 based on information input from the first device. The first device has a transmission / reception unit 130 that transmits the inference request input from the third device to a second device having the large-scale language model 200 and receives an inference result from the second device.

[0069] The search expansion generation unit 120 generates an inference request influenced by the first information included in the first search result, based on the first search result including the first information obtained by searching the first database 400 included in the second device, and the second search result including the second information obtained by searching the second database 500 included in the second device. Note that since the search expansion generation unit 120 is included in the third device, the search performed by the search expansion generation unit 120 is executed via the first device.

[0070] In the information processing system 1000 according to the other embodiment of the present invention described above, at least two of the first device, the second device, and the third device may be integrated.

[0071] An information processing system 1000 according to another embodiment of the present invention includes a plurality of computer devices forming a peer-to-peer network. At least one of the computer devices included in the information processing system 1000 includes a search expansion generating unit 120 that generates an inference request to be input to a large-scale language model 200 that has undergone machine learning, based on input information. At least one of the computer devices included in the information processing system 1000 also includes a transmitting / receiving unit 130 that transmits an inference request to the large-scale language model 200 and receives an inference result from the large-scale language model 200.

[0072] At least one of the computer devices included in the information processing system 1000 has an input information generation unit that generates information to be input to the search expansion generation unit 120. At least one of the computer devices included in the information processing system 1000 has a display signal generation unit 140 that generates a display signal including at least a part of the inference result. The search expansion generation unit 120 generates an inference request influenced by the first information included in the first search result, based on a first search result including the first information obtained by searching a first database 400 in which first information linked to an individual identified by personal identification information is stored, and a second search result including the second information obtained by searching a second database 500 in which second information consisting of non-personal information is stored.

[0073] From the viewpoint of increasing the level of information security, the control unit 100C may have an inference request confirmation unit for confirming that the inference request generated by the search extension generation unit 120 does not contain information that can identify an individual. In this case, only the inference request that has been confirmed by the inference request confirmation unit as not containing information that can identify an individual is transmitted from the transmission / reception unit 130 to the large-scale language model 200. The transmission / reception unit 130 may have the function of the inference request confirmation unit. [Explanation of symbols]

[0074] 1000: Information processing systems 100: Information processing device 100C: Control section 110: Input / output section 120: Search extension generation unit 130: Transmitter / receiver 140:Display signal generation section 150: Teacher Data Creation Department 200: Large-scale language models 300: Personal information database 310: Individual health check information 320: Receipt information 400: First database 410: Health Risk Data 420: Stress visualization data 430: Health risk fluctuation data 500: Second database 510: Medical Papers 520: Health and Productivity Management Information 530: Nutrition and diet related information 540: Information related to motor function 550: Doctor's findings 600: Information Processing Department 610: Health Risk Prediction Engine 620: Stress Visualization Engine 630: Health risk fluctuation prediction engine

Claims

1. A control unit having a search expansion generation unit that generates an inference request to be input to a large-scale language model that has been machine-learned based on input information including personal identification information that can identify an individual; a transceiver unit for transmitting the inference request to the large-scale language model and receiving an inference result from the large-scale language model; An information processing device comprising: The search extension generator, generating the inference request influenced by the individual identifying information based on a first search result obtained by searching a first database in which first information linked to the individual identified by the individual identifying information is stored, and a second search result obtained by searching a second database in which second information consisting of non-personal information is stored; performing a search of the first database before searching the second database, and setting search conditions for the second database based on the results of the first search; An information processing device comprising:

2. The information processing apparatus according to claim 1 , wherein the inference request generated by the search extension generating unit does not include the personally identifying information.

3. The information processing apparatus according to claim 1 , wherein the search conditions of the second database do not include the individual identifying information.

4. a personal-related information database in which personal-related information useful for searching the first information linked to the personal identification information is stored; The information processing apparatus according to claim 1 , wherein the search extension generating unit sets a first search condition for obtaining the first search result based on the personal-related information obtained by searching the personal-related information database.

5. The information processing device according to claim 4 , wherein the first search result does not include the individual identifying information and does not include the individual related information.

6. The information processing apparatus according to claim 4 , wherein the individual-related information includes at least one of medical checkup information and social insurance medical fee statement information.

7. The information processing apparatus according to claim 1 , wherein the first information includes at least one selected from the group consisting of health risk data, stress visualization data, and health risk fluctuation data of the individual.

8. The information processing device according to claim 1 , wherein the second information includes at least one selected from the group consisting of medical papers, health management related information, nutrition and diet related information, motor function related information, and doctor's findings related information.

9. The information processing device according to claim 1 , wherein the large-scale language model is managed in a closed environment with limited access.

10. The information processing apparatus according to claim 1 , wherein the first database is managed in a closed environment in which access is restricted.

11. The information processing apparatus according to claim 1 , further comprising: a display signal generating unit that receives the inference result from the large-scale language model and generates a display signal including at least a part of the inference result.

12. The information processing apparatus according to claim 1 , wherein the control unit further comprises a teacher data creation unit that creates teacher data based on the inference request, the inference result, and corrections to the inference result.

13. An information processing system comprising a first device having an input / output function, a second device having a data storage function, and a third device having a control unit, the first device being capable of communicating with the second device and the third device, The second device is A model storage device in which a large-scale language model that has undergone machine learning is stored; a first database in which first information associated with an individual identified by the individual identification information input to the first device is stored; a second database in which second information consisting of non-personal information is stored; having the control unit of the third device has a search expansion generating unit that generates an inference request to be input to the large-scale language model based on information input from the first device; the first device includes a transceiver unit that transmits the inference request input from the third device to the second device having the large-scale language model and receives an inference result from the second device; The search extension generator, generating the inference request influenced by the first information included in the first search result based on a first search result including the first information obtained by searching the first database and a second search result including the second information obtained by searching the second database; performing a search of the first database before searching the second database, and setting search conditions for the second database based on the results of the first search; An information processing system comprising:

14. The information processing system according to claim 13 , wherein at least two of the first device, the second device and the third device are integral.

15. An information processing system having a plurality of computer devices forming a peer-to-peer network, At least one of the plurality of computer devices has a search expansion generation unit that generates an inference request to be input to a machine-learned large-scale language model based on input information, At least one of the plurality of computer devices has a transceiver unit that transmits the inference request to the large-scale language model and receives an inference result from the large-scale language model; At least one of the plurality of computer devices has an input information generating unit that generates information to be input to the search extension generating unit; At least one of the plurality of computer devices has a display signal generating unit that generates a display signal including at least a part of the inference result; The search extension generator, generating the inference request influenced by the first information included in the first search result based on a first search result including the first information obtained by searching a first database in which first information linked to an individual identified by personal identification information is stored, and a second search result including the second information obtained by searching a second database in which second information consisting of non-personal information is stored; performing a search of the first database before searching the second database, and setting search conditions for the second database based on the results of the first search; An information processing system comprising:

Citation Information

Patent Citations

  • Diagnostic support device, diagnostic support system, diagnostic support method, and diagnostic support program

    JP2022190008A

  • Animal Health Management System

    JP7260943B1

  • Information processing device, information processing method, and program

    WO2023210217A1

  • Similar document search device, similar document search method, and similar document search program

    JP7300237B2

  • Medical Support Devices

    JP7454090B1