Business support system and business support method

The model evaluation system objectively assesses LLM expertise using multiple evaluation methods, improving user trust and system efficiency by providing clear expertise scores and retention processes.

JP2025110027APending Publication Date: 2025-07-28HITACHI LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024003712
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-15
Publication Date
2025-07-28

AI Technical Summary

Technical Problem

Existing business support systems using large language models (LLMs) lack an effective method to objectively evaluate the expertise of their outputs, placing a heavy burden on users to assess the appropriateness and accuracy of the LLM's specialized content, which can reduce system utilization efficiency.

Method used

A model evaluation system that includes a storage device for business-related data, a natural language model, and a control device to calculate a specialty metric, enabling objective evaluation of the LLM's expertise through multiple angles such as textbook evaluation, oral question assessment, reality evaluation, and self-evaluation, with integrated scoring to determine the LLM's expertise.

Benefits of technology

Enables accurate and efficient evaluation of the LLM's expertise, enhancing user trust and system effectiveness by providing a clear expertise score and retention processes to improve the LLM's performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025110027000001_ABST
    Figure 2025110027000001_ABST
Patent Text Reader

Abstract

To appropriately evaluate the expertise of a natural language model.SOLUTION: A business support system 1 is provided with a storage device for storing a database in which natural language data including descriptions about a field of business are registered, and a control device for executing expertise metric calculation processing in which the natural language data extracted from the database are input to a natural language model having the natural language data as input and data corresponding to the natural language data as output, output data corresponding to the input natural language data are acquired, values of parameters representing expertise in the field of business of the natural language model are calculated based on the acquired output data, and information on the calculated values of the parameters is output to an output device.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a business support system and a business support method.

Background Art

[0002] In specialized operations carried out in many fields, the performance of such activities strongly depends on the skills of experts such as behavioral patterns, experience, talent, and know-how.

[0003] For example, salespersons in the financial industry conduct various sales activities, and the performance of such activities strongly depends on characteristics such as the individual behavioral patterns, experience, talent, and know-how of the salespersons. Another example is that the parameter design when constructing a network for a specific industry depends on the extensive experience of skilled engineers. Another example is that in systems engineering for constructing an IT system (IT: Information Technology) for a specific industry, this activity not only depends on only the manualized text but also on the experience of the engineers. Another example is that the performance of the activities of engineers who perform maintenance on specialized measuring instruments also depends on the experience of the engineers.

[0004] Some organizations have tried to share the knowledge and know-how of successful experts throughout the organization, but it is generally difficult for others to appropriately utilize specialized knowledge and know-how, and if this is done without an appropriate advisor, business efficiency and results often decline.

[0005] In recent years, with the progress of artificial intelligence technology, it has been considered to introduce large language models (LLMs) as advisor software into business support systems. Here, an LLM is a natural language processing AI model (AI: Artificial Intelligence) trained using a large amount of text data, and it is a machine learning model that has the ability to receive information including text as input and output information including text.

[0006] FIG. 34 is a diagram showing an example of the configuration of the current business support system 3 using a large language model. This business support system 3 includes a preprocessing unit 121 and a postprocessing unit 122. The preprocessing unit 121 receives a business-related question from the user and passes it to the large language model system 130 equipped with an LLM. The postprocessing unit 122 receives the information of the answer (advice) to the question generated by the large language model system 130 and returns it to the user.

[0007] Here, since there is a possibility that the LLM of the large language model system 130 does not have appropriate expertise or generates uncertain answers, it is necessary to evaluate the appropriateness of the information provided by the LLM.

[0008] However, in order to evaluate the appropriateness of the information provided by the LLM, it is necessary for the user who evaluates the LLM to have advanced expertise and judgment ability. In particular, when the information provided by the LLM includes specialized content, in order to evaluate the accuracy of the output of the LLM, it is required that the user himself has a deep understanding of that specialized field. This places a heavy burden on the user and may consequently reduce the utilization efficiency and effectiveness of the system using the LLM.

[0009] Here, Patent Document 1 proposes a method of adding information to a question sentence to generate a prompt so that the LLM can generate a more appropriate answer. Specifically, it is disclosed that by adding information to the question sentence to generate a prompt, the LLM can generate a more appropriate answer.

[0010] In addition, Patent Document 2 discloses a technology related to a persona chatbot control method and system for maintaining a consistent dialogue experience and the flow of dialogue. Specifically, it includes steps of receiving a user utterance, adding it to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, inputting the encoded prompt into a language model, and generating a chatbot utterance as a response to the user utterance. Further, it is disclosed that a dialogue topic detector is used to determine a dialogue topic related to the user utterance, and a dialogue scene search model is used to obtain a dialogue scene related to the determined dialogue topic.

Prior Art Documents

Patent Documents

[0011]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0012] Since the technology of Patent Document 1 has a character limit for the input to the LLM, it improves the appropriateness of answers including expertise by giving predetermined additional information to the LLM. However, Patent Document 1 does not disclose a configuration for objectively showing whether the LLM has appropriate expertise (evaluating the accuracy of the LLM's output).

[0013] On the other hand, Patent Document 2 discloses a checking mechanism for a chatbot using an LLM, but the check is for the consistency of the chatbot's character and not for expertise.

[0014] Thus, the current situation is that an effective method for confirming the expertise of the LLM's answers has not yet been developed.

[0015] The present invention has been made in view of such circumstances, and an object thereof is to provide a model evaluation system and a model evaluation method capable of appropriately evaluating the expertise of a natural language model.

Means for Solving the Problems

[0016] One aspect of the present invention for solving the above problems includes a storage device that stores a database registering natural language data including descriptions related to business fields, and a natural language model that takes natural language data as input and outputs data corresponding to the natural language data. By inputting the natural language data extracted from the database, output data corresponding to the input natural language data is obtained, and based on the obtained output data, the value of a parameter representing the expertise of the natural language model in the business field is calculated. A business support system comprising a control device that executes a specialty metric calculation process for outputting information on the calculated parameter value to an output device.

Effects of the Invention

[0017] According to the present invention described above, the expertise of a natural language model can be appropriately evaluated. Configurations, effects, and the like other than those described above will be clarified by the description of the following embodiments.

Brief Description of the Drawings

[0018]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Mode for Carrying Out the Invention

[0019] Hereinafter, embodiments of the present invention will be described in detail based on the drawings.

[0020] FIG. 1 is a diagram showing an example of the configuration of the business support system 1 according to this embodiment. Business support System 1 is configured to include each information processing system (a system consisting of one or more information processing devices) of a large language model system 130, a model evaluation system 120, and a user interface system 100. The information processing systems are connected to each other by a wired or wireless communication network 110 such as, for example, the Internet, a LAN (Local Area Network), a WAN (Wide Area Network), or a dedicated line.

[0021] The large language model system 130 stores a plurality of large language models (LLMs) such as BERT or GPT-3 for each business field (for each specialized field of business). This large language model (hereinafter also referred to as an advisor LLM) is a trained model that uses data (such as text data) described in natural language as input data and outputs data (such as text data) described in natural language corresponding to the input data. For example, when a prompt of a question sentence described in natural language is input, the advisor LLM outputs the text of an answer sentence corresponding to the question sentence.

[0022] The type of the advisor LLM is not particularly limited. For example, it is a neural network having an input layer to which input data is input, one or more intermediate layers (hidden layers) that extract and output feature amounts of the input data, and an output layer that outputs output data from the feature amounts. Examples of the neural network include an RNN (Recurrent Neural Network) and a CNN (Convolution Neural Network). In addition to the neural network, it is not precluded from adopting a model to which an SVM (Support Vector Machine), a Bayesian network, a regression tree, or the like is applied.

[0023] The user interface system 100 is an information processing system used by a user who conducts business and the like using an advisor LLM. The user interface system 100 has a user interface that accepts input of question information (characters, voice, etc.) from the user, and converts the input into natural language data that can be input to the advisor LLM. The user interface system 100 is, for example, a voice dialogue system 101, a mobile terminal 103, or a personal computer 104 (PC: Personal Computer). Further, the user interface system 100 may be an XR goggle 102 that includes a display device for displaying the real space or virtual space and a microphone device for accepting voice input from the user, and realizes augmented reality (AR) or virtual reality (VR).

[0024] The model evaluation system 120 mediates the processing of the large language model system 130 and evaluates the expertise of the advisor LLM.

[0025] FIG. 2 is a diagram for explaining the functions of the model evaluation system 120 in more detail. The model evaluation system 120 has functional units such as a preprocessing unit 121, a postprocessing unit 122, an expertise metric calculator 123, and an expertise retention unit 124.

[0026] The preprocessing unit 121 performs predetermined preprocessing on the question data received from the user interface system 100, and transmits the preprocessed data to the large language model system 130.

[0027] The postprocessing unit 122 receives the answer data for the above question generated by the large language model system 130, performs predetermined postprocessing on the received answer data, and transmits the postprocessed data to the user interface system 100.

[0028] The above is the same as each function of the business support system shown in FIG. 34. The business support system 1 of the present embodiment further includes the following functional units.

[0029] That is, the expertise metric calculator 123 periodically calculates an expertise score, which is an index value representing the expertise of the advisor LLM of the large language model system 130, and presents the calculation result to the user of the user interface system 100.

[0030] When the expertise retention device 124 determines that the expertise score calculated by the expertise metric calculator 123 is low, it performs a process (retention process) to enhance the expertise of the advisor LLM of the large language model system 130.

[0031] (Expertise metric calculator) FIG. 3 is a diagram showing an example of the configuration of the expertise metric calculator 123. The expertise metric calculator 123 evaluates the expertise of the LLM from a plurality of different angles, and based on these evaluation results, outputs an expertise score S.

[0032] Specifically, the expertise metric calculator 123 includes a specialty field setter 400 that receives a specialty field setting from the user, and a hyperparameter setter 401 that receives a hyperparameter setting from the user.

[0033] In addition, the expertise metric calculator 123 includes a textbook evaluation module 402 that calculates a first score S1 based on the set specialty field and hyperparameters, an oral question evaluation module 403 that calculates a second score S2 based on the set specialty field and hyperparameters, a reality evaluation module 404 that calculates a third score S3 based on the set specialty field and hyperparameters, and an answer self-evaluation module 405 that calculates a fourth score S4 based on the set specialty field and hyperparameters. Details of the first score S1 to the fourth score S4 will be described later.

[0034] In addition, the expertise metric calculator 123 includes an LLM interface 407 that transmits and receives data of questions for calculating each score and data of answers corresponding to the questions between each module of the textbook evaluation module 402, the oral question evaluation module 403, the reality evaluation module 404, and the answer self-evaluation module 405, and the large language model system 130.

[0035] In addition, the expertise metric calculator 123 includes an integrated score calculator 406 that calculates an expertise score S by integrating the scores calculated by each module of the textbook evaluation module 402, the oral question evaluation module 403, the reality evaluation module 404, and the answer self-evaluation module 405.

[0036] (Textbook Evaluation Module) FIG. 4 is a block diagram showing an example of the configuration of the textbook evaluation module 402. The textbook evaluation module 402 calculates a first score S1 to evaluate whether the advisor LLM can understand and appropriately use specific expertise.

[0037] The textbook evaluation module 402 has a specialty setting device 600 that receives a specialty setting from the user, and from a textbook database 601 that stores texts (such as texts in textbooks) explaining each matter related to each specialty, obtains texts related to the set specialty, and generates data of a combination of a question q1 (a multiple-choice question) that allows the user to select a correct answer from a plurality of options with a certain word or text in the obtained text as the correct answer and any other word or text as the incorrect answer and its correct answer t1 (hereinafter referred to as multiple-choice data), and stores the data in a test database 603.

[0038] In addition, the textbook evaluation module 402 includes a multiple-choice test selector 604 that extracts a predetermined number N of multiple-choice tests from the test database 603 and sends the questions in the extracted multiple-choice tests to the large language model system 130, an answer correctness determiner 605 that determines the correctness of the answers to the questions received from the large language model system 130, and an individual score calculator 606 that calculates a first score S1 (correct answer rate) based on the determination result of the correctness.

[0039] (Oral question evaluation module 403) FIG. 5 is a block diagram showing an example of the configuration of the oral question evaluation module 403. The oral question evaluation module 403 presents questions in the field of expertise in the form of oral questions to the advisor LLM, and calculates a second score S2 by having another LLM (specialized LLM) in the field of expertise grade the answers.

[0040] First, the oral question evaluation module 403 includes a specialized LLM database 902 that stores each specialized LLM, and a fine-tuning device 907 that trains or fine-tunes each specialized LLM using a predetermined dataset and registers it in the specialized LLM database 902. Note that the specialized LLM may be an LLM smaller than the advisor LLM.

[0041] In addition, the oral examination evaluation module 403 includes a specialty setting device 900 that accepts a specialty setting from the user, a specialty LLM selector 903 that selects a specialty LLM for the set specialty from the specialty LLM database 902, and a free-response question database 901 that is a database storing data of a combination of a question (free-response question) answered in text about matters related to each specialty and the text of the correct answer to the question (free-response question data). A free-response question selector 906 that acquires the question q2 for the set specialty and its correct answer (model answer t2) from the free-response question database 901 and sends the question q2 to the large language model system 130, a specialty LLM answer grader 904 that grades by comparing the answer a2 to the question q2 received from the large language model system 130 with the model answer t2 output by the selected specialty LLM as the answer to the question q2, and an individual score calculator 905 that calculates a second score S2 based on the grading result.

[0042] (Reality evaluation module 404) FIG. 6 is a block diagram showing an example of the configuration of the reality evaluation module 404. The reality evaluation module 404 collates the answer to a question asking about a fact occurring or existing in the business of the specialty field output by the advisor LLM with the correct answer information stored in an external information processing system different from the advisor LLM, and calculates a third score S3 based on the degree of coincidence to evaluate whether the advisor LLM has a correct understanding of the reality in the specialty field.

[0043] The reality evaluation module 404 includes a specialty setting device 1200 that accepts specialty and hyperparameter settings from the user, an external system selector 1204 that selects an external system 1205 related to the set specialty from a plurality of external systems 1205 that store information on facts occurring in each specialty (for example, detailed information on companies in the field, personal information of customers, etc.), and an external grounding test generator 1202 that generates data (external grounding test data) in the form of a combination of a question q3 of a question (external grounding test) asking about facts in the set specialty and words or sentences (external grounding data t3) representing the facts that are the correct answers, and registers the data in an external grounding test database 1203.

[0044] Note that the facts targeted by the question q3 of the external grounding test are current or past facts. This fact may include, for example, personal information or financial information of a specific operator, etc., and may include facts that are generally not known to those other than a specific user.

[0045] Furthermore, the reality evaluation module 404 has an external grounding test selector 1201 that acquires external grounding test data of the question q3 and the external grounding data t3 from the external grounding test database 1203 and sends the question q3 to the large language model system 130, an answer correctness determination device 1206 that compares the answer a3 to the question q3 received from the large language model system 130 with the external grounding data t3 and grades it, and an individual score calculator 1207 that calculates a third score S3 based on the grading result.

[0046] (Answer self-evaluation module 405) FIG. 7 is a block diagram showing an example of the configuration of the answer self-evaluation module 405. The answer self-evaluation module 405 evaluates specialty and calculates a fourth score S4 by giving the advisor LLM a perspective to introspectively verify its own specialty.

[0047] The self-evaluation module 405 for answers includes a specialization setter 1500 that accepts the setting of the field of expertise and hyperparameters from the user, and a self-awareness prompt database 1502 in which data (self-awareness prompt data for expertise) of questions (expertise self-awareness tests) asking about the degree of proficiency in the field of expertise to the advisor LLM itself is registered. A self-awareness prompt selector 1501 that selects a question q4 (self-awareness prompt for expertise) related to the set field of expertise from the self-awareness prompt database 1502 and sends the selected self-awareness prompt for expertise to the large language model system 130, and an individual score calculator 1503 that calculates a fourth score S4 based on the answer to the self-awareness prompt for expertise received from the large language model system 130.

[0048] The self-awareness prompt for expertise takes, for example, the form of "Please answer the confidence level of your expertise in {field of expertise} on a metric from 0 to 100."

[0049] (Integrated score calculator 406) Next, FIGS. 8 and 9 are block diagrams showing an example of the configuration of the integrated score calculator 406. The integrated score calculator 406 integrates the first score S1, the second score S2, the third score S3, and the fourth score S4 to calculate an expertise score S. Here, the calculation formula for the expertise score S incorporates weighting parameters w (w1, w2, w3, w4) associated with the first score S1, the second score S2, the third score S3, and the fourth score S4. The parameters w (w1, w2, w3, w4) are parameters representing the reliability of the expertise evaluation by multiple-choice questions, free-response questions, external grounding tests, and self-awareness tests for expertise, and are set for each field of expertise of each test.

[0050] Specifically, first, as shown in FIG. 8, the integrated score calculator 406 is configured to determine (learn) the parameter w. That is, the integrated score calculator 406 includes a parameter updater 2202 that initializes the parameter w, and the first score S1, the second score S2, the third score S3, and the fourth score S4 calculated in the past, respectively obtained from each of the textbook evaluation module 402, the oral question evaluation module 403, the reality evaluation module 404, and the answer self-evaluation module 405, and a learning-based integrator 2200 that calculates a specialty score S (for learning) based on the current parameter w, a human-based specialty metric storage unit 2204 that stores a human evaluation score Sh, which is an estimated value of the specialty score S independently evaluated and determined by an expert or the like for the specialty of the advisor LLM, a comparator 2201 that compares the calculated specialty score S and the human evaluation score Sh and calculates the divergence (loss) between the two, and a parameter storage unit 2203 that stores the parameter w whose value has been adjusted (learned) by the parameter updater 2202 based on the loss in association with the corresponding specialty field.

[0051] After that, the learning-based integrator 2200 of the integrated score calculator 406 calculates the specialty score S using the learned parameter w. That is, as shown in FIG. 9, the learning-based integrator 2200 of the integrated score calculator 406 obtains the specialty field and hyperparameters respectively set by the specialty field setter 400 and the parameter setter 2300, and based on the first score S1, the second score S2, the third score S3, and the fourth score S4 respectively obtained from each of the textbook evaluation module 402, the oral question evaluation module 403, the reality evaluation module 404, and the answer self-evaluation module 405, and the learned parameter w, calculates the specialty score S.

[0052] (Specialty Retention Device 124) FIG. 10 is a block diagram showing an example of the configuration of the expertise retention device 124. The expertise retention device 124 checks the expertise of the advisor LLM using questions (retention examples) provided to the advisor LLM to maintain the expertise of the advisor LLM. Here, it is assumed that there is an explicit correct answer in the retention examples.

[0053] The expertise retention device 124 includes a specialty setting unit 1800 that receives the setting of the specialty field and hyperparameters from the user, and a retention example selector 1801 that acquires data (retention examples) of a combination of a question θ and a correct answer t related to the set specialty field from a retention example database 1802 in which a plurality of retention examples are registered, and transmits the question θ to the large language model system 130. A score calculator 1803 that receives an answer p corresponding to the question θ from the large language model system 130, compares the answer p with the correct answer t, and calculates a retention score Sa that is an index value representing the degree of expertise of the advisor LLM, and a retention information prompter 1804 that generates a retention information prompt that is a prompt for the retention example and transmits it to the large language model system 130.

[0054] Here, FIG. 11 is a diagram showing an example of the hardware included in the model evaluation system 120. The model evaluation system 120 includes a control device 91 such as a CPU (Central Processing Unit), a main memory device 92 such as a RAM (Random Access Memory) and a ROM (Read Only Memory), an auxiliary storage device 93 such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive), an input device 94 composed of a keyboard, a mouse, a touch panel, etc., an output device 95 for performing screen display composed of a monitor (display), etc., and a communication device 96 composed of a NIC (Network Interface Card), a wireless communication module, a USB (Universal Serial Interface) module, a serial communication module, etc. Note that other information processing devices in the business support system 1 also include similar hardware.

[0055] The functions of the model evaluation system 120 are realized by the hardware of the model evaluation system 120 or by the control device 91 of the model evaluation system 120 reading and executing each program stored in the main memory device 92 or the auxiliary storage device 93. Also, these programs are stored in a storage device such as a secondary storage device, a non-volatile semiconductor memory, a hard disk drive, an SSD, etc., or a recording medium readable by the model evaluation system 120 such as an IC card, an SD card, a DVD, etc. Note that the model evaluation system 120 may be realized using virtual information processing resources provided using virtualization technology, process space separation technology, etc., such as a virtual server provided by a cloud system, for example. Also, all or part of the functions provided by these devices may be realized by a service provided by a cloud system via an API (Application Programming Interface), etc. Next, the processing performed in the business support system 1 will be described.

[0056] FIG. 12 is a flowchart for explaining an example of business support processing. The business support processing is started when the model evaluation system 120 receives business question data for the advisor LLM from the user interface system 100.

[0057] First, the preprocessing unit 121 performs predetermined preprocessing on the question data received from the user interface system 100, and transmits the preprocessed question data to the large language model system 130 (S300).

[0058] Then, the advisor LLM of the large language model system 130 generates answer data for the question data (S301). The large language model system 130 transmits the generated answer data to the model evaluation system 120. The model evaluation system 120 transmits the received answer data (for example, business support advice) to the user interface system 100.

[0059] At this time, the model evaluation system 120 executes the following processes S302 to S306. That is, first, the expertise metric calculator 123 executes a score calculation process S302 for calculating the expertise score S of the advisor LLM based on the answer data received from the large language model system 130.

[0060] The expertise metric calculator 123 determines whether the number of times of the retention process performed so far is greater than a predetermined threshold number N (S303). If the number of times of the retention process is greater than the predetermined threshold number N (S303: Yes), the expertise metric calculator 123 executes the process of S306. If the number of times of the retention process is not greater than the predetermined threshold number N (S303: No), the expertise metric calculator 123 executes the process of S304.

[0061] In S304, the expertise metric calculator 123 determines whether the expertise score S calculated in the score calculation process S302 is greater than a predetermined threshold S'. If the expertise score S calculated in the score calculation process S302 is greater than the predetermined threshold S' (S304: Yes), the expertise metric calculator 123 executes the process of S306. If the expertise score S calculated in the score calculation process S302 is not greater than the predetermined threshold S' (S304: No), the expertise metric calculator 123 executes the process of S305.

[0062] In S305, the expertise retention device 124 executes a retention process S305 to enhance the expertise of the advisory LLM. After that, the process of S300 is performed.

[0063] In S306, the expertise metric calculator 123 causes the expertise score S calculated in S302 to be displayed on the screen of the user interface system 100 together with the business support advice received from the large language model system 130. Thus, the business support process ends.

[0064] FIG. 13 is a diagram showing an example of a screen displayed on the user interface system 100.

[0065] This screen 1300 displays the question text G02 input by the user and the answer text G03 (business support advice) from the large language model system 130 for this question in a dialogue format together with the graphic G01 representing the user and the graphic G04 representing the large language model system 130.

[0066] And on this screen 1300, corresponding to (in the vicinity of) the answer text G03, the specialized field G05 set by the user and the expertise score G06 calculated for that specialized field are displayed.

[0067] This enables the user to intuitively and visually understand the expertise of the advisor LLM, which helps in judging the reliability of its advice. And it can assist the user in making business decisions and contribute to improving business efficiency.

[0068] Note that a detailed button may be provided on this screen 1300, and when the user clicks it, the evaluation results of each expertise evaluation module can be individually confirmed.

[0069] Note that this screen 1300 may output a warning (message, voice, etc.) to the user when the expertise score S does not reach a certain standard.

[0070] Regarding the display of the expertise score on this screen 1300, the expertise scores of multiple specialized fields may be displayed in a layout that enables visual comparison of the expertise scores using a figure in a predetermined form (for example, a graph such as a radar chart).

[0071] Also, when the user interface system 100 is the XR goggles 102, this screen 1300 may display the content of the expertise score or output voice to the goggles. When the user interface system 100 is the mobile terminal 103, the content of this screen 1300 may be output as voice. This allows the user to confirm the expertise of the advisor LLM even while performing other business operations.

[0072] In addition, this screen 1300 may display the expertise score in various display modes according to its value. For example, a sentence corresponding to the value of the expertise score may be added to the response sentence G03 and displayed on the screen 1300, or the response sentence G03 may be changed according to the value of the expertise score and displayed on the screen 1300. For example, when the expertise score is within a predetermined normal value (range), it may be displayed as "This is ○○", and when the expertise score is within a predetermined low value (range), it may be displayed as "Probably, this may be ○○". Also, a graphic such as an emoji corresponding to the value of the expertise score may be displayed. Further, audio corresponding to the value of the expertise score may be output. For example, audio with a tone or intonation corresponding to the value of the expertise score may be output. Hereinafter, the details of each process in the business support process will be described.

[0073] <Expertise Score Calculation Process> FIG. 14 is a flowchart for explaining the details of the score calculation process S302.

[0074] The expertise field setter 400 displays a predetermined input screen and accepts the setting of the test expertise field from the user (S501). Also, the hyperparameter setter 401 displays a predetermined input screen and accepts the setting of the hyperparameters in the test from the user (S502). Hyperparameters are, for example, the number of questions for each test, the database or system to be used in the test, which will be described later. The hyperparameters and the expertise field may be set in the first score calculation process S503, the second score calculation process S504, the third score calculation process S505, and the fourth score calculation process S506 as described later.

[0075] Subsequently, the textbook evaluation module 402 executes a first score calculation process S503 for calculating a first score S1 based on the expertise field set in S501 and the hyperparameters set in S502.

[0076] In addition, the oral examination evaluation module 403 executes a second score calculation process S504 for calculating a second score S2 based on the specialized field set in S501 and the hyperparameters set in S502.

[0077] In addition, the reality evaluation module 404 executes a third score calculation process S505 for calculating a third score S3 based on the specialized field set in S501 and the hyperparameters set in S502.

[0078] In addition, the answer self-evaluation module 405 executes a fourth score calculation process S506 for calculating a fourth score S4 based on the specialized field set in S501 and the hyperparameters set in S502.

[0079] The integrated score calculator 406 executes an integrated score calculation process S507 for calculating a specialty score S based on the scores calculated in the first score calculation process S503, the second score calculation process S504, the third score calculation process S505, and the fourth score calculation process S506. Hereinafter, the details of the first score calculation process S503, the second score calculation process S504, the third score calculation process S505, the fourth score calculation process S506, and the integrated score calculation process S507 will be described.

[0080] <First Score Calculation Process> FIG. 15 is a flowchart for explaining the details of the first score calculation process S503.

[0081] First, the textbook evaluation module 402 displays a predetermined input screen and accepts from the user the setting of the specialized field and hyperparameters related to the multiple-choice questions (S800).

[0082] The textbook evaluation module 402 reads the textbook database 601 related to the specialized field set in S800 (S801).

[0083] The textbook evaluation module 402 generates multiple-choice questions based on the textbook database 601 read in S801 (S802).

[0084] For example, the textbook evaluation module 402 generates a plurality of data of questions q1 having a plurality of answer options and correct answers t1 thereto (multiple-choice data) based on all the data registered in the textbook database 601, and stores them in the test database 603.

[0085] The textbook evaluation module 402 extracts multiple-choice questions to be presented to the user from the test database 603 under the conditions indicated by the hyperparameters set in S800 (S803).

[0086] For example, the textbook evaluation module 402 extracts N pieces of multiple-choice data indicated by the hyperparameters set in S800 from the test database 603.

[0087] The textbook evaluation module 402 transmits one of the data of the multiple-choice question q1 extracted in S803 to the large language model system 130 (S804). Then, the advisory LLM of the large language model system 130 outputs an answer a1 which is the output for the input value question q1. The large language model system 130 returns the output answer a1 to the textbook evaluation module 402.

[0088] The textbook evaluation module 402 checks whether the conditions indicated by the hyperparameters set in S800 are met (for example, checks whether N multiple-choice questions q1 indicated by the hyperparameters have been sent) (S805). If the conditions indicated by the hyperparameters set in S800 are met (S805: Yes), the textbook evaluation module 402 executes the process of S806. If the conditions indicated by the hyperparameters set in S800 are not met (S805: No), the textbook evaluation module 402 repeats the process of S804 to send the next multiple-choice question q1.

[0089] The textbook evaluation module 402 calculates the first score S1 by comparing the answers a1 and the correct answers t1 for N multiple-choice questions q1. For example, for each question q1, the textbook evaluation module 402 gives +1 point if the answer a1 and the correct answer t1 match, and +0 point if they do not match. Also, the textbook evaluation module 402 may use the correct answer rate as the first score S1.

[0090] Here, FIG. 16 is a diagram showing an example of a multiple-choice question q1. This multiple-choice question q1 has a question text 1601 and a plurality of options 1602 (words or texts) related to the question.

[0091] Also, FIG. 17 is a diagram showing an example of the correct answer t1 or the answer a1 for the multiple-choice question q1. This correct answer t1 or answer a1 is the option (word or text) that is the correct answer or the answer among the options in the multiple-choice question q1.

[0092] <Second score calculation process> Next, FIG. 18 is a flowchart for explaining the details of the second score calculation process S504.

[0093] First, the oral question evaluation module 403 displays a predetermined input screen and accepts from the user the setting of the specialized field and hyperparameters related to the free-response question (S1100).

[0094] The oral interview evaluation module 403 selects a specialized LLM related to the specialized field set in S1100 from the specialized LLM database 902 and loads the specialized LLM to calculate the second score S2 (S1101).

[0095] In addition, the oral interview evaluation module 403 obtains one piece of free-response question data (question q2 and model answer t2) related to the free-response question related to the specialized field set in S1100 from the free-response question database 901 by specialized field (S1102).

[0096] The oral interview evaluation module 403 sends the question q2 of the free-response question data obtained in S1102 to the large language model system 130 (S1103). Then, the advisory LLM of the large language model system 130 outputs an answer a2, which is the output for the input value of the question q2. The large language model system 130 returns the output answer a2 to the oral interview evaluation module 403.

[0097] The oral interview evaluation module 403 calculates the score S2_sub related to the question q2 of the free-response question data by using the specialized LLM loaded in S1101 (S1104).

[0098] For example, the oral interview evaluation module 403 obtains the score S2_sub as an output value by inputting question data asking to what extent the answer a2 to the question q2 of the free-response question data approximates the model answer t2 to the specialized LLM.

[0099] The oral examination evaluation module 403 checks whether the conditions indicated by the hyperparameters set in S1100 are met (for example, checks whether the questions q2 of N2 free-response questions indicated by the hyperparameters have been sent) (S1105). If the conditions indicated by the hyperparameters set in S1100 are met (S1105: Yes), the oral examination evaluation module 403 executes the process of S1106. If the conditions indicated by the hyperparameters set in S1100 are not met (S1105: No), the oral examination evaluation module 403 repeats the process of S1102 to send the next free-response question q2.

[0100] In S1106, the oral examination evaluation module 403 calculates the second score S2 by summing the scores S2_sub calculated in S1103. Note that the oral examination evaluation module 403 may use the average value of the scores S2_sub as the second score S2.

[0101] Here, FIG. 19 is a diagram showing an example of the question q2 of the free-response question. This question q2 of the free-response question has a text of a question that allows a free answer.

[0102] Also, FIG. 20 is a diagram showing an example of the model answer t2 or the answer a2 for the question q2 of the free-response question. This model answer t2 or answer a2 is a text or word that is the correct answer or an answer for the question q2 of the free-response question.

[0103] <Third score calculation process> FIG. 21 is a flowchart for explaining the details of the third score calculation process S505.

[0104] First, the reality evaluation module 404 displays a predetermined input screen and accepts from the user the setting of the specialized field and hyperparameters related to the external ground test (S1400).

[0105] The reality evaluation module 404 generates an external grounding test by generating external grounding test data composed of the question q3 and the external grounding data t3 (correct answer) from the external system 1205 related to the specialized field set in S1400 (S1401).

[0106] The reality evaluation module 404 sends the question q3 in the external grounding test data generated in S1401 to the large language model system 130 (S1402). Then, the advisory LLM of the large language model system 130 outputs an answer a3 which is the output for the input value of the question q3. The large language model system 130 returns the output answer a3 to the reality evaluation module 404.

[0107] The reality evaluation module 404 determines the correctness of the external grounding test by comparing the external grounding data t3 in the external grounding test data with the answer a3 (S1403). For example, the reality evaluation module 404 stores it as a correct answer when the external grounding data t3 and the answer a3 match, and stores it as a wrong answer when the external grounding data t3 and the answer a3 do not match. Also, for example, the reality evaluation module 404 may calculate a score by calculating the degree of match between the external grounding data t3 and the answer a3.

[0108] The reality evaluation module 404 checks whether the conditions indicated by the hyperparameters set in S1400 are satisfied (for example, checks whether the question q3 of N3 external grounding tests indicated by the hyperparameters has been sent) (S1404). If the conditions indicated by the hyperparameters set in S1400 are satisfied (S1404: Yes), the oral question evaluation module 403 executes the process of S1405. If the conditions indicated by the hyperparameters set in S1400 are not satisfied (S1404: No), the oral question evaluation module 403 repeats the process of S1401 to send the question q3 of the next external grounding test.

[0109] In S1405, the reality evaluation module 404 calculates the third score S3 by aggregating the correct / incorrect answers or scores identified in S1103. For example, the reality evaluation module 404 sets the number of correct answers or the correct answer rate as the third score S3. Also, for example, the reality evaluation module 404 sets the total value of the scores calculated in S1403 as the third score S3.

[0110] Here, FIG. 22 is a diagram showing an example of the question q3 of the external grounding test. This question q3 of the external grounding test has a question sentence that asks about the content of the fact. For example, if the target external system 1205 is a business support system, the question q3 is a question that asks about specific data that the current customer has, such as age or gender. For example, if the target external system 1205 is a system that handles measuring instruments or the like, the question is about the settings of the parameters, model numbers, current, voltage, or ambient conditions such as temperature that the measuring instrument has.

[0111] Also, FIG. 23 is a diagram showing an example of the correct answer t3 or the answer a3 to the question q3 of the external grounding test. This correct answer t3 or answer a3 is a word or sentence that is the correct answer or answer to the question q3 of the external grounding test.

[0112] <Fourth Score Calculation Process> FIG. 24 is a flowchart for explaining the details of the fourth score calculation process S506.

[0113] First, the answer self-evaluation module 405 displays a predetermined input screen and accepts from the user the setting of the specialized field and hyperparameters related to the self-awareness test of expertise (S1700).

[0114] The answer self-evaluation module 405 obtains the self-awareness of expertise prompts (question q4) related to the above-set field of expertise from the expertise self-awareness prompt database 1502. Then, the answer self-evaluation module 405 sends the self-awareness of expertise prompt (question q4) to the large language model system 130 (S1701). After that, the advisory LLM of the large language model system 130 outputs an answer a4 that is the output for the input value of the self-awareness of expertise prompt (question q4). The large language model system 130 returns the output answer a4 to the answer self-evaluation module 405.

[0115] The answer self-evaluation module 405 calculates the fourth score S4 based on the answer a4 (S1703). For example, when the answer a4 is a numerical value, the answer self-evaluation module 405 may use that numerical value as the fourth score S4, or when the answer a4 is a text, it may perform semantic analysis of the content and use it as the fourth score S4.

[0116] Here, FIG. 25 is a diagram showing an example of the self-awareness of expertise prompt (question q4) of the self-awareness of expertise test. This question q4 has a text that asks about the expertise of the advisory LLM with a numerical value between 0 and 100.

[0117] Also, FIG. 26 is a diagram showing an example of the answer a4 for the self-awareness of expertise prompt (question q4) of the self-awareness of expertise test. This answer a4 is a numerical value that is the answer to the question q4.

[0118] <Integrated score calculation process> FIG. 27 is a flowchart for explaining the details of the integrated score calculation process S507.

[0119] The integrated score calculator 406 executes a parameter learning process S5071 for learning the value of the parameter w for calculating the expertise score S. Also, the integrated score calculator 406 executes a score calculation execution process S5702 for calculating the expertise score S based on the parameter w. Note that whether to execute the parameter learning process S5071 for the second and subsequent times may be optional.

[0120] Hereinafter, the details of the parameter learning process S5071 and the score calculation execution process S5702 will be described.

[0121] <Parameter Learning Process> FIG. 28 is a flowchart for explaining the details of the parameter learning process S5071.

[0122] The integrated score calculator 406 receives from the user the setting of the specialized field and parameters related to the score for learning the value of the parameter w (for example, the initial value of the parameter w, the value of the human evaluation score Sh) (S2400). The human evaluation score Sh is, for example, a score separately set by the user or the like for learning the parameter w.

[0123] The integrated score calculator 406 initializes the value of the parameter w and generates, for example, the following formula (1) for calculating the expertise score S using the parameter w (S2401).

[0124] Expertise score S = First score S1 × w1 + Second score S2 × w2 + Third score S3 × w3 + Fourth score S4 × w4 ··· Formula (1)

[0125] Here, w1 is the parameter w (weight value) related to the first score S1, w2 is the parameter w (weight value) related to the second score S2, w3 is the parameter w (weight value) related to the third score S3, and w4 is the parameter w (weight value) related to the fourth score S4. Note that the formula for calculating the expertise score S using the parameter w shown here is an example, and other formulas may be used as long as they reflect the weights of each score in the expertise score S. As the formulation method, weighted addition, weighted average, or integration by a neural network can be considered.

[0126] The integrated score calculator 406 calculates the expertise score S by substituting the score values (first score S1, second score S2, third score S3, fourth score S4) for learning the value of the parameter w into the formula for calculating the expertise score S (S2402).

[0127] The integrated score calculator 406 calculates the degree of deviation between the two by comparing the calculated expertise score S with the corresponding human evaluation score Sh and calculating the loss Loss = L(S, Sh) (S2403).

[0128] The integrated score calculator 406 changes the parameter w (w1, w2, w3, w4) so that the loss Loss becomes smaller (S2404).

[0129] When the loss Loss is smaller than a predetermined threshold (S2405: Yes), the integrated score calculator 406 executes the process of S2406. When the loss Loss is larger than the predetermined threshold (S2405: Yes), the integrated score calculator 406 repeats the processes after S2402 using the changed parameter w.

[0130] In S2406, the integrated score calculator 406 stores the current parameter w in association with the field of expertise related to the expertise score S.

[0131] <Score calculation execution process S5702> FIG. 29 is a flowchart for explaining the details of the score calculation execution process S5702.

[0132] The integrated score calculator 406 receives from the user the setting of the field of expertise and parameters related to the expertise score S to be calculated (S2500).

[0133] The integrated score calculator 406 reads the parameter w determined in the parameter learning process S5071 and the scores (first score S1, second score S2, third score S3, fourth score S4) necessary for calculating the expertise score S (S2501).

[0134] The integrated score calculator 406 calculates the expertise score S using the read parameter w (for example, calculated by Equation (1)) (S2502).

[0135] <Expertise retention process> FIG. 30 is a flowchart for explaining the details of the specialty retention process S305.

[0136] The specialty retention device 124 receives from the user the setting of the specialty field and parameters to be subject to the retention process (S1900). Here, it is assumed that the specialty retention device 124 acquires from the retention example database 1802 the retention examples (data of a combination of questions θ = (θ1, θ2,... θn) and correct answers t = (t1, t2,... tn)) related to the set specialty field.

[0137] The specialty retention device 124 transmits the question θ in the retention example to the large language model system 130 (S1901). Then, the advisory LLM of the large language model system 130 outputs an answer p (p1, p2,... pn) which is the output for the input value question θ. The large language model system 130 returns the output answer p to the specialty retention device 124.

[0138] The specialty retention device 124 calculates a retention score Sa by comparing the correct answer t and the answer p in the retention example (S1902). For example, the specialty retention device 124 sets the number of cases where the correct answer t and the answer p match as the retention score Sa.

[0139] The specialty retention device 124 determines whether the retention score Sa is higher than a predetermined threshold (S1903). If the retention score Sa is higher than the predetermined threshold (S1903: Yes), the specialty retention process S302 ends. If the retention score Sa is not higher than the predetermined threshold (S1903: No), the specialty retention device 124 executes the process of S1904.

[0140] In S1904, the expertise retention device 124 transmits the retention example problems (question θ and correct answer t) obtained in S1900 to the large language model system 130. Then, the large language model system 130 performs machine learning of the advisor LLM using the retention example problems as learning data. Thereby, the expertise of the advisor LLM can be maintained.

[0141] Note that the expertise retention device 124 may transmit the correctly answered retention example problems to the large language model system 130 as learning data, or may transmit all the retention example problems to the large language model system 130 as learning data in order to make the large language model system 130 learn that the incorrect answers in the incorrectly answered retention example problems are incorrect answers.

[0142] <Modification Example of Expertise Retention Device> In the above, the configuration and processing of the expertise retention device 124 when there is explicit correct answer data in the retention example problems have been described. Here, the expertise retention device 124 when there is no explicit correct answer data in the retention example problems will be described.

[0143] The retention example problems (questions) here are, for example, about the optimal value (parameter θtmp) of the parameter θ related to the operation of the systems of each facility or equipment such as factory equipment, plants, or communication networks. The parameter θ is, for example, a parameter set for factory equipment, an operation condition (temperature, pressure, etc.) set for a plant, or a device setting parameter (communication speed, communication route, buffer size, etc.) of a communication network.

[0144] Also, the advisor LLM here can take as input the expertise score and the parameters for calculating the expertise score in addition to the retention example problems, and outputs parameters that can calculate an expertise score higher than the input expertise score.

[0145] FIG. 31 is a diagram showing an example of the configuration of the expertise retention device 124 when there is no explicit correct answer data in the retention example problem. This expertise retention device 124 includes a specialty setting device 2000 that receives settings of a specialty field and hyperparameters from a user, and a parameter prediction prompt database 2001 that is a database of prompts for questions (retention example problems) in each specialty field where there is no explicit answer. A parameter prediction prompt selector 2002 that acquires a prompt for a retention example problem related to the set specialty field, a parameter prediction prompt presenter 2003 that transmits the acquired prompt to the large language model system 130, a score calculator 2004 that acquires a parameter θtmp corresponding to the prompt from the large language model system 130 and requests an external system 2009 to calculate the expertise score of the acquired parameter θtmp and acquires the expertise score Sa, a score level determiner 2005 that determines the expertise score Sa, a top score selector 2006 that repeats the calculation of the parameter θtmp until the value of the expertise score Sa becomes sufficiently high, and a retention information presenter 2007 that transmits the calculated parameter θtmp to the large language model system 130.

[0146] The external system 2009 is an information processing system that calculates a value of a parameter indicating how appropriate the parameter θ output from the advisor LLM is. In this embodiment, the parameter is assumed to be an expertise score. However, the score calculated by the external system 2009 may be a different type of score other than the expertise score. The external system 2009 is composed of, for example, a simulator that reproduces the behavior of the system or a predetermined pre-trained model.

[0147] FIG. 32 is a flowchart for explaining the details of the expertise retention process S305A when there is no explicit correct answer data in the retention example problem.

[0148] The expertise retention device 124 receives from the user the setting of the field of expertise and parameters to be subject to the retention process (S2100). For example, the expertise retention device 124 acquires information on retention examples related to the set field of expertise from the parameter prediction prompt database 2001.

[0149] The expertise retention device 124 searches for an external system 2009 among the plurality of external systems 2009 that can calculate the score of the facility or installation set in S2100 (S2101). For example, the expertise retention device 124 transmits inquiry information including hyperparameters to each external system 2009, and receives information indicating whether the external system 2009 can calculate the expertise score S.

[0150] If there is no external system 2009 that can calculate the expertise score S (S2101: No), the expertise retention process S305A ends. If there is an external system 2009 that can calculate the expertise score S (S2101: Yes), the expertise retention device 124 executes the process of S2102.

[0151] In S2102, the expertise retention device 124 generates information for requesting the calculation of the expertise score S. For example, the expertise retention device 124 generates API information API(θ) for calling an external system 2009 that calculates the expertise score S from the parameters θ = (θ1, θ2, … θn).

[0152] The expertise retention device 124 transmits a prompt of a retention example including the parameter θ to the large language model system 130 to predict a parameter θtmp (such that the expertise score S is optimized, i.e., made higher) (S2103). The advisor LLM of the large language model system 130 outputs the parameter θtmp and transmits it to the expertise retention device 124.

[0153] The expertise retention device 124 sets the parameter θtmp in the API(θ) and sends it to the external system 2009 searched in S2101, thereby obtaining the expertise score Sa from the external system 2009 (S2104).

[0154] The expertise retention device 124 determines whether the expertise score S is greater than a predetermined threshold (S2105). If the expertise score S is greater than the predetermined threshold (S2105: Yes), the expertise retention device 124 executes S2106. If the expertise score S is not greater than the predetermined threshold (S2105: No), the expertise retention device 124 executes the process of S2107.

[0155] In S2106, the expertise retention device 124 sends the parameter θtmp to the large language model system 130. Thereafter, the large language model system 130 performs machine learning of the advisor LLM using the parameter θtmp and its expertise score Sa as learning data for optimal parameters.

[0156] In S2107, the expertise retention device 124 repeats the process of S2103 to input the expertise score Sa as a past performance value, the corresponding parameter θtmp, and the current parameter θ to the advisor LLM. As a result, a new parameter θtmp can be obtained, which is expected to increase the expertise score S.

[0157] In this way, when there is no explicit correct answer data for the prompt, the expertise retention device 124 estimates whether the advisor LLM correctly understands the system based on the expertise score provided by the external system 2009. By strengthening the ability to predict the optimal parameter set, expertise retention is performed. That is, the large language model system 130 utilizes the feedback from the external system 2009 to understand the operation of the system and hone the ability to predict the optimal parameter set, thereby maintaining and enhancing its expertise.

[0158] <Modification Example of the Configuration of Business Support System 1> In the third score calculation process S505, the model evaluation system 120 performed an external grounding test on the advisor LLM.

[0159] FIG. 33 is a diagram showing an example of another configuration of the business support system 2 when performing an external grounding test. In this business support system 2, the large language model system 130 is communicably connected to a predetermined external database 140.

[0160] The external database 140 manages information (for example, past business daily reports, sales manuals, customer relationship management systems (CRM)) that stores facts related to the external grounding test.

[0161] When the advisor LLM of the large language model system 130 determines that it is necessary to obtain information on the fact that the prompt for the external grounding test question has been input, it accesses the external database 140 to obtain the fact and generate response data.

[0162] In this way, by providing a configuration in which the external database 140 is provided separately from the large language model system 130, the advisor LLM can be constructed without incorporating and learning sensitive facts such as personal information, and unnecessary diffusion of information can be prevented in the external grounding test.

[0163] As described above, the business support system 1 of the present embodiment inputs questions (natural language data) in a specific specialized field extracted from various databases or evaluation modules into a natural language model (advisor LLM that takes a natural language question prompt as input and outputs a corresponding answer) related to a predetermined specialized field, obtains response data corresponding to the input questions, calculates a specialization score S representing the specialization of the advisor LLM in the specific specialized field based on the obtained response data, and outputs information on the calculated specialization score S to an output device.

[0164] In this way, by calculating and outputting the score of the answer obtained for a question in a specific field input to the advisor LLM in that field, the user can know to what extent the advisor LLM has expertise in that field. As a result, the user can directly and quantitatively evaluate the reliability of the answer provided by the advisor LLM. And thereby, the user can confirm the reliability of the advice from the advisor LLM and reduce the uncertainty when taking actions based on that advice.

[0165] In this way, according to the business support system 1 of the present embodiment, the expertise of the natural language model can be appropriately evaluated.

[0166] Further, the business support system 1 of the present embodiment uses multiple-choice data (textbook database 601), which is data of a combination of a question including multiple answer options regarding matters in a specialized field and the correct answer option for that question, and determines whether the answer obtained by inputting the multiple-choice data question to the advisor LLM matches the correct answer, and calculates the expertise score S.

[0167] Thereby, it is possible to appropriately evaluate whether the advisor LLM grasps the basic matters (knowledge, etc.) of the specialized field.

[0168] Furthermore, the business support system 1 of the present embodiment uses free-response question data (free-response question database 901 by specialized field), which is data of a combination of a question that answers in text regarding matters related to a specialized field and the text of the correct answer for that question, and determines the degree to which the answer obtained by inputting the free-response question data question to the advisor LLM matches the correct answer, and calculates the expertise score S.

[0169] Thereby, it is possible to evaluate in detail whether the advisor LLM organically grasps the content (abstract knowledge, etc.) of the specialized field.

[0170] In addition, the business support system 1 of this embodiment uses external grounding test data (external grounding test database 1203), which is data of a combination of questions asking about facts in a specialized field and correct answers to those facts, and calculates a specialization score S by determining whether the words or sentences of the answer obtained by inputting a question of the external grounding test data to the advisor LLM match the correct answer.

[0171] Thereby, it is possible to evaluate whether the advisor LLM grasps the current situation that should be known in the specialized field as facts.

[0172] In addition, the business support system 1 of this embodiment uses specialization self - recognition prompt data (specialization self - recognition prompt database 1502), which is data of questions asking the advisor LLM about the degree to which the advisor LLM is proficient in a specialized field, and calculates a specialization score S based on the words or sentences of the answer obtained by inputting the specialization self - recognition prompt data to the advisor LLM.

[0173] Thereby, the specialization of the advisor LLM can be evaluated from a subjective aspect.

[0174] In addition, when there is a correct answer to a question, if the specialization score S is equal to or lower than a predetermined threshold, the business support system 1 of this embodiment uses the data of questions and answers in the specialized field as learning data for the advisor LLM and makes the advisor LLM learn.

[0175] In this way, when the specialization of the advisor LLM is low, it is possible to use the data of questions and answers to strengthen and update the specialization of the advisor LLM while adapting to the needs of the user.

[0176] Further, when there is no correct answer to a question and the expertise score S is below a predetermined threshold, the business support system 1 of this embodiment inputs a question in a specialized field to the advisor LLM and transmits the answer data (parameter θtmp) for optimizing the expertise score S, which is obtained thereby, to the external system 209. Then, when the received expertise score S is not greater than the predetermined threshold, a retention process is executed to obtain new answer data by inputting the answer data (parameter θtmp) and the received score Sa into the advisor LLM as input data.

[0177] In this way, when the expertise of the advisor LLM is low, by inputting the answer (parameter θtmp) with a low expertise score S obtained from the advisor LLM and its score Sa into the advisor LLM, the advisor LLM can output answer data (parameter θ) with a more suitable score in the future.

[0178] Also, when it is determined that the expertise score S is below a predetermined threshold, the business support system 1 of this embodiment displays a warning on the screen.

[0179] Thereby, the user can know that the expertise of the advisor LLM is low.

[0180] Further, for each of a plurality of specialized fields, the business support system 1 of this embodiment calculates the expertise score S based on the corresponding advisor LLM and displays a graphic (such as a radar chart) representing the comparison of the calculated expertise scores S on the screen.

[0181] Thereby, the user can compare the expertise of each advisor LLM and determine which advisor LLM in which specialized field is suitable or which advisor LLM in which specialized field should be improved.

[0182] In addition, the business support system 1 of the present embodiment generates question data and the like to be input to the advisor LLM based on the voice acquired from the XR goggles 102.

[0183] As a result, the user can ask questions to the advisor LLM while performing other tasks.

[0184] In addition, the business support system 1 of the present embodiment displays a message or a graphic corresponding to the expertise score S on the screen.

[0185] As a result, the user can quickly grasp the degree of expertise of the advisor LLM.

[0186] Although the embodiments of the present invention have been described above, the present invention is not limited to the above embodiments, and can be implemented using any components without departing from the gist thereof. The embodiments and modifications described above are merely examples, and the present invention is not limited to these contents as long as the features of the invention are not impaired. Also, although various embodiments and modifications have been described above, the present invention is not limited to these contents. Other aspects conceivable within the scope of the technical idea of the present invention are also included in the scope of the present invention.

[0187] For example, a part of the hardware provided in each device of the present embodiment may be provided in other devices.

[0188] Also, each program of each device may be provided in other devices, a certain program may be composed of a plurality of programs, or a plurality of programs may be integrated into one program.

[0189] Also, the storage format of the information described in each example is an example, and may be appropriately changed to other formats or the like.

[0190] In addition, the business support system 1 of this embodiment can be used for various operations. For example, a salesman in the financial industry can receive advice on formulating business strategies and customer interaction strategies from the advisor LLM. Also, an engineer engaged in network design can receive specialized advice on design guidelines or methods or predicted performance from the advisor LLM. Further, a system engineer building an IT system can receive specialized advice on its design guidelines or setting methods from the advisor LLM. Moreover, an engineer performing maintenance on specialized equipment such as semiconductor manufacturing equipment can receive specialized advice on maintenance guidelines or work procedures from the advisor LLM. Also, in the medical field, inexperienced doctors and nurses can receive specialized advice on diagnosis or medical treatment from the advisor LLM. Additionally, an engineer optimizing a production line in a manufacturing factory can receive specialized advice on equipment placement or scheduling from the advisor LLM.

[0191] Also, in this embodiment, it is assumed that the business support process starts when the question data is received from the user interface system 100, but it may be executed at other timings. For example, the business support process may perform only the processes after S302 when a predetermined timing (predetermined time interval, predetermined time) or a designated input from the user is received, regardless of whether the question data has been received from the user interface system 100.

Explanation of Signs

[0192] 1 Business support system, 130 Large language model system, 123 Specialization metric calculator, 124 Specialization retention device

Claims

1. A storage device that stores a database registering natural language data including a description related to a business field, and a control device that executes a specialty metric calculation process of inputting natural language data extracted from the database into a natural language model that takes natural language data as input and outputs data corresponding to the natural language data, obtaining output data corresponding to the input natural language data, calculating a value of a parameter representing the specialty in the business field of the natural language model based on the obtained output data, and outputting information on the calculated value of the parameter to an output device A business support system comprising the above.

2. The storage device stores multiple-choice data in the database, which is data of a combination of a question including multiple answer options related to matters in the business field and an answer option for the correct answer to the question, In the specialty metric calculation process, the control device inputs the question data in the multiple-choice data into the natural language model to obtain output data of an answer corresponding to the question, and executes a textbook evaluation process of calculating the value of the parameter by determining whether the obtained output data matches the correct answer in the multiple-choice data. The business support system according to Claim 1.

3. The storage device stores free-response question data in the database, which is data of a combination of a question that answers a matter related to the business field in a text and a text of the correct answer to the question, In the specialty metric calculation process, the control device inputs the question data in the free-response question data into the natural language model to obtain output data of a text of an answer corresponding to the question, and executes an oral question evaluation process of calculating a value of a parameter representing the degree to which the text indicated by the obtained output data matches the text of the correct answer in the free-response question data. The business support system according to Claim 1.

4. The storage device stores external grounding test data in the database, which is data of a combination of a question asking about facts in the business field and the correct answers to the facts. In the expertise metric calculation process, the control device inputs the question data in the external grounding test data into the natural language model to obtain output data of an answer corresponding to the question, and calculates the value of the parameter by determining whether the word or sentence indicated by the obtained output data matches the word or sentence representing the correct answer in the external grounding test data, and executes a reality evaluation process. The business support system according to claim 1.

5. The storage device stores, in the database, expertise self - recognition prompt data which is data of a question asking about the degree to which the natural language model is proficient in the field of the business. In the expertise metric calculation process, the control device inputs the expertise self - recognition prompt data into the natural language model to obtain output data of a word or sentence of an answer corresponding to the expertise self - recognition prompt data, and executes an answer self - evaluation process of calculating the value of the parameter based on the word or sentence indicated by the obtained output data. The business support system according to claim 1.

6. The control device determines whether the calculated value of the parameter is less than or equal to a predetermined threshold value. When the calculated value of the parameter is less than or equal to the predetermined threshold value, the control device executes a retention process of causing the natural language model to learn by using the question and answer data in the field of the business as learning data for the natural language model. The business support system according to claim 1.

7. The control device determines whether the calculated value of the parameter is less than or equal to a predetermined threshold value. When the calculated value of the parameter is less than or equal to the predetermined threshold value, the control device inputs the natural language data in the field of the business and the calculated value of the parameter into the natural language model to obtain output data for improving the value of the parameter corresponding to the natural language data, transmits the obtained output data to a predetermined information processing system to receive a score of the obtained output data, determines whether the received score is greater than a predetermined threshold value, and when it is determined that the received score is not greater than the predetermined threshold value, executes a retention process of inputting the data including the output data, the natural language data, and the received score into the natural language model to obtain new output data. The business support system according to claim 1.

8. The control device in the expertise metric calculation process, determines whether the value of the calculated parameter is less than or equal to a predetermined threshold value, and when it is determined that the value of the calculated parameter is less than or equal to the predetermined threshold value, outputs a warning to an output device The business support system according to claim 1

9. The storage device stores the natural language model for each of a plurality of fields The control device in the expertise metric calculation process, calculates the parameter for each of the plurality of fields based on the corresponding natural language model, and outputs a graph representing a comparison of the values of the calculated parameters The business support system according to claim 1

10. Further includes a display device for displaying a real or virtual space, and a device for receiving voice input The control device, in the expertise metric calculation process, acquires the voice input from the device to the device, and generates data of the natural language input to the natural language based on the acquired voice The business support system according to claim 1

11. The control device in the expertise metric calculation process, outputs a message or a graph corresponding to the value of the calculated parameter to an output device The business support system according to claim 1

12. A business support method by an information processing device including a storage device storing a database registering natural language data including a description related to a business field, and a control device, wherein the control device inputs the natural language data extracted from the database into a natural language model that takes the natural language data as an input and outputs data corresponding to the natural language data, thereby obtaining output data corresponding to the input natural language data, and based on the obtained output data, executes an expertise metric calculation process for calculating a value of a parameter representing the expertise in the business field of the natural language model, and outputs information on the value of the calculated parameter to an output device Business support method

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A

  • Text generation device and text generation method

    JP7313757B1