Information processing device, information processing method, and program

The information processing apparatus optimizes generative AI services by integrating communication control, business process selection, and result evaluation to automate business processes, enhancing accuracy and reliability while hiding AI complexity from users.

JP2026055783APending Publication Date: 2026-03-31WINGARC 1ST
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Conventional generative AI technologies require repetitive trial and error for achieving desired results, and users need to be aware of the AI processes, which is inefficient and cumbersome.

Method used

An information processing apparatus that integrates communication control, business process function selection, prompt embedding, parallel processing, and result evaluation to optimize generative AI services, allowing users to perform business processes without direct interaction with AI, ensuring reliability and flexibility.

Benefits of technology

Enables seamless business operations using generative AI behind the scenes, improving processing accuracy, reliability, and continuity by utilizing multiple AI services, and adapting to user proficiency and business requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026055783000001_ABST
    Figure 2026055783000001_ABST
Patent Text Reader

Abstract

The goal is to effectively utilize generative AI behind the scenes of business operations, enabling users to perform tasks without being aware of the generative AI. [Solution] The LLM gateway unit 51 controls communication with multiple vendors, each providing a generation AI service. The business processing function selection unit 52 selects a business processing function to be executed based on a business processing request received from the business terminal 6. The prompt embedding unit 54 embeds a prompt optimized for the selected business processing function into the processing block (processing request). The parallel processing unit 56 simultaneously (in parallel) sends the processing request, including the prompt, to multiple generation AI services. The optimal solution selection unit 57 evaluates the processing results returned from the multiple generation AI services and selects the optimal processing result based on the evaluation results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, an information processing method, and a program.

Background Art

[0002] Conventionally, there is a technique for extracting information from images and documents using generative AI (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in recent years, when using generative AI in business, there are characteristics in the way of writing prompts, and there are many cases where Try&Error is repeated until the desired result is achieved. There is a demand for solving such problems, but the conventional technologies including the technology of Patent Document 1 cannot sufficiently meet such demands.

[0005] The present invention has been made in view of such a situation, and an object thereof is to make full use of generative AI behind the scenes of business so that users can execute business processes without being aware of generative AI.

Means for Solving the Problems

[0006] To achieve the above object, an information processing apparatus according to an aspect of the present invention includes communication control means for controlling communication with a plurality of vendors each providing a generative AI service, business process function selection means for selecting a business process function to be executed from among a plurality of business process functions related to the business document based on reading conditions of the business document preset according to business purposes. A prompt embedding means that incorporates a prompt for a generation AI service optimized for the processing content of the selected business processing function into the processing request, A parallel processing means that transmits the processing request to the multiple vendors in parallel, A processing result selection means that evaluates each processing result returned from the aforementioned multiple vendors based on pre-set evaluation indicators and selects the optimal processing result based on the results of the evaluation, It is equipped with.

[0007] An information processing method and program according to one aspect of the present invention are, respectively, a method and a program corresponding to an information processing apparatus according to one aspect of the present invention. [Effects of the Invention]

[0008] According to the present invention, it becomes possible to utilize generative AI in the background of business operations, enabling users to perform business processes without being aware of the generative AI. [Brief explanation of the drawing]

[0009] [Figure 1] This figure shows an overview of the service that can be realized by an information processing system to which a server according to one embodiment of the information processing device of the present invention is applied. [Figure 2] This figure shows an example of the configuration of an information processing system to which a server according to one embodiment of the information processing device of the present invention is applied. [Figure 3] Figure 2 is a block diagram showing an example of the server hardware configuration in the information processing system. [Figure 4] This is a functional block diagram showing an example of the functional configuration of the server in Figure 3 that constitutes the information processing system in Figure 2. [Figure 5] This figure shows an overview of the LLM function technology stack provided by the information processing device of the present invention. [Figure 6] This diagram shows the characteristics of each processing level in prompt-based and model-embedded business responses. [Figure 7]It is a diagram showing an example of a prompt instruction screen (processing block setting screen) specialized for causing an AI to perform OCR processing on an image. [Figure 8] It is a diagram showing an example of an income statement in information extraction processing from financial statements. [Figure 9] It is a diagram showing an example of a balance sheet in information extraction processing from financial statements. [Figure 10] It is a diagram showing an example of a setting screen in financial statement processing. [Figure 11] It is a diagram showing an example of an editing screen after reading a financial statement. [Figure 12] It is a diagram showing an example of a convenience store receipt. [Figure 13] It is a diagram showing an example of a hotel receipt. [Figure 14] It is a diagram showing an example of an aircraft receipt. [Figure 15] It is a diagram showing an example of a setting screen in transportation expense receipt processing. [Figure 16] It is a diagram showing an example of an editing screen for batch processing of multiple receipts. [Figure 17] It is a diagram showing an example of an invoice including both header information and detail information. [Figure 18] It is a diagram showing an example of a setting screen in invoice processing. [Figure 19] It is a diagram showing an example of a detail data acquisition screen for an invoice. [Figure 20] It is a diagram showing the system configuration in the case of invoiceAgent cooperation. [Figure 21] It is a diagram showing the system configuration in the case of MotionBoard cooperation. [Figure 22] It is a diagram showing a list of business categories considered as current business application patterns. [Figure 23] It is a diagram showing an example of a configuration in which a server realizes cooperation with multiple generative AI services. [Figure 24] It is a diagram showing an automatic switching function in case of a failure of a generative AI service. [Figure 25]This diagram shows the comparison and selection function for processing results obtained from multiple generation AI services. [Modes for carrying out the invention]

[0010] Embodiments of the present invention will be described below with reference to the drawings.

[0011] First, with reference to Figure 1, an overview of the service (hereinafter referred to as "this service") that can be realized by an information processing system (see Figure 2, described later) to which a server according to one embodiment of the information processing device of the present invention is applied will be explained. Figure 1 shows an overview of the service that can be realized by an information processing system to which a server according to one embodiment of the information processing device of the present invention is applied.

[0012] This service utilizes generative AI behind the scenes of business operations, supporting users (e.g., users) to perform business processes without being aware of the generative AI.

[0013] Specifically, Figure 1 shows, for example, how employee A operates the work terminal 6 to request OCR processing for image data such as an invoice (hereinafter referred to as "invoice image"). When employee A sends a business processing request from business terminal 6 (S1), server 1, which provides AI business processing services, receives the processing request. Server 1 then sends processing requests (requests) simultaneously to multiple AI service vendors (first AI vendor 2 (e.g., OpenAI®), second AI vendor 3 (e.g., Azure OpenAI), third AI vendor 4 (e.g., Google Vertex AI), fourth AI vendor 5 (e.g., Amazon Bedrock)) based on the received processing requests (S2).

[0014] Each generation AI service extracts text information from the invoice image using OCR processing and returns the result to Server 1. Server 1 compares the results returned from multiple generation AI services, selects the most appropriate (most likely) result, and returns it to the business terminal 6 (S3).

[0015] In this way, users can easily perform business processes that extract necessary information from image data of business documents without having to be aware of selecting a generation AI service or writing prompts. Furthermore, by using multiple generation AI services in parallel and obtaining multiple processing results, the reliability of the processing results is improved by selecting and using the result that is most suitable for the task from among the multiple processing results, and business operations can continue even if a specific service fails.

[0016] Although not shown in Figure 1, this service may also allow for setting processing levels in stages. In other words, Server 1 can select from three levels: a general-purpose processing level that uses prompts entered by the user through free-form text input (e.g., Level 0 in Figure 6); a basic function level that uses prompts combining fixed options and free-form text input for basic business processing functions such as extracting information from electronic data of business documents (e.g., Level 1 in Figure 6); and a specific business level that incorporates prompts optimized for specific tasks (e.g., Level 2 in Figure 6). This enables flexible processing tailored to the user's skill level and business requirements.

[0017] Although not shown in Figure 1, this service may also be configured to perform processing specifically tailored to the forms of a particular service provider. In other words, Server 1 can incorporate prompts, AI vendors, models, etc., that are specialized for the format and description style of documents issued by specific service providers (e.g., Level 2 in Figure 6), such as airline invoices, taxi dispatch service receipts, and railway company ticket receipts, into encapsulated processing blocks. This significantly improves the accuracy of recognizing documents from specific service providers.

[0018] Although not shown in Figure 1, this service may also provide a graphical user interface. In other words, Server 1 allows users to select business processing functions via a GUI, and automatically incorporates the optimal prompt even if the user does not know how to write prompts. This will allow users without knowledge of prompt engineering to utilize advanced generative AI processing.

[0019] Although not shown in Figure 1, this service may also automate error handling. In other words, if any of the generation AI services fails, Server 1 can automatically switch to another service and continue processing. This improves service availability and ensures business continuity.

[0020] Although not shown in Figure 1, this service may be configured to be expandable. In other words, Server 1 is equipped with an expandable interface that allows for the addition of new generative AI services, and can automatically adapt existing business processing functions to the added services. This will make it easier to adapt to new AI services that may emerge in the future.

[0021] Although not illustrated in Figure 1, this service may also provide business templates. In other words, Server 1 provides templates for various business categories, such as electronic bookkeeping compliance, daily report creation, inquiry management, utilization of internal company know-how, competitor research, order processing, inspection, receipt slip processing, document data entry, meeting efficiency improvement, and information / report distribution. Users can start processing tasks simply by selecting the appropriate template. This will allow for immediate application to business operations.

[0022] The evaluation metrics may include the degree of agreement between multiple processing results, the priority of the generating AI service, processing speed, cost, majority vote, or a combination thereof. For example, as shown in Figure 25, if three generating AI services return the same result "T001" and one service returns a different result "T00i", "T001" with the higher degree of agreement will be selected. Also, if a priority is set, the result of the service with the highest priority will be selected.

[0023] Next, with reference to Figure 2, we will describe the configuration of an information processing system to which an information processing system that realizes the provision of the above-mentioned service, i.e., an information processing system to which a server according to one embodiment of the information processing device of the present invention is applied. Figure 2 shows an example of the configuration of an information processing system to which a server according to one embodiment of the information processing device of the present invention is applied.

[0024] The information processing system shown in Figure 2 is configured to include a server 1, a first AI vendor 2, a second AI vendor 3, a third AI vendor 4, a fourth AI vendor 5, and a business terminal 6. Server 1, the first AI vendor 2, the second AI vendor 3, the third AI vendor 4, the fourth AI vendor 5, and the business terminal 6 are interconnected via a network N such as the Internet.

[0025] Server 1 is an information processing device managed by the service provider of this service (Figure 1). Server 1 performs various processes to realize this service while communicating as needed with the first AI vendor 2, the second AI vendor 3, the third AI vendor 4, the fourth AI vendor 5, and the business terminal 6. The first AI vendor 2 is a system that provides generative AI services. The second AI vendor 3 is a generative AI service system provided by Microsoft (registered trademark) and others. The third AI vendor 4 is a generative AI service system provided by Google® and others. The fourth AI vendor 5 is a generative AI service system provided by Amazon® and others. The business terminal 6 is an information processing device operated by employee A, and consists of a smartphone, tablet, personal computer, etc.

[0026] Figure 3 is a block diagram showing an example of the server hardware configuration in the information processing system shown in Figure 2.

[0027] Server 1 comprises a CPU (Central Processing Unit) 11, ROM (Read Only Memory) 12, RAM (Random Access Memory) 13, a bus 14, an input / output interface 15, an input unit 16, an output unit 17, a storage unit 18, a communication unit 19, and a drive 20.

[0028] The CPU 11 executes various processes according to the program recorded in the ROM 12 or the program loaded from the storage unit 18 into the RAM 13. RAM13 also stores data and other information necessary for the CPU11 to perform various processes.

[0029] The CPU 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output interface 15 is also connected to this bus 14. An input / output interface 15 is connected to an input unit 16, an output unit 17, a storage unit 18, a communication unit 19, and a drive 20.

[0030] The input unit 16 is configured, for example, with a keyboard, and is used to input various types of information. The output unit 17 consists of a display such as an LCD and a speaker, and outputs various information as images and sounds. The memory unit 18 is composed of DRAM (Dynamic Random Access Memory) and stores various types of data. The communication unit 19 communicates with other devices (for example, the first AI vendor 2, second AI vendor 3, third AI vendor 4, fourth AI vendor 5, and business terminal 6 in Figure 2) via a network N including the Internet.

[0031] The drive 20 is appropriately equipped with removable media 21, such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory. Programs read from the removable media 21 by the drive 20 are installed in the storage unit 18 as needed. Furthermore, the removable media 21 can store various types of data stored in the storage unit 18, just as the storage unit 18 does.

[0032] Although not shown in the diagram, the first AI vendor 2, second AI vendor 3, third AI vendor 4, fourth AI vendor 5, and business terminal 6 in Figure 2 can also have a configuration that is basically the same as the hardware configuration shown in Figure 3. Therefore, the explanation of the hardware configuration of the first AI vendor 2, second AI vendor 3, third AI vendor 4, fourth AI vendor 5, and business terminal 6 will be omitted.

[0033] Through the cooperation of various hardware and software components that make up the information processing system in Figure 2, including Server 1 in Figure 3, various processes for providing the service in Figure 1 can be executed.

[0034] Figure 4 is a functional block diagram showing an example of the functional configuration of the server in Figure 3 within the information processing system shown in Figure 2.

[0035] As shown in Figure 4, the CPU 11 of server 1 functions as follows: LLM gateway unit 51, business processing function selection unit 52, business optimization unit 53, prompt embedding unit 54, AI service switching unit 55, parallel processing unit 56, and optimal solution selection unit 57. Furthermore, one area of ​​the storage unit 18 of server 1 is provided with prompt information DB71, model information DB72, business pattern information DB73, AI service information DB74, and priority information DB75.

[0036] The prompt information DB71 stores optimized prompt templates corresponding to each business processing function, as well as prompts specific to the forms of particular service providers. The Model Information DB72 stores a list of models provided by each generation AI service, the characteristics of each model, and the correspondence between business processing functions and models. The business pattern information DB73 stores a list of available business processing functions, the processing level (Level 0 to 2) for each business, templates for each business category, and the processing parameters required for each business. The AI ​​service information DB74 stores the API endpoint, authentication information, service operating status, and supported processing capabilities for each generated AI service. The priority information DB75 stores the priority, cost information, processing speed performance, and reliability score for each generated AI service.

[0037] The LLM gateway unit 51 controls communication with multiple generation AI services (first AI vendor 2, second AI vendor 3, third AI vendor 4, and fourth AI vendor 5). The LLM gateway unit 51 obtains API endpoint information for each generation AI service from the AI ​​service information DB 74 and establishes a connection via the communication unit 19.

[0038] The business processing function selection unit 52 selects the business processing function to be executed based on the business processing request received from the business terminal 6. Specifically, the business processing function selection unit 52 obtains a list of available business processing functions from the business pattern information DB 73 and selects the function that best suits the user's request. There are three levels of selectable business processing functions: general processing level (Level 0), basic function level (Level 1), and specific business level (Level 2).

[0039] The business optimization unit 53 sets the optimal processing parameters for the selected business processing function. Specifically, the business optimization unit 53 obtains detailed information about the relevant business from the business pattern information DB 73 and makes the necessary settings for processing that business. For example, if invoice processing is selected, it sets the extraction pattern for invoice-specific items (registration number, date, amount, etc.).

[0040] The prompt integration unit 54 integrates prompts for the generated AI service, optimized for the selected business processing function, into a block (data unit of processing request). Specifically, the prompt integration unit 54 obtains a prompt template for the relevant business from the prompt information DB 71 and generates a complete prompt by combining it with the data to be processed. It also selects the most suitable model for each generated AI service from the model information DB 72 and integrates it into the block along with the prompt.

[0041] The AI ​​service switching unit 55 monitors the availability of the generated AI service and automatically switches to another service in the event of a failure. The AI ​​service switching unit 55 obtains the priority of each service from the priority information DB 75 and instructs the LLM gateway unit 51 to attempt to connect to the service in order from the highest priority service.

[0042] The parallel processing unit 56 simultaneously (in parallel) sends processing requests with prompts embedded by the prompt embedding unit 54 to multiple generation AI services. The parallel processing unit 56 executes the same processing request to each generation AI service in parallel and receives the responses asynchronously. Responses from generation AI services that have encountered errors are excluded, and only the results from generation AI services that have responded successfully are passed to the optimal solution selection unit 57.

[0043] The optimal solution selection unit 57 evaluates the processing results returned from multiple generation AI services and selects the optimal result. Specifically, the optimal solution selection unit 57 evaluates each processing result based on pre-set evaluation indicators. These evaluation indicators include the degree of agreement between multiple processing results, the priority of the generation AI services, processing speed, and cost. For example, if three services return the same result and one service returns a different result, the unit selects the result with the highest degree of agreement. The selected result is sent back to the business terminal 6.

[0044] In this way, users can easily perform business processes that extract necessary information (such as text data for each item) from document images (business documents) without having to be aware of the selection of generation AI services or the writing of prompts. Furthermore, by using multiple generation AI services in parallel, the reliability of the processing results is improved, and business operations can continue even if a particular service fails.

[0045] Furthermore, by setting processing levels in stages, flexible processing becomes possible according to the user's proficiency and business requirements.

[0046] Furthermore, by performing processing specifically tailored to the forms of particular service providers, the accuracy of recognizing those forms is significantly improved.

[0047] Furthermore, by providing a graphical user interface, even users without knowledge of prompt engineering can utilize advanced generative AI processing.

[0048] Furthermore, automating error handling improves service availability and ensures business continuity.

[0049] Furthermore, by making the screen settings expandable, it will be easy to adapt to new AI services that may emerge in the future.

[0050] Furthermore, providing business templates enables immediate application to business processes.

[0051] Figure 5 shows an overview of the LLM function technology stack provided by a server according to one embodiment of the information processing device of the present invention. As shown in Figure 5, Server 1 provides a hierarchical technology stack for utilizing generative AI behind the scenes of business operations. This technology stack has three processing levels, and each level enables optimal utilization of generative AI according to the user's proficiency and business requirements.

[0052] The lowest layer, Level 0 (general-purpose processing level), is a primitive LLM layer where users can freely input prompts and directly access foundational generative AI services such as ChatGPT®, Gemini®, and Llama®. At this level, users can understand the detailed operation of the generative AI and freely write custom prompts to achieve highly flexible processing.

[0053] The intermediate layer, Level 1 (basic functionality level), functions as a multimodal interpretation layer, providing basic processing functions such as text extraction from images, information extraction from videos, and format generation. At this level, semi-standardized prompts are used, allowing users to utilize basic generative AI functions without being aware of specific prompt descriptions. Integration with products such as SVF®, invoiceAgent®, and Motionboard® is also achieved at this level.

[0054] The top-tier Level 2 (Specific Task Level) provides a concrete functionality layer, offering processing fully optimized for specific tasks such as invoice data extraction and construction site condition analysis. At this level, prompts are fully integrated, allowing users to focus solely on their task objectives and execute processes without being aware of the generative AI running in the background.

[0055] Figure 6 is a diagram that shows in detail the characteristics of each processing level in prompt and model-embedded business responses. As shown in Figure 6, the structure is such that as the functional level moves from general-purpose to specific, the presence of generative AI in the user's awareness gradually decreases, and their ability to concentrate on their work improves. This gradual design allows users with varying levels of proficiency to utilize generative AI in an optimal way according to their respective abilities and requirements.

[0056] At Level 0, users have free access to the generative AI, with complete control over both prompting and model selection. At this stage, users with expertise in generative AI can leverage maximum flexibility to achieve their own unique processing. The prompting function allows users to directly instruct the generative AI and obtain the desired results.

[0057] At Level 1, users can customize certain aspects of single-document OCR processing, achieving a basic level of efficiency. While model selection is automated at this stage, users can customize some prompts. This allows even users without specialized knowledge to perform processing utilizing basic generation AI functions.

[0058] At Levels 2-3, for specific tasks such as ANA invoice processing, users can fully concentrate on what they want to do, and the generated AI is provided as a pre-optimized function with integrated models and prompts. At this highest level, users can focus solely on the essential tasks of their work and do not need to be aware of any technical details.

[0059] Figure 7 shows an example of a prompt specification screen (processing block setting screen) specifically designed for instructing the generating AI to perform OCR processing on images. As shown in Figure 7, Server 1 provides the optimal processing blocks for each business purpose that the generating AI is to perform, and completely encapsulates the vendor, model, and prompts of the generating AI suitable for that processing, eliminating the need for the user to individually configure them. In this configuration screen, the user can focus only on the settings items necessary for their business, with the technical complexity hidden in the background.

[0060] Specifically, the basic settings area contains items directly related to business operations, such as error settings, output value confirmation, block number, and memo field. In the acquisition item settings, you can intuitively specify the information you want to extract, such as the business company and date, from a selection of options. Furthermore, multilingual support can be easily achieved by setting the prompt language.

[0061] This allows users to intuitively configure settings based on their desired functionality without having to be aware of the technical details of the generating AI. In the background, Server 1 automatically selects the optimal generating AI service, optimizes prompts, adjusts processing parameters, and more, to most efficiently fulfill the user's requests.

[0062] Figures 8 and 9 are diagrams that illustrate in detail examples of information extraction from financial statements. Figure 8 shows an income statement, and Figure 9 shows a balance sheet, demonstrating the effectiveness of the present invention in extracting information from these complex financial documents.

[0063] The income statement in Figure 8 starts with sales of 300,000 yen and includes financial information with a complex hierarchical structure, such as cost of goods sold (beginning inventory of 50,000 yen, purchases of goods for the period of 160,000 yen, etc.), selling, general and administrative expenses (personnel expenses of 20,000 yen, depreciation expenses of 5,000 yen, etc.), non-operating income and expenses, extraordinary gains and losses, and corporate taxes.

[0064] Figure 9 shows the detailed financial situation of the balance sheet, including assets (current assets: cash and deposits of 300 yen, notes receivable of 30 yen, etc., fixed assets: tangible fixed assets of 600 yen, etc.) and liabilities and equity (current liabilities: notes payable of 50 yen, etc., fixed liabilities: corporate bonds of 100 yen, etc., equity: shareholders' equity of 750 yen, etc.).

[0065] Server 1 enables the accurate extraction of necessary financial information from these highly structured financial documents simply by configuring image recognition and specifying the items to be read. This significantly streamlines the previously manual process of entering financial data, simultaneously reducing human error and improving processing speed.

[0066] Figure 10 is a detailed diagram showing the settings screen for financial statement processing. As shown in Figure 10, the settings screen 100 has an image recognition setting area 101 and a reading item setting area 102, providing an intuitive and easy-to-use user interface.

[0067] In the image recognition settings area 101, you can efficiently configure the types of documents to be processed. In the form fields, you can select the type of financial statement, such as "Income Statement" or "Balance Sheet," and in the input method field, you can specify the processing method, such as "OCR (Optical Character Recognition)." In the OCR settings, you can clearly define the scope of processing by selecting the format (e.g., Income Statement) and setting the maximum number of items (e.g., 10). Furthermore, you can customize the appearance of the user interface through the button design settings.

[0068] In the reading item setting area 102, you can set specific financial items to be extracted in detail. In the character recognition settings, hints (available to businesses) are provided for format selection, and key financial indicators such as sales revenue, gross profit, operating profit, ordinary profit, pre-tax net profit, and net profit can be individually specified as acquisition items. For each item, it is possible to set details such as the item name in the file to be scanned, the display name in the app, and hints to improve scanning accuracy.

[0069] Even more importantly, in addition to these standard forms and selection of extraction items, users can also directly specify the items they want to obtain using OCR by entering text. For example, they can specify additional information in natural language, such as "I want to process the image while acquiring it," "I want to extract detailed information on a specific account," or "I want to acquire the calculation formula as well." This feature allows for flexible handling of special requests that cannot be addressed by standard processing.

[0070] Figure 11 shows a detailed view of the editing screen after reading the data. As shown in Figure 11, the editing screen 110 effectively combines the profit and loss statement display area 111 and the read information area 112, providing an environment in which the user can intuitively check and edit the processing results.

[0071] In the income statement display area 111, the original image of the income statement being processed is clearly displayed, allowing the user to visually compare the original document with the extracted results. This image display allows for immediate verification of the accuracy of the OCR processing and corrections to be made as needed. Furthermore, by understanding the positional relationships of items in the image, it is possible to efficiently detect any missing or misrecognized data.

[0072] In the read information area 112, the numerical values ​​of each financial item extracted by OCR processing are displayed in a structured format. Key financial indicators such as sales of 300,000 yen, gross profit of 120,000 yen, operating profit of 30,000 yen, ordinary profit of 21,000 yen, net profit before tax of 20,000 yen, and net profit of 16,000 yen are displayed in an organized manner that retains the structure of the original income statement.

[0073] Furthermore, header information such as the company name "XYZ Co., Ltd." is accurately extracted and displayed, making document identification and management easier. Users can review these extracted results and edit them directly as needed, utilizing them as final business data. The editing function allows for the rapid correction of minor misrecognitions that occur during OCR processing, enabling the efficient acquisition of highly accurate financial data.

[0074] Figures 12 to 14 are representative examples illustrating the diversity of various types of receipts. Figure 12 shows a receipt from a convenience store, including a transaction record from April 18, 2025, at the Arai 1-chome convenience store. This receipt includes details of the items purchased, such as a seaweed rice ball for 167 yen and grilled beef skirt steak for 297 yen, a total transaction amount of 464 yen, tax information including 34 yen in consumption tax, and detailed transaction information such as the amount of points used and the points member ID. It includes characteristic elements of convenience store receipts, such as a vertical format, small font size, and diverse product information.

[0075] Figure 13 shows a hotel receipt, which includes information specific to the lodging industry, such as the accommodation fee of 6,500 yen, the receipt date of June 2, 2024, the reservation order number IN119537685, and the name of the facility used (YYYYYYY). This receipt clearly shows the breakdown of charges, including the accommodation fee of 6,500 yen, breakfast charges, etc., and the payment method, such as credit card payment, representing a standard receipt format in the lodging industry.

[0076] Figure 14 shows a receipt for an airline ticket purchase, showing a ticket price of 62,050 yen, an issue date of October 18, 2023, and including airline industry-specific information such as the flight number and reservation number. This receipt contains structured, detailed transaction information specific to the airline industry, including fare details, timetables, and flight change information.

[0077] Server 1 can provide highly accurate information extraction processing through a unified interface for diverse receipts with different industries, formats, and information structures. Optimized prompts and model selection, taking into account the specific characteristics of each industry, achieve high recognition accuracy even for industry-specific terminology and formats.

[0078] Figure 15 is a detailed diagram showing the settings screen for processing travel expense receipts. As shown in Figure 15, the settings screen 150 effectively combines the image recognition setting area 151 and the reading item setting area 152 to provide functions specifically tailored for travel expense settlement operations.

[0079] In the image recognition setting area 151, form fields such as title input, description input, input method selection, and character recognition settings can be configured. A particularly important feature is the ability to pre-define and select various types of transportation-related documents, such as travel expense receipts, airline ticket receipts, and accommodation receipts. Furthermore, the ability to set a maximum number of fields allows for efficient processing of large volumes of receipts.

[0080] In the reading item setting area 152, you can systematically set the specific items necessary for expense reimbursement. Select "Expense Receipt" as the format, and hints (available to businesses) are provided to help you select this format. The items to be retrieved include company name (please retrieve the official name of the transportation company), date (please retrieve month, day, time, and date information), amount (please retrieve only numbers), and description number (please retrieve using hyphens), etc., comprehensively covering the main items necessary for expense reimbursement.

[0081] Even more importantly, in addition to these standard settings, users can specify special requests using natural language text input. For example, they can specify requests that cannot be handled by standard settings, such as "I want to get detailed information about the travel section," "I want to check whether a discount is applicable," or "I want to perform character recognition while improving image quality," by describing them in natural language. This flexibility allows the system to accommodate a wide range of travel expense reimbursement requests.

[0082] Figure 16 is a detailed diagram showing the editing screen for processing multiple receipts at once. As shown in Figure 16, the editing screen 110 allows for the simultaneous processing of multiple receipts of different types (convenience store, airline tickets, accommodation, etc.), and enables the extraction and display of necessary information such as company name, date, and amount from each receipt in a single batch. This batch processing function makes it possible to efficiently process a variety of receipts that previously had to be processed individually, using a unified workflow.

[0083] Specifically, diverse transaction information has been accurately extracted, such as 464 yen from a convenience store on April 18, 2025; 62,050 yen from XXXX Airlines on April 16, 2025; and 6,500 yen from a hotel named YYYYYYY on May 31, 2024.

[0084] Furthermore, "edit" and "re-analyze" functions are provided for each receipt, allowing users to individually review and modify the extracted results. This ensures consistent results through unified processing, even when receipts from different industries and formats are mixed together.

[0085] The batch processing function allows users to significantly reduce the effort required for individual settings when processing large volumes of receipts, dramatically improving operational efficiency. Furthermore, the consistent and accurate processing results ensure smooth execution of subsequent expense reimbursement processes and integration with accounting systems.

[0086] Figure 17 shows a detailed example of a complex invoice that includes both header and item information. As shown in Figure 17, this invoice includes basic transaction information as header information, such as the issue date May 28, 2025, issuer "Sample Co., Ltd." (with company seal), recipient "Client Co., Ltd.", Mr. Taro Yamada, General Affairs Department Accounting, and total amount 672,216 yen (including consumption tax, etc.).

[0087] Further detailed information is provided in a structured table format, including product name (SVF, invoiceAgent, Dr.Sum®, MotionBoard, dejiren®), quantity (1, 2, 3, 4, 5), unit price (11,111, 22,222, 33,333, 44,444, 55,555), and total amount (¥11,111, ¥44,444, ¥99,999, ¥177,776, ¥277,775).

[0088] The document also includes tax calculation information such as a subtotal of 611,105 yen, consumption tax (10%) of 61,111 yen, and a total of 672,216 yen, as well as payment details such as the recipient bank, "○○○ Bank ○○ Branch, Ordinary Account No. 0000000, Sample Co., Ltd."

[0089] Server 1 enables accurate extraction of both header and detail information from invoices with such complex structures, simply by using image recognition and setting reading items, and obtaining them as structured data. This significantly automates invoice processing tasks that were previously performed manually, simultaneously improving processing accuracy and efficiency.

[0090] Figure 18 is a detailed diagram showing the settings screen for invoice processing. As shown in Figure 18, the settings screen 180 effectively integrates the image recognition setting area 181 and the header / detail reading area 182, providing advanced settings functions to handle complex invoice processing.

[0091] In the image recognition settings area 181, you can enter the form name and application website, and in addition to basic settings such as error settings, output value confirmation, block number, and input settings, you can also configure detailed OCR (optical character recognition) settings. In particular, specifying a specific generation AI model such as OpenAI ChatGPT gpt4o ensures optimal processing performance. Detailed formatting settings for forms allow for flexible handling of various invoice formats.

[0092] The header / detail reading area 182 allows for detailed item settings specifically for invoice processing. For character recognition settings, guidance such as "Enter text" and "Please enter the format" allows users to intuitively configure the settings.

[0093] In the retrieval items, you can set detailed header information such as customer name (please accurately retrieve the customer name and billing name), transaction date (please retrieve the transaction date, including the date information from the invoice, including the main body, month and day, date, and recording time), and total amount (please retrieve only the total amount).

[0094] Furthermore, by checking a checkbox to retrieve details, you can enable the acquisition of detailed information, and individually set detail items such as item name, quantity, unit price, and amount as item reading items. By setting detailed extraction instructions for each item, such as using symbols or only extracting numbers, highly accurate information extraction can be achieved.

[0095] Importantly, in addition to these standard settings, users can specify special requests in natural language via text input. For example, complex requests that cannot be handled by standard processing, such as "I want to capture the image including the seal impression," "I want to improve the recognition accuracy of the handwritten portion," or "I want to perform calculation verification for specific items," can be specified by describing them in natural language.

[0096] Figure 19 is a detailed diagram showing the screen for retrieving invoice details. As shown in Figure 19, the detailed data acquisition screen 190 displays the extracted basic information and detailed information, along with the original invoice image, in a highly structured table format, allowing for editing of each item. This screen allows users to intuitively understand the processing results and efficiently review and edit them.

[0097] The invoice image on the left clearly displays the invoice addressed to "Client Co., Ltd." dated May 28, 2025, allowing for visual confirmation of its correspondence with the original document. The invoice includes header information such as issuer information, recipient information, and total amount of 672,216 yen, as well as detailed item information such as product names (SVF, invoiceAgent, Dr.Sum, MotionBoard, dejiren, etc.) and corresponding quantities, unit prices, and amounts.

[0098] In the extraction results display on the right, the basic information tab and the details tab are separated, allowing for systematic management of each piece of information. In the details tab, the columns for serial number, item name, quantity, unit price, and amount are clearly separated, and detailed information such as 1st row: SVF, 1, 11111, 11111, 2nd row: invoiceAgent, 2, 22222, 44444, 3rd row: Dr.Sum, 3, 33333, 99999, 4th row: MotionBoard, 4, 44444, 177776, 5th row: dejiren, 5, 55555, 277775 is accurately extracted and displayed.

[0099] Furthermore, editing functionality is provided for each line item, allowing users to individually review and modify the extracted results. New line items can also be added using the "Add" button, enabling flexible line item management. This advanced line item processing functionality significantly streamlines the process of extracting information from complex invoices and reduces manual data entry errors.

[0100] Figure 20 is a diagram that shows the system configuration in the case of integration with invoiceAgent in detail. As shown in Figure 20, the operational image and the value that can be provided by utilizing generation AI_OCR in AI x server integration are comprehensively illustrated. In this integrated system, the manual input workload for compliance with the Electronic Bookkeeping Act can be significantly reduced by integrating with the non-standard OCR function that generation AI excels at.

[0101] The system flow begins with the generation AI being used as the entry point. In this environment, non-standard forms such as faxed receipts and purchase orders with varying formats for each trading partner are collected through invoiceAgent's automatic / manual archiving function. These diverse document formats possess a non-standard nature and complexity that makes them difficult to process with conventional OCR technology.

[0102] Next, Server 1 operates as an API external request function and executes the conversion process of the collected document images. At this stage, the OCR request (processing request) is sent to each vendor of the generation AI service along with a prompt, and character recognition is performed by the generation AI through advanced multimodal processing. Due to the generation AI's natural language understanding capabilities, it can achieve high-precision recognition even for handwritten characters, complex layouts, and various fonts, which were difficult to handle with conventional OCR.

[0103] The OCR processing results from the generating AI, that is, the data converted from images to text, are processed on Server 1 (associating text with each item in the form) and then returned to InvoiceAgent. At this time, the extracted data is automatically incorporated into the appropriate business flow through the automatic custom property update and automatic sorting function.

[0104] This system integration will significantly reduce the human workload involved in complying with the Electronic Bookkeeping Law, simultaneously improving processing accuracy and optimizing operational efficiency. In particular, by highly automating tasks that previously relied on manual processes for handling non-standard forms, it will provide significant value in reducing the burden on companies complying with the Electronic Bookkeeping Law.

[0105] Figure 21 is a diagram showing a detailed system configuration when using MotionBoard integration. As shown in Figure 21, Case 5, a dashboard analysis example, illustrates the use of generative AI in dashboard analysis. In this system, Server 1 periodically acquires MotionBoard dashboards and sends them to stakeholders along with the analysis results during scheduled distribution, thereby realizing data-driven decision support.

[0106] The specific workflow begins with the closing operations dashboard, which visualizes various business metrics (team performance charts, period-based trend analysis, goal achievement rates, etc.). This dashboard includes various types of graphs, tables, and metrics, providing an integrated display of complex business information.

[0107] Next, Server 1 retrieves this dashboard information and performs advanced analysis processing using a generating AI. Based on the analysis results, value-added services such as information notification to stakeholders, flow creation, and automatic report generation are provided.

[0108] A more important feature is that, based on the generated AI analysis results, Server 1 automatically issues alert notifications to the responsible personnel. For example, along with quantitative analysis results such as "Team A: Actual results 2,861, achievement rate 58%, 20% compared to the previous month" and "Team B: Actual results 3,157, achievement rate 63%, 12% compared to the previous month," a message is delivered that includes specific improvement suggestions such as "Team A's achievement rate is below the target. We recommend considering additional measures."

[0109] Such advanced analytics and notification systems automate dashboard monitoring tasks that previously relied on manual processes, enabling timely decision-making support and business improvement suggestions. Furthermore, regular distribution of analytical reports facilitates overall organizational visibility and streamlines the improvement cycle.

[0110] Figure 22 is a diagram showing a comprehensive list of business categories currently being considered as application patterns for business operations. As shown in Figure 22, Server 1 can provide optimized functional blocks corresponding to 11 major business categories across diverse business domains. Each business category addresses processing requests specific to a particular industry or business process, and each comes pre-configured with the most suitable models and prompts.

[0111] For electronic bookkeeping compliance operations (Pri: high), features such as electronic bookkeeping compliance processing for faxed documents, automatic OCR data reading, and document transmission from smartphones ensure reliable compliance with legal requirements. For daily report creation / activity recording, features such as form-free daily report creation and daily report creation from on-site photos and audio recordings support the efficiency of on-site operations.

[0112] In inquiry management, the quality of customer support operations is improved by managing information such as summarizing customer inquiry information and system update schedules. In terms of leveraging internal know-how, it is possible to handle questions using internal know-how information (electronic files, CSV / Excel®, etc.) and build a knowledge base using RAG functionality.

[0113] In competitive analysis, the system supports strategic decision-making by providing pricing support based on competitors' flyers and automatically collecting market analysis information. In order processing, it enables sophisticated procurement operations through a fax-free ordering system (EDI), master checks based on parts information (RAG), and order handling.

[0114] In inspection operations, the system facilitates the automation of the quality control process through the reading of card-type forms and data checking functions. In receipt processing, the system improves the accuracy of medical and pharmaceutical operations through OCR for receipts at pharmacies, location information, and identity verification.

[0115] In document data entry, the digitization and storage of data from documents significantly improves the efficiency of traditional manual data entry. In meeting efficiency improvements, meeting minutes and speaker separation are implemented, and pre-meeting management reviews and understanding are conducted via chat, simultaneously shortening meeting times and improving their quality.

[0116] Information and report distribution allows for the quantification of BI analysis results in text format, and the construction of a data-driven management foundation through trend analysis and other methods. For these business categories, it can process diverse data formats such as text, audio, images, and electronic files without requiring AI to be aware of the generation process, achieving processing optimized for the specific characteristics and business requirements of each industry. Furthermore, it incorporates related technological elements from previously filed patents, enabling the provision of a comprehensive business automation solution.

[0117] Figure 23 shows a detailed configuration example in which a server enables collaboration with multiple generation AI services. As shown in Figure 23, Server 1 enables collaboration with multiple generation AI services, and the configuration of the automation mechanism is shown, taking into consideration the ability for users to use the generation AI without being aware of it. In this configuration, users can call multiple generation AIs uniformly from a single block, creating an environment where they do not need to be aware of the technical characteristics or availability status of each generation AI service at all.

[0118] The upper layers of the system house major generative AI service providers such as OpenAI, Azure OpenAI, Google Vertex AI, and Amazon Bedrock, and also ensure scalability for the addition of new services in the future under "Others." Although each generative AI service has different API specifications, authentication methods, and processing performance characteristics, Server 1 absorbs these differences and provides a unified interface.

[0119] Server 1 centrally manages communication with these multiple generation AI services and continuously monitors the operational status, processing speed, cost, etc., of each service. When a processing request is received from a user, Server 1 automatically performs tasks such as selecting the optimal service, optimizing prompts, and controlling parallel processing, allowing the user to obtain the processing results necessary for their work without having to be aware of any technical details.

[0120] At the lower layer, basic functions such as "sending a conversation" and "ending a conversation" provided by the generative AI service are abstracted and reconfigured as higher-level functions tailored to business objectives. This abstraction ensures compatibility between different generative AI services and enables transparent migration and parallel use between services.

[0121] Furthermore, this integrated architecture allows for easy addition of new generative AI services through the expansion interface of Server 1, enabling service expansion while maintaining compatibility with existing business processing functions.

[0122] Figure 24 is a diagram illustrating in detail the automatic switching function in the event of a failure in the generation AI service. As shown in Figure 24, the mechanism for handling failures and ensuring business continuity is illustrated, assuming that the processing results are the same for all generating AIs. In the previous state, where only a specific cloud service was contracted, if the generating AI service stopped due to a failure or other reason, business interruption was unavoidable.

[0123] However, a system utilizing Server 1 can ensure business continuity by continuously monitoring other services and automatically switching to available services. This automatic switching system has pre-assigned priorities for each AI generation service, implementing a hierarchical prioritization system such as Amazon Bedrock: priority 1, Azure OpenAI: priority 2, Google Vertex AI: priority 3.

[0124] The specific failure response flow involves first attempting to connect to Amazon Bedrock, which is the highest priority. If this service is running normally, processing requests will be executed as scheduled. However, if Amazon Bedrock experiences a service failure or becomes unreachable, Server 1 will automatically switch to Azure OpenAI, which is the next highest priority.

[0125] In the diagram, OpenAI (two instances) is indicated by an "X" mark, showing that these services are unavailable. In this situation, Server 1 automatically distributes processing to available services, avoiding a "business interruption" and enabling continuous business execution through "automatic switching" and "prioritization" from a "service connection unavailable" state.

[0126] This advanced fault tolerance capability enables the construction of robust business systems that leverage the redundancy of multiple services, rather than relying on a single generation AI service. Furthermore, automatic switching, including prioritization, allows for the optimization of the balance between cost-effectiveness and processing quality while ensuring business continuity.

[0127] Figure 25 is a diagram that shows in detail the comparison and selection function for processing results obtained from multiple generation AI services. As shown in Figure 25, in situations where it is difficult to predict which generation AI will produce the best answer, Server 1 simultaneously sends processing requests to multiple generation AIs, comprehensively evaluates the results, and selects the optimal answer—a sophisticated mechanism.

[0128] The specific processing flow involves Server 1 first sending the same processing request (for example, OCR processing of an invoice) in parallel to four generative AI services: OpenAI, Azure OpenAI, Google Vertex AI, and Amazon Bedrock. Each service performs the processing independently, generating results based on different recognition algorithms and training data.

[0129] As an example of processing results, the case is shown where OpenAI returned "Registration Number: T001", Azure OpenAI returned "Registration Number: T001", Google Vertex AI returned "Registration Number: T00i", and Amazon Bedrock returned "Registration Number: T001". In this case, three of the four results matched as "T001", and one differed from the others, showing "T00i".

[0130] Server 1's optimal solution selection function analyzes this result pattern and performs a logical evaluation, stating, "Since the three results match, T001 is judged to have the highest probability, and the result is returned as T001." This evaluation process implements a sophisticated selection algorithm that considers not only a simple majority rule principle, but also a combination of factors such as the past performance of each generating AI service, the characteristics of the data being processed, and reliability scores.

[0131] Furthermore, this parallel processing and results comparison system automatically corrects temporary performance degradations and recognition errors in individual AI generation services, maintaining high overall processing accuracy. Additionally, by utilizing diverse processing results obtained from multiple services, it is possible to identify processing problems and areas for improvement that would be difficult to discover using a single service.

[0132] This approach of simultaneously requesting processing from multiple generative AIs can effectively support users in obtaining the correct answer, significantly reducing the traditional trial-and-error process and improving both operational efficiency and processing accuracy.

[0133] Although one embodiment of the present invention has been described above, the present invention is not limited to the embodiments described above, and any modifications, improvements, etc. that can achieve the objectives of the present invention are considered to be included in the present invention.

[0134] Furthermore, the system configuration shown in Figure 2 and the hardware configuration of Server 1 shown in Figure 3 are merely illustrative examples for achieving the objectives of the present invention and are not particularly limited.

[0135] Furthermore, the functional block diagram shown in Figure 4 is merely illustrative and not particularly limiting. In other words, it is sufficient that the information processing system in Figure 2 has the functionality to execute the various processes described above as a whole, and the functional blocks and databases used to realize this functionality are not particularly limited to the example in Figure 4.

[0136] Furthermore, the location of the functional blocks and database is not limited to Figure 4 and can be any location. For example, at least a portion of the functional blocks and database located on Server 1 may be provided on the First AI Vendor 2, Second AI Vendor 3, Third AI Vendor 4, Fourth AI Vendor 5, Business Terminal 6, or other information processing devices not shown.

[0137] Furthermore, the series of processes described above can be executed by hardware or by software. Furthermore, a single functional block may consist of hardware alone, software alone, or a combination of both.

[0138] When a series of processes are executed by software, the programs that make up that software are installed on a computer or other device from a network or storage medium. The computer may be a computer that is built into dedicated hardware. Furthermore, a computer can be any computer capable of performing various functions by installing various programs, such as a server, a general-purpose smartphone, or a personal computer.

[0139] Such recording media containing programs may consist not only of removable media (not shown) distributed separately from the main unit of the device to provide the program to the user, but also of recording media provided to the user in a state where they are pre-installed in the main unit of the device.

[0140] In this specification, the step of describing a program to be recorded on a recording medium includes not only processes that are performed chronologically in that order, but also processes that are not necessarily performed chronologically, but are executed in parallel or individually.

[0141] In summary, the information processing device to which the present invention applies only needs to have the following configuration, and can take various forms. That is, the information processing device to which the present invention is applied (for example, Server 1 in Figures 2 to 4) is: (1) A communication control means (for example, the LLM gateway unit 51 in Figure 4) that controls communication with multiple vendors (for example, the first AI vendor 2, the second AI vendor 3, the third AI vendor 4, and the fourth AI vendor 5 in Figure 1) that each provide a generation AI service, A business processing function selection means (for example, the business processing function selection unit 52 in Figure 4) selects a business processing function to be executed from among multiple business processing functions related to the business document based on pre-set reading conditions for the business document according to the business purpose (for example, the form of the business document and the items to be read), A prompt embedding means (e.g., prompt embedding unit 54 in Figure 4) that incorporates a prompt for a generated AI service optimized for the processing content of the selected business processing function into a processing request (e.g., processing block), A parallel processing means (for example, the parallel processing unit 56 in Figure 4) that transmits the processing request to the multiple vendors in parallel, A processing result selection means (for example, the optimal solution selection unit 57 in Figure 4) evaluates the processing results returned from the aforementioned multiple vendors based on pre-set evaluation indicators and selects the optimal processing result (for example, a processing result suitable for the user's business) based on the results of the evaluation. Having that will suffice.

[0142] In this way, it becomes possible to ensure business continuity, enable use without being aware of the generated AI, and improve the reliability of processing results.

[0143] (2) Furthermore, the business processing function selection means (for example, the business processing function selection unit 52 in Figure 4) A general-purpose processing level that uses prompts entered by the user's free-form text (e.g., Level 0 in Figure 6), For basic business processing functions that extract information from the electronic data of the aforementioned business documents, a basic function level (e.g., Level 1 in Figure 6) is used, which combines standardized options and the aforementioned free-form text prompts. A specific task level (e.g., Level 2 in Figure 6) incorporates prompts optimized for that specific task, From these three levels, select the level according to the user's instructions or the reading conditions. It is possible.

[0144] This enables flexibility through gradual optimization, allowing processing to be tailored to the user's skill level and business requirements.

[0145] (3) Furthermore, the business processing function selection means (for example, the business processing function selection unit 52 in Figure 4) If, at the aforementioned specific business level, the user selects a business processing function specialized for documents issued by a specific service provider (for example, the convenience store receipt in Figure 12, the hotel receipt in Figure 13, and the airline receipt in Figure 14), The prompt embedding means (for example, the prompt embedding unit 54 in Figure 4) incorporates a service-specific prompt that includes information regarding the form format, items to be recorded, and location of recording for the specific service provider. The processing result selection means (for example, the optimal solution selection unit 57 in Figure 4) determines the validity of each processing result based on the form format of the specified service provider and selects a processing result based on the result of the validity determination. It is possible.

[0146] This enables high-precision handling of specific service forms and significantly improves the accuracy of recognition of forms from specific service providers.

[0147] (4) Furthermore, the business processing function selection means (for example, the business processing function selection unit 52 in Figure 4) A graphical user interface (for example, the settings screen 100 in Figure 10 and the settings screen 150 in Figure 15) is provided, which allows the user to select any of the multiple business processing functions through user operation. The prompt embedding means (for example, the prompt embedding unit 54 in Figure 4) incorporates the prompt corresponding to the business processing function selected by the user's operation on the graphical user interface into the processing request. It is possible.

[0148] This eliminates the need for users to know how to write prompts, allowing them to easily perform business processes through GUI operations.

[0149] (5) Furthermore, the parallel processing means (for example, the parallel processing unit 56 in Figure 4) If an error response is received from any vendor (for example, if the service connection is unavailable as shown in Figure 24), the AI ​​service generated by that vendor is excluded, and processing continues using the processing results from other vendors. The processing result selection means (for example, the optimal solution selection unit 57 in Figure 4) makes the processing results from vendors that have responded normally the subject of the evaluation. It is possible.

[0150] This allows operations to continue even in the event of a failure in the generation AI service, and enables the automatic processing of trial and error.

[0151] (6) Furthermore, the communication control means (for example, the LLM gateway unit 51 in Figure 4) It features an expandable interface that allows the user to add new generative AI services through user operations, The prompt integration means is, The user's actions on the extended interface apply prompts corresponding to existing business processing functions to newly added generation AI services. It is possible.

[0152] This makes it easier to support new AI services and allows for future expansion.

[0153] (7) Furthermore, the business processing function selection means (for example, the business processing function selection unit 52 in Figure 4) It provides a graphical user interface where templates for different business categories (for example, the electronic bookkeeping compliance tasks shown in Figure 22, daily report creation / activity record, customer inquiry management, etc.) can be selected and arranged. It is possible.

[0154] This allows users to execute business processes simply by selecting a business template, without needing to know how to write prompts. [Explanation of Symbols]

[0155] 1...Server, 2...1st AI Vendor, 3...2nd AI Vendor, 4...3rd AI Vendor, 5...4th AI Vendor, 6...Business Terminal, 11...CPU, 12...ROM, 13...RAM, 14...Bus, 15...Input / Output Interface, 16...Input Unit, 17...Output Unit, 18...Storage Unit, 19...Communication Unit, 20...Drive, 21...Removable Media, 51... ··LLM Gateway Unit, 52··Business Processing Function Selection Unit, 53··Business Optimization Unit, 54··Prompt Embedding Unit, 55··AI Service Switching Unit, 56··Parallel Processing Unit, 57··Optimal Solution Selection Unit, 71··Prompt Information DB, 72··Model Information DB, 73··Business Pattern Information DB, 74··AI Service Information DB, 75··Priority Information DB, A··Business Personnel, N··Network

Claims

1. A communication control means for controlling communication with multiple vendors that each provide a generation AI service, A business processing function selection means that selects a business processing function to be executed from among multiple business processing functions related to a business document, based on pre-set reading conditions for the business document according to the business purpose, A prompt embedding means for incorporating a prompt for a generation AI service optimized for the processing content of the selected business processing function into a processing request, A parallel processing means that transmits the processing request to the multiple vendors in parallel, A processing result selection means that evaluates each processing result returned from the aforementioned multiple vendors based on pre-set evaluation indicators and selects the optimal processing result based on the results of the evaluation, An information processing device equipped with the following features.

2. The aforementioned business processing function selection means is A general-purpose processing level that uses prompts entered by the user's free-form text, A basic function level that uses a combination of predefined options and free-form text prompts for basic business processing functions that extract information from the electronic data of the aforementioned business documents, Task-specific levels with prompts optimized for specific tasks, From these three levels, select the level according to the user's instructions or the reading conditions. The information processing apparatus according to claim 1.

3. The aforementioned business processing function selection means is If, at the aforementioned specific business level, the user selects a business processing function specialized for documents issued by a specific service provider, The prompt integration means is, The system incorporates a service-specific prompt that includes information regarding the form format, items to be entered, and location of entries for the aforementioned specific service provider. The processing result selection means is Based on the form format of the specified service provider, the validity of each processing result is determined, and a processing result is selected based on the result of the validity determination. The information processing apparatus according to claim 2.

4. The aforementioned business processing function selection means is A graphical user interface is provided that allows the user to select any of the multiple business processing functions through user operation. The prompt integration means is, The prompt corresponding to the business processing function selected by the user's operation on the graphical user interface is incorporated into the processing request. The information processing apparatus according to claim 1.

5. The parallel processing means is If an error response is received from any vendor, the AI ​​service generated by that vendor will be excluded, and processing will continue using the processing results from other vendors. The processing result selection means makes the processing results from vendors that have responded normally the subject of the evaluation. The information processing apparatus according to claim 1.

6. The aforementioned communication control means is It includes an expandable interface that allows the user to add new generation AI services through user operation, The prompt integration means is, The user's operation on the extended interface applies prompts corresponding to existing business processing functions to the newly added generation AI service. The information processing apparatus according to claim 1.

7. The aforementioned business processing function selection means provides a graphical user interface in which templates for each business category are arranged in a selectable manner. The information processing apparatus according to claim 1.

8. An information processing method performed by an information processing device, A communication control step that controls communication with vendors providing multiple generation AI services, A business process function selection step in which, based on pre-set reading conditions for a business document according to the business purpose, a business process function to be executed is selected from among multiple business process functions related to the business document, and A prompt embedding step that incorporates a prompt for a generation AI service optimized for the processing content of the selected business processing function into the processing request, A parallel processing request step of sending the processing request to the multiple vendors in parallel, A processing result selection step involves evaluating each processing result returned from the multiple vendors in response to the processing request based on pre-set evaluation indicators, and selecting the optimal processing result based on the results of the evaluation. Information processing methods including

9. On the computer, A communication control step that controls communication with vendors providing multiple generation AI services, A business process function selection step in which, based on pre-set reading conditions for a business document according to the business purpose, a business process function to be executed is selected from among multiple business process functions related to the business document, and A prompt embedding step that incorporates a prompt for a generation AI service optimized for the processing content of the selected business processing function into the processing request, A parallel processing request step of sending the processing request to the multiple vendors in parallel, A processing result selection step involves evaluating each processing result returned from the multiple vendors in response to the processing request based on pre-set evaluation indicators, and selecting the optimal processing result based on the results of the evaluation. A program that executes control processes, including those mentioned above.

Citation Information

Patent Citations

  • Using generative artificial intelligence to supplement automated information extraction

    JP2025026286A