Processing apparatus, processing method, and processing program

The processing system addresses the challenge of suboptimal translation by using generative AI models and IOWN technology for real-time, context-aware translation, enhancing efficiency and usability in interpretation services.

JP7857339B2Active Publication Date: 2026-05-12NTT DOCOMO BUSINESS INC
View PDF 9 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NTT DOCOMO BUSINESS INC
Filing Date
2024-05-22
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Conventional interpretation systems lack the ability to provide optimal translation according to the situation, making it difficult to appropriately utilize translated content.

Method used

A processing system utilizing generative AI models like Tsuzumi and ChatGPT, combined with a low-latency IOWN network, to translate voice or text data in real-time, tailored to the specific context and needs of a meeting, including language, accuracy, and industry-specific requirements.

Benefits of technology

Enables efficient, real-time simultaneous interpretation among multiple users, improving business efficiency and reducing costs by eliminating translation delays and allowing for convenient audio output, while enhancing the appropriateness and usability of translated content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007857339000001
    Figure 0007857339000001
  • Figure 0007857339000002
    Figure 0007857339000002
  • Figure 0007857339000003
    Figure 0007857339000003
Patent Text Reader

Abstract

To enable appropriate use of translation content.SOLUTION: A server device 10 includes: an input control unit 13 which inputs voice data or text data received from a user terminal to a translation generation model with prompts set thereon for translating input voice data or text data into a designated language in a natural context; and a storage unit 63 which stores, in a storage unit 70, information on knowledge about tasks extracted based on translation data which is translated voice data or text data output from the translation generation model in accordance with the input by the input control unit 13.SELECTED DRAWING: Figure 38
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a processing device, a processing method, and a processing program.

Background Art

[0002] Conventionally, as an interpretation service using human resources, there has been a service such as entrusting interpretation to specialized contractors for each type of interpretation service, and the interpreter contracted by the contractor provides the result to the entrusting party.

[0003] And in recent years, as an interpretation system, various systems using IT technology have been provided, such as an automatic translation system between languages and a system that converts voice and text using speech recognition technology and speech synthesis technology.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, since conventional interpretation systems are only for automating translation processing, it is difficult to appropriately utilize the translated content.

[0006] The present invention has been made in view of the above, and an object thereof is to provide a processing device, a processing method, and a processing program that can provide an optimal translation according to the situation.

Means for Solving the Problems

[0007] To solve the above-mentioned problems and achieve the objective, the processing apparatus of the present invention is characterized by comprising: an input control unit that inputs audio data or text data received from a user terminal to a translation generation model that is set to prompt a prompt that commands the input audio data or text data to be translated into a specified language in a natural context; and a storage unit that stores in the storage unit information related to business knowledge extracted based on the translated data, which is the translated audio data or text data output from the translation generation model in response to the input by the input control unit. [Effects of the Invention]

[0008] According to the present invention, it is possible to make appropriate use of translated content. [Brief explanation of the drawing]

[0009] [Figure 1] Figure 1 shows an example of the configuration of a processing system according to the first embodiment. [Figure 2] Figure 2 is a diagram illustrating the overview of IOWN technology. [Figure 3] Figure 3 illustrates an example of the use of the processing system according to the first embodiment. [Figure 4] Figure 4 is a diagram illustrating the processing flow of the processing system. [Figure 5] Figure 5 is a diagram illustrating the overview of the processing system. [Figure 6] Figure 6 shows an example of pre-meeting information. [Figure 7] Figure 7 is a diagram illustrating step S11 shown in Figure 5. [Figure 8] Figure 8 is a diagram illustrating step S12 shown in Figure 5. [Figure 9] Figure 9 is a diagram illustrating step S13 shown in Figure 5. [Figure 10] Figure 10 is a diagram illustrating step S14 shown in Figure 5. [Figure 11] Figure 11 is a diagram illustrating step S15 shown in Figure 5. [Figure 12] Figure 12 is a diagram for explaining step S16 shown in Figure 5. [Figure 13] Figure 13 is a diagram for explaining step S17 shown in Figure 5. [Figure 14] Figure 14 is a diagram for explaining the processing of the processing system. [Figure 15] Figure 15 is a diagram showing another example of the prompt creation process. [Figure 16] Figure 16 is a diagram showing another example of the prompt creation process. [Figure 17] Figure 17 is a diagram showing another example of the prompt creation process. [Figure 18] Figure 18 is a diagram for explaining another usage example of the processing system. [Figure 19] Figure 19 is a diagram for explaining another usage example of the processing system. [Figure 20] Figure 20 is a diagram for explaining the outline of the task management support process of the processing system. [Figure 21] Figure 21 is an example of a sequence diagram showing the processing procedure of the processing method according to the first embodiment. [Figure 22] Figure 22 is an example of a sequence diagram showing the processing procedure of another processing method according to the first embodiment. [Figure 23] Figure 23 is a diagram for comparing the conventional translation service with the simultaneous interpretation service provided by the processing system according to the first embodiment. [Figure 24] Figure 24 is a diagram for comparing the conventional translation service with the simultaneous interpretation service provided by the processing system according to the first embodiment. [Figure 25] Figure 25 is a diagram for explaining the processing of the processing system according to the second embodiment. [Figure 26] Figure 26 is a diagram showing the configuration of the processing system according to the second embodiment. [Figure 27] Figure 27 is a diagram for explaining an example of the processing by the processing system according to the second embodiment. [Figure 28]Figure 28 is a diagram illustrating an example of processing by the processing system according to the second embodiment. [Figure 29] Figure 29 is a diagram illustrating an example of processing by the processing system according to the second embodiment. [Figure 30] Figure 30 is a sequence diagram showing the processing procedure according to the second embodiment. [Figure 31] Figure 31 is a sequence diagram showing the processing procedure according to the second embodiment. [Figure 32] Figure 32 is a sequence diagram showing the processing procedure according to the second embodiment. [Figure 33] Figure 33 is a diagram illustrating the processing of the processing system according to the third embodiment. [Figure 34] Figure 34 shows the configuration of the processing system according to the third embodiment. [Figure 35] Figure 35 illustrates an example of processing performed by the processing system according to the third embodiment. [Figure 36] Figure 36 is a sequence diagram showing the processing procedure according to the third embodiment. [Figure 37] Figure 37 is a diagram illustrating the processing of the processing system according to the fourth embodiment. [Figure 38] Figure 38 shows the configuration of the processing system according to the fourth embodiment. [Figure 39] Figure 39 is a table diagram showing an example of information regarding the knowledge of a business that has undergone a predetermined correspondence process according to the fourth embodiment. [Figure 40] Figure 40 is a diagram illustrating an example of processing by the processing system according to the fourth embodiment. [Figure 41] Figure 41 is a diagram illustrating an example of processing by the processing system according to the fourth embodiment. [Figure 42] Figure 42 is a sequence diagram showing the processing procedure according to the fourth embodiment. [Figure 43] Figure 43 is a sequence diagram showing the processing procedure according to the fourth embodiment. [Figure 44]Figure 44 is a diagram illustrating the processing of the processing system according to the fifth embodiment. [Figure 45] Figure 45 shows the configuration of the processing system according to the fifth embodiment. [Figure 46] Figure 46 is a table diagram showing an example of training data according to the fifth embodiment. [Figure 47] Figure 47 is a sequence diagram showing the processing procedure according to the fifth embodiment. [Figure 48] Figure 48 shows an example of a computer in which a server device is realized when a program is executed. [Modes for carrying out the invention]

[0010] Hereinafter, one embodiment of the present invention will be described in detail with reference to the drawings. However, the present invention is not limited by this embodiment. Furthermore, in the drawings, the same parts are denoted by the same reference numerals.

[0011] <First Embodiment> [Processing System] The configuration of the processing system according to the first embodiment will now be described. The processing system according to the first embodiment provides a simultaneous interpretation service in a meeting between multiple users, using generative AI (Artificial Intelligence) (generative model) to translate the voice data or text data of other users into the language specified by each user, and outputting the translated voice data or text data to a predetermined terminal.

[0012] Figure 1 is a diagram showing an example configuration of a processing system according to the first embodiment. As shown in Figure 1, the processing system 100 according to the first embodiment includes user terminals 20A to 20D used by users A to D, who are users of the simultaneous interpretation service, earphone microphones 30A to 30D attached to users A to D and communicating with the user terminals 20A to 20D, and a cloud server device 10. The server device 10 communicates with the generation AI server 40 and the generation AI server 50. When referring to the user terminals 20A to 20D collectively, they are referred to as user terminal 20. The number of user terminals 20 can be two or more, and is not limited to four.

[0013] The generation AI server 40 is equipped with the generation AI Tsuzumi (Registered Trademark) 41 (First Generation Model). Tsuzumi 41 is a natural language processing model finely tuned for specific fields. These specific fields include, for example, medicine, semiconductors, IT (Information Technology), academia, factories (plants), law, and office services. Tsuzumi 41 is built with an emphasis on low power consumption and has a faster processing speed than ChatGPT 51 (described later).

[0014] The generation AI server 50 is equipped with the generation AI ChatGPT (Registered Trademark) 51 (Second Generation Model). ChatGPT 51 is a large-scale natural language processing model, which is slower but more accurate than Tsuzumi 41. Tsuzumi 41 and ChatGPT 51 process the input voice or text data in natural language according to the set prompts, generate voice or text data, and output it. Note that the above generation AI is just one example, and servers equipped with multiple other generation AIs may be provided.

[0015] User terminals 20A to 20D are terminal devices capable of inputting and outputting voice and text data, and communicating with the server device 10. User terminals 20A to 20D may also input and output voice data via earphone microphones 30A to 30D worn by each user A to D.

[0016] User terminals 20A to 20D are, for example, PCs (Personal Computers), notebook PCs, tablet devices, smartphones, etc. User terminals 20A to 20D communicate with the server device 10 regarding simultaneous interpretation processing by launching the simultaneous interpretation application. At this time, user terminals 20A to 20D accept input of pre-meeting information about the meeting the user will participate in, based on the user's actions before the meeting, and transmit the pre-meeting information to the server device 10.

[0017] The server device 10 is implemented, for example, by loading a predetermined program into a computer including ROM (Read Only Memory), RAM (Random Access Memory), CPU (Central Processing Unit), etc., and having the CPU execute the predetermined program. The server device 10 also has a communication interface for sending and receiving various information with other devices connected via a network, etc. (for example, user terminals 20A to 20D, generation AI servers 40, 50).

[0018] The server device 10 provides a simultaneous interpretation service in meetings between multiple users, using Tsuzumi41 or ChatGPT51 to translate the voice or text of other users into the language specified by each user, and outputting the translated voice or text data to a designated terminal.

[0019] Specifically, the server device 10 selects either Tsuzumi41 or ChatGPT51 based on pre-meeting information about the meeting the user is participating in. The server device 10 then adds content based on the pre-meeting information to the selected generating AI and sets prompts that instruct it to translate the input voice or text data into a specified language in a natural context.

[0020] This allows the server device 10 to obtain the most appropriate translation from the generating AI, tailored to the situation, and provide it to the user.

[0021] Furthermore, the server device 10 communicates with user terminals 20A to 20D and the generation AI servers 40 and 50 via a low-latency communication network related to IOWN (Innovative Optical and Wireless Network) (hereinafter referred to as the IOWN network 60).

[0022] [Overview of IOWN technology] Here, we will explain IOWN technology. Figure 2 is a diagram illustrating the overview of IOWN technology. As shown in Figure 2, IOWN technology consists of three main technological fields: "All-Photonics Network (APN)", "Digital Twin Computing (DTC)", and "Cognitive Foundation (CF)".

[0023] [All-Photonics Network] The APN related to IOWN technology is a technology that enables the construction of high-speed networks by processing all network transmission functions in the optical domain. Specifically, the APN related to IOWN technology is a technology that realizes low power consumption, high quality, large capacity, and low latency communication based on optical-based (photonics-based) technologies such as "optoelectronic convergence technology," "high-capacity optical transmission system and device technology," "optical Ising machine," and "optical lattice clock network."

[0024] [Digital Twin Computing] DTC related to IOWN technology is a technology that maps individual objects from the real world onto a virtual space using the vast amount of data collected by devices connected to the aforementioned APN.

[0025] Traditional digital twin frameworks are used, for example, to map individual objects such as automobiles and robots onto a virtual space, perform analyses and predictions on them, and then project the results of those analyses and predictions back onto the real world.

[0026] On the other hand, DTC related to IOWN technology is a technology that develops the conventional concept of digital twins by freely combining digital twins of various industries and objects and people to perform calculations, thereby accurately reproducing combinations of multiple objects, such as people and automobiles in a city. Furthermore, DTC related to IOWN technology enables not only the representation of a person's external appearance but also the digital representation of their internal aspects such as consciousness and thoughts. This is achieved by combining technologies such as "speech recognition," "speech synthesis," and "understanding of emotions and intentions" to collect information and construct a digital twin environment.

[0027] Thus, DTC related to IOWN technology is a technology that enables the creation of digital twins that do not exist in the real world by combining multiple entities that are single in the real world and replicating them as digital twins in a virtual space, or by exchanging or merging some of the components between multiple digital twins.

[0028] [Cognitive Foundation] IOWN technology's Cloud Computing (CF) is a technology that centrally manages, configures, integrates, manages, and operates ICT (Information and Communication Technology) resources at different layers, such as cloud computing, edge computing, network services, and user equipment. Specifically, IOWN technology's CF treats diverse targets as a group of virtualized ICT resources and uses multi-orchestration functionality as a hub to optimally integrate multiple resources at different layers.

[0029] Furthermore, as shown in Figure 2, IOWN technology integrates the aforementioned APN, DTC, and network services provided by service providers to deliver high value-added services.

[0030] For example, as shown in Figure 2(1), IOWN technology provides a technology for transmitting information collected via APN to other terminal devices at high speed and with low latency. Also, as shown in Figure 2(2), IOWN technology provides a technology for transmitting large amounts of information collected from terminal devices and outputting analysis results and other information from the service provider's services at high speed and with low latency in information analysis services provided by service providers. Furthermore, as shown in Figure 2(3), IOWN technology provides a technology for constructing a digital twin environment using information acquired from surveillance cameras, automobile sensors, etc., by transmitting large amounts of information at high speed and with low latency, and for performing future predictions and outputting the prediction results to the user.

[0031] Based on the IOWN technology described above, it is believed that the large-capacity, high-speed, and low-latency information transmission infrastructure will facilitate the construction of digital twin environments and the collaboration between different digital twin environments.

[0032] The processing system 100 communicates via a high-capacity, high-speed, and low-latency information transmission infrastructure based on the IOWN technology described above. For example, when coordinating with the processing system 100, an APN is used to realize a low-latency simultaneous interpretation service. Specifically, the processing system 100 can provide a simultaneous interpretation service that outputs translated audio or text data in real time, even when audio or text data to be translated is input, and even when there are three or more users.

[0033] [Server equipment] Returning to Figure 1, the server device 10 will be described. The server device 10 includes a pre-meeting information acquisition unit 11 (acquisition unit), a prompt creation unit 12 (creation unit), an input control unit 13, an output control unit 14, and a summary creation unit 15.

[0034] The pre-meeting information acquisition unit 11 acquires pre-meeting information about the meeting that the user will participate in by communicating with user terminals 20A to 20D.

[0035] The prompt creation unit 12 selects one of several generating AIs based on the pre-meeting information. The prompt creation unit 12 selects either Tsuzumi41 or ChatGPT51 based on the pre-meeting information. The prompt creation unit 12 creates a prompt (command) that instructs the selected generating AI to translate the input voice data or text data into a specified language in a natural context, and adds content based on the pre-meeting information.

[0036] The prompt generation unit 12 uses a generating AI (Tsuzumi 41 or ChatGPT 51) to determine, based on pre-meeting information, the following: the specified language (e.g., the user's language), the accuracy of the meeting, the speed of responses, the meeting status, the industry related to the meeting, and / or whether the meeting is a specific field or a precision-focused meeting. The prompt generation unit 12 is not limited to the above and may change the content of its determination depending on the industry, field, members, and situation. The server device 10 may also determine, according to predetermined rules, the language used by user A, the accuracy of the meeting, the speed of responses, the meeting status, the industry related to the meeting, and / or whether the meeting is a specific field or a precision-focused meeting.

[0037] The prompt generation unit 12 selects either Tsuzumi 41 or ChatGPT 51 based on the content determined using the generation AI. For example, if the meeting is in a specific field (e.g., medicine) and the response speed is set to a relatively fast level where speed is important, the prompt generation unit 12 will select Tsuzumi 41. On the other hand, if accuracy is important, the prompt generation unit 12 will select ChatGPT 51.

[0038] The prompt creation unit 12 then adds the content determined by the generation AI to the prompt. Based on the determined content, the prompt creation unit 12 selects a generation AI suitable for the meeting, adjusts the prompt to be used when instructing that generation AI, and then issues a translation instruction to the selected generation AI. The prompt creation unit 12 outputs the created prompt to the input control unit 13.

[0039] The input control unit 13 provides the prompt created by the prompt creation unit 12 to the generation AI selected by the prompt creation unit 12, and inputs the voice data or text data entered from the user terminals 20A to 20D.

[0040] The output control unit 14 outputs the voice data or text data output from the generated AI selected by the prompt creation unit 12 to a predetermined terminal (user terminal 20A to 20D).

[0041] When the meeting ends, the summary creation unit 15 instructs the generation AI selected by the prompt creation unit 12 to create a summary of the meeting content in the language specified by each user A to D. The summary creation unit 15 then sends each summary created by the generation AI to each user's terminal 20A to 20D. The sent summaries are used by users A to D for meeting review, task management, etc. In addition to automatically creating summaries of meeting content, the generation AI can also automatically generate Q&A and save it to a database, and it is also possible to provide chatbot functionality.

[0042] [Examples of application] Figure 3 illustrates an example of using the processing system 100 according to the first embodiment. Figure 3 shows a meeting between user A who speaks Japanese, user B who speaks English, user C who speaks Chinese, and user D who speaks French as an example.

[0043] In this case, the prompt creation unit 12 creates a prompt that instructs the selected generating AI to translate the input voice data or text data into the language used by each user. Specifically, voice data spoken by other users B, C, and D that is output to user A is to be translated into Japanese. Voice data spoken by other users A, C, and D that is output to user B is to be translated into English. Voice data spoken by other users A, B, and D that is output to user C is to be translated into Chinese. Voice data spoken by other users A, B, and C that is output to user D is to be translated into French.

[0044] As a result, the generating AI translates the spoken voice of each user A, B, C, and D into the language used by each user A, B, C, and D. The output control unit 14 transmits the translated voice data or text data into the language used by each user A to D to each user terminal 20A to 20D used by each user.

[0045] As a result, each translated audio data can be output via each user's user terminal 20A-20D, or it may be output from the earphone microphones 30A-30D worn by each user A, B, C, and D. For example, the speech of users B, C, and D is translated into Japanese and output from user A's earphone microphone 30A.

[0046] [Process flow] The processing flow of the processing system 100 in the case of Figure 3 will be explained. Figure 4 is a diagram illustrating the processing flow of the processing system 100. Here, the example will be that user A's voice data is translated for other users B to D.

[0047] User A launches the application on user terminal 20A (referred to as "App" in the diagram) and enters a summary of the meeting as pre-meeting information before the meeting (Step S1). From the pre-meeting information, the generating AI (Tsuzumi41 or ChatGPT51) determines the specified language (the language used by User A), the industry related to the meeting, the meeting situation, the accuracy of the meeting, the response time (degree of response speed), and / or whether the meeting is a specific field meeting or an accuracy-focused meeting.

[0048] The server device 10 uses a generation AI (Tsuzumi 41 or ChatGPT 51) to summarize user A's pre-meeting information and inputs it as a prompt (instruction) for user A to the generation AI (Tsuzumi 41 or ChatGPT 51). Based on the pre-meeting information and the generation AI's determination of user A's language, the accuracy of the meeting, the degree of response speed, the industry related to the meeting, the meeting situation, and / or whether the meeting is in a specific field or is a precision-oriented meeting, the server device 10 selects a generation AI suitable for the meeting, adjusts the prompt for instructing the generation AI, and then issues a translation instruction to the selected generation AI (step S2). Furthermore, the server device 10 may, in accordance with predetermined rules, determine, for example, the language used by user A, the accuracy of the meeting, the speed of responses, the status of the meeting, the industry related to the meeting, and / or whether the meeting is in a specific field or is a meeting where accuracy is important. Based on the determination, it may select a generation AI suitable for the meeting, adjust the prompts used when instructing the generation AI, and then issue translation instructions to the selected generation AI.

[0049] The server device 10 sets the prompt for user A via the API (Application Programming Interface) and inputs the voice data (which may also be text data) sent from user terminal 20A into the selected generating AI (step S4).

[0050] Based on pre-meeting information provided by User A (the conversation organizer or administrator), the system determines the languages ​​used by Users B through D and inputs this information into the generating AI (Tsuzumi41 or ChatGPT51) for each of Users B through D.

[0051] The generating AI (Tsuzumi41 or ChatGPT51) translates the input voice data of user A (step S3) into the languages ​​used by users B, C, and D according to each user's prompt (step S5).

[0052] Specifically, the generating AI translates User A's Japanese voice data into English for User B. The generating AI translates User A's Japanese voice data into Chinese for User C. The generating AI translates User A's Japanese voice data into French for User D.

[0053] The generating AI returns each translated audio data (which may also be text data) to the server device 10 (step S6). For example, the generating AI converts the translated data into speech using synthesized speech.

[0054] Then, the server device 10 sends and outputs to user B's user terminal 20B an audio file containing user A's Japanese audio data translated into English (step S7-1). The server device 10 sends and outputs to user C's user terminal 20C an audio file containing user A's Japanese audio data translated into Chinese (step S7-2). The server device 10 sends and outputs to user D's user terminal 20D an audio file containing user A's Japanese audio data translated into French (step S7-3). Note that each translation data may also be in text format.

[0055] Then, when voice data (or text data) is input from users B to D, the server device 10 similarly uses a generating AI to translate it into each user's language, and outputs the translated voice data (or text data) from the corresponding user terminal 20.

[0056] The server device 10 can speed up each of the above steps by communicating via the low-latency IOWN network 60, enabling simultaneous interpretation even among three or more users.

[0057] [Process Overview] The processing overview of the processing system 100 will be explained. Figure 5 is a diagram illustrating the processing overview of the processing system 100.

[0058] As shown in Figure 5, the user launches the application on user terminal 20 and selects a meeting room (Figure 5(1)).

[0059] Next, before the meeting starts, users enter their name and a brief description of the meeting, thus inputting pre-meeting information (Figure 5 (2), (2-1)).

[0060] The server device 10 acquires pre-meeting information and uses a generating AI to determine from the pre-meeting information the specified language (e.g., the user's language), the accuracy of the meeting, the degree of response speed, the status of the meeting, the industry related to the meeting, and / or whether the meeting is a specific field meeting or an accuracy-focused meeting.

[0061] To create a prompt, the server device 10 instructs the generating AI to determine the user's language from the pre-meeting information (step S11) and to create a summary of the meeting information (step S12). Then, the server device 10 instructs the generating AI to determine, based on the created summary, the accuracy of the meeting, the speed of responses, the status of the meeting, the industry related to the meeting, and / or whether the meeting is in a specific field or is a meeting where accuracy is important (steps S13 to S17).

[0062] The server device 10 selects either Tsuzumi41 or ChatGPT51 based on the information determined using the generating AI. Furthermore, the server device 10 adds this determined information to the prompt and sets the created prompt to the selected generating AI (Figure 5 (3)). For example, the instruction to Tsuzumi41 is, "For the next conversation, directed to the ID number, translate from Japanese, in a casual office setting."

[0063] Once the prompt settings are complete, the conversation begins (Figure 5 (4)). A simultaneous interpretation service is provided to users A, B, C, and D through translation by Tsuzumi41 or ChatGPT51 according to the set prompts (steps S21, S22) (Figure 5 (5)). In step S22, user X is a collective term for users B, C, and D.

[0064] [Create prompt] Next, we will explain the prompt creation process. Figure 6 shows an example of pre-meeting information. As shown in Figure 6, for example, the user inputs pre-meeting information C11, which is, "Today I will have a meeting with Mr. / Ms. A to discuss whether or not to continue the service in the future. It is an important meeting."

[0065] Figures 7 to 13 illustrate steps S11 to S17 shown in Figure 5. Figure 14 illustrates the processing of the processing system 100.

[0066] When the server device 10 receives pre-meeting information C11 from the user terminal 20, it inputs the pre-meeting information C11 into the generating AI, as shown in Figure 7, and has the AI ​​determine the user's language of use (step S11). The generating AI returns that the user's language of use is Japanese (answer A11).

[0067] As shown in Figure 8, the server device 10 causes the generating AI to create a summary A12 of the pre-meeting information C11 (step S12).

[0068] As shown in Figure 9, the server device 10 instructs the generating AI to determine from the summary A12 whether the accuracy of the meeting is between 60% and 90% (step S13). The generating AI returns that the accuracy of the meeting is 90% (answer A13).

[0069] As shown in Figure 10, the server device 10 causes the generating AI to determine from the summary A12 whether the response speed is fast, standard, or slow (step S14). The generating AI returns that the response speed is standard (response A14).

[0070] As shown in Figure 11, the server device 10 instructs the generating AI to determine from the summary A12 whether the meeting situation is casual, formal, or emergency (step S15). The generating AI returns that the meeting situation is formal (answer A15).

[0071] As shown in Figure 12, the server device 10 instructs the generating AI to determine from the summary A12 whether the industry related to the meeting is medical, semiconductor, service, academic, or factory (step S16). Note that the industry to be determined by the generating AI is not limited to the above. The generating AI returns that the industry related to the meeting is service (answer A16).

[0072] As shown in Figure 13, the server device 10 instructs the generating AI to determine from the summary A12 and answers A13-A16 whether the generating AI suitable for translation is Tsuzumi41, which is specific to a particular field and prioritizes speed, or ChatGPT51, which prioritizes accuracy (step S17). In addition to automatic selection by the generating AI, the user can also select ChatGPT51 or Tsuzumi41 through settings. The generating AI returns that ChatGPT51 is the generating AI suitable for translation (answer A17). Based on this, the server device 10 selects ChatGPT51 as the generating AI suitable for the meeting.

[0073] Then, as shown in Figure 14, the server device 10 sets prompt P2, including answers A11, A13~A16, to ChatGPT51, and the conversation between users begins (Figure 14 (4)). When the users start talking, the server device 10 starts the simultaneous interpretation service using the generation AI (Figure 14 (5)).

[0074] [Other examples of prompt creation] Figures 15 to 17 show other examples of the prompt creation process.

[0075] For example, let's consider the case of pre-meeting information C2 in Figure 15, which states, "I will now speak with a patient. I will explain the nature of the illness and the treatment method." In this case, the generating AI determines that the language used is Japanese, and based on the summary it creates, it returns answer A2, which indicates that the meeting belongs to the medical field, the meeting accuracy is 90%, the response speed is slow, the meeting situation is formal, and Tsuzumi41 is the generating AI best suited for translation. A prompt containing this answer A2 is then set for Tsuzumi41.

[0076] Next, let's explain the case of pre-meeting information C3 shown in Figure 16, which reads, "Today I will have a meeting with Mr. / Ms. A to discuss whether or not to continue with IT services in the future. It is an important meeting." In this case, the generating AI determines that the language used is Japanese, and based on the summary it has created, it returns answer A3, which indicates that the meeting belongs to IT services, the accuracy of the meeting is 85%, the response speed is standard, the meeting situation is formal, and the generating AI suitable for translation is Tsuzumi41. A prompt containing this answer A3 is then set for Tsuzumi41.

[0077] Next, let's explain the case of pre-meeting information C4, which is "Meeting with a client. Very important," as shown in Figure 17. In this case, the generating AI determines that the language used is Japanese, and based on the summary it has created, it returns answer A4, which indicates that the meeting belongs to office services, the accuracy of the meeting is 90%, the response speed is standard, the meeting situation is formal, and the generating AI suitable for translation is Tsuzumi41. A prompt containing this answer A4 is then set for Tsuzumi41.

[0078] Figures 18 and 19 illustrate other examples of the use of the processing system 100. Figure 18 shows an example where voice data or text data from users A, B, and C are translated into English, which is a common language for users A, B, and C, and the translated voice data or text data is output to terminal 20E.

[0079] In this case, the prompt creation unit 12 creates a prompt that instructs the selected generating AI to translate the input voice data or text data into each user's common language. The output control unit 14 transmits the voice data or text data corresponding to the common language (English) to a predetermined terminal (for example, terminal 20E). Alternatively, the output control unit 14 may transmit the voice data or text data corresponding to the common language (English) to the user terminal used by each user. At this time, the server device 10 may automatically recognize the user using the AI ​​and start translation using the generating AI.

[0080] Figure 19 illustrates an example where voice data from user A, who speaks Japanese, is translated into English, the language of user B, and the translated voice data (text data) is output from user terminal 20B in speaker mode. In this case, the prompt creation unit 12 creates a prompt that instructs the selected generating AI to translate the input Japanese voice data or text data from user A (first user) into the language used by user B (second user) (e.g., English). The output control unit 14 transmits the voice data or text data translated into user B's language (English) to user terminal 20B. For example, this can be applied to conversations in places such as police stations, pharmacies, front desks, banks, ticket counters, and shops.

[0081] [Task management support] Next, the task management support processing by the server device 10 will be described. Figure 20 is a diagram illustrating the overview of the task management support processing of the processing system 100.

[0082] After the meeting ends, the server device 10 has the selected generation AI for the meeting automatically create a summary of the meeting content in the language specified by each user A to D (Figure 20 (6-1)).

[0083] The server device 10 requests the selected generative AI for the meeting to create a meeting summary (step S26), create Action Items (A / I) (step S27), and condense the summary into a single slide (step S28). This slide may include images as well as a text summary. The server device 10 then has the generative model AI save the content summary, A / I, and slide to the cloud and an application (Figure 20 (6-2)). By reviewing these, the user can review the meeting, manage tasks, etc. (Figure 20 (6)).

[0084] [Simultaneous interpretation processing] Next, the processing procedure of the processing method according to the first embodiment will be described. Figure 21 is an example of a sequence diagram showing the processing procedure of the processing method according to the first embodiment. Figure 21 shows the processing procedure for simultaneous interpretation, in which voice data or text data from user A is translated into user B's language and output from user B's user terminal 20B, and also translated into user C's language and output from user C's user terminal 20C.

[0085] As shown in Figure 21, for example, when the application is launched on user terminal 20A (step S31), communication between server device 10 and user terminal 20A outputs information about meeting rooms that user A can join to user terminal 20A (step S32). User A operates user terminal 20A and selects a meeting room to join (steps S33, S34), and then inputs a summary of the meeting as pre-meeting information before the meeting (step S35). User terminal 20A transmits the pre-meeting information it has received to server device 10 (step S36).

[0086] The server device 10 uses a generating AI to determine the language used by user A from the pre-meeting information (step S37). The server device 10 uses a generating AI to create a summary of the pre-meeting information (step S38). The server device 10 uses a generating AI to determine, based on the summary, the accuracy of the meeting, the speed of responses, the status of the meeting, the industry related to the meeting, and / or whether the meeting is a specific field meeting or a meeting where accuracy is important (step S39).

[0087] Based on the judgment in step S39, the server device 10 selects either Tsuzumi41 or ChatGPT51, adds the judgment to the prompt, and sets the created prompt to the selected generating AI (steps S40 to S42). The example in Figure 21 shows the case where Tsuzumi41 is selected and the prompt is set. If ChatGPT51 is selected, the prompt is set to ChatGPT51, and ChatGPT51 is instructed to perform the translation process described later (step S48).

[0088] When the server device 10 receives a prompt setting notification (step S43), it sends a message to the user terminal 20A indicating that a conversation can be started (step S44).

[0089] Then, when voice data or text data entered by user A is transmitted from user terminal 20A (steps S45, S46), server device 10 inputs the voice data or text data entered by user A into Tsuzumi 41 (step S47) and has it translated into the languages ​​used by users B and C respectively (step S48).

[0090] The server device 10 transmits the audio data or text output from Tsuzumi 41, translated into the language used by user B (step S49), to user terminal 20B (step S50) and causes it to output (step S53). The server device 10 transmits the audio data or text output from Tsuzumi 41, translated into the language used by user C (step S51), to user terminal 20C (step S52) and causes it to output (step S54).

[0091] The same applies not only to user A, but also when voice data or text data from users A, B, and C are input to the server device 10. The server device 10 should input the voice data or text data from users A, B, and C into Tsuzumi 41, and then output the voice data or text data output from Tsuzumi 41, translated into the languages ​​used by users A, B, and C, respectively, to the user terminals 20A to 20C of each user A to C.

[0092] [Task management support processing] Next, the task management support process will be described. Figure 22 is an example of a sequence diagram showing the processing procedure of another processing method according to the first embodiment.

[0093] As shown in Figure 22, the server device 10 determines whether the meeting by users A to C has ended (step S61). If the meeting has not ended (step S61: No), the process returns to step S61. If the meeting has ended (step S61: Yes), the server device 10 has the generation AI selected for the meeting create a summary of the meeting content in the language used by each user A to D (steps S62, S63).

[0094] The server device 10 stores each summary output from the generating AI (steps S64, S65) and transmits the corresponding language summary to each user terminal 20A~20C (steps S66~S68).

[0095] [Effects of the first embodiment] Figures 23 and 24 are diagrams comparing a conventional translation service with a simultaneous interpretation service provided by the processing system according to the first embodiment.

[0096] As shown in Figure 23, conventional translation services that translate Japanese user voice data and output it as text data take time to output the translation. In contrast, in the simultaneous interpretation service of the first embodiment, each user uses their own terminal, which results in extra communication between the terminal and the server. However, because the high-capacity, high-speed, and low-latency IOWN network 60 is used, this delay can be almost ignored, and as a result, the conversation can be accelerated. Note that the network used is not limited to the IOWN network 60; other networks may also be used.

[0097] Therefore, for example, in a conventional translation service, the first dialogue can be completed by time t1, while in the simultaneous interpretation service of the first embodiment, the second dialogue can be completed.

[0098] Thus, according to the first embodiment, since there is no waiting time for translation, the stress of dialogue can be eliminated. As a result, according to the first embodiment, business efficiency can be improved and costs can be reduced. In addition, in the first embodiment, since the translated audio data is output, the user only needs to listen to the audio data and does not need to read the text, making it highly convenient. Of course, the input and output may be either audio data or text data.

[0099] Furthermore, the processing system 100 communicates via the high-capacity, high-speed, and low-latency IOWN network 60. Therefore, the processing system 100 can provide a simultaneous interpretation service that outputs translated audio or text data in real time, even among three or more users.

[0100] In the processing system 100, for example, even if four people are participating in a meeting, the voice data of one user is translated into the language used by each user and output from each user's user terminal 20. In the simultaneous interpretation service of the first embodiment, simultaneous interpretation can be achieved even among three or more users.

[0101] Furthermore, in the first embodiment, simultaneous interpretation is performed by selecting either Tsuzumi41 or ChatGPT51 based on pre-meeting information about the meeting the user is participating in, thereby enabling appropriate translation tailored to the meeting.

[0102] Furthermore, in the first embodiment, the server device 10 adds content based on pre-meeting information to the selected generating AI and sets prompts that instruct it to translate the input voice data or text data into a specified language in a natural context. Therefore, according to the first embodiment, the most appropriate translation for the situation can be obtained and provided to the user.

[0103] In the first embodiment, Tsuzumi41 and ChatGPT51 were described as example generative AIs, but other generative AIs may be used, and it is not limited to two generative AIs; any of three or more generative AIs may be selected.

[0104] From here, several embodiments that are alternative forms of the first embodiment described above will be explained. Note that the processing described in the first embodiment can be used for the second to fifth embodiments described thereafter.

[0105] <Second Embodiment> First, a second embodiment realized by the processing system 100 according to the first embodiment will be described. The second embodiment is an embodiment in which input voice data or text data (hereinafter sometimes simply referred to as "dialogue data") is translated by the server device 10 included in the processing system 100, and then re-translated based on translation commands from a user or the like.

[0106] (Processing by the processing system) First, the processing of the processing system 100 according to the second embodiment will be explained using Figure 25. Figure 25 is a diagram illustrating the processing of the processing system 100 according to the second embodiment.

[0107] First, the user launches the application on user terminal 20 and selects a meeting room (Figure 25 (1)). Next, the user enters pre-meeting information by entering their name and meeting summary before the meeting starts (Figure 25 (2)).

[0108] The server device 10 selects a predetermined large language model (LLM) based on the content determined using the generating AI. In the second to fifth embodiments described below, the server device 10 will be described using the example of selecting either Tsuzumi41 or ChatGPT51, but the large language models used by the server device 10 are not limited to Tsuzumi41 and ChatGPT51.

[0109] Furthermore, the server device 10 adds the determined information to the prompt and sets the created prompt to the selected generating AI (Figure 25 (3)). Once the prompt setting is complete, the conversation begins (Figure 25 (4)).

[0110] The server device 10 provides simultaneous interpretation services to the user by translating using Tsuzumi41 or ChatGPT51 according to the set prompt (Figure 5(5)). The processes from (1) to (5) in Figure 25 are realized by the processes according to the first embodiment described above.

[0111] The server device 10 displays the results of the translation (translated content) in step (5) of Figure 25 to the users of the service (Figure 25 (6)). As an example, the server device 10 displays the translated content, which has been translated into the language specified by each user, on each user's chat screen in a group chat format (Figure 25 (6-1)).

[0112] For example, if user E, who speaks Japanese, enters "こんにちは" (konnichiwa), the server device 10 will display "こんにちは" (Konnichiwa) on user E's screen (Figure 25 (6-2)). The server device 10 will also display the English translation of "こんにちは" (Konnichiwa) on user F's screen (Figure 25 (6-3)). The server device 10 will also display the Chinese translation of "こんにちは" (Konnichiwa) on user G's screen (Figure 25 (6-4)).

[0113] Here, the server device 10 performs a retranslation of the target translation result based on the retranslation command received from the user (Figure 25 (7)). For example, when the user issues a retranslation command because the conversation is not making sense due to an unintended translation, the server device 10 checks the translated text and performs a retranslation.

[0114] On the other hand, the server device 10 receives a quote reply as a response based on a quote reply command received from the user, and associates the utterance specified by the user with the utterance (Figure 25 (8)). For example, when the user requests a quote reply command from the original user because the user does not understand the meaning of the translated content, the server device 10 processes a quote reply associated with the target utterance.

[0115] The server device 10 then displays the results of the re-translation process described above on the chat screen of the user who issued the re-translation command to the server device 10 (Figure 25 (9)). The server device 10 also displays the results of the quote reply process on each user's chat screen in a group chat format (Figure 25 (9)).

[0116] (Processing system 100) Next, the configuration of the processing system 100 according to the second embodiment will be described. Figure 26 is a diagram showing the configuration of the processing system 100 according to the second embodiment. As shown in Figure 26, the processing system 100 consists of a server device 10, a user terminal 20, an earphone microphone 30, a generation AI server 40, and a generation AI server 50.

[0117] Since the processing system 100 according to the second embodiment has the same configuration as the processing system 100 according to the first embodiment described above, the explanation of the overlapping parts will be omitted.

[0118] (Server device 10) The server device 10 according to the second embodiment includes a pre-meeting information acquisition unit 11 (acquisition unit), a prompt creation unit 12 (creation unit), an input control unit 13, an output control unit 14, a summary creation unit 15, a translation unit 61, and a storage unit 70.

[0119] (Storage unit 70) The storage unit 70 is implemented using semiconductor memory elements such as RAM (Random Access Memory) and flash memory, or storage devices such as hard disks and optical discs. The storage unit 70 stores data and programs used for various processes by the server device 10, as well as various data acquired as the server device 10 operates.

[0120] Specifically, the memory unit 70 temporarily or permanently stores information such as dialogue data, pre-meeting information, created prompts, translation data, and created summary results.

[0121] (Input control unit 13) The input control unit 13 inputs the dialogue data received from the user terminal to a translation generation model that has been set with prompts instructing it to translate the input dialogue data into a specified language in a natural context.

[0122] For example, the input control unit 13 inputs dialogue data in a first language entered by a user of the simultaneous interpretation service via a user terminal into a translation generation model that is set to prompt for translation into a second language specified for each other user of the simultaneous interpretation service.

[0123] (Output control unit 14) The output control unit 14 outputs to the user terminal 20 the first translated data translated in response to the input from the input control unit 13, and the second translated data translated from the first language to the second language by the translation unit 61 described later. An example of the processing of the output control unit 14 will be explained in the section "An Example of Processing According to the Second Embodiment" below.

[0124] (Translation Section 61) The translation unit 61 takes the first translated data, which is translated dialogue data output from the translation generation model in response to input from the input control unit 13, and inputs it into the translation generation model to translate the first translated data into second translated data in a language different from the language of the first translated data. An example of the processing of the translation unit 61 will be described later in the section "An Example of Processing According to the Second Embodiment".

[0125] (An example of processing according to the second embodiment) From here, an example of processing by the processing system 100 according to the second embodiment will be explained using Figures 27 to 29. Figures 27 to 29 are diagrams illustrating an example of processing by the processing system 100 according to the second embodiment. Figure 27 is a diagram illustrating an example of the first example, the "re-translation function". Figure 28 is a diagram illustrating an example of the second example, the "quote reply function". Figure 29 is a diagram illustrating an example of the third example, the "translation alert function".

[0126] (Example 1) First, we will explain the first example, the "re-translation function," using Figure 27. In the example in Figure 27, User E is a user who uses Japanese. User F is a user who uses English. User G is a user who uses Chinese.

[0127] First, user E makes the utterance "Wait until it arrives" to the processing system 100. However, since user E's utterance "Wait until it arrives" is not unique, it can be interpreted in various ways. For example, the utterance "Wait until it arrives" could be translated as meaning "Wait in the car," meaning that the user wants the system to wait in the car. Alternatively, the utterance "Wait until it arrives" could be translated as meaning "Wait until it arrives," meaning that the system wants the system to wait until user E arrives at the specified location.

[0128] Therefore, the server device 10 performs a re-translation process to verify whether the translated result displayed to the user using the simultaneous interpretation service is what the user intended.

[0129] First, the server device 10 inputs the dialogue data entered by the user into a translation generation AI, which then translates it into translation data in the first language. For example, the server device 10 translates user E's utterance, "Wait until the car arrives," into English and Chinese. The server device 10 then displays the translated data in Japanese (Figure 27 (1-1)), English (Figure 27 (1-2)), and Chinese (Figure 27 (1-3)) on each user's terminal.

[0130] Here, as mentioned above, if the utterance "Wait until it arrives" is translated as "Wait until it arrives" instead of "Wait in the car" as intended by user E, a misunderstanding may occur in the conversation between user E, who made the utterance, and users F and G, who reviewed the translation data. Therefore, users can confirm whether their utterance has been translated to what they intended by issuing a re-translation command to the server device 10.

[0131] Specifically, when the server device 10 receives a translation command from a user to translate the first translation data, which has been translated into a first language, into a second language, which is a different language, the server device 10 inputs the first translation data into a translation generative model. Next, the server device 10 displays the second translation data, which has been output from the translation generative model in response to the input of the first translation data and has been translated into the second language, alongside the first translation data.

[0132] For example, when the server device 10 receives a retranslation command from user E, it performs a retranslation of the first translation data corresponding to "Wait until it arrives (Figure 27 (1-1))", which is the translation result into English (Figure 27 (1-2)) and the translation result into Chinese (Figure 27 (1-3)).

[0133] As a specific example, the server device 10 translates the English translation result (Figure 27 (1-2)) back into the user E's specified language (Japanese). Then, as shown in Figure 27 (2-1), the server device 10 displays the re-translated result, "Wait until I arrive," to user E.

[0134] Furthermore, the server device 10 re-translates the translation result into Chinese (Figure 27 (1-3)) into the language specified by user E (Japanese). Then, as shown in Figure 27 (2-2), the server device 10 displays the re-translated result, "Wait for me," to user E.

[0135] Through the process described above, user E can learn that their utterance, "Wait until it arrives," has been translated into English and Chinese as "Wait until it arrives," rather than "Wait in the car."

[0136] Furthermore, the server device 10 can perform retranslation processing when it receives a retranslation command from another user who speaks the same language as the user who made the initial statement. For example, the process described above was explained using the case where the retranslation command was made by user E, who said, "Wait until it arrives." However, the server device 10 can perform retranslation processing based on retranslation commands from other users besides user E.

[0137] Furthermore, the server device 10 can display the results of speech recognition to the user when the user inputs speech via voice. If the user confirms the results of the speech recognition and then inputs a re-translation command, the server device 10 can execute a re-translation process based on that command. As a result, the server device 10 allows the user to confirm whether the translation of their own voiced utterance has been translated into another language as intended by the user.

[0138] (Second example) Next, we will explain the second example, the "quote reply function," using Figure 28. In the example in Figure 28, User E is a user who uses Japanese, and User F is a user who uses English.

[0139] Figure 28 shows a scene in which user E and user F engage in a dialogue based on the translation provided by the processing system 100, and the dialogue concerns "a request for the creation of release decision documents and the schedule for their creation." It should be noted that the translation of the conversation in the second example is assumed to be performed based on the translation process according to the first embodiment described above, and the same applies hereafter.

[0140] For example, the word "Ririhan" mentioned above means "release decision," but it is a unique neologism used by User E and the organization to which User E belongs. Therefore, the server device 10 may translate the Japanese dialogue data "Prepare the materials for Ririhan" as "Prepare the materials for Ririhan." However, "Ririhan" is not a translation of "Ririhan," but simply a word created by transliterating the spoken sound "ririhan" into Roman letters, so User F may not understand the meaning of the word.

[0141] Therefore, the server device 10 can facilitate questions about unfamiliar words among users by receiving a quoted reply as a response associated with an utterance that a given user cannot understand, and displaying it to other users.

[0142] Specifically, the server device 10 receives utterances associated with words specified by the user for the first translation data output from the generative model for translation. The server device 10 then outputs the words specified by the user in association with the received utterances. For example, the server device 10 implements a quote reply function through the following process.

[0143] The server device 10 translates user E's utterance, "Prepare the materials for Ririhan (Figure 28 (1-1))", to "Prepare the materials for Ririhan (Figure 28 (1-2))".

[0144] Server device 10 receives a click of the quote reply button (Figure 28 (2-1)) by user F. Then, server device 10 receives user F's "What is Ririhan? (Figure 28 (2-2))" in correspondence with "Prepare the materials for Ririhan (Figure 28 (1-2))".

[0145] The server device 10 translates user F's utterance, "What is Ririhan?" (Figure 28 (2-2)), to "What is 'Ririhan'?" and displays it on user E's user terminal (Figure 28 (2-3)).

[0146] Server device 10 receives the click of the quote reply button (Figure 28 (3-1)) by user E. Then, in response to the question "What is 'Lilihan'?" (Figure 28 (2-3)), server device 10 receives user E's response "It's about release determination" (Figure 28 (3-2)).

[0147] Server device 10 translates user E's utterance, "I mean release judgment (Figure 28 (3-2))", to "Ririhan means Release judgment," and displays it on user F's user terminal (Figure 28 (3-3)).

[0148] Furthermore, in the second example, the server device 10 can also display the original sentence of a quoted utterance alongside the translated text. For example, the server device 10 can display the original "What is Ririhan?" spoken by user F alongside the utterance "What is Ririhan?" (Figure 28 (2-3)) (Figure 28 (4-1)). On the other hand, when user F inputs "What is Ririhan?" (Figure 28 (2-2)), the server device 10 can display the translated result of that utterance, "What is Ririhan?", alongside it (Figure 28 (4-2)).

[0149] Through the above-described process, the server device 10 facilitates the exchange of questions and answers between users, even when accurate translation of neologisms or technical terms is difficult, thereby enabling smooth communication in different languages.

[0150] (Third example) Next, we will explain the third example, the "translation alert function," using Figure 29. In the example in Figure 29, a conversation is shown between User H and User I, who each use different languages. In Figure 29, it is assumed that User H and User I will engage in dialogue based on the translated content provided by the processing system 100.

[0151] Figure 29 shows a scene where User H and User I are having a conversation about the progress of their work. Specifically, User H asks, "How is the progress on the XX task? (Figure 29 (1-1))" to which User I replies, "It's not finished yet. (Figure 29 (1-2))"

[0152] Here, user H is about to type and send a reprimanding statement to user I, such as, "You're always late with your work. Shouldn't you reconsider your time management?" (Figure 29 (2-1)). However, actions such as "pointing out someone's mistakes in front of others," as described above, can be considered insulting in certain cultures and may be undesirable.

[0153] Therefore, if the processing system 100 detects an undesirable utterance, it issues an alert to the speaker. Specifically, the server device 10 translates the dialogue using a generative model for translation that has been set with prompts instructing the input dialogue data to correct any rude expressions towards the other party. For example, the server device 10 implements the translation alert function through the following process. The "rude expressions" mentioned above refer to expressions that are disrespectful to the other party, such as expressions that include meanings such as rudeness, insult, or disrespect.

[0154] As an example of processing, the server device 10 receives dialogue data from user H, such as "Your work is always running late. Shouldn't you review your time management? (Figure 29 (2-1))." The server device 10 then inputs the received dialogue data into a generative model for translation, which is configured with prompts that instruct the model to correct expressions that may be considered rude to the other party.

[0155] The server device 10 displays the corrections output from the translation generation model, which has the above prompt set, as an alert. Specifically, the server device 10 issues an alert to user H such as, "In certain cultures, it is not customary to point out mistakes in front of others. Are you sure you want to send this?" (Figure 29 (2-2)). In addition, the server device 10 can directly correct the dialogue data entered by the user instead of issuing an alert.

[0156] In addition to the example of expression shown in Figure 29, the server device 10 can also issue alerts and perform corrections regarding content that should not be explicitly mentioned, such as history, political criticism, and religious criticism that the other party would not want to discuss, as well as behaviors that are considered rude in certain cultures, such as asking about height, and expressions that are not unique.

[0157] Furthermore, the server device 10 can issue alerts and perform corrections not only for explicit content but also for implicit content such as sarcasm in a specific context, such as "You're really smart" (expressions that are rude to the person you're talking to). In addition, the server device 10 can issue alerts and perform corrections for dialogue data that has different meanings depending on the context, such as expressions that have multiple meanings, such as "I see" meaning "I understand" and "I understand (when the person is actually surprised and does not understand)," as well as proverbs, idioms, honorifics, etc. (expressions that are not unique).

[0158] In addition, the server device 10 can use a generative model for translation that has been set to prompt the input dialogue data to perform at least one of the following actions: correction of homonyms, correction of grammatical errors, correction of kanji conversion errors, and correction of speech recognition errors.

[0159] For example, the server device 10 can identify and correct errors in words that have the same or similar pronunciation but different meanings, such as "Sato" and "Sugar," if there is a possibility of misspelling, homonym usage errors, or speech recognition errors based on the context. Also, for example, the server device 10 can identify and correct grammatical errors, such as "What is Lilihan?"

[0160] (Procedure for processing according to the second embodiment) From here, the processing procedure by the processing system 100 according to the second embodiment will be explained using Figures 30 to 32. Figures 30 to 32 are sequence diagrams showing the processing procedure according to the second embodiment.

[0161] First, the processing procedure related to the "re-translation function" will be explained using Figure 30. As shown in Figure 30, for example, user terminal 20A receives dialogue data input from user A (S101). Subsequently, user terminal 20A transmits the received dialogue data to the generation AI server 40 (S102).

[0162] The generating AI server 40 performs translation processing using the received dialogue data (S103). Then, the generating AI server 40 transmits the first translated data, which is the translated dialogue data, to user terminals 20A, 20B, and 20C (S104, S105, and S106).

[0163] User terminals 20A, 20B, and 20C, having received the first translation data, output the translation results to each user (users A, B, and C) (S107, S108, and S109).

[0164] Here, the user terminal 20A receives a translation command from user A (S110). Subsequently, the user terminal 20A sends the first translation data to the generation AI server 40 (S111).

[0165] The generation AI server 40 performs translation processing using the received first translation data (S112). Then, the generation AI server 40 transmits the second translation data, which is the first translation data retranslated, to user terminals 20A, 20B, and 20C (S113, S114, and S115).

[0166] The user terminals 20A, 20B, and 20C that received the first translation data output the translation results to each user (users A, B, and C) (S116, S117, and S118).

[0167] Furthermore, the same processing is performed not only for user A, but also when dialogue data from users A, B, and C is input to the server device 10. The server device 10 inputs the dialogue data from users A, B, and C to Tsuzumi 41, and then outputs the dialogue data, which has been translated into the languages ​​used by users A, B, and C respectively, to the user terminals 20A to 20C of each user A to C. The same applies to the following sequence diagrams.

[0168] Next, the processing procedure related to the "quote reply function" will be explained using Figure 31. Note that the steps S201 to S209 shown in Figure 31 are the same as the steps S101 to S109 shown in Figure 30, so the explanation will be omitted.

[0169] User terminal 20A receives a quote reply command from user A (S210). Subsequently, user terminal 20A sends an associated utterance based on the quote reply command to the generation AI server 40 (S211).

[0170] The generating AI server 40 performs translation processing using the received associated utterance (S212). Then, the generating AI server 40 transmits the translated associated utterance to user terminals 20A, 20B, and 20C (S213, S214, and S215).

[0171] User terminals 20A, 20B, and 20C, having received the first translation data, output the translation results to each user (users A, B, and C) (S216, S217, and S218).

[0172] Next, the processing procedure related to the "translation alert function" will be explained using Figure 32. As shown in Figure 32, for example, user terminal 20A receives dialogue data input from user A (S301). Subsequently, user terminal 20A transmits the received dialogue data to the generation AI server 40 (S302).

[0173] The generating AI server 40 performs translation processing using the received dialogue data (S303). Then, the generating AI server 40 transmits the translated dialogue data to the server device 10 (S304).

[0174] If no alert is issued (No in S305), the server device 10 sends the translation data received from the generation AI server 40 to the user terminal 20A (S306). On the other hand, if an alert is issued (Yes in S305), the server device 10 sends a predetermined translation alert to the user terminal 20A (S307).

[0175] The user terminal 20A then outputs the received translation data or translation alert to the user (S308).

[0176] (effect) The effects of the processing system 100 according to the second embodiment will now be explained. The input control unit 13 of the processing system 100 inputs the dialogue data received from the user terminal to a translation generation model that has been set with a prompt that commands the input dialogue data to be translated into a specified language in a natural context. The translation unit 61 of the processing system 100 inputs the first translated data, which is the translated dialogue data output from the translation generation model in response to the input from the input control unit 13, into the translation generation model and translates it into second translated data in a language different from the language of the first translated data. The output control unit 14 of the processing system 100 outputs the second translated data translated by the translation unit 61 to the user terminal. Therefore, the processing system 100 according to the second embodiment has the effect of making it possible to provide the optimal translation according to the situation.

[0177] When the translation unit 61 receives a translation command to translate the first translation data, which has been translated into a first language, into a second language, it inputs the first translation data into a translation generative model. The output control unit 14 then displays the second translation data, which has been output from the translation generative model in response to the input from the translation unit 61 and translated into the second language, alongside the first translation data.

[0178] Therefore, the processing system 100 has the effect of making it easier for users to check how their own speech is being translated when using the simultaneous interpretation service.

[0179] The output control unit 14 receives utterances associated with words specified by the user for the first translation data output from the generation model for translation, and outputs the words and utterances in correspondence.

[0180] Therefore, the processing system 100 enables the translation of utterances associated with the incomprehensible utterances when a user of the simultaneous interpretation service cannot understand the translated utterances. Thus, by achieving optimal translation according to the situation, the processing system 100 effectively facilitates smoother communication between users who speak different languages.

[0181] The translation unit 61 can use a generative model for translation that has been set with prompts instructing it to perform at least one of the following actions on the input dialogue data: correcting homonyms, correcting grammatical errors, correcting kanji conversion errors, correcting speech recognition errors, correcting non-unique expressions, and correcting expressions that are rude to the conversation partner.

[0182] Therefore, the processing system 100 not only translates the content of speech, but also has the effect of suppressing communication problems that occur in communication between users from different cultures.

[0183] Furthermore, the processing system 100 enables a reduction and improvement in the efficiency of computer processing by performing re-translation of already translated data as a series of processes, whereas previously it was necessary to verify the translated content using other translation devices.

[0184] <Third Embodiment> Next, a third embodiment, which is another form of the processing system 100 according to the first embodiment, will be described. The third embodiment is an embodiment in which an agreement document is generated using information such as decisions made at a meeting and tasks to be implemented, extracted from a summary of translation data created by the processing system 100, and notified to the user, and agreement on the decisions made and tasks to be implemented is accepted from the user. In the following sections, "agreement document" refers to a document (screen) used to reach an agreement with participating users regarding the summary, decisions made, tasks to be implemented, etc., of the implementation meeting.

[0185] (Processing by the processing system) First, the processing of the processing system 100 according to the third embodiment will be explained using Figure 33. Figure 33 is a diagram illustrating the processing of the processing system 100 according to the third embodiment.

[0186] First, the processing system 100 translates the dialogue data input by the user. Note that the series of processes related to the translation process shown in (1) to (5) of Figure 33 are the same as the processes in the second embodiment, so their explanation is omitted.

[0187] Next, the server device 10 instructs the selected generation AI for the meeting to automatically create a summary of the meeting content in the language specified by each user (Figure 33 (6)).

[0188] The server device 10 receives a summary as input to a generation model for extraction, which has predetermined prompts set. The server device 10 automatically determines the information to be extracted from the meeting content and the appropriate prompts for that information, and extracts information regarding the meeting's conclusions, decisions, and agreed-upon items such as tasks to be implemented (Figure 33 (7)). The predetermined prompts mentioned above may be, for example, prompts like those shown in Figure 33 (7-1).

[0189] The server device 10 uses the extracted information regarding the agreements to generate an agreement document in a predetermined format (Figure 33 (8)). The server device 10 then notifies the target user of the generated agreement document (Figure 33 (9)).

[0190] As part of the consensus-building process, the server device 10 accepts input from users indicating "agreement" or "disagreement" with the notified agreement document. The server device 10 then stores information identifying the user who agrees or disagrees with the agreement document in its storage unit (Figure 33 (10)).

[0191] (Processing system 100) Next, the configuration of the processing system 100 according to the third embodiment will be described. Figure 34 is a diagram showing the configuration of the processing system 100 according to the third embodiment. As shown in Figure 34, the processing system 100 consists of a server device 10, a user terminal 20, an earphone microphone 30, a generation AI server 40, and a generation AI server 50.

[0192] Since the processing system 100 according to the third embodiment has the same configuration as the processing system 100 according to the second embodiment described above, the explanation of the overlapping parts will be omitted.

[0193] (Server device 10) The server device 10 according to the third embodiment includes a pre-meeting information acquisition unit 11 (acquisition unit), a prompt creation unit 12 (creation unit), an input control unit 13, an output control unit 14, a summary creation unit 15, a consensus building unit 62, and a storage unit 70.

[0194] (Storage unit 70) The memory unit 70 temporarily or permanently stores information such as dialogue data, pre-meeting information, generated prompts, translation data, generated summary results, information on agreed matters, written agreements, and the results of the agreement-making process.

[0195] For example, the memory unit 70 stores the conclusions, decisions, and tasks to be implemented of the meeting, extracted from the generated summary, as information regarding the agreed-upon matters.

[0196] Furthermore, for example, the storage unit 70 stores information identifying the user who entered an agreement to the agreement document, information indicating the agreement, and the agreement document in association with each other. Also, the storage unit 70 stores information identifying the user who entered a disagreement to the agreement document, information indicating the disagreement, and the agreement document in association with each other.

[0197] (Summary creation section 15) The summary generation unit 15 uses the translated dialogue data output from the translation generation model in response to input from the input control unit 13 to summarize in a specific language, and then translates it into other languages. The summary generation unit 15 then summarizes the meeting content in the user-specified language for each translation. An example of the processing of the summary generation unit 15 will be explained in the section "An Example of Processing According to the Third Embodiment" below.

[0198] (Consensus Building Department 62) The consensus building unit 62 accepts input from the user indicating agreement or disagreement based on the summary created by the summary creation unit 15. An example of the processing performed by the consensus building unit 62 will be described later in the section "An Example of Processing According to the Third Embodiment".

[0199] (An example of processing according to the third embodiment) From here, an example of processing by the processing system 100 according to the third embodiment will be explained using Figure 35. Figure 35 is a diagram illustrating an example of processing by the processing system 100 according to the third embodiment. Note that Figure 35 is a diagram illustrating an example of the "consensus building function".

[0200] (An example of a consensus-building process) As described above, the processing system 100 according to the third embodiment generates an agreement document using a summary of a meeting, etc., and notifies the user of the agreement document to facilitate consensus building among the users who participated in the meeting.

[0201] Server device 10 generates an agreement document that includes "date and time," "participants," "summary," and "decisions / Action Items" (Figure 35 (1)). Specifically, server device 10 inputs user-specific translation data output from a translation generation model to an extraction generation model that has been set with a prompt instructing it to extract information about the agreements related to the meeting contained in the user's natural language dialogue data. Next, server device 10 uses the extracted information about the agreements to generate an agreement document that includes a summary of the meeting content, decisions, and at least one of the Action Items. Finally, server device 10 sends the generated agreement document to the user.

[0202] "Date and time" is information indicating the date and time the meeting took place. For example, as shown in (2) of Figure 35, the server device 10 generates an agreement document that includes information such as "2023 / 11 / 02 10:00~11:00" as the date and time.

[0203] "Participating members" is information that identifies the users who participated in the meeting. For example, as shown in (3) of Figure 35, the server device 10 generates an agreement document that includes information such as "User J (you), User K, User L" as participating members.

[0204] A "summary" is information that summarizes the content of the dialogue in a meeting. For example, as shown in (4) of Figure 35, the server device 10 generates an agreement document that includes information such as "User J requested User K and User L to perform sorting work" as a summary.

[0205] "Decisions / Action Items" are information such as decisions made, to-dos, and policies at a meeting. Specifically, as shown in (5) of Figure 35, the server device 10 generates an agreement document that includes information such as "User K will create release decision materials by August 12th" and "Everyone will work to reduce costs in order to achieve the business plan."

[0206] The server device 10 then accepts input from the user indicating "agree" or "disagree" after they have received notification of the generated agreement document.

[0207] Specifically, the server device 10 displays the generated agreement document and an input button for accepting agreement to the agreement document on the user's terminal. Next, when the user presses the button to respond with agreement, the server device 10 accepts the user's agreement to the agreement document. Then, the server device 10 stores information identifying the user who has agreed to the contents of the agreement document in the storage unit 70, associating it with the agreement document.

[0208] On the other hand, the server device 10 displays the generated agreement document and an input button for accepting disagreement with the agreement document on the user's terminal. Next, when the user presses the button to respond with disagreement, the server device 10 accepts the user's disagreement with the agreement document. The server device 10 then stores information identifying the user who disagreed with the contents of the agreement document in the storage unit 70, associating it with the agreement document.

[0209] Furthermore, the consensus-building unit 62 can display on the user's terminal an agreement document relating to the entire meeting, or agreement documents for each predetermined level of detail, along with a button for responding with the agreement.

[0210] For example, server device 10 accepts "agreement / disagreement" for each "decision / action item". Specifically, server device 10 accepts the press of a button (Figure 35 (5-1)) indicating agreement for a decision / action item such as "User K will create release decision documents by 8 / 12". On the other hand, server device 10 accepts the press of a button (Figure 35 (5-2)) indicating disagreement for a decision / action item such as "User K will create release decision documents by 8 / 12".

[0211] Furthermore, in addition to the "Decision Items / Action Items" mentioned above, the server device 10 can perform consensus-building processing for each item, such as date and time, participating members, and summary.

[0212] Furthermore, the server device 10 can display the conversation log screen or conversation list screen related to the agreement document by pressing the buttons displayed at the top of the agreement document screen ((6) and (7) in Figure 35).

[0213] (Procedure for processing according to the third embodiment) From here, the processing procedure of the processing system 100 according to the third embodiment will be explained using Figure 36. Figure 36 is a sequence diagram showing the processing procedure according to the third embodiment.

[0214] As shown in Figure 36, for example, the server device 10 sends a request to the generation AI server 40 to create a summary (S401). Subsequently, the generation AI server 40 creates a summary of the target meeting based on the summary creation request (S402). Then, the generation AI server 40 sends the created summary to the server device 10 (S403).

[0215] The server device 10 stores the received summary in the storage unit 70 (S404). Then, the server device 10 sends a command to the generation AI server 40 to extract information about the agreed-upon items from the summary (S405, S406).

[0216] The generating AI server 40 extracts information regarding the agreed-upon items from the summary based on the extraction command (S407). Then, the generating AI server 40 transmits the extracted information regarding the agreed-upon items to the server device 10 (S408).

[0217] The server device 10 generates an agreement document using the information regarding the received agreement (S409). Then, the server device 10 sends the agreement document to each user terminal 20A, 20B, and 20C (S410, S411, and S412).

[0218] User terminals 20A, 20B, and 20C, which have received the agreement document, output the translation results to each user (users A, B, and C) (S413, S414, and S415).

[0219] The server device 10 receives responses from each user indicating agreement or disagreement with the written agreement (S416, S417, and S418). Then, the server device 10 performs an agreement formation process based on the responses received from each user (S419).

[0220] (effect) The effects of the processing system 100 according to the third embodiment will now be explained. The input control unit 13 of the processing system 100 inputs the dialogue data received from the user terminal to a translation generation model that has been set with prompts instructing it to translate the input dialogue data into a specified language in a natural context. The summary creation unit 15 of the processing system 100 uses the translated dialogue data output from the translation generation model in response to the input from the input control unit 13 to summarize the contents of the meeting in a language specified by each user. The consensus building unit 62 of the processing system 100 receives input from the user indicating agreement or disagreement based on the summary created by the summary building unit 15. Therefore, the processing system 100 according to the third embodiment has the effect of enabling appropriate consensus building on the interpretation content.

[0221] The summary creation unit 15 receives user-specific translation data output from the translation generation model in response to input from the input control unit 13, into an extraction generation model that has been set with a prompt instructing it to extract information about the agreements related to the meeting contained in the dialogue data, which is the user's natural language dialogue data. Next, the summary creation unit 15 uses the information about the agreements to generate an agreement document that includes a summary of the meeting content, decisions made, and at least one of the Action Items. Then, the agreement formation unit 62 sends the generated agreement document to the user.

[0222] Therefore, the processing system 100 has the effect of making it possible to easily reach an agreement by generating an agreement document used for consensus building using dialogue data between users.

[0223] The consensus formation unit 62 displays the generated agreement document and an input button for accepting agreement to the agreement document on the user's terminal. Next, when the user presses the button to respond with agreement, the consensus formation unit 62 accepts the user's agreement to the agreement document. Then, the consensus formation unit 62 stores information identifying the user who has agreed to the contents of the agreement document in the storage unit 70, associating it with the agreement document.

[0224] Therefore, the processing system 100 has the effect of making it possible to easily reach an agreement regarding a meeting based on an automatically generated agreement document.

[0225] The consensus-building unit 62 displays on the user's terminal an agreement document relating to the entire meeting, or agreement documents for each predetermined level of detail, along with a button for responding with the agreement.

[0226] Therefore, the processing system 100 has the effect of making it easy to reach agreements according to the situation, such as reaching an agreement on a part of the agreement document or on the entire agreement document.

[0227] Furthermore, the processing system 100 eliminates the need to input translation data into a separate device from the translation processing device to create a summary, and then use the created summary to reach an agreement individually. Instead, it performs the process of reaching an agreement using the agreement document generated from the summary as a series of processes, thereby reducing and streamlining the computer's processing.

[0228] <Fourth Embodiment> Next, a fourth embodiment, which is another form of the processing system 100 according to the first embodiment, will be described. The fourth embodiment is an embodiment in which the processing system 100 extracts information related to business knowledge from user-to-user dialogue data such as translation data, and presents the information related to the business knowledge to the user in a predetermined format. In the following sections, "information related to business knowledge" includes business knowledge and documented information accumulated in a predetermined organization.

[0229] (Processing by the processing system) First, the processing of the processing system 100 according to the fourth embodiment will be explained using Figure 37. Figure 37 is a diagram illustrating the processing of the processing system 100 according to the fourth embodiment.

[0230] First, the processing system 100 translates the dialogue data input by the user. Note that the series of processes related to the translation process shown in (1) to (5) of Figure 37 are the same as the processes in the second embodiment, so their explanation is omitted.

[0231] Next, the server device 10 extracts information related to business knowledge from the pre-stored translation data (Figure 37 (6)).

[0232] The server device 10 performs a predetermined mapping process on the extracted business knowledge information (Figure 37 (7)). Specifically, the server device 10 inputs user dialogue data and translation data into a generation model with predetermined prompts set, and performs a predetermined mapping process such as mapping questions to their answers. The predetermined prompts mentioned above may be, for example, prompts like those shown in Figure 37 (7-1).

[0233] The server device 10 stores information about the business knowledge for which the mapping process has been performed in its storage unit (Figure 37(8)). The server device 10 then outputs the stored information about the business knowledge to the user in a predetermined format (Figure 37(9)).

[0234] (Processing system 100) Next, the configuration of the processing system 100 according to the fourth embodiment will be described. Figure 38 is a diagram showing the configuration of the processing system 100 according to the fourth embodiment. As shown in Figure 38, the processing system 100 consists of a server device 10, a user terminal 20, an earphone microphone 30, a generation AI server 40, and a generation AI server 50.

[0235] Since the processing system 100 according to the fourth embodiment has the same configuration as the processing system 100 according to the second embodiment described above, the explanation of the overlapping parts will be omitted.

[0236] (Server device 10) The server device 10 according to the third embodiment includes a pre-meeting information acquisition unit 11 (acquisition unit), a prompt creation unit 12 (creation unit), an input control unit 13, an output control unit 14, a summary creation unit 15, a storage unit 63, a display unit 64, an inquiry reception unit 65, and a storage unit 70.

[0237] (Storage unit 70) The memory unit 70 temporarily or permanently stores information such as dialogue data, pre-meeting information, created prompts, translation data, and information related to business knowledge on which a predetermined correspondence process has been performed.

[0238] Here, an example of information relating to business knowledge that has undergone a predetermined correspondence process and is stored in the memory unit 70 will be explained using Figure 39. Figure 39 is a table diagram showing an example of information relating to business knowledge that has undergone a predetermined correspondence process according to the fourth embodiment.

[0239] As shown in Figure 39, the storage unit 70 stores the items "Key," "lang," "question," and "answer" in association with each other, to identify the record in which the knowledge of the business that has undergone a predetermined correspondence process is stored.

[0240] In the above-mentioned "lang" item, information indicating the language is stored. Also, in the "question" item, information regarding the inquiry content included in the dialogue data or translation data is stored. Further, in the "answer" item, information regarding the answer content to the inquiry included in the dialogue data or translation data is stored.

[0241] For example, the storage unit 70 stores by associating key "1", lang "JPN", question "What does 'Ririhan' mean?", and answer "Release judgment". The above-mentioned information means that for the inquiry "What does 'Ririhan' mean?" made in Japanese (JPN), the answer "Release judgment" was given.

[0242] Also, for example, the storage unit 70 stores by associating key "1", lang "ENG", question "What does the Ririhan mean?", and answer "Ririhan means release judgement.". The above-mentioned information means that for the inquiry "What does the Ririhan mean?" made in English (ENG), the answer "Ririhan means release judgement." was given.

[0243] Furthermore, the storage unit 70 can store by associating information regarding the business insights extracted from the translation data translated into different languages between different languages.

[0244] For example, as described above, the questions and answers in lang "JPN" are associated with key "1". On the other hand, the questions and answers in lang "ENG" are associated with key "1". Therefore, the storage unit 70 can store by associating the questions and answers in lang "JPN" and the questions and answers in lang "ENG" identified by the common key "1".

[0245] Note that the storage unit 70 can store information regarding knowledge of various types of operations, not limited to the inquiries and responses regarding the "Riri Judgment" described above. For example, the storage unit 70 can store information such as "explanation of the meaning of terms" and "information regarding overall operations such as the purpose of an operation, operation procedures, operation instructions for tools and systems, and troubleshooting" as information regarding knowledge of operations.

[0246] (Storage unit 63) The storage unit 63 stores, in the storage unit 70, information regarding knowledge of operations, which is extracted based on translation data that is translated dialogue data output from a generation model for translation in response to an input by the input control unit 13.

[0247] Specifically, the storage unit 63 inputs, for each user, translation data to a generation model for extraction in which a prompt for instructing extraction of questions and answers corresponding to the questions from dialogue data in the natural language of the user is set.

[0248] Next, the storage unit 63 associates the questions extracted by the generation model for extraction with the answers corresponding to the questions, and stores the result in the storage unit 70 as information regarding knowledge of operations. For example, the storage unit 63 stores, in the storage unit 70, by associating them respectively, inquiries regarding the meaning of unclear terms and the corresponding answers.

[0249] Specifically, the storage unit 63 inputs, for each user, translation data to a generation model for extraction in which a prompt for instructing extraction of information regarding operation education from dialogue data in the natural language of the user is set. Next, the storage unit 63 stores, in the storage unit 70, the information regarding operation education extracted by the generation model for extraction as information regarding knowledge of operations.

[0250] For example, the storage unit 63 stores, in the storage unit 70, as information regarding operation education, information in which inquiries regarding information regarding overall operations such as the purpose of an operation, operation procedures, operation instructions for tools and systems, and troubleshooting and the corresponding answers are associated, and individual explanatory contents.

[0251] Furthermore, the storage unit 63 can store information in the memory unit 70 that associates the user's knowledge with the user themselves. Specifically, the storage unit 63 inputs user-specific translation data into a generative model for extraction, which is configured with prompts that instruct the model to extract information identifying the user who made the utterance from the user's natural language dialogue data, and information relating to the user's business knowledge. Next, the storage unit 63 associates the user identification information extracted by the generative model with the information relating to the user's business knowledge and stores it in the memory unit 70.

[0252] For example, the storage unit 63 can associate the speaker of the dialogue data that forms the basis of the insights extracted from the generative model for extraction with information about business knowledge, such as "Who is knowledgeable about XX?", and store it in the memory unit 70.

[0253] (Display section 64) The display unit 64 displays the business knowledge information stored in the storage unit 63, associating questions with their corresponding answers. An example of the processing of the display unit 64 will be described later in the section "An Example of Processing According to the Fourth Embodiment".

[0254] (Inquiry reception desk 65) The inquiry reception unit 65 uses the business knowledge information stored in the memory unit 70 to generate a response to the user's inquiry and outputs it to the user. An example of the processing of the inquiry reception unit 65 will be described later in the section "An Example of Processing According to the Fourth Embodiment".

[0255] (An example of processing according to the fourth embodiment) From here, an example of processing by the processing system 100 according to the fourth embodiment will be explained using Figures 40 and 41. Figures 40 and 41 are diagrams illustrating an example of processing by the processing system 100 according to the fourth embodiment. Figure 40 is a diagram illustrating "output of a Q&A list" as the first example. Figure 41 is a diagram illustrating an example of "automatic response to inquiries" as the second example.

[0256] (Example 1) First, we will explain the first example, "Output of Q&A List," using Figure 40. In the example in Figure 40, User E is a user who uses Japanese, and User F is a user who uses English.

[0257] In the first example, the server device 10 displays a Q&A list, which is information relating to business knowledge stored in the storage unit 70 in association with the processing according to the fourth embodiment described above, on each user's user terminal in a language specified for each user.

[0258] Specifically, the server device 10 displays a "Q&A list" that associates questions and answers with each predetermined category. For example, as shown in Figure 40 (1), the server device 10 displays a list (Figure 40 (1-1)) that associates the category "Explanation of the meaning of terms," ​​identified by No. "1," with the question "What does 'Release Judgment' mean?" and the answer "Release Judgment."

[0259] Furthermore, the server device 10 can display a translated version of the Q&A list (Figure 40 (2-1)) for English-speaking users, as shown in Figure 40 (2).

[0260] The "Q&A List" mentioned above is merely an example, and the server device 10 can display information related to business knowledge in other formats. Furthermore, the "Explanation of Terminology" classification mentioned above is merely an example, and the server device 10 can display Q&A in other classifications without limitation.

[0261] Through the above-described processing, users E and F who use different languages can refer to information regarding business knowledge such as Q&A lists in their respective languages.

[0262] (Second example) Next, using FIG. 41, the "automatic response to inquiries", which is the second example, will be described. In an example of FIG. 41, user E is a user who uses Japanese. Also, user F is a user who uses English.

[0263] In the server device 10 in the second example, a prompt for instructing the generation of a response to a predetermined inquiry input by the user is set for a generation model for inquiry response that is based on information regarding business knowledge stored in the storage unit 70, and the inquiry by the user is input. Then, the server device 10 outputs to the user the response to the inquiry by the user generated by the generation model for inquiry response.

[0264] Specifically, as shown in (1) of FIG. 41, the server device 10 receives an inquiry from the user in a chat format or the like. Then, the server device 10 generates a response corresponding to the inquiry from the user and outputs it to the user.

[0265] For example, the server device 10 receives an inquiry of "What is a release judgment?" from user E ((1-1) of FIG. 41). Next, the server device 10 generates a response corresponding to the inquiry from user E based on the business knowledge stored in the storage unit 70. Then, the server device 10 outputs the generated response of "It means'release judgment'" to user E who made the inquiry ((1-2) of FIG. 41).

[0266] Also, regarding the above-described automatic response to inquiries, the server device 10 can perform an automatic response to inquiries in English for users who use English ((2) of FIG. 41).

[0267] For example, server device 10 receives a query from user F, "What does the Ririhan mean?" (Figure 41 (2-1)). Next, server device 10 generates an answer to user F's query based on the business knowledge stored in memory unit 70. Then, server device 10 outputs the generated answer, "Ririhan means release judgment," to user E who made the query (Figure 41 (2-2)).

[0268] Through the process described above, users E and F, who use different languages, can ask questions about unclear points in the language they use and efficiently obtain answers. Furthermore, the server device 10 can output answers to users while considering (correcting, etc.) topics that should not be explicitly mentioned, such as history, political criticism, or religious criticism that the other party may not want to discuss, as well as behaviors that are considered rude in certain cultures, such as asking about height, and non-unique expressions.

[0269] It should be noted that the "automated response to inquiries" described above is merely one example, and the server device 10 can display information related to business knowledge in other formats. Furthermore, the automated response to inquiries related to "reporting decisions" described above is merely one example, and the server device 10 can automatically respond to other types of inquiries without limitation.

[0270] (Processing procedure) From here, the processing procedure by the processing system 100 according to the fourth embodiment will be described. Figures 42 and 43 are sequence diagrams showing the processing procedure according to the fourth embodiment.

[0271] First, the processing procedure for outputting the "Q&A list" will be explained using Figure 42. As shown in Figure 42, for example, user terminal 20A receives dialogue data input from user A (S501). Subsequently, user terminal 20A transmits the received dialogue data to the generation AI server 40 (S502).

[0272] The generating AI server 40 performs translation processing using the received dialogue data (S503). Then, the generating AI server 40 transmits the translated data to the server device 10 (S504).

[0273] Upon receiving the translation data, the server device 10 extracts information related to business knowledge from the translation data (S505). Next, the server device 10 performs a mapping process on the extracted information related to business knowledge (S506). Then, the server device 10 stores the information related to business knowledge that has undergone the mapping process in the storage unit 70 (S507).

[0274] Here, when user terminal 20A receives a command from user A to view the Q&A list, server device 10 transmits information on business knowledge (Q&A list) to user terminal 20A (S508). User terminal 20A then outputs the received information on business knowledge (Q&A list) (S509).

[0275] Next, the processing procedure related to the output of "automated inquiry response" will be explained using Figure 43. Note that the steps from S601 to S607 shown in Figure 43 are the same as the steps from S501 to S507 shown in Figure 42, so the explanation will be omitted.

[0276] Here, user terminal 20A receives the inquiry input from user A (S608). Then, user terminal 20A transmits the inputted inquiry content to server device 10 (S609).

[0277] The server device 10 generates a response corresponding to the inquiry received from the user terminal 20A (S610). Then, the server device 10 sends the response corresponding to the inquiry to the user terminal 20A (S611). Then, the user terminal 20A outputs the response corresponding to the received inquiry (S612).

[0278] (effect) The effects of the processing system 100 according to the fourth embodiment will now be explained. The input control unit 13 of the processing system 100 inputs the dialogue data received from the user terminal to a translation generation model that has been set with prompts that instruct the input dialogue data to be translated into a specified language in a natural context. The storage unit 63 of the processing system 100 stores in the storage unit 70 information related to business knowledge extracted based on the translated dialogue data output from the translation generation model in response to the input from the input control unit 13. Therefore, the processing system 100 according to the fourth embodiment has the effect of enabling appropriate utilization of the interpreted content.

[0279] The storage unit 63 inputs user-specific translation data into a generative model for extraction, which is configured with prompts that instruct the model to extract questions and corresponding answers from the user's natural language dialogue data. The storage unit 63 then associates the questions and corresponding answers extracted by the generative model for extraction and stores them in the memory unit 70 as information related to business knowledge.

[0280] Furthermore, the storage unit 63 inputs user-specific translation data to a generative model for extraction, which is configured with prompts that instruct the model to extract information related to job training from the user's natural language dialogue data. The storage unit 63 then stores the information related to job training extracted by the generative model for extraction in the storage unit 70 as information related to job knowledge.

[0281] Therefore, the processing system 100 has the effect of automatically extracting information for generating a list of Q&A and for automatically responding to inquiries from user interaction data in the course of business and storing it in a database.

[0282] The display unit 64 displays information related to business knowledge stored in the storage unit 63, associating questions with their corresponding answers.

[0283] Therefore, the processing system 100 enables the display of a Q&A list to users using information on business knowledge that is automatically extracted from user interaction data during business operations and stored in a database. As a result, the processing system 100 has the effect of enabling effective utilization of the organization's knowledge accumulated through business operations.

[0284] The inquiry reception unit 65 inputs the user's inquiry to a generation model for inquiry response, which is configured with prompts that instruct the model to generate an answer to a predetermined inquiry entered by the user, based on the business knowledge information stored in the memory unit 70. The inquiry reception unit 65 outputs the answer to the user's inquiry, generated by the generation model for inquiry response, to the user.

[0285] Therefore, the processing system 100 can automatically respond to business inquiries from users by using information on business knowledge that is automatically extracted from user interaction data during business operations and stored in a database. As a result, the processing system 100 has the effect of enabling effective utilization of the organization's knowledge accumulated through business operations.

[0286] Furthermore, the server device 10 according to the fourth embodiment enables the use of information related to business knowledge that can be handled in cases of different languages, cultures, etc. For example, when cultural backgrounds differ, the ease of understanding terminology and explanations may differ. Therefore, conventionally, when generating a Q&A list from business conversation logs, etc., it is necessary to understand the ease or difficulty of understanding the content in each language and adjust the description accordingly.

[0287] Specifically, let's assume that a conversation took place between users speaking Japanese, English, and Chinese, where: 1. One user asked a question about the procedure for a task and answered it; 2. The user speaking English asked a follow-up question about the meaning of term 'a' used in the answer; and 3. The user speaking Chinese requested a further explanation of step 'b' of the procedure for the task.

[0288] When generating a Q&A list from the above-mentioned conversation (dialogue data), since term a is unfamiliar to English-speaking users, the English Q&A list needs to be adjusted to include an explanation of term a. On the other hand, since the explanation of process b is difficult for Chinese-speaking users to understand, the description needs to be adjusted to be expressed in a way that is easy for Chinese-speaking users to understand.

[0289] Here, the processing system 100 stores translated data, which is dialogue data translated into a language specified by a user using a different language, for each language. The processing system 100 can then use the translated data stored for each language to extract information related to business knowledge, generate a list of Q&A, and automatically respond to inquiries. As a result, the processing system 100 enables smoother intercultural communication. Furthermore, in addition to directly creating knowledge in each language specified by the user, the processing system 100 can achieve further optimization of the knowledge for each language based on the re-translation function and quote reply function described in the second embodiment.

[0290] Furthermore, while conventional processing systems 100 were required to store information related to business knowledge, format it into a predetermined format, and output it to the user, separate from the device that performs translation processing, the processing system 100 extracts information related to business knowledge from translation data and outputs it to the user in a predetermined format as a series of processes, thereby enabling a reduction and improvement in the efficiency of computer processing.

[0291] <Fifth Embodiment> Next, a fifth embodiment, which is another form of the processing system 100 according to the first embodiment, will be described. The fifth embodiment is an embodiment in which the processing system 100 generates training data to be used to train a generative model using dialogue data, translation data, etc., stored by the processing system 100, and further trains the generative model using the generated training data and dictionary files of specialized terms etc., which are defined in advance by the user. In the fifth embodiment, the generative model that the processing system 100 is trained on may include a speech recognition model, etc.

[0292] (Processing by the processing system) First, the processing of the processing system 100 according to the fifth embodiment will be explained using Figure 44. Figure 44 is a diagram illustrating the processing of the processing system 100 according to the fifth embodiment.

[0293] First, the processing system 100 translates the dialogue data input by the user. Note that the series of processes related to the translation process shown in (1) to (5) of Figure 44 are the same as the processes in the second embodiment, so their explanation is omitted.

[0294] First, the server device 10 extracts dialogue history data that meets predetermined conditions, such as the user being unable to understand the dialogue, from the pre-stored dialogue data and translation data between users (Figure 44 (6)). Specifically, the server device 10 inputs the dialogue data and translation data between users into a generation model for extraction, which has predetermined prompts set, and performs the process of extracting dialogue history data, which is the history data of when the user did not understand the meaning of the words and made an inquiry or a quoted reply. The predetermined prompts mentioned above may be, for example, prompts like those shown in Figure 44 (6-1).

[0295] The server device 10 generates training data using the extracted dialogue history data (Figure 44 (7)). Then, the server device 10 trains a generative model for translation using the generated training data (Figure 44 (8)).

[0296] (Processing system 100) Next, the configuration of the processing system 100 according to the fifth embodiment will be described. Figure 45 is a diagram showing the configuration of the processing system 100 according to the fifth embodiment. As shown in Figure 45, the processing system 100 consists of a server device 10, a user terminal 20, an earphone microphone 30, a generation AI server 40, and a generation AI server 50.

[0297] Since the processing system 100 according to the fifth embodiment has the same configuration as the processing system 100 according to the second embodiment described above, the explanation of the overlapping parts will be omitted.

[0298] (Server device 10) The server device 10 according to the fifth embodiment includes a pre-meeting information acquisition unit 11 (acquisition unit), a prompt creation unit 12 (creation unit), an input control unit 13, an output control unit 14, a summary creation unit 15, a storage unit 63, a learning data generation unit 66, a learning unit 67, and a storage unit 70.

[0299] (Storage unit 70) The memory unit 70 temporarily or permanently stores information such as dialogue data, pre-meeting information, created prompts, translation data, dialogue history data, and learning data received by the server device 10.

[0300] Here, an example of learning data stored by the memory unit 70 will be explained using Figure 46. Figure 46 is a table diagram showing an example of learning data according to the fifth embodiment. As shown in Figure 46, the memory unit 70 stores items such as "Key" to identify individual records of the learning data, "lang", "unknown word", and "correct answer" in association with each other.

[0301] As mentioned above, the "lang" field stores information indicating the language. The "unknown words" field stores information about parts of the dialogue or translation data where the intended meaning was not conveyed, or parts that were mistranslated, etc., for each language. The "correct answer" field stores information about expressions that accurately convey the intended meaning even between different languages.

[0302] For example, the memory unit 70 stores by associating key "1", lang "JPN", unknown word "Ririhan", and correct answer "Release Judgment". The above-described information means that for "Ririhan" in Japanese (JPN), the content that is accurately transmitted between different languages is "Release Judgment".

[0303] Also, for example, the memory unit 70 stores by associating key "1", lang "ENG", unknown word "Ririhan", and correct answer "release judgement". The above-described information means that for "Ririhan" in English (ENG), the content that is accurately transmitted between different languages is "release judgement".

[0304] Furthermore, the memory unit 70 can store information regarding findings of operations extracted from translation data in different languages by associating them between different languages.

[0305] For example, as described above, the unknown words and correct answers in lang "JPN" are associated with key "1". On the other hand, the unknown words and correct answers in lang "ENG" are associated with key "1". Therefore, the memory unit 70 can store by associating the unknown words and correct answers in lang "JPN" and the unknown words and correct answers in lang "ENG" that are identified by the common key "1".

[0306] Note that the above-described association is merely an example, and the memory unit 70 can store without being limited to the type, content, data format, etc. of the data as long as it is data that can be used as learning data. For example, in addition to learning data related to words that the user cannot understand, the memory unit 70 can store learning data related to coined words, technical terms, etc. used in a specific field or area.

[0307] Also, the memory unit 70 can store predetermined term data such as a dictionary file of technical terms, etc. preset by the user.

[0308] (Storage section 63) The storage unit 63 stores information that will be used as learning data, such as user interaction data and translation data, in the memory unit 70.

[0309] Specifically, the storage unit 63 receives utterances associated with words specified by the user from the translation data output from the translation generation model, and stores the words and utterances in the memory unit 70 in correspondence.

[0310] Furthermore, the storage unit 63 associates the first user's inquiry about the wording contained in the dialogue data with the second user's response to the inquiry and stores it in the memory unit 70. As an example, the storage unit 63 associates the first utterance (inquiry) that is the target of the quote reply with the second utterance (response) corresponding to the first utterance with the dialogue data (translation data) related to the quote reply function of the second embodiment, as shown in Figure 28, and stores it in the memory unit 70.

[0311] Furthermore, for example, the storage unit 63 stores in the memory unit 70 the words that have been retranslated in the dialogue data (translated data) that has undergone retranslation processing according to the second embodiment, as shown in Figure 27.

[0312] Furthermore, the storage unit 63 can store dialogue data specified by the user in the memory unit 70. Specifically, if the user realizes that a translation is incorrect, the storage unit 63 can receive a storage command along with the correct translation and store the incorrect translation and the correct translation in association.

[0313] (Training data generation unit 66) The learning data generation unit 66 generates learning data for training the translation generative model using the translated dialogue data output from the translation generative model in response to the input from the input control unit 13.

[0314] The learning data generation unit 66 generates data as learning data, which is the correspondence between words stored by the storage unit 63 and quoted replies (utterances). The learning data generation unit 66 also generates data as learning data, which is the correspondence between queries stored by the storage unit 63 and the answers to those queries.

[0315] For example, the learning data generation unit 66 generates learning data by associating information such as parts where the intent was not conveyed or parts that were mistranslated, extracted from dialogue data or translation data for each language, with correct data.

[0316] (Learning Section 67) The learning unit 67 trains a generative model for translation using the learning data generated by the learning data generation unit 66. Specifically, the learning unit 67 uses at least one of the learning data generated by the learning data generation unit 66 and predetermined terminology data such as dictionary files of specialized terms set in advance by the user to train the generative model for translation to improve the accuracy of translation, such as reducing mistranslations, replacing specialized terms or coined words in translation, and excluding discriminatory translations.

[0317] Here, we will describe an example of a learning process in which a learning process is performed to improve the translation accuracy of a generative model for translation by substituting technical terms or coined words for translation. The learning unit 67 uses the learning data stored in the memory unit 70 to train the target generative model for translation.

[0318] For example, the word "Ririhan" is a word whose meaning is only accurately understood among specific users. Regarding "Ririhan," data from past user conversations has been linked to the expression "release judgment," and this is stored in the memory unit 70. Furthermore, data has been linked to the expression "release judgment," which is the direct translation of "Ririhan," and this is also stored in the memory unit 70.

[0319] The learning unit 67 uses the above-described learning data to train a generation model for translation so that when "リリ判" is input by a Japanese user, it is translated as "release judgement" instead of "Ririha" for an English user.

[0320] (Procedure of the process) Hereinafter, the procedure of the process by the processing system 100 according to the fifth embodiment will be described with reference to FIG. 47. FIG. 47 is a sequence diagram showing the procedure of the process according to the fifth embodiment.

[0321] As shown in FIG. 47, for example, the user terminal 20A receives an input of dialogue data from the user A (S701). Subsequently, the user terminal 20A transmits the received dialogue data to the generation AI server 40 (S702).

[0322] The generation AI server 40 executes a translation process using the received dialogue data (S703). Then, the generation AI server 40 transmits the translated translation data to the server device 10 (S704).

[0323] The server device 10 that has received the translation data extracts dialogue history data that satisfies a predetermined condition (S705). Next, the server device 10 generates learning data using the extracted dialogue history data (S706). Then, the server device 10 performs learning using the generated learning data (S707, S708).

[0324] Here, when the end condition of learning is not satisfied (No in S709), the server device 10 continues learning. On the other hand, when the end condition of learning is satisfied (Yes in S709), the server device 10 ends learning.

[0325] Note that the above-described end condition of learning is not particularly limited. For example, the server device 10 may determine that the learning condition is satisfied when the number of learning times for the target generation model reaches a predetermined number of learning times or when the model accuracy of the generation model exceeds a predetermined threshold value.

[0326] (effect) The effects of the processing system 100 according to the fifth embodiment will now be explained. The input control unit 13 of the processing system 100 inputs the dialogue data received from the user terminal to a translation generation model that has been set with prompts instructing it to translate the input dialogue data into a specified language in a natural context. The learning data generation unit 66 of the processing system 100 generates learning data for training the translation generation model using the translated dialogue data output from the translation generation model in response to the input from the input control unit 13. Then, the learning unit 67 of the processing system 100 trains the translation generation model using the learning data generated by the learning data generation unit 66. Therefore, the processing system 100 according to the fifth embodiment has the effect of improving the accuracy of the interpretation content.

[0327] The storage unit 63 receives utterances associated with words specified by the user for the first translation data output from the generative model for translation, and stores the words and utterances in the memory unit 70 in correspondence. The learning data generation unit 66 then generates the correspondence data between the words stored in the storage unit 63 and the quoted replies as learning data.

[0328] The storage unit 63 associates the first user's inquiry about the wording contained in the dialogue data with the second user's response to the inquiry and stores it in the memory unit 70. The learning data generation unit 66 then generates learning data by associating the inquiry with the response stored in the storage unit 63.

[0329] Therefore, the processing system 100 has the effect of making it easy to train the generative model by automatically generating training data from business dialogue data that is used for additional training to optimize the generative model for each organization and improve translation accuracy.

[0330] The learning unit 67 uses the learning data generated by the learning data generation unit 66 to learn how to improve the accuracy of translation by the generative model for translation.

[0331] Therefore, the processing system 100 makes it easy to perform a learning process to optimize the generative model for each organization by using learning data that is automatically generated from business dialogue data. As a result, the processing system 100 has the effect of improving the accuracy of the interpreted content by optimizing the generative model for each organization.

[0332] <Variation> The following describes some modifications implemented by the processing system 100 according to this embodiment.

[0333] (Data, etc.) The simultaneous interpretation service, re-translation (processing / function), quoted reply (processing / function), translation alert (function), information regarding agreements, agreement documents, decisions, Action Items, Q&A lists, automated responses to inquiries, learning data, names of functional parts of the processing system 100, steps, processes, names of steps or processes, etc., used in the description of the above embodiment are merely examples and can be changed at will.

[0334] (Regarding sequence diagrams) The steps of the process described using the sequence diagram may be rearranged as long as they do not contradict each other, and some steps may be omitted. Furthermore, conjunctions such as "next," "continue," "furthermore," "at this time," and "on this occasion" in the explanation of the sequence diagram do not limit the order or timing of the execution of the processes in the sequence diagram.

[0335] As described above, the server device 10 can use a generative model for translation that has been configured with prompts to instruct it to perform at least one of the following actions on the input dialogue data: correction of homonyms, correction of grammatical errors, correction of kanji conversion errors, correction of speech recognition errors, correction of non-unique expressions, and correction of expressions that are rude to the dialogue partner. However, it is not limited to this. For example, in addition to the generative model for translation, the server device 10 can use a generative model for correction that performs the correction processes described above. In that case, the server device 10 can perform the correction process using a generative model for correction that has been configured with prompts to perform the above actions: correction of homonyms, correction of grammatical errors, correction of kanji conversion errors, correction of speech recognition errors, correction of non-unique expressions, and correction of expressions that are rude to the dialogue partner.

[0336] [System Configuration of This Embodiment] The server device 10 is a functional concept and does not necessarily have to be physically configured as shown in the diagram. In other words, the specific forms of distribution and integration of the functions of the server device 10 are not limited to those shown in the diagram, and all or part of it can be configured by functionally or physically distributing or integrating it in any unit according to various loads and usage conditions.

[0337] Furthermore, each process performed in the server device 10 may be implemented, in whole or in part, by a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and a program that is analyzed and executed by the CPU and GPU. Alternatively, each process performed in the server device 10 may be implemented as hardware using wired logic.

[0338] Furthermore, among the processes described in this embodiment, all or part of the processes described as being performed automatically can be performed manually. Alternatively, all or part of the processes described as being performed manually can be performed automatically by known methods. In addition, the processing procedures, control procedures, specific names, and information including various data and parameters described above and illustrated can be modified as appropriate unless otherwise specified.

[0339] [program] Figure 48 shows an example of a computer in which a server device 10 is realized when a program is executed. The computer 1000 has, for example, memory 1010 and a CPU 1020. The computer 1000 also has a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.

[0340] Memory 1010 includes ROM 1011 and RAM 1012. ROM 1011 stores, for example, a boot program such as the BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to the hard disk drive 1090. The disk drive interface 1040 is connected to the disk drive 1100. For example, a removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to, for example, the mouse 1110 and the keyboard 1120. The video adapter 1060 is connected to, for example, the display 1130.

[0341] The hard disk drive 1090 stores, for example, an OS (Operating System) 1091, an application program 1092, a program module 1093, and program data 1094. That is, the programs that define each process of the server device 10 are implemented as program modules 1093 containing code executable by the computer 1000. The program modules 1093 are stored, for example, in the hard disk drive 1090. For example, a program module 1093 for performing processes similar to those in the functional configuration of the server device 10 is stored in the hard disk drive 1090. Note that the hard disk drive 1090 may be replaced by an SSD (Solid State Drive).

[0342] Furthermore, the configuration data used in the processing of this embodiment described above is stored as program data 1094 in, for example, memory 1010 or hard disk drive 1090. The CPU 1020 then reads the program module 1093 and program data 1094 stored in memory 1010 or hard disk drive 1090 into RAM 1012 as needed and executes them.

[0343] Furthermore, the program module 1093 and program data 1094 are not limited to being stored in the hard disk drive 1090; for example, they may be stored in a removable storage medium and read by the CPU 1020 via a disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (LAN (Local Area Network), WAN (Wide Area Network), etc.). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via a network interface 1070.

[0344] Although this embodiment applying the invention made by the present inventors has been described above, the present invention is not limited by the description and drawings that constitute part of the disclosure of the present invention in this embodiment. That is, all other embodiments, examples, and operational techniques made by those skilled in the art based on this embodiment are included in the scope of the present invention. [Explanation of Symbols]

[0345] 10 Server devices 11. Pre-meeting Information Acquisition Department 12. Prompt Creation Section 13 Input Control Unit 14 Output Control Unit 15 Summary Preparation Section 20, 20A~20D User Terminals 30A~30D Earphone Microphone 40,50 Generation AI Server 61 Translation Department 62 Consensus Building Department 63 Storage Unit 64 Display section 65 Inquiry Reception Department 66. Training Data Generation Unit 67 Learning Department 70 Memory section

Claims

1. An input control unit receives audio or text data from a user terminal and inputs it to a generation model for translation, which is configured with prompts that instruct the model to translate incoming audio or text data into a specified language in a natural context. A storage unit stores in a memory unit information related to business knowledge, which is extracted based on the translated audio data or text data output from the translation generation model in response to the input from the input control unit, The system includes an inquiry response unit which inputs an inquiry into an inquiry response generation model that has been set to prompt the generation model to generate an answer to a predetermined inquiry input by the user based on information about the business knowledge stored in the memory unit, translates the answer to the user's inquiry generated by the inquiry response generation model according to the language used by the user, and outputs it to the user. The aforementioned storage unit is Information relating to business knowledge, such as inquiries and corresponding answers concerning at least one of the following: business objectives, business procedures, instructions for using tools and systems, and troubleshooting, or information relating to business training that is the content of individual explanations, is stored in the memory unit. A processing apparatus characterized by the following:

2. The aforementioned storage unit is The user-specific translation data is input to a generative model for extraction, which is configured with prompts that instruct the model to extract questions and corresponding answers from the voice data or text data, which are the user's natural language dialogue data. The questions extracted by the aforementioned generation model for extraction and the corresponding answers are associated and stored in the storage unit as information relating to the knowledge of the business. The apparatus according to claim 1.

3. The aforementioned storage unit is The following steps are performed: A generative model for extraction, which is configured with prompts that instruct the model to extract information related to job training from the voice data or text data, which are natural language dialogue data of the user, is input with the translation data for each user. The information regarding the training of the business extracted by the aforementioned generation model is stored in the memory unit as information regarding the knowledge of the business. The apparatus according to claim 1.

4. The aforementioned storage unit is The following steps are performed: A generative model for extraction, which is configured with prompts that instruct the model to extract information identifying the user who made the utterance from the voice data or text data, which are natural language dialogue data of the user, and information relating to the user's business knowledge, is input with the translation data for each user. The information identifying the user extracted by the extraction generation model is associated with the information relating to the user's business knowledge and stored in the storage unit. The apparatus according to claim 1.

5. The system further includes a display unit that displays information relating to the knowledge of the business stored in the storage unit, associating questions with answers to those questions. The apparatus according to any one of claims 1 to 3.

6. Based on pre-meeting information about meetings that pre-configured users will participate in, it is determined whether the meeting is in a specific field or a meeting where accuracy is paramount. Based on the determined results, select either the first generative model, which is a natural language processing model finely tuned for a specific field, or the second generative model, which is a large-scale natural language processing model. The selected generation model further includes a creation unit that creates prompts that instruct the model to translate the input audio data or text data into the language used by each user. The apparatus according to any one of claims 1 to 3.

7. The aforementioned processing apparatus is The user terminal and each server device equipped with the generation model communicate via a communication network related to IOWN (Innovative Optical and Wireless Network). The apparatus according to any one of claims 1 to 3.

8. A method of processing that a computer will execute, An input control step involves inputting audio or text data received from a user terminal to a generative translation model that has been configured with prompts instructing it to translate incoming audio or text data into a specified language in a natural context. A storage step involves storing information relating to business knowledge, extracted based on the translated audio data or text data output from the translation generation model in response to the input from the input control step, in a storage unit. A processing method comprising: an inquiry response step, in which an inquiry is input to an inquiry response generation model, which is set to have a prompt that instructs the generation model to generate an answer to a predetermined inquiry input by the user based on information regarding the business knowledge stored in the memory unit; and an inquiry is input to the user, the answer to the user's inquiry generated by the inquiry response generation model is translated according to the language used by the user and output to the user; The aforementioned storage process is, Information relating to business knowledge, such as inquiries and corresponding answers concerning at least one of the following: business objectives, business procedures, instructions for using tools and systems, and troubleshooting, or information relating to business training that is the content of individual explanations, is stored in the memory unit. Processing method.

9. An input control step involves inputting audio or text data received from a user terminal into a generative translation model that has been configured with prompts instructing it to translate incoming audio or text data into a specified language in a natural context. A storage step involves storing information relating to business knowledge, extracted based on the translated audio data or text data output from the translation generation model in response to the input from the input control step, in a storage unit. A processing program that causes a computer to execute an inquiry response process, which includes: inputting an inquiry from the user into an inquiry response generation model that has been set with prompts instructing it to generate an answer to a predetermined inquiry input by the user based on information about the business knowledge stored in the memory unit; translating the answer to the user's inquiry generated by the inquiry response generation model according to the language used by the user; and outputting it to the user. The storage step is, Information relating to business knowledge, such as inquiries and corresponding answers concerning at least one of the following: business objectives, business procedures, instructions for using tools and systems, and troubleshooting, or information relating to business training that is the content of individual explanations, is stored in the memory unit. Processing program.