Processing apparatus, processing method, and processing program

The processing device employs context-aware generative AI models over a low-latency IOWN network to deliver optimal translations in real-time, addressing the limitations of conventional systems by enhancing communication efficiency and reducing stress in multi-user settings.

JP2025177017AActive Publication Date: 2025-12-05NTT DOCOMO BUSINESS INC
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2024083481
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-22
Publication Date
2025-12-05
Estimated Expiration
2044-05-22

AI Technical Summary

Technical Problem

Conventional interpretation systems fail to provide optimal translations based on situational context, leading to suboptimal language conversion in various scenarios.

Method used

A processing device utilizing a generative AI model, such as Tsuzumi or ChatGPT, to translate voice or text data into specified languages with context-aware prompts, supported by a low-latency IOWN network for real-time simultaneous interpretation among multiple users.

Benefits of technology

Enables context-dependent, real-time translation services that enhance communication efficiency and reduce stress during multi-user interactions by providing accurate and timely language conversions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025177017000001_ABST
    Figure 2025177017000001_ABST
Patent Text Reader

Abstract

To provide an appropriate translation depending on the situation.SOLUTION: A server device 10 includes: an input control unit 13 which inputs voice data or text data received from a user terminal to a translation generation model with prompts set thereon for translating input voice data or text data into a designated language in a natural context; a translation unit 61 which inputs first translation data to the translation generation model to obtain second translation data in a language different from that of the first translation data, the first translation data being translated voice data or text data output from the translation generation model in accordance with the input by the input control unit 13; and an output control unit 14 which outputs the second translation data translated by the translation unit 61 to the user terminal.SELECTED DRAWING: Figure 26
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a processing device, a processing method, and a processing program. [Background technology]

[0002] Conventionally, manual interpretation services have involved requesting interpretation from specialized companies for each type of interpretation service, and having interpreters under contract with those companies provide the interpretation to the client.

[0003] In recent years, various systems using IT technology have been provided as interpretation systems, such as automatic translation systems between various languages ​​and systems that convert voice and text using voice recognition technology and voice synthesis technology. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Publication No. 2022-003441 [Patent Document 2] Japanese Patent Application Publication No. 2019-139663 Summary of the Invention [Problem to be solved by the invention]

[0005] However, conventional interpretation systems only automate the translation process and do not provide optimal translation depending on the situation.

[0006] The present invention has been made in view of the above, and has as its object to provide a processing device, a processing method, and a processing program that can provide an optimal translation depending on the situation. [Means for solving the problem]

[0007] In order to solve the above-mentioned problems and achieve the object, the processing device of the present invention is characterized by having an input control unit that inputs voice data or text data received from a user terminal to a generation model for translation that is set with a prompt that instructs the input voice data or text data to be translated into a specified language in a natural context, a translation unit that inputs first translation data, which is the translated voice data or text data output from the generation model for translation in response to input by the input control unit, into the generation model for translation and translates it into second translation data in a language different from the language of the first translation data, and an output control unit that outputs the second translation data translated by the translation unit to the user terminal. [Effects of the Invention]

[0008] According to the present invention, it is possible to provide an optimum translation depending on the situation. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of a processing system according to the first embodiment. [Figure 2] FIG. 2 is a diagram showing an overview of the IOWN technology. [Figure 3] FIG. 3 is a diagram illustrating an example of use of the processing system according to the first embodiment. [Figure 4] FIG. 4 is a diagram illustrating the flow of processing in the processing system. [Figure 5] FIG. 5 is a diagram illustrating an outline of the processing of the processing system. [Figure 6] FIG. 6 is a diagram illustrating an example of pre-meeting information. [Figure 7] FIG. 7 is a diagram illustrating step S11 shown in FIG. [Figure 8] FIG. 8 is a diagram illustrating step S12 shown in FIG. [Figure 9] FIG. 9 is a diagram illustrating step S13 shown in FIG. [Figure 10]FIG. 10 is a diagram illustrating step S14 shown in FIG. [Figure 11] FIG. 11 is a diagram illustrating step S15 shown in FIG. [Figure 12] FIG. 12 is a diagram illustrating step S16 shown in FIG. [Figure 13] FIG. 13 is a diagram illustrating step S17 shown in FIG. [Figure 14] FIG. 14 is a diagram illustrating the processing of the processing system. [Figure 15] FIG. 15 is a diagram illustrating another example of the prompt generation process. [Figure 16] FIG. 16 is a diagram illustrating another example of the prompt generation process. [Figure 17] FIG. 17 is a diagram illustrating another example of the prompt generation process. [Figure 18] FIG. 18 is a diagram illustrating another example of use of the processing system. [Figure 19] FIG. 19 is a diagram illustrating another example of use of the processing system. [Figure 20] FIG. 20 is a diagram for explaining an outline of the task management support process of the processing system. [Figure 21] FIG. 21 is an example of a sequence diagram showing a processing procedure of the processing method according to the first embodiment. [Figure 22] FIG. 22 is an example of a sequence diagram illustrating a processing procedure of another processing method according to the first embodiment. [Figure 23] FIG. 23 is a diagram comparing a conventional translation service with the simultaneous interpretation service provided by the processing system according to the first embodiment. [Figure 24] FIG. 24 is a diagram comparing a conventional translation service with the simultaneous interpretation service provided by the processing system according to the first embodiment. [Figure 25] FIG. 25 is a diagram illustrating the processing of the processing system according to the second embodiment. [Figure 26] FIG. 26 is a diagram showing the configuration of a processing system according to the second embodiment. [Figure 27]FIG. 27 is a diagram illustrating an example of processing by the processing system according to the second embodiment. [Figure 28] FIG. 28 is a diagram illustrating an example of processing by the processing system according to the second embodiment. [Figure 29] FIG. 29 is a diagram illustrating an example of processing by the processing system according to the second embodiment. [Figure 30] FIG. 30 is a sequence diagram showing a processing procedure according to the second embodiment. [Figure 31] FIG. 31 is a sequence diagram showing a processing procedure according to the second embodiment. [Figure 32] FIG. 32 is a sequence diagram showing a processing procedure according to the second embodiment. [Figure 33] FIG. 33 is a diagram illustrating the processing of the processing system according to the third embodiment. [Figure 34] FIG. 34 is a diagram showing the configuration of a processing system according to the third embodiment. [Figure 35] FIG. 35 is a diagram illustrating an example of processing by the processing system according to the third embodiment. [Figure 36] FIG. 36 is a sequence diagram showing a processing procedure according to the third embodiment. [Figure 37] FIG. 37 is a diagram illustrating the processing of the processing system according to the fourth embodiment. [Figure 38] FIG. 38 is a diagram showing the configuration of a processing system according to the fourth embodiment. [Figure 39] FIG. 39 is a table diagram showing an example of information related to business knowledge for which a predetermined association process has been performed according to the fourth embodiment. [Figure 40] FIG. 40 is a diagram illustrating an example of processing by the processing system according to the fourth embodiment. [Figure 41] FIG. 41 is a diagram illustrating an example of processing by the processing system according to the fourth embodiment. [Figure 42] FIG. 42 is a sequence diagram showing a processing procedure according to the fourth embodiment. [Figure 43]FIG. 43 is a sequence diagram showing a processing procedure according to the fourth embodiment. [Figure 44] FIG. 44 is a diagram illustrating the processing of the processing system according to the fifth embodiment. [Figure 45] FIG. 45 is a diagram showing the configuration of a processing system according to the fifth embodiment. [Figure 46] FIG. 46 is a table illustrating an example of learning data according to the fifth embodiment. [Figure 47] FIG. 47 is a sequence diagram showing a processing procedure according to the fifth embodiment. [Figure 48] FIG. 48 is a diagram illustrating an example of a computer that implements a server device by executing a program. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited to this embodiment. In addition, in the description of the drawings, the same parts are designated by the same reference numerals.

[0011] First Embodiment [Processing System] The configuration of a processing system according to the first embodiment will be described. The processing system according to the first embodiment provides a simultaneous interpretation service in a conference between multiple users, using a generative AI (artificial intelligence) (generative model) to translate the voice data or text data of other users into a language designated by each user, and outputting the translated voice data or text data to a specified terminal.

[0012] FIG. 1 is a diagram showing an example of the configuration of a processing system according to the first embodiment. As shown in FIG. 1, the processing system 100 according to the first embodiment includes user terminals 20A to 20D used by users A to D, who are users of the simultaneous interpretation service, earphone microphones 30A to 30D worn by the users A to D and which communicate with the user terminals 20A to 20D, and a cloud server device 10. The server device 10 communicates with a generation AI server 40 and a generation AI server 50. The user terminals 20A to 20D are collectively referred to as user terminals 20. The number of user terminals 20 may be two or more and is not limited to four.

[0013] The generative AI server 40 is equipped with Tsuzumi (registered trademark) 41 (first generative model), a generative AI. Tsuzumi 41 is a natural language processing model fine-tuned to specific fields, such as medicine, semiconductors, IT (Information Technology), academia, factories (plants), law, and office services. Tsuzumi 41 was built with an emphasis on low power consumption and has a faster processing speed than ChatGPT 51 (described below).

[0014] The generation AI server 50 is equipped with ChatGPT (registered trademark) 51 (second generation model), which is a generation AI. ChatGPT 51 is a large-scale natural language processing model that is slower than Tsuzumi 41 but more accurate. Tsuzumi 41 and ChatGPT 51 perform natural language processing on input voice data or text data according to set prompts, and generate and output voice data or text data. Note that the above generation AI is just one example, and additional servers equipped with multiple other generation AIs may be provided.

[0015] The user terminals 20A to 20D are terminal devices that can input and output voice data and text data, and communicate with the server device 10. The user terminals 20A to 20D may input and output voice data via earphone microphones 30A to 30D worn by each of the users A to D.

[0016] The user terminals 20A to 20D are, for example, personal computers (PCs), notebook PCs, tablet terminals, smartphones, etc. The user terminals 20A to 20D start a simultaneous interpretation application to communicate with the server device 10 regarding the simultaneous interpretation process. At this time, the user terminals 20A to 20D accept input of pre-conference information regarding the conference in which the user will participate, through an operation by the user before the conference, and transmit the pre-conference information to the server device 10.

[0017] The server device 10 is realized by loading a predetermined program into a computer or the like including, for example, a ROM (Read Only Memory), a RAM (Random Access Memory), a CPU (Central Processing Unit), etc., and having the CPU execute the predetermined program. The server device 10 also has a communication interface for transmitting and receiving various information to and from other devices (for example, user terminals 20A to 20D, generation AI servers 40, 50) connected via a network or the like.

[0018] In a conference between multiple users, the server device 10 uses Tsuzumi41 or ChatGPT51 to provide a simultaneous interpretation service in which the speech or text of other users is translated into a language specified by each user, and the translated speech or text data is output to a specified terminal.

[0019] Specifically, the server device 10 selects either Tsuzumi41 or ChatGPT51 based on pre-meeting information about the meeting the user will be attending. Then, the server device 10 adds content based on the pre-meeting information to the selected generation AI and sets a prompt to instruct the AI ​​to translate input voice data or text data into a specified language with a natural context.

[0020] This allows the server device 10 to obtain the most suitable translation for the situation from the generating AI and provide it to the user.

[0021] Furthermore, the server device 10 communicates with the user terminals 20A to 20D and the generation AI servers 40 and 50 via a low-latency communication network related to the IOWN (Innovative Optical and Wireless Network) (hereinafter referred to as an IOWN network 60).

[0022] [IOWN Technology Overview] Here, we will explain the IOWN technology. Figure 2 is a diagram showing an overview of the IOWN technology. As shown in Figure 2, the IOWN technology consists of three main technology areas: "All-Photonics Network (APN)," "Digital Twin Computing (DTC)," and "Cognitive Foundation (CF) (registered trademark)."

[0023] [All Photonics Network] The APN related to IOWN technology is a technology that enables the construction of high-speed networks by processing all network transfer functions in the optical domain. Specifically, the APN related to IOWN technology is a technology that realizes low-power, high-quality, large-capacity, and low-latency communications based on optical-based (photonics-based) technologies such as "photonics-electronic convergence technology," "large-capacity optical transmission system and device technology," "optical Ising machine," and "optical lattice clock network."

[0024] [Digital Twin Computing] DTC, which is related to IOWN technology, is a technology that maps individual objects in the real world onto a virtual space using the vast amount of data collected by devices connected to the APN mentioned above.

[0025] Conventional digital twin frameworks are used by mapping individual objects, such as automobiles and robots, into a virtual space, performing analysis and predictions on them, and then mapping the results of the analysis and predictions back onto the real world.

[0026] On the other hand, DTC related to IOWN technology expands on the conventional concept of digital twins, freely combining digital twins of various industries, objects, and people to perform calculations, thereby reproducing with high accuracy the combination of multiple objects, such as people and automobiles in a city. Furthermore, DTC related to IOWN technology enables not only the expression of a person's external appearance, but also the digital expression of their internal state, such as consciousness and thoughts, by combining technologies that enable "speech recognition," "speech synthesis," "understanding of emotions and intentions," etc. to collect information and build a digital twin environment.

[0027] In this way, DTC related to IOWN technology is a technology that enables the creation of digital twins that do not exist in the real world by combining multiple entities that are single in the real world and replicating them as digital twins in a virtual space, or by exchanging or merging some of the components between multiple digital twins.

[0028] [Cognitive Foundation] CF related to IOWN technology is a technology that centrally performs the deployment, configuration, linkage, management, and operation of ICT (Information and Communication Technology) resources at different layers, from the cloud to edge computers, network services, user equipment, etc. Specifically, CF related to IOWN technology treats various targets as a group of virtualized ICT resources, and optimally integrates multiple resources at different layers using multi-orchestration functions as a hub.

[0029] Furthermore, as shown in Figure 2, IOWN technology provides high-value-added services by linking the above-mentioned APN, DTC, and network services provided by operators.

[0030] For example, as shown in Figure 2 (1), IOWN technology provides a technology for transmitting information collected via APN to other terminal devices at high speed and with low latency. Also, as shown in Figure 2 (2), IOWN technology provides a technology for collecting large amounts of information from terminal devices and outputting information such as analysis results from the service provided by the operator at high speed and with low latency in services such as information analysis. Also, as shown in Figure 2 (3), IOWN technology provides a technology for transmitting large amounts of information at high speed and with low latency, using information obtained from surveillance cameras, automobile sensors, etc. to build a digital twin environment, make future predictions, and output the prediction results to the user.

[0031] It is believed that the high-capacity, high-speed, low-latency information transmission infrastructure based on the IOWN technology described above will advance the construction of digital twin environments and the linkage between different digital twin environments.

[0032] The processing system 100 communicates via a large-capacity, high-speed, low-latency information transmission infrastructure based on the above-mentioned IOWN technology. For example, when linking with the processing system 100, a low-latency simultaneous interpretation service is realized using an APN. Specifically, the processing system 100 can provide a simultaneous interpretation service that outputs translated speech data or text data in real time, even when speech data or text data to be translated is input, even when there are three or more users.

[0033] [Server device] 1, the server device 10 will be described. The server device 10 includes a pre-conference information acquisition unit 11 (acquisition unit), a prompt creation unit 12 (creation unit), an input control unit 13, an output control unit 14, and a summary creation unit 15.

[0034] The pre-conference information acquisition unit 11 acquires pre-conference information related to the conference in which the user will participate through communication with the user terminals 20A to 20D.

[0035] The prompt creation unit 12 selects one of a plurality of generation AIs based on the pre-meeting information. The prompt creation unit 12 selects one of Tsuzumi41 and ChatGPT51 based on the pre-meeting information. The prompt creation unit 12 creates a prompt (command) that instructs the selected generation AI to translate input voice data or text data into a specified language with a natural context, and adds content based on the pre-meeting information.

[0036] The prompt creation unit 12 uses a generation AI (Tsuzumi41 or ChatGPT51) to determine, based on pre-meeting information, for example, a specified language (e.g., the user's language), the accuracy of the meeting, the response speed, the situation of the meeting, the industry related to the meeting, and / or whether the meeting is a specific field or an accuracy-oriented meeting. The prompt creation unit 12 may change the content of the determination depending on the industry, field, members, and situation, without being limited to the above. Furthermore, the server device 10 may determine, according to predetermined rules, the language used by user A, the accuracy of the meeting, the response speed, the situation of the meeting, the industry related to the meeting, and / or whether the meeting is a specific field or an accuracy-oriented meeting.

[0037] The prompt creation unit 12 selects either Tsuzumi41 or ChatGPT51 based on the content determined using the generation AI. For example, if the conference is in a specific field (e.g., medicine), the response speed is set to a relatively fast level, and speed is emphasized, the prompt creation unit 12 selects Tsuzumi41. On the other hand, if accuracy is emphasized, the prompt creation unit 12 selects ChatGPT51.

[0038] The prompt creation unit 12 then adds the content determined using the generation AI to the prompt. Based on the determination, the prompt creation unit 12 selects a generation AI suitable for the conference, adjusts the prompt to be used to instruct the generation AI, and then issues a translation instruction to the selected generation AI. The prompt creation unit 12 outputs the created prompt to the input control unit 13.

[0039] The input control unit 13 gives the prompt created by the prompt creating unit 12 to the generated AI selected by the prompt creating unit 12, and inputs the voice data or text data input from the user terminals 20A to 20D.

[0040] The output control unit 14 outputs the voice data or text data output from the generated AI selected by the prompt creating unit 12 to a predetermined terminal (user terminals 20A to 20D).

[0041] When the conference ends, the summary creation unit 15 causes the generation AI selected by the prompt creation unit 12 to create a summary of the conference contents in the language specified by each of the users A to D. The summary creation unit 15 transmits each summary created by this generation AI to the user terminals 20A to 20D of each user. The transmitted summaries are used by the users A to D for reviewing the conference, task management, etc. By using the generation AI, in addition to creating an automatic summary of the conference contents, the generation AI can also generate automatic Q&A and store it in a database, and it is also possible to provide ChatBot functionality.

[0042] [Application example] 3 is a diagram illustrating an example of use of the processing system 100 according to the first embodiment. In FIG. 3, a conference is shown as an example between user A who speaks Japanese, user B who speaks English, user C who speaks Chinese, and user D who speaks French.

[0043] In this case, the prompt creation unit 12 creates a prompt that instructs the selected generation AI to translate the input voice data or text data into the language used by each user. Specifically, the voice data uttered by other users B, C, and D and to be output to user A is translated into Japanese. The voice data uttered by other users A, C, and D and to be output to user B is translated into English. The voice data uttered by other users A, B, and D and to be output to user C is translated into Chinese. The voice data uttered by other users A, B, and C and to be output to user D is translated into French.

[0044] As a result, the generation AI translates the speech of each of users A, B, C, and D into the language used by each of users A, B, C, and D. The output control unit 14 transmits the speech data or text data translated into the language used by each of users A to D to each of user terminals 20A to 20D used by each user.

[0045] As a result, the translated voice data is output via the user terminals 20A to 20D of the respective users, and may also be output from the earphone microphones 30A to 30D worn by the respective users A, B, C, and D. For example, the speech of users B, C, and D is translated into Japanese and output from the earphone microphone 30A of user A.

[0046] [Processing flow] The process flow of the processing system 100 in the case of Fig. 3 will be described. Fig. 4 is a diagram illustrating the process flow of the processing system 100. Here, the case where speech data of user A is translated for other users B to D will be described as an example.

[0047] User A starts an application (represented as an app in the figure) on user terminal 20A and inputs a summary of the meeting as pre-meeting information before the meeting (step S1). From the pre-meeting information, the generation AI (Tsuzumi41 or ChatGPT51) determines the specified language (language used by user A), the industry related to the meeting, the situation of the meeting, the accuracy of the meeting, the response time (degree of response speed), and / or whether the meeting is a specific field or an accuracy-oriented meeting.

[0048] The server device 10 uses the generation AI (Tsuzumi41 or ChatGPT51) to summarize user A's pre-meeting information and inputs the summarized information into a prompt (instruction) for user A to the generation AI (Tsuzumi41 or ChatGPT51). The server device 10 selects a generation AI appropriate for the meeting based on the language used by user A, the accuracy of the meeting, the degree of response speed, the industry related to the meeting, the situation of the meeting, and / or whether the meeting is in a specific field or an accuracy-oriented meeting, which have been determined by the generation AI using the pre-meeting information, and adjusts the prompt when instructing the generation AI, and then issues a translation instruction to the selected generation AI (step S2). In addition, the server device 10 may determine, according to predetermined rules, for example, the language used by user A, the accuracy of the meeting, the degree of response speed, the situation of the meeting, the industry related to the meeting, and / or whether the meeting is in a specific field or a meeting that emphasizes accuracy, and based on the determination, select a generation AI that is suitable for the meeting, adjust the prompts when instructing the generation AI, and then issue translation instructions to the selected generation AI.

[0049] The server device 10 sets a prompt for user A via an API (Application Programming Interface) and inputs the voice data (or text data) sent from the user terminal 20A to the selected generation AI (step S4).

[0050] From the pre-conference information provided by User A (the conversation organizer or administrator), the languages ​​used by Users B to D are determined and input into the prompts for each of Users B to D to the generation AI (Tsuzumi41 or ChatGPT51).

[0051] The generation AI (Tsuzumi41 or ChatGPT51) translates the input voice data of user A (step S3) into the languages ​​used by users B, C, and D according to each user's prompts (step S5).

[0052] Specifically, the generation AI translates user A's Japanese speech data into English for user B. The generation AI translates user A's Japanese speech data into Chinese for user C. The generation AI translates user A's Japanese speech data into French for user D.

[0053] The generation AI returns each piece of translated voice data (which may be text data) to the server device 10 (step S6). For example, the generation AI vocalizes the translated data using a synthesized voice.

[0054] Then, the server device 10 transmits voice data obtained by translating the Japanese voice data of user A into English to the user terminal 20B of user B, and causes it to be output (step S7-1). The server device 10 transmits voice data obtained by translating the Japanese voice data of user A into Chinese to the user terminal 20C of user C, and causes it to be output (step S7-2). The server device 10 transmits voice data obtained by translating the Japanese voice data of user A into French to the user terminal 20D of user D, and causes it to be output (step S7-3). Each piece of translation data may be text data.

[0055] Then, when voice data (which may be text data) is input from users B to D, the server device 10 similarly uses the generation AI to translate the voice data into the language used by each user, and outputs the translated voice data (which may be text data) from the corresponding user terminal 20.

[0056] The server device 10 can speed up each of the above steps by performing communication via the low-delay IOWN network 60, and can realize simultaneous interpretation even among three or more users.

[0057] [Processing Overview] An outline of the processing performed by the processing system 100 will now be described. FIG.

[0058] As shown in FIG. 5, the user starts an application on the user terminal 20 and selects a conference room ((1) in FIG. 5).

[0059] Next, before the start of the meeting, the user inputs pre-meeting information by inputting his / her name and a meeting summary ((2), (2-1) in FIG. 5).

[0060] The server device 10 acquires pre-meeting information and, using a generation AI, determines from the pre-meeting information the specified language (e.g., the language used by the user), the accuracy of the meeting, the degree of response speed, the situation of the meeting, the industry related to the meeting, and / or whether the meeting is a specific field or an accuracy-oriented meeting.

[0061] To create the prompt, the server device 10 causes the generation AI to determine the language used by the user from the pre-meeting information (step S11) and create a summary of the meeting information (step S12).Then, the server device 10 causes the generation AI to determine the accuracy of the meeting, the degree of response speed, the situation of the meeting, the industry related to the meeting, and / or whether the meeting is a specific field or an accuracy-oriented meeting based on the created summary (steps S13 to S17).

[0062] The server device 10 selects either Tsuzumi41 or ChatGPT51 based on the content determined using the generation AI. Furthermore, the server device 10 adds this determined content to the prompt and sets the created prompt to the selected generation AI ((3) in FIG. 5). For example, the instruction to Tsuzumi41 is, "Upcoming conversation, ID number, interpret Japanese, casual office setting."

[0063] Once the prompt settings are complete, the conversation begins ((4) in Figure 5). Tsuzumi 41 or ChatGPT 51 translates according to the set prompts (steps S21 and S22), and simultaneous interpretation services are provided to users A, B, C, and D ((5) in Figure 5). Note that in step S22, user X is a collective term for users B, C, and D.

[0064] [Create Prompt] Next, the prompt creation process will be described. Fig. 6 is a diagram showing an example of pre-meeting information. As shown in Fig. 6, for example, the user inputs pre-meeting information C11 such as "Today I will be meeting with Mr. A to discuss whether to continue the service in the future. This is an important meeting."

[0065] 7 to 13 are diagrams for explaining steps S11 to S17 shown in Fig. 5. Fig. 14 is a diagram for explaining the processing of the processing system 100.

[0066] When the server device 10 receives the pre-meeting information C11 from the user terminal 20, the server device 10 inputs the pre-meeting information C11 to the generation AI and determines the user's preferred language (step S11), as shown in Fig. 7. The generation AI returns that the user's preferred language is Japanese (answer A11).

[0067] As shown in FIG. 8, the server device 10 causes the generation AI to create a summary A12 of the pre-meeting information C11 (step S12).

[0068] 9, the server device 10 causes the generation AI to determine whether the accuracy of the meeting is between 60% and 90% from the summary A12 (step S13). The generation AI returns that the accuracy of the meeting is 90% (answer A13).

[0069] 10, the server device 10 causes the generation AI to determine whether the response speed is fast, normal, or slow from the summary A12 (step S14). The generation AI returns that the response speed is normal (answer A14).

[0070] 11, the server device 10 causes the generation AI to determine whether the situation of the meeting is casual, formal, or emergency from the summary A12 (step S15). The generation AI returns that the situation of the meeting is formal (answer A15).

[0071] 12, the server device 10 causes the generation AI to determine from the summary A12 whether the industry related to the conference is medical, semiconductor, service, academic, or factory (step S16). Note that the industry determined by the generation AI is not limited to the above. The generation AI returns that the industry related to the conference is service (answer A16).

[0072] As shown in FIG. 13, the server device 10 causes the generation AI to determine, from the summary A12 and the answers A13 to A16, whether the generation AI suitable for translation is Tsuzumi41, which is in a specific field and prioritizes speed, or ChatGPT51, which prioritizes accuracy (step S17). In addition to automatic selection by the generation AI, the user can also select ChatGPT51 or Tsuzumi41 by setting it themselves. The generation AI returns that the generation AI suitable for translation is ChatGPT51 (answer A17). In response to this, the server device 10 selects ChatGPT51 as the generation AI suitable for the conference.

[0073] 14, the server device 10 sets the prompt P2 including the answers A11, A13 to A16 in the ChatGPT 51, and the conversation between the users starts ((4) in FIG. 14). Once the users start talking, the server device 10 starts a simultaneous interpretation service using the generation AI ((5) in FIG. 14).

[0074] [Other examples of creating prompts] 15 to 17 are diagrams showing another example of the prompt creation process.

[0075] For example, consider the case of pre-meeting information C2 in Figure 15, which states, "I will now talk to the patient. I will explain the details of the illness and the treatment." In this case, the generation AI determines that the language used is Japanese, and based on the created summary, returns answer A2, which indicates that the meeting belongs to the medical field, the accuracy of the meeting is 90%, the response speed is slow, and the situation of the meeting is formal, and the generation AI suitable for translation is Tsuzumi41. A prompt including this answer A2 is set to Tsuzumi41.

[0076] Next, let us consider the case of pre-meeting information C3 in Figure 16, which states, "Today I will be meeting with Mr. A to discuss whether to continue providing IT services in the future. This is an important meeting." In this case, the generation AI determines that the language used is Japanese, and based on the created summary, returns answer A3, which indicates that the meeting belongs to the IT service category, the accuracy of the meeting is 85%, the response speed is standard, and the situation of the meeting is formal, and the generation AI suitable for translation is Tsuzumi41. A prompt including this answer A3 is set to Tsuzumi41.

[0077] Next, let us consider the case of pre-meeting information C4 in Figure 17, which is "Meeting with a client. Very important." In this case, the generation AI determines that the language used is Japanese, and based on the created summary, returns answer A4, which indicates that the meeting belongs to office services, the accuracy of the meeting is 90%, the response speed is standard, the situation of the meeting is formal, and the generation AI suitable for translation is Tsuzumi41. A prompt including this answer A4 is set to Tsuzumi41.

[0078] 18 and 19 are diagrams illustrating other examples of use of the processing system 100. Fig. 18 shows an example in which voice data or text data of users A, B, and C is translated into English, which is a common language of users A, B, and C, and the translated voice data or text data is output to terminal 20E.

[0079] In this case, the prompt creation unit 12 creates a prompt that instructs the selected generation AI to translate the input voice data or text data into the common language of each user. The output control unit 14 transmits the voice data or text data corresponding to the common language (English) to a predetermined terminal (e.g., terminal 20E). Alternatively, the output control unit 14 may transmit the voice data or text data corresponding to the common language (English) to the user terminal used by each user. At this time, the server device 10 may automatically recognize the user using the AI ​​and start interpretation using the generation AI.

[0080] FIG. 19 illustrates an example in which voice data of Japanese-speaking user A is translated into English, which is the language of user B, and the translated voice data (text data) is output from user terminal 20B in speaker mode. In this case, prompt creation unit 12 creates a prompt that instructs the selected generation AI to translate the input Japanese voice data or text data of user A (first user) into the language (e.g., English) used by user B (second user). Output control unit 14 transmits the voice data or text data translated into user B's language (English) to user terminal 20B. This is applicable to conversations at, for example, police stations, pharmacists, front desks, banks, ticket counters, kiosks, etc.

[0081] [Task management support] Next, a description will be given of the task management support process by the server device 10. Fig. 20 is a diagram for explaining an outline of the task management support process by the processing system 100.

[0082] After the conference is over, the server device 10 causes the generation AI selected for the conference to automatically create a summary of the contents of the conference in the language designated by each of the users A to D ((6-1) in FIG. 20).

[0083] The server device 10 requests the generation AI selected for the meeting to create a summary of the meeting (step S26), create an Action Item (A / I) (step S27), and compile the summary on one slide (step S28). Note that this slide may contain not only a simple text summary but also an image. The server device 10 then saves the content summary, A / I, and slide created by the generation model AI in the cloud and in an application ((6-2) in Figure 20). By checking these, the user can review the meeting, manage tasks, etc. ((6) in Figure 20).

[0084] [Simultaneous interpretation processing] Next, the processing procedure of the processing method according to the first embodiment will be described. Fig. 21 is an example of a sequence diagram showing the processing procedure of the processing method according to the first embodiment. Fig. 21 shows the processing procedure of simultaneous interpretation processing in which voice data or text data from user A is translated into the language used by user B and output from user terminal 20B of user B, and also translated into the language used by user C and output from user terminal 20C of user C.

[0085] 21, for example, when an application is started on user terminal 20A (step S31), information about conference rooms in which user A can participate is output to user terminal 20A through communication between server device 10 and user terminal 20A (step S32). User A operates user terminal 20A to select a conference room to participate in (steps S33 and S34), and then inputs a conference outline as pre-conference information before the conference (step S35). User terminal 20A transmits the input pre-conference information to server device 10 (step S36).

[0086] The server device 10 uses the generation AI to determine the language used by user A from the pre-meeting information (step S37). The server device 10 uses the generation AI to create a summary of the pre-meeting information (step S38). The server device 10 uses the generation AI to determine, based on the summary, the accuracy of the meeting, the degree of response speed, the situation of the meeting, the industry related to the meeting, and / or whether the meeting is a specific field or an accuracy-oriented meeting (step S39).

[0087] Based on the determination made in step S39, the server device 10 selects either Tsuzumi41 or ChatGPT51, adds the determination to the prompt, and sets the created prompt to the selected generation AI (steps S40 to S42). The example in Fig. 21 shows a case where Tsuzumi41 is selected and a prompt is set. When ChatGPT51 is selected, the prompt is set to ChatGPT51, and ChatGPT51 is caused to perform the translation process (step S48) described below.

[0088] When the server device 10 receives the prompt setting notification (step S43), it transmits to the user terminal 20A a message indicating that conversation can be started (step S44).

[0089] Then, when the voice data or text data input by user A is sent from user terminal 20A (steps S45, S46), server device 10 inputs the voice data or text data input by user A to Tsuzumi41 (step S47) and has it translated into the languages ​​used by users B and C, respectively (step S48).

[0090] The server device 10 transmits the voice data or text output from Tsuzumi 41 translated into the language used by user B (step S49) to user terminal 20B (step S50) and causes it to be output (step S53). The server device 10 transmits the voice data or text output from Tsuzumi 41 translated into the language used by user C (step S51) to user terminal 20C (step S52) and causes it to be output (step S54).

[0091] The same applies not only to user A, but also to cases where voice data or text data from users A, B, and C is input to server device 10. Server device 10 inputs the voice data or text data from users A, B, and C to Tsuzumi 41, and then outputs the voice data or text data translated into the languages ​​used by users A, B, and C output from Tsuzumi 41 to user terminals 20A to 20C of users A to C.

[0092] [Task management support processing] Next, the task management support process will be described. Fig. 22 is an example of a sequence diagram showing the processing procedure of another processing method according to the first embodiment.

[0093] 22, the server device 10 determines whether the conference of users A to C has ended (step S61). If the conference has not ended (step S61: No), the process returns to step S61. If the conference has ended (step S61: Yes), the server device 10 causes the generation AI selected for the conference to create a summary of the contents of the conference in the language used by each of users A to D (steps S62 and S63).

[0094] The server device 10 stores each summary output from the generation AI (steps S64, S65), and transmits the summary in the corresponding language to each of the user terminals 20A to 20C (steps S66 to S68).

[0095] [Advantages of the first embodiment] 23 and 24 are diagrams comparing a conventional translation service with the simultaneous interpretation service provided by the processing system according to the first embodiment.

[0096] As shown in Figure 23, in conventional translation services that translate a user's Japanese voice data and output it as text data, it takes time for the translation to be output. In contrast, in the simultaneous interpretation service of the first embodiment, each user uses their own terminal, which results in extra communication between the terminal and the server. However, since the IOWN network 60, which has high capacity, high speed, and low latency, is used, this delay can be almost ignored, resulting in faster dialogue. Note that the network used is not limited to the IOWN network 60, and other networks may also be used.

[0097] Therefore, for example, in the simultaneous interpretation service of the first embodiment, the second dialogue can be completed by the time t1 when the first dialogue ends in the conventional translation service.

[0098] As described above, according to the first embodiment, there is no need to wait for translation, which can alleviate stress during conversations. As a result, according to the first embodiment, it is possible to improve business efficiency and reduce costs. Furthermore, in the first embodiment, translated voice data is output, so the user only needs to listen to the voice data and does not need to read the text, which is highly convenient. Of course, the input and output may be either voice data or text data.

[0099] Furthermore, the processing system 100 communicates via the large-capacity, high-speed, low-latency IOWN network 60. Therefore, the processing system 100 can provide a simultaneous interpretation service that outputs translated audio data or text data in real time, even between three or more users.

[0100] In the processing system 100, even if, for example, four people are holding a conference, the voice data of one user is translated into each user's language and output from each user's user terminal 20. The simultaneous interpretation service in the first embodiment can realize simultaneous interpretation even among three or more multiple users.

[0101] Furthermore, in the first embodiment, simultaneous interpretation is performed by selecting either Tsuzumi41 or ChatGPT51 based on pre-conference information regarding the conference in which the user is participating, so that appropriate translation can be performed according to the conference.

[0102] In the first embodiment, the server device 10 adds content based on the pre-conference information to the selected generation AI and sets a prompt to instruct the AI ​​to translate the input voice data or text data into a specified language with a natural context. Therefore, according to the first embodiment, the optimal translation according to the situation can be obtained and provided to the user.

[0103] In the first embodiment, Tsuzumi41 and ChatGPT51 were used as examples of the generation AIs to be used, but other generation AIs may also be used, and the number of generation AIs is not limited to two, but may be any one of three or more generation AIs.

[0104] From here, we will explain several embodiments that are variations of the first embodiment described above. Note that the processing according to the first embodiment described above can be used for the second to fifth embodiments described below.

[0105] <Second embodiment> First, a second embodiment implemented by the processing system 100 according to the first embodiment will be described. The second embodiment is an embodiment in which input voice data or text data (hereinafter, sometimes simply referred to as "dialogue data") is translated by the server device 10 included in the processing system 100, and the translated dialogue data (hereinafter, sometimes simply referred to as "translated data") is retranslated based on a translation command from a user or the like.

[0106] (Processing by processing system) First, the processing of the processing system 100 according to the second embodiment will be described with reference to Fig. 25. Fig. 25 is a diagram illustrating the processing of the processing system 100 according to the second embodiment.

[0107] First, the user starts the application on the user terminal 20 and selects a conference room ((1) in FIG. 25). Next, before the conference starts, the user inputs pre-conference information by inputting their name and conference summary ((2) in FIG. 25).

[0108] The server device 10 selects a predetermined large-scale language model (LLM) based on the content determined using the generation AI. Note that in the following second to fifth embodiments, the server device 10 selects either Tsuzumi41 or ChatGPT51, but the large-scale language models used by the server device 10 are not limited to Tsuzumi41 and ChatGPT51.

[0109] Furthermore, the server device 10 adds the determined content to the prompt and sets the created prompt to the selected generation AI ((3) in FIG. 25). Then, when the prompt setting is complete, the conversation starts ((4) in FIG. 25).

[0110] The server device 10 provides the user with a simultaneous interpretation service by translating using Tsuzumi41 or ChatGPT51 in accordance with the set prompt ((5) in FIG. 5). Note that the processes from (1) to (5) in FIG. 25 are realized by the process according to the first embodiment described above.

[0111] The server device 10 displays the results of the interpretation (interpretation content) in step (5) of Fig. 25 to the users using the service ((6) of Fig. 25). As an example, the server device 10 displays the interpretation content translated into a language designated by each user on each user's chat screen in a group chat format ((6-1) of Fig. 25).

[0112] For example, when "Hello" is input by user E who speaks Japanese, server device 10 displays "Hello" ((6-2) in FIG. 25) on the screen of user E. Server device 10 also displays the result of translating "Hello" into English ((6-3) in FIG. 25) on the screen of user F who speaks English. Server device 10 also displays the result of translating "Hello" into Chinese ((6-4) in FIG. 25) on the screen of user G who speaks Chinese.

[0113] Here, the server device 10 retranslates the target translation result based on a retranslation command received from the user ((7) in FIG. 25). For example, if a command for retranslation is issued by the user when an event occurs in which the conversation is translated into unintended content, causing the conversation to become disjointed, the server device 10 checks the translated text and retranslates it.

[0114] On the other hand, the server device 10 receives an utterance associated with the utterance specified by the user as a quote reply based on a quote reply instruction received from the user ((8) in FIG. 25). For example, when a quote reply instruction is received from the user when the user does not understand the meaning of the translated content and checks the content with the user who originally made the utterance, the server device 10 performs quote reply processing associated with the target utterance.

[0115] Then, the server device 10 displays the result of the above-mentioned re-translation process on the chat screen of the user who instructed the server device 10 to re-translate ((9) in FIG. 25). In addition, the server device 10 displays the result of the quote reply process on the chat screen of each user in a group chat format ((9) in FIG. 25).

[0116] (Processing system 100) Next, the configuration of the processing system 100 according to the second embodiment will be described. Fig. 26 is a diagram showing the configuration of the processing system 100 according to the second embodiment. As shown in Fig. 26, the processing system 100 is made up of a server device 10, a user terminal 20, an earphone microphone 30, a generation AI server 40, and a generation AI server 50.

[0117] The processing system 100 according to the second embodiment has the same configuration as the processing system 100 according to the first embodiment described above, and therefore, a description of overlapping parts will be omitted.

[0118] (Server device 10) The server device 10 according to the second embodiment includes a pre-conference information acquisition unit 11 (acquisition unit), a prompt creation unit 12 (creation unit), an input control unit 13, an output control unit 14, a summary creation unit 15, a translation unit 61, and a memory unit 70.

[0119] (Storage unit 70) The storage unit 70 is realized by a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, or a storage device such as a hard disk or an optical disk. The storage unit 70 stores data and programs used for various processes by the server device 10, and various data acquired by the operation of the server device 10.

[0120] Specifically, the storage unit 70 temporarily or permanently stores information such as dialogue data, pre-conference information, created prompts, translation data, and created summary results.

[0121] (Input control unit 13) The input control unit 13 inputs dialogue data received from a user terminal to a generative model for translation in which a prompt is set to instruct the translation of the input dialogue data into a specified language with a natural context.

[0122] For example, the input control unit 13 inputs dialogue data in a first language input by a user using the simultaneous interpretation service via a user terminal into a translation generation model in which a prompt is set to translate the dialogue data into a second language specified for each other user using the simultaneous interpretation service.

[0123] (Output control unit 14) The output control unit 14 outputs the first translation data translated in response to the input by the input control unit 13 and the second translation data translated from the first language into the second language by the translation unit 61 described below to the user terminal 20. An example of the processing by the output control unit 14 will be described later in the section "Example of processing according to the second embodiment."

[0124] (Translation Department 61) The translation unit 61 inputs first translation data, which is dialogue data translated from the generative model for translation in response to input from the input control unit 13, into the generative model for translation and translates it into second translation data in a language different from the language of the first translation data. Note that an example of processing by the translation unit 61 will be described later in the section "Example of processing according to the second embodiment."

[0125] (Example of processing according to the second embodiment) Next, an example of processing by the processing system 100 according to the second embodiment will be described with reference to Figs. 27 to 29. Figs. 27 to 29 are diagrams illustrating an example of processing by the processing system 100 according to the second embodiment. Fig. 27 is a diagram illustrating an example of a "retranslation function" as a first example. Fig. 28 is a diagram illustrating an example of a "quote reply function" as a second example. Fig. 29 is a diagram illustrating an example of a "translation alert function" as a third example.

[0126] (First example) First, a first example of the "retranslation function" will be described with reference to Fig. 27. In the example of Fig. 27, User E is a user who speaks Japanese, User F is a user who speaks English, and User G is a user who speaks Chinese.

[0127] First, user E utters "Wait for me until I get there" to the processing system 100. Here, user E's utterance "Wait for me until I get there" is not unique and may be interpreted in various ways. For example, the utterance "Wait for me until I get there" may be translated as "Wait for me in the car," which means that the user wants the user to wait in the car. Furthermore, the utterance "Wait for me until I get there" may be translated as "Wait for me until I get there," which means that the user wants the user to wait until the user E arrives at a specified location.

[0128] Therefore, the server device 10 executes a retranslation process for checking whether the translation result displayed to the user who uses the simultaneous interpretation service is what the user intended.

[0129] First, the server device 10 inputs dialogue data input by a user into a generation AI for translation, and translates it into translation data in a first language. For example, the server device 10 translates an utterance by user E, "Wait until I arrive," into English and Chinese. Then, the server device 10 displays the translation data in Japanese ((1-1) in FIG. 27), English ((1-2) in FIG. 27), and Chinese ((1-3) in FIG. 27) on the user terminal of each user.

[0130] Here, as described above, if the utterance "Wait until I get there" is translated as "Wait until I get there" instead of "Wait in the car" as intended by user E, a discrepancy may arise in the conversation between user E who made the utterance and users F and G who checked the translation data. Therefore, by issuing a retranslation command to server device 10, the user can check whether their utterance has been translated into the content they intended.

[0131] Specifically, when the server device 10 receives a translation command from a user to translate first translation data translated into a first language into a different second language, the server device 10 inputs the first translation data into a generative model for translation. Next, the server device 10 displays the second translation data translated into the second language and output from the generative model for translation in response to the input of the first translation data, side by side with the first translation data.

[0132] For example, when the server device 10 receives a retranslation command from user E, it performs retranslation of the first translation data corresponding to "Wait until I come ((1-1) in Figure 27)," which is the translation result into English ((1-2) in Figure 27) and the translation result into Chinese ((1-3) in Figure 27).

[0133] As a specific example, the server device 10 retranslates the translation result into English ((1-2) in FIG. 27) into the language specified by user E (Japanese). Then, the server device 10 displays the retranslated result, "Wait until I come," to user E, as shown in (2-1) in FIG.

[0134] The server device 10 also retranslates the translation result into Chinese ((1-3) in FIG. 27) into the language specified by user E (Japanese). Then, the server device 10 displays the retranslated result, "Wait for me," to user E, as shown in (2-2) in FIG.

[0135] Through the above-described processing, User E can know that his / her utterance "Wait until I get there" is translated into English and Chinese as "Wait until I get there" rather than "Wait in the car."

[0136] The server device 10 can execute the retranslation process when it receives a retranslation command from another user who speaks the same language as the user who made the utterance. For example, the above-described process has been described as an example of a case in which a retranslation command is issued by user E who uttered "Wait until I get there," but the server device 10 can execute the retranslation process based on a retranslation command from a user other than user E.

[0137] Additionally, when a user inputs an utterance by voice, the server device 10 can display the result of the voice recognition to the user. Then, when the user confirms the result of the voice recognition and inputs a retranslation command, the server device 10 can execute a retranslation process based on the retranslation command. As a result, the server device 10 allows the user to confirm whether the translation result of the user's utterance into another language is as the user intended.

[0138] (Second example) Next, a second example of the "quote reply function" will be described with reference to Fig. 28. In the example of Fig. 28, User E is a user who speaks Japanese, and User F is a user who speaks English.

[0139] 28 shows a scene in which user E and user F are having a conversation based on the interpretation content translated by the processing system 100, and the conversation content is about "a request to create release determination materials and a schedule for the creation." Note that the translation of the conversation in the second example is assumed to be performed based on the translation processing according to the first embodiment described above, and the same applies hereinafter.

[0140] For example, the above-mentioned word "Ririhan" means "release decision," but is a coined word unique to user E or the organization to which user E belongs. Therefore, the server device 10 may translate the Japanese dialogue data "Prepare the materials for Ririhan" as "Prepare the materials for Ririhan." However, since "Ririhan" is not a translation of "Ririhan," but is simply a word in which the spoken sounds of "Ririhan" are transcribed into Roman letters, user F may not understand the meaning of the word.

[0141] Therefore, when a specific user cannot understand a certain utterance, the server device 10 accepts a reply associated with the utterance as a quoted reply and displays it to other users, thereby facilitating questions between users about unknown words.

[0142] Specifically, the server device 10 receives an utterance associated with a word specified by a user for the first translation data output from the generative model for translation. The server device 10 then outputs the received utterance in association with the word specified by the user. For example, the server device 10 realizes the quote reply function by the following process.

[0143] The server device 10 translates the user E's utterance "Prepare the materials for Ririhan ((1-1) in FIG. 28)" into "Prepare the materials for Ririhan ((1-2) in FIG. 28)."

[0144] The server device 10 accepts the press of the quote reply button ((2-1) in FIG. 28) by user F. Then, the server device 10 accepts "What is Ririhan? ((2-2) in FIG. 28)" from user F in association with "Prepare the materials for Ririhan ((1-2) in FIG. 28)."

[0145] The server device 10 translates the user F's utterance "What is Ririhan?" ((2-2) in FIG. 28) into "What is 'Ririhan'?" and displays it on the user terminal of the user E ((2-3) in FIG. 28).

[0146] The server device 10 accepts the press of the quote reply button ((3-1) in FIG. 28) by user E. Then, the server device 10 accepts user E's "It means release decision ((3-2) in FIG. 28)" in association with "What is 'Relihan'? ((2-3) in FIG. 28)".

[0147] The server device 10 translates the user E's utterance "It means release judgment ((3-2) in FIG. 28)" into "Ririhan means Release judgment." and displays it on the user terminal of the user F ((3-3) in FIG. 28).

[0148] Furthermore, in the second example, the server device 10 can also display the original sentence of the quoted utterance together with the original sentence before translation. For example, in response to "What is 'Ririhan'? ((2-3) of FIG. 28)", the server device 10 can display the original sentence "What is Ririhan?" uttered by user F alongside the original sentence ((4-1) of FIG. 28). On the other hand, when user F inputs "What is Ririhan? ((2-2) of FIG. 28)", the server device 10 can display the translation result of the utterance, "What is 'Ririhan'?" alongside the original sentence ((4-2) of FIG. 28).

[0149] Through the above-described processing, the server device 10 facilitates the exchange of questions and answers between users, even when accurate translation of coined words, technical terms, etc. is difficult, thereby enabling smooth communication in different languages.

[0150] (Third example) Next, a third example of the "translation alert function" will be described with reference to Fig. 29. Note that the example in Fig. 29 shows a conversation between user H and user I, who each speak a different language. Note that in Fig. 29, it is assumed that user H and user I are engaged in a dialogue based on the interpretation content translated by the processing system 100.

[0151] 29 shows a scene in which user H and user I are having a conversation about the progress of work. Specifically, in response to user H's question, "How is the progress of work XX? ((1-1) in FIG. 29)," user I replies, "It's not finished yet. ((1-2) in FIG. 29)."

[0152] Here, user H is about to input and send a utterance to scold user I, such as "You're always late with your work. Shouldn't you reconsider your time management?" ((2-1) in Figure 29). However, for example, pointing out someone's mistakes in public, as mentioned above, may be considered an insult in certain cultures and may not be desirable.

[0153] Therefore, when an undesirable utterance is detected, the processing system 100 issues an alert to the speaker. Specifically, the server device 10 translates the dialogue using, as a translation generation model, a generation model in which a prompt is set that instructs the dialogue partner to correct rude expressions in the input dialogue data. For example, the server device 10 realizes the translation alert function by the following process. Note that the above-mentioned "rude expressions" are expressions that are disrespectful to the other party, and include, for example, expressions that include meanings such as rudeness, insult, and disrespect.

[0154] As an example of processing, the server device 10 receives dialogue data such as "You're always late with your work. Shouldn't you reconsider your time management? ((2-1) in FIG. 29)" from user H. Then, the server device 10 inputs the received dialogue data into a translation generation model in which a prompt is set to instruct the other party to correct any expressions that include rude expressions.

[0155] The server device 10 displays, as an alert, the correction content output from the generative model for translation for which the above-mentioned prompt is set. Specifically, the server device 10 issues an alert to the user H, such as "In certain cultures, it is not advisable to point out mistakes in public. Are you sure you want to send this?" ((2-2) in FIG. 29). Note that the server device 10 can directly correct the dialogue data input by the user, rather than issuing an alert.

[0156] In addition to the examples of expressions shown in FIG. 29, the server device 10 can issue alerts or make corrections regarding content that should not be mentioned, such as history, political criticism, or religious criticism, which the other person does not want to be touched on, as well as words and actions that are considered rude in a particular culture, such as asking about someone's height, and non-unique expressions.

[0157] Furthermore, the server device 10 can issue an alert or correct not only explicit content but also implicit content (such as rude expressions toward the conversation partner) such as sarcasm in a specific context, such as "You're really smart." The server device 10 can also issue an alert or correct dialogue data (such as non-unique expressions) that have different meanings depending on the context, such as expressions that have multiple meanings, such as "I see," which means "I understand," and "I understand (when the person actually does not understand and is surprised)," and proverbs, idioms, honorifics, etc.

[0158] In addition, the server device 10 can use as a generation model for translation a generation model that has a prompt set to instruct the server device 10 to perform at least one of the following operations for the input dialogue data: correction of homonyms, correction of grammatical errors, correction of kanji conversion errors, and correction of speech recognition errors.

[0159] For example, when there is a possibility of a typographical error, an incorrect use of a homonym, or a speech recognition error based on the context for words that have the same or similar pronunciation but different meanings, such as "Sato" and "sugar," the server device 10 can point out and correct the error. Also, for example, when the server device 10 detects a grammatical error such as "What is Lilihan?", it can point out and correct the error.

[0160] (Processing procedure according to the second embodiment) From here, the processing procedure by the processing system 100 according to the second embodiment will be described with reference to Fig. 30 to Fig. 32. Fig. 30 to Fig. 32 are sequence diagrams showing the processing procedure according to the second embodiment.

[0161] First, the processing procedure related to the "retranslation function" will be described with reference to Fig. 30. As shown in Fig. 30, for example, the user terminal 20A accepts input of dialogue data from the user A (S101). Subsequently, the user terminal 20A transmits the accepted dialogue data to the generation AI server 40 (S102).

[0162] The generation AI server 40 executes translation processing using the received dialogue data (S103), and then transmits the translated dialogue data, which is the first translation data, to the user terminals 20A, 20B, and 20C (S104, S105, and S106).

[0163] User terminals 20A, 20B, and 20C that have received the first translation data output the translation results to each user (users A, B, and C) (S107, S108, and S109).

[0164] Here, the user terminal 20A accepts a translation command from the user A (S110). Subsequently, the user terminal 20A transmits the first translation data to the generation AI server 40 (S111).

[0165] The generation AI server 40 executes translation processing using the received first translation data (S112), and then transmits second translation data obtained by retranslating the first translation data to the user terminals 20A, 20B, and 20C (S113, S114, and S115).

[0166] User terminals 20A, 20B, and 20C that have received the first translation data output the translation results to each user (users A, B, and C) (S116, S117, and S118).

[0167] Note that similar processing is performed not only for user A but also when dialogue data from users A, B, and C is input to server device 10. Server device 10 inputs the dialogue data from users A, B, and C to Tsuzumi 41, and then outputs the dialogue data translated into the languages ​​used by users A, B, and C output from Tsuzumi 41 to user terminals 20A to 20C of users A to C. Note that the same applies to subsequent sequence diagrams.

[0168] Next, a processing procedure related to the "quote reply function" will be described with reference to Fig. 31. Note that steps S201 to S209 shown in Fig. 31 are the same as steps S101 to S109 shown in Fig. 30, and therefore description thereof will be omitted.

[0169] The user terminal 20A receives a quote reply command from the user A (S210). Subsequently, the user terminal 20A transmits an associated utterance based on the quote reply command to the generation AI server 40 (S211).

[0170] The generation AI server 40 executes a translation process using the received associated utterance (S212), and then transmits the translated associated utterance to the user terminals 20A, 20B, and 20C (S213, S214, and S215).

[0171] User terminals 20A, 20B, and 20C that have received the first translation data output the translation results to each user (users A, B, and C) (S216, S217, and S218).

[0172] Next, the processing procedure related to the "translation alert function" will be described with reference to Fig. 32. As shown in Fig. 32, for example, the user terminal 20A accepts input of dialogue data from the user A (S301). Subsequently, the user terminal 20A transmits the accepted dialogue data to the generation AI server 40 (S302).

[0173] The generation AI server 40 executes translation processing using the received dialogue data (S303), and then transmits the translated dialogue data to the server device 10 (S304).

[0174] If an alert is not to be issued (No in S305), the server device 10 transmits the translation data received from the generation AI server 40 to the user terminal 20A (S306). On the other hand, if an alert is to be issued (Yes in S305), the server device 10 transmits a predetermined translation alert to the user terminal 20A (S307).

[0175] Then, user terminal 20A outputs the accepted translation data or a translation alert to the user (S308).

[0176] (effect) Next, the effects of the processing system 100 according to the second embodiment will be described. The input control unit 13 of the processing system 100 inputs dialogue data received from a user terminal to a generative model for translation, which is configured with a prompt instructing the translation of the input dialogue data into a specified language with a natural context. The translation unit 61 of the processing system 100 inputs first translation data, which is translated dialogue data output from the generative model for translation in response to the input by the input control unit 13, into the generative model for translation, and translates the first translation data into second translation data in a language different from the language of the first translation data. The output control unit 14 of the processing system 100 outputs the second translation data translated by the translation unit 61 to the user terminal. Therefore, the processing system 100 according to the second embodiment has the effect of being able to provide an optimal translation according to the situation.

[0177] When the translation unit 61 receives a translation command to translate the language of first translation data translated into a first language into a different second language, the translation unit 61 inputs the first translation data into a generative model for translation. The output control unit 14 then displays the second translation data translated into the second language and output from the generative model for translation in response to the input by the translation unit 61 alongside the first translation data.

[0178] Therefore, the processing system 100 has the effect of making it easy for a user to check how his or her own speech has been translated when using a simultaneous interpretation service.

[0179] The output control unit 14 receives an utterance associated with a word designated by a user for the first translation data output from the generative model for translation, and outputs the word and the utterance in association with each other.

[0180] Therefore, when a user using a simultaneous interpretation service cannot understand the translated utterance, the processing system 100 enables the user to make an utterance associated with the utterance that the user cannot understand. Therefore, the processing system 100 realizes the optimal translation according to the situation, thereby achieving the effect of facilitating smoother communication between users who speak different languages.

[0181] The translation unit 61 can use as a generation model for translation a generation model that has a prompt set to instruct the translation unit 61 to perform at least one of the following operations for the input dialogue data: correction of homonyms, correction of grammatical errors, correction of kanji conversion errors, correction of speech recognition errors, correction of non-unique expressions, and correction of expressions that are rude to the dialogue partner.

[0182] Therefore, the processing system 100 not only translates the content of the speech, but also has the effect of preventing communication problems from occurring in communication between users with different cultures.

[0183] Furthermore, whereas in the past it was necessary to check the translation content using other translation devices, the processing system 100 performs re-translation of already translated translation data as a series of processes, thereby reducing computer processing and improving efficiency.

[0184] <Third embodiment> Next, a third embodiment, which is another form of the processing system 100 according to the first embodiment, will be described. In the third embodiment, an agreement document is generated using information on decisions made in a meeting, tasks to be performed, etc. extracted from a summary of translation data created by the processing system 100, and notified to users, and consensus building on the decisions made, tasks to be performed, etc. is accepted from the users. Note that the "agreement document" in the following sections refers to a document (screen) used to build consensus with participating users on the summary, decisions made, tasks to be performed, etc. of the meeting.

[0185] (Processing by processing system) First, the processing of the processing system 100 according to the third embodiment will be described with reference to Fig. 33. Fig. 33 is a diagram for explaining the processing of the processing system 100 according to the third embodiment.

[0186] First, the processing system 100 translates dialogue data input by a user. Note that the series of steps related to the translation process shown in (1) to (5) in Fig. 33 are the same as those in the second embodiment, and therefore their explanation will be omitted.

[0187] Next, the server device 10 causes the generation AI selected for the conference to automatically create a summary of the contents of the conference in the language designated by each user ((6) in FIG. 33).

[0188] The server device 10 inputs a summary into a generative model for extraction that has a predetermined prompt set, automatically determines the information to be extracted from the contents of the meeting and the appropriate prompt for that information, and extracts information related to the agreements such as the conclusion of the meeting, decisions, and tasks to be performed ((7) in FIG. 33). Note that the predetermined prompt may be, for example, a prompt such as that shown in (7-1) in FIG. 33.

[0189] The server device 10 generates a written agreement in a predetermined format using the extracted information on the agreement items ((8) in FIG. 33). Then, the server device 10 notifies the target user of the generated written agreement ((9) in FIG. 33).

[0190] As an agreement formation process, the server device 10 accepts the user's input of "agreement" or "disagreement" with respect to the notified agreement document. Then, the server device 10 stores information identifying the user who agrees or disagrees with the agreement document in the storage unit ((10) in FIG. 33).

[0191] (Processing system 100) Next, the configuration of the processing system 100 according to the third embodiment will be described. Fig. 34 is a diagram showing the configuration of the processing system 100 according to the third embodiment. As shown in Fig. 34, the processing system 100 is made up of a server device 10, a user terminal 20, an earphone microphone 30, a generation AI server 40, and a generation AI server 50.

[0192] The processing system 100 according to the third embodiment has the same configuration as the processing system 100 according to the second embodiment described above, and therefore, a description of overlapping parts will be omitted.

[0193] (Server device 10) The server device 10 according to the third embodiment includes a pre-meeting information acquisition unit 11 (acquisition unit), a prompt creation unit 12 (creation unit), an input control unit 13, an output control unit 14, a summary creation unit 15, a consensus building unit 62, and a memory unit 70.

[0194] (Storage unit 70) The memory unit 70 temporarily or permanently stores information such as dialogue data, pre-meeting information, created prompts, translation data, created summary results, information on agreed matters, agreement documents, and results of the consensus building process.

[0195] For example, the storage unit 70 stores the conclusions, decisions, tasks to be performed, etc. of the meeting extracted from the created summary as information on the agreements.

[0196] For example, the storage unit 70 stores information identifying a user who has made an input relating to agreement into the agreement document, information indicating agreement, and the agreement document in association with each other. The storage unit 70 also stores information identifying a user who has made an input relating to non-agreement into the agreement document, information indicating non-agreement, and the agreement document in association with each other.

[0197] (Summary Creation Section 15) The summary creation unit 15 uses translation data, which is dialogue data translated from a translation generation model in response to input from the input control unit 13, to create a summary in one specific language and then translate the summary into another language. The summary creation unit 15 then summarizes the contents of the meeting in the translated language specified for each user. An example of the processing of the summary creation unit 15 will be described later in the section "Example of processing according to the third embodiment."

[0198] (Consensus Building Department 62) The consensus building unit 62 accepts input of agreement or disagreement from the user based on the summary created by the summary creation unit 15. An example of the processing by the consensus building unit 62 will be described later in the section "Example of processing related to the third embodiment."

[0199] (Example of processing according to the third embodiment) From here, an example of processing by the processing system 100 according to the third embodiment will be described with reference to Fig. 35. Fig. 35 is a diagram illustrating an example of processing by the processing system 100 according to the third embodiment. Note that Fig. 35 is a diagram illustrating an example of a "consensus building function".

[0200] (An example of consensus building processing) As described above, the processing system 100 according to the third embodiment realizes the function of generating an agreement document using a summary of a meeting, etc., notifying the users of the agreement document, and forming an agreement among the users who participated in the meeting.

[0201] The server device 10 generates a written agreement including "date and time," "participating members," "summary," "decisions / action items," etc. ((1) in Figure 35). Specifically, the server device 10 inputs translation data for each user output from the generative model for translation to a generative model for extraction, which is set with a prompt that instructs the extraction of information related to the agreed matters related to the meeting included in the dialogue data in the user's natural language. Next, the server device 10 uses the extracted information related to the agreed matters to generate a written agreement including at least one of a summary of the contents of the meeting, decisions, and action items. The server device 10 then transmits the generated written agreement to the user.

[0202] "Date and time" is information indicating the date and time when the conference was held. For example, as shown in (2) of Fig. 35, the server device 10 generates a written agreement including information such as "2023 / 11 / 02 10:00-11:00" as the date and time.

[0203] "Participating members" is information for identifying users who participated in the conference. For example, as shown in (3) of Fig. 35, the server device 10 generates a written agreement including information such as "User J (you), User K, and User L" as participating members.

[0204] The "summary" is information summarizing the content of the conversation in the conference. For example, as shown in (4) of Fig. 35, the server device 10 generates a written agreement including information such as "User J has requested User K and User L to perform the sorting work" as a summary.

[0205] "Decision Items / Action Items" is information on decisions made in a meeting, ToDos, policies, etc. Specifically, for "Decision Items / Action Items," for example, as shown in (5) of Fig. 35, the server device 10 generates a written agreement including information such as "User K will prepare release decision materials by August 12th" and "All members will work on cost reduction to achieve the business plan" as the "Decision Items / Action Items."

[0206] The server device 10 then accepts input of "agree" or "disagree" from the user who has been notified of the generated written agreement.

[0207] Specifically, the server device 10 displays the generated agreement document and an input button for accepting agreement to the agreement document on the user's terminal. Next, when the user presses the button for responding to agreement, the server device 10 accepts the user's agreement to the agreement document. Then, the server device 10 stores information identifying the user who has agreed to the contents of the agreement document in the storage unit 70 in association with the agreement document.

[0208] On the other hand, the server device 10 displays the generated agreement document and an input button for accepting disagreement with the agreement document on the user's terminal. Next, when the user presses the button for responding to disagreement, the server device 10 accepts the user's disagreement with the agreement document. Then, the server device 10 stores information identifying the user who disagreed with the contents of the agreement document in the storage unit 70 in association with the agreement document.

[0209] Furthermore, the consensus building unit 62 can display on the user's terminal a written agreement relating to the entire conference or a written agreement for each predetermined granularity, in association with a button for responding to the agreement.

[0210] For example, the server device 10 accepts "agreement / disagreement" for each item of "decision matter / action item." Specifically, the server device 10 accepts pressing of a button ((5-1) in FIG. 35) which means agreement with a decision matter / action item such as "User K will prepare release determination materials by August 12." On the other hand, the server device 10 accepts pressing of a button ((5-2) in FIG. 35) which means disagreement with a decision matter / action item such as "User K will prepare release determination materials by August 12."

[0211] Furthermore, the server device 10 can perform consensus building processing for each item, such as date and time, participating members, summary, etc., in addition to the above-mentioned "decision items / action items."

[0212] Furthermore, the server device 10 can display a conversation log screen or a conversation list screen related to the agreement document by pressing the buttons ((6) and (7) in FIG. 35) displayed at the top of the agreement document screen.

[0213] (Processing procedure according to the third embodiment) From here, the processing procedure by the processing system 100 according to the third embodiment will be described with reference to Fig. 36. Fig. 36 is a sequence diagram showing the processing procedure according to the third embodiment.

[0214] 36, for example, the server device 10 transmits a request to create a summary to the generation AI server 40 (S401). Subsequently, the generation AI server 40 generates a summary for the target conference based on the summary creation request (S402). Then, the generation AI server 40 transmits the created summary to the server device 10 (S403).

[0215] The server device 10 stores the received summary in the storage unit 70 (S404). Then, the server device 10 transmits a command to the generation AI server 40 to extract information related to the agreement from the summary (S405, S406).

[0216] The generation AI server 40 extracts information about the agreed-upon matters from the summary based on the extraction command (S407), and then transmits the extracted information about the agreed-upon matters to the server device 10 (S408).

[0217] The server device 10 generates a written agreement using the received information on the terms of the agreement (S409), and then transmits the written agreement to each of the user terminals 20A, 20B, and 20C (S410, S411, and S412).

[0218] Upon receiving the written agreement, user terminals 20A, 20B, and 20C output the translation results to each user (users A, B, and C) (S413, S414, and S415).

[0219] The server device 10 receives responses from each user indicating agreement or disagreement with the written agreement (S416, S417, and S418). Then, the server device 10 executes a consensus building process based on the responses received from each user (S419).

[0220] (effect) Next, the effects of the processing system 100 according to the third embodiment will be described. The input control unit 13 of the processing system 100 inputs dialogue data received from a user terminal to a generative model for translation, which is configured with a prompt commanding that the input dialogue data be translated into a specified language with a natural context. The summary creation unit 15 of the processing system 100 summarizes the contents of the meeting in the specified language for each user using translation data, which is the translated dialogue data output from the generative model for translation in response to the input by the input control unit 13. The consensus building unit 62 of the processing system 100 accepts input of agreement or disagreement from users based on the summary created by the summary creation unit 15. Therefore, the processing system 100 according to the third embodiment advantageously enables appropriate consensus building on the content to be interpreted.

[0221] The summary creation unit 15 inputs translation data for each user output from the translation generation model in response to input from the input control unit 13 to a generation model for extraction, which has a prompt set to instruct extraction of information about the agreed-upon matters included in the dialogue data, which is dialogue data in the user's natural language. Next, the summary creation unit 15 uses the information about the agreed-upon matters to generate a written agreement including at least one of a summary of the contents of the conference, decisions, and action items. The consensus building unit 62 then sends the generated written agreement to the user.

[0222] Therefore, the processing system 100 has the effect of enabling easy consensus building by generating a written agreement to be used for consensus building using dialogue data between users.

[0223] The agreement building unit 62 displays the generated agreement document and an input button for accepting agreement to the agreement document on the user's terminal. Next, when the user presses the button for responding to agreement, the agreement building unit 62 accepts the user's agreement to the agreement document. Then, the agreement building unit 62 stores information identifying the user who agreed to the contents of the agreement document in the storage unit 70 in association with the agreement document.

[0224] Therefore, the processing system 100 has the effect of making it possible to easily reach an agreement about a conference based on an automatically generated written agreement.

[0225] The consensus building unit 62 displays on the user's terminal a written agreement relating to the entire conference or a written agreement for each predetermined granularity, in association with a button for responding to the agreement.

[0226] Therefore, the processing system 100 has the effect of making it easy to reach an agreement depending on the situation, such as reaching an agreement on a part of the agreement document or on the entire agreement document.

[0227] Furthermore, whereas in the past, translation data had to be input into a device separate from the device that performed the translation process, a summary had to be created, and consensus had to be reached individually using the created summary, the processing system 100 performs processes such as reaching an agreement using an agreement document generated from the summary as a series of processes, thereby reducing computer processing and making it more efficient.

[0228] <Fourth embodiment> Next, a fourth embodiment, which is another form of the processing system 100 according to the first embodiment, will be described. In the fourth embodiment, the processing system 100 extracts information about business knowledge from dialogue data between users, such as translation data, and presents the information about the business knowledge to the user in a predetermined format. In the following sections, "information about business knowledge" includes knowledge about business accumulated in a predetermined organization, written information, and the like.

[0229] (Processing by processing system) First, the processing of the processing system 100 according to the fourth embodiment will be described with reference to Fig. 37. Fig. 37 is a diagram for explaining the processing of the processing system 100 according to the fourth embodiment.

[0230] First, the processing system 100 translates the dialogue data input by the user. Note that the series of steps related to the translation process shown in (1) to (5) in Fig. 37 are the same as those in the second embodiment, and therefore their explanation will be omitted.

[0231] Next, the server device 10 extracts information related to business knowledge from the translation data stored in advance ((6) in FIG. 37).

[0232] The server device 10 performs a predetermined association process on the extracted information related to business knowledge ((7) in FIG. 37). Specifically, the server device 10 inputs dialogue data between users and translation data into a generative model in which a predetermined prompt is set, and performs a predetermined association process such as associating a question with an answer to the question. Note that the above-mentioned predetermined prompt may be, for example, a prompt as shown in (7-1) in FIG. 37.

[0233] The server device 10 stores the information about the business knowledge for which the association process has been performed in the storage unit ((8) in FIG. 37). Then, the server device 10 outputs the stored information about the business knowledge to the user in a predetermined format ((9) in FIG. 37).

[0234] (Processing system 100) Next, the configuration of the processing system 100 according to the fourth embodiment will be described. Fig. 38 is a diagram showing the configuration of the processing system 100 according to the fourth embodiment. As shown in Fig. 38, the processing system 100 is made up of a server device 10, a user terminal 20, an earphone microphone 30, a generation AI server 40, and a generation AI server 50.

[0235] The processing system 100 according to the fourth embodiment has a similar configuration to the processing system 100 according to the second embodiment described above, and therefore, a description of overlapping parts will be omitted.

[0236] (Server device 10) The server device 10 according to the third embodiment includes a pre-meeting information acquisition unit 11 (acquisition unit), a prompt creation unit 12 (creation unit), an input control unit 13, an output control unit 14, a summary creation unit 15, a storage unit 63, a display unit 64, an inquiry reception unit 65, and a memory unit 70.

[0237] (Storage unit 70) The storage unit 70 temporarily or permanently stores information such as dialogue data, pre-conference information, created prompts, translation data, and information related to business knowledge that has undergone predetermined association processing.

[0238] An example of information related to business knowledge for which a predetermined association process has been performed and which is stored in the storage unit 70 will now be described with reference to Fig. 39. Fig. 39 is a table diagram illustrating an example of information related to business knowledge for which a predetermined association process has been performed according to the fourth embodiment.

[0239] As shown in Figure 39, the storage unit 70 stores the following items in association with each other: "Key," which identifies a record in which information about business knowledge for which a predetermined association process has been performed is stored; "lang," "question," and "answer."

[0240] The "lang" field stores information indicating the language. The "question" field stores information about the inquiry contained in the dialogue data or translation data. The "answer" field stores information about the answer to the inquiry contained in the dialogue data or translation data.

[0241] For example, the storage unit 70 stores the key "1", the lang "JPN", the question "What does 'Ririban' mean?", and the answer "Release decision" in association with each other. The above information means that the answer "Release decision" was given in response to the inquiry "What does 'Ririban' mean?" made in Japanese (JPN).

[0242] Furthermore, for example, the storage unit 70 stores the key "1", the lang "ENG", the question "What does the Ririhan mean?", and the answer "Ririhan means release judgment" in association with each other. The above information means that the answer "Ririhan means release judgment" was given in response to the inquiry "What does the Ririhan mean?" made in English (ENG).

[0243] Furthermore, the storage unit 70 can store information relating to business knowledge extracted from translation data translated into different languages ​​in association with each other across different languages.

[0244] For example, as described above, questions and answers for the lang "JPN" are associated with key "1." On the other hand, questions and answers for the lang "ENG" are associated with key "1." Therefore, the storage unit 70 can store questions and answers for the lang "JPN" and questions and answers for the lang "ENG," identified by the common key "1," in association with each other.

[0245] The storage unit 70 can store information related to various types of business knowledge, not limited to inquiries and answers related to the "Release Date" described above. For example, the storage unit 70 can store information related to business knowledge such as "explanations of the meanings of terms," ​​"information related to the purpose of the business, business procedures, operating instructions for tools and systems, troubleshooting, etc."

[0246] (Storage section 63) The storage unit 63 stores in the memory unit 70 information about business knowledge extracted based on translation data, which is translated dialogue data output from the generative model for translation in response to input by the input control unit 13.

[0247] Specifically, the storage unit 63 inputs translation data for each user into a generation model for extraction that has a prompt set to instruct the extraction of questions and corresponding answers from the user's dialogue data in natural language.

[0248] Next, the storage unit 63 associates the questions extracted by the generative model for extraction with the answers corresponding to the questions, and stores them as information about business knowledge in the memory unit 70. For example, the storage unit 63 associates inquiries about the meanings of unknown terms with the corresponding answers, etc., and stores them in the memory unit 70.

[0249] Specifically, the storage unit 63 inputs translation data for each user to a generative model for extraction, which is set with a prompt that commands extraction of information related to business training from the user's dialogue data in natural language. Next, the storage unit 63 stores the information related to business training extracted by the generative model for extraction in the memory unit 70 as information related to business knowledge.

[0250] For example, the storage unit 63 stores in the memory unit 70 information that matches inquiries and corresponding answers regarding general business information such as the purpose of the business, business procedures, instructions for using tools and systems, troubleshooting, etc., as well as individual explanation contents, etc., as information regarding business training.

[0251] Furthermore, the storage unit 63 can store information associating the knowledge possessed by a user with the user in the memory unit 70. Specifically, the storage unit 63 inputs translation data for each user to a generative model for extraction in which a prompt is set to instruct the extraction of information identifying the user who made the utterance from the dialogue data in the user's natural language and information related to the user's business knowledge. Next, the storage unit 63 associates the information identifying the user extracted by the generative model for extraction with the information related to the user's business knowledge, and stores the information in the memory unit 70.

[0252] For example, the storage unit 63 can associate the speaker of the dialogue data that is the source of the knowledge extracted from the generative model for extraction with information about business knowledge such as ``Who is knowledgeable about XX?'' and store it in the memory unit 70.

[0253] (Display section 64) The display unit 64 displays the information related to business knowledge stored by the storage unit 63 in association with questions and answers to the questions. An example of the processing of the display unit 64 will be described later in the section "Example of processing related to the fourth embodiment."

[0254] (Inquiry Reception Department 65) The inquiry receiving unit 65 generates a response to an inquiry from a user using information related to business knowledge stored in the storage unit 70, and outputs the response to the user. Note that an example of the processing of the inquiry receiving unit 65 will be described later in the section "Example of processing related to the fourth embodiment."

[0255] (Example of processing according to the fourth embodiment) Next, an example of processing by the processing system 100 according to the fourth embodiment will be described with reference to Fig. 40 and Fig. 41. Fig. 40 and Fig. 41 are diagrams illustrating an example of processing by the processing system 100 according to the fourth embodiment. Note that Fig. 40 is a diagram illustrating "output of a Q&A list" as a first example. Also, Fig. 41 is a diagram illustrating an example of "automatic response to inquiries" as a second example.

[0256] (First example) First, the first example, "output of Q&A list", will be described with reference to Fig. 40. In the example of Fig. 40, User E is a user who speaks Japanese, and User F is a user who speaks English.

[0257] In the first example, the server device 10 displays a list of Q&As, which are information related to business knowledge that has been associated and stored in the memory unit 70 by the processing related to the fourth embodiment described above, on each user's user terminal in a language specified for each user.

[0258] Specifically, the server device 10 displays questions and answers associated with each other for each predetermined category as a "Q&A list." For example, as shown in (1) of Fig. 40, the server device 10 displays a list ((1-1) of Fig. 40) in which the category "Explanation of the Meaning of Terms" identified by No. "1" is associated with the question "What does "Release Decision" mean and the answer "Release Decision."

[0259] Furthermore, the server device 10 can display the above-mentioned Q&A list translated into English ((2-1) in FIG. 40) for a user who speaks English, as shown in (2) in FIG.

[0260] It should be noted that the above-mentioned "Q&A list" is merely an example, and the server device 10 can display information related to business knowledge in other formats. Also, the above-mentioned classification "explanation of term meanings" is merely an example, and the server device 10 can display Q&As of other classifications without any limitations.

[0261] Through the above-described process, user E and user F, who each speak different languages, can refer to information related to business knowledge, such as a Q&A list, for each language they use.

[0262] (Second example) Next, a second example of "automatic response to inquiries" will be described with reference to Fig. 41. In the example of Fig. 41, User E is a user who speaks Japanese, and User F is a user who speaks English.

[0263] In the second example, the server device 10 inputs a user's inquiry to a generative model for inquiry handling, which has set therein a prompt that commands the generation of an answer to a specific inquiry input by the user, based on information related to business knowledge stored in the storage unit 70. The server device 10 then outputs to the user an answer to the user's inquiry that has been generated by the generative model for inquiry handling.

[0264] Specifically, the server device 10 accepts an inquiry from a user in a chat format or the like, as shown in (1) of Fig. 41. Then, the server device 10 generates a response corresponding to the inquiry from the user and outputs it to the user.

[0265] For example, the server device 10 receives an inquiry from user E, "What is release judgment?" ((1-1) in FIG. 41). Next, the server device 10 generates a response to the inquiry from user E based on the business knowledge stored in the storage unit 70. Then, the server device 10 outputs the generated response, "It means 'release judgment'," to user E who made the inquiry ((1-2) in FIG. 41).

[0266] Furthermore, with regard to the automatic response to the above-mentioned inquiry, the server device 10 can automatically respond to an inquiry in English for a user who speaks English ((2) in FIG. 41).

[0267] For example, the server device 10 receives an inquiry from user F, "What does the Ririhan mean?" ((2-1) in FIG. 41). Next, the server device 10 generates a response to the inquiry from user F, based on the business knowledge stored in the storage unit 70. Then, the server device 10 outputs the generated response, "Ririhan means release judgment." to user E who made the inquiry ((2-2) in FIG. 41).

[0268] Through the above-described process, users E and F, who speak different languages, can inquire about unclear points and efficiently obtain answers in each language they use. Note that the server device 10 can output answers to users by taking into consideration (correcting, etc.) content that should not be mentioned clearly, such as history, political criticism, or religious criticism, which the other person does not want to be touched on, words or actions that are considered rude in a particular culture, such as asking about someone's height, and inconsistent expressions.

[0269] Note that the above-described "automatic response to inquiries" is merely an example, and the server device 10 can display information related to business knowledge in other formats. Also, the above-described automatic response to inquiries related to "reli-ban" is merely an example, and the server device 10 can automatically respond to other types of inquiries without any limitations.

[0270] (Processing Procedure) Hereinafter, a description will be given of the processing procedure by the processing system 100 according to the fourth embodiment. Figures 42 and 43 are sequence diagrams showing the processing procedure according to the fourth embodiment.

[0271] First, the processing procedure for outputting the "Q&A list" will be described with reference to Fig. 42. As shown in Fig. 42, for example, the user terminal 20A accepts input of dialogue data from the user A (S501). Subsequently, the user terminal 20A transmits the accepted dialogue data to the generation AI server 40 (S502).

[0272] The generation AI server 40 executes translation processing using the received dialogue data (S503), and then transmits the translated data to the server device 10 (S504).

[0273] The server device 10, which has received the translation data, extracts information relating to business knowledge from the translation data (S505). Next, the server device 10 performs a correlation process on the extracted information relating to business knowledge (S506). Then, the server device 10 stores the correlation-processed information relating to business knowledge in the storage unit 70 (S507).

[0274] Here, when the user terminal 20A receives a command to view the Q&A list from the user A, the server device 10 transmits information about business knowledge (Q&A list) to the user terminal 20A (S508). Then, the user terminal 20A outputs the received information about business knowledge (Q&A list) (S509).

[0275] Next, a processing procedure for outputting "automatic response to inquiry" will be described with reference to Fig. 43. Note that steps S601 to S607 shown in Fig. 43 are the same as steps S501 to S507 shown in Fig. 42, and therefore description thereof will be omitted.

[0276] Here, the user terminal 20A accepts an input of an inquiry from the user A (S608), and then the user terminal 20A transmits the input inquiry content to the server device 10 (S609).

[0277] The server device 10 generates a response corresponding to the inquiry received from the user terminal 20A (S610). Then, the server device 10 transmits the response corresponding to the inquiry to the user terminal 20A (S611). Then, the user terminal 20A outputs the response corresponding to the received inquiry (S612).

[0278] (effect) Next, the effects of the processing system 100 according to the fourth embodiment will be described. The input control unit 13 of the processing system 100 inputs dialogue data received from a user terminal to a generative model for translation, which is configured with a prompt that commands the input dialogue data to be translated into a specified language with a natural context. The storage unit 63 of the processing system 100 stores, in the memory unit 70, information related to business knowledge extracted based on translation data, which is translated dialogue data output from the generative model for translation in response to input by the input control unit 13. Therefore, the processing system 100 according to the fourth embodiment has the effect of enabling appropriate utilization of interpretation content.

[0279] The storage unit 63 inputs translation data for each user to a generative model for extraction, which is set with prompts that instruct the extraction of questions and corresponding answers from the dialogue data in the user's natural language. The storage unit 63 then associates the questions and corresponding answers extracted by the generative model for extraction with each other, and stores them in the memory unit 70 as information related to business knowledge.

[0280] The storage unit 63 also inputs translation data for each user to a generative model for extraction, which is configured with a prompt that commands the extraction of information related to business training from the user's natural language dialogue data. The storage unit 63 then stores the information related to business training extracted by the generative model for extraction in the memory unit 70 as information related to business knowledge.

[0281] Therefore, the processing system 100 has the effect of automatically extracting information for generating a Q&A list and automatically responding to inquiries from dialogue data between users in the course of business and storing the information in a database.

[0282] The display unit 64 displays the information about business knowledge stored in the storage unit 63 in association with questions and answers to the questions.

[0283] Therefore, the processing system 100 can display a Q&A list to a user using information about business knowledge that is automatically extracted from dialogue data between users in business and accumulated in a database, thereby achieving the effect of enabling effective utilization of organizational knowledge accumulated through business execution.

[0284] The inquiry receiving unit 65 inputs an inquiry from a user to a generative model for inquiry handling, which has set therein a prompt that commands the generation of an answer to a specific inquiry input by the user, based on information related to business knowledge stored in the storage unit 70. The inquiry receiving unit 65 outputs to the user an answer to the user's inquiry that has been generated by the generative model for inquiry handling.

[0285] Therefore, the processing system 100 can automatically respond to business inquiries from users by using information on business knowledge that is automatically extracted from dialogue data between users in business and stored in a database. As a result, the processing system 100 has the effect of making it possible to effectively utilize organizational knowledge that is accumulated through business execution.

[0286] Furthermore, the server device 10 according to the fourth embodiment enables the utilization of information on business knowledge corresponding to different languages, different cultures, etc. For example, when cultural backgrounds differ, the ease of understanding of terms and explanations may differ. Therefore, in the past, when generating a Q&A list from business conversation logs, etc., it was necessary to understand the ease / difficulty of understanding of the content for each language and adjust the description content accordingly.

[0287] Specifically, let's say a conversation takes place between users who speak Japanese, English, and Chinese: "1. A question is asked about the work procedure and answered," "2. The English user asks again about the meaning of term a used in the answer," and "3. The Chinese user requests another explanation about step b of the work procedure."

[0288] When generating a Q&A list from the above-mentioned conversation (dialogue data), since term a is unfamiliar to English-speaking users, the description content must be adjusted to include an explanation of term a in the English Q&A list. On the other hand, since the explanation of step b is difficult for Chinese-speaking users to understand, the description content must be adjusted to be more suitable for Chinese-speaking users.

[0289] Here, the processing system 100 stores, for each language, translation data, which is dialogue data translated into a language specified by a user who speaks a different language. The processing system 100 can then extract information related to business knowledge using the translation data stored for each language, generate a Q&A list, and automatically respond to inquiries. Therefore, the processing system 100 enables smoother communication between different cultures. Furthermore, the processing system 100 not only directly creates knowledge in each language specified by the user, but also optimizes knowledge for each language based on the retranslation function and quote reply function described in the second embodiment.

[0290] Furthermore, whereas processing system 100 previously required a device that performed translation processing to store information related to business knowledge, convert it into a predetermined format, and output it to the user, the processing system extracts information related to business knowledge from translation data and outputs it to the user in a predetermined format as a series of processes, thereby enabling computer processing to be reduced and made more efficient.

[0291] <Fifth embodiment> Next, a fifth embodiment, which is another form of the processing system 100 according to the first embodiment, will be described. The fifth embodiment is an embodiment in which training data used for training a generative model is generated using dialogue data, translation data, etc. stored by the processing system 100, and the generative model is additionally trained using the generated training data and a dictionary file of technical terms, etc., predefined by the user. Note that in the fifth embodiment, the generative model to be trained by the processing system 100 may include a speech recognition model, etc.

[0292] (Processing by processing system) First, the processing of the processing system 100 according to the fifth embodiment will be described with reference to Fig. 44. Fig. 44 is a diagram for explaining the processing of the processing system 100 according to the fifth embodiment.

[0293] First, the processing system 100 translates dialogue data input by a user. Note that the series of steps related to the translation process shown in (1) to (5) in Fig. 44 are the same as those in the second embodiment, and therefore will not be described here.

[0294] First, the server device 10 extracts dialogue history data that meets a predetermined condition, such as the user being unable to understand, from pre-stored dialogue data between users and translation data ((6) in FIG. 44). Specifically, the server device 10 inputs the dialogue data between users and translation data into a generative model for extraction, to which a predetermined prompt is set, and performs a process of extracting dialogue history data, which is history data in which a user makes an inquiry because they are unable to understand the meaning of a sentence, or makes a quote reply. Note that the above-mentioned predetermined prompt may be, for example, a prompt as shown in (6-1) in FIG. 44.

[0295] The server device 10 generates learning data using the extracted dialogue history data ((7) in FIG. 44). Then, the server device 10 uses the generated learning data to train a generative model for translation ((8) in FIG. 44).

[0296] (Processing system 100) Next, the configuration of the processing system 100 according to the fifth embodiment will be described. Fig. 45 is a diagram showing the configuration of the processing system 100 according to the fifth embodiment. As shown in Fig. 45, the processing system 100 is made up of a server device 10, a user terminal 20, an earphone microphone 30, a generation AI server 40, and a generation AI server 50.

[0297] The processing system 100 according to the fifth embodiment has the same configuration as the processing system 100 according to the second embodiment described above, and therefore, a description of overlapping parts will be omitted.

[0298] (Server device 10) The server device 10 according to the fifth embodiment includes a pre-meeting information acquisition unit 11 (acquisition unit), a prompt creation unit 12 (creation unit), an input control unit 13, an output control unit 14, a summary creation unit 15, a storage unit 63, a learning data generation unit 66, a learning unit 67, and a memory unit 70.

[0299] (Storage unit 70) The storage unit 70 temporarily or permanently stores information such as the dialogue data received by the server device 10, pre-conference information, created prompts, translation data, dialogue history data, and learning data.

[0300] An example of the training data stored in the storage unit 70 will now be described with reference to Fig. 46. Fig. 46 is a table diagram showing an example of training data according to the fifth embodiment. As shown in Fig. 46, the storage unit 70 stores the following items in association with each other: "Key" for identifying an individual record of the training data, "lang", "unknown word", and "correct answer".

[0301] The "lang" field above stores information indicating the language. The "unknown word" field stores information about parts of the sentence that were not conveyed in each language extracted from the dialogue data or translation data, as well as mistranslations. The "correct answer" field stores information about expressions that accurately convey the meaning across different languages.

[0302] For example, the storage unit 70 stores the key "1", the lang "JPN", the unknown word "Ririhan", and the correct answer "Release Judgment" in association with each other. The above information means that for "Ririhan" in Japanese (JPN), the correct meaning that can be conveyed between different languages ​​is "Release Judgment".

[0303] Furthermore, for example, the storage unit 70 stores the key "1", the lang "ENG", the unknown word "Ririhan", and the correct answer "release judgment" in association with each other. The above information means that for "Ririhan" in English (ENG), the correct meaning that can be conveyed between different languages ​​is "release judgment".

[0304] Furthermore, the storage unit 70 can store information relating to business knowledge extracted from translation data in different languages ​​in association with each other in different languages.

[0305] For example, as described above, the unknown words and correct answers for the lang "JPN" are associated with key "1." On the other hand, the unknown words and correct answers for the lang "ENG" are associated with key "1." Therefore, the storage unit 70 can store the unknown words and correct answers for the lang "JPN" identified by the common key "1" in association with the unknown words and correct answers for the lang "ENG."

[0306] Note that the above-described associations are merely examples, and the storage unit 70 can store any data that can be used as training data, without any limitations on the type, content, data format, etc. For example, the storage unit 70 can store training data related to coined words and technical terms used in specific fields or areas, in addition to training data related to words that the user cannot understand.

[0307] The storage unit 70 can also store predetermined terminology data such as dictionary files of technical terms and the like that are set in advance by the user.

[0308] (Storage section 63) The storage unit 63 stores information used as learning data for dialogue data between users and translation data in the storage unit 70.

[0309] Specifically, the storage unit 63 receives an utterance associated with a word designated by the user for the translation data output from the generative model for translation, and stores the word and the utterance in association with each other in the memory unit 70.

[0310] Furthermore, the storage unit 63 associates an inquiry about a wording included in the dialogue data by the first user with a response to the inquiry by the second user, and stores the associated data in the memory unit 70. As an example, the storage unit 63 associates a first utterance (inquiry) that is the subject of a quote reply with a second utterance (response) that corresponds to the first utterance for dialogue data (translation data) related to the quote reply function of the second embodiment, as shown in FIG.

[0311] Furthermore, for example, the storage unit 63 stores in the storage unit 70 the words that have been subjected to the re-translation process for the dialogue data (translation data) that has been subjected to the re-translation process according to the second embodiment, as shown in FIG.

[0312] Furthermore, the storage unit 63 can store dialogue data specified by the user in the memory unit 70. Specifically, if the user himself / herself realizes that the translation is incorrect, the storage unit 63 can accept a storage command along with the correct translation and store the incorrect translation in association with the correct translation.

[0313] (Learning data generation unit 66) The learning data generation unit 66 generates learning data for training the generation model for translation using translation data, which is translated dialogue data output from the generation model for translation in response to input by the input control unit 13.

[0314] The learning data generation unit 66 generates, as learning data, data associating words and quoted replies (utterances) stored by the storage unit 63. The learning data generation unit 66 also generates, as learning data, data associating inquiries and responses to the inquiries stored by the storage unit 63.

[0315] For example, the learning data generation unit 66 generates learning data by associating information on parts where the intention was not conveyed for each language extracted from the dialogue data or translation data, mistranslated parts, etc. with correct answer data.

[0316] (Study Section 67) The learning unit 67 uses the training data generated by the training data generation unit 66 to train a generative model for translation. Specifically, the learning unit 67 uses at least one of the training data generated by the training data generation unit 66 and predetermined terminology data such as a dictionary file of technical terms and the like preset by the user to perform training so as to improve the accuracy of translation by the generative model for translation, such as reducing mistranslations, replacing technical terms or neologisms with new words, and excluding discriminatory translations.

[0317] Here, an example of a learning process for improving the translation accuracy of a generative model for translation by replacing technical terms or neologisms will be described. The learning unit 67 uses the learning data stored in the storage unit 70 to learn the generative model for translation.

[0318] For example, the word "Ririhan" is a word whose intention can be accurately conveyed only between specific users. For this "Ririhan," data is stored in the storage unit 70, which is associated with the expression that conveys the intention of "Ririhan" based on past dialogue data between users, and the expression is "release judgment." Furthermore, data is stored in the storage unit 70, which is associated with the expression that conveys the intention of "Ririhan," which is a direct translation of "Ririhan," and the expression is "release judgment."

[0319] Using the above-mentioned learning data, the learning unit 67 trains a generative model for translation so that when a Japanese user inputs "Ririhan," it translates it to "release judgment" for an English user instead of "Ririha."

[0320] (Processing Procedure) From here, the processing procedure by the processing system 100 according to the fifth embodiment will be described with reference to Fig. 47. Fig. 47 is a sequence diagram showing the processing procedure according to the fifth embodiment.

[0321] As shown in Fig. 47, for example, the user terminal 20A receives input of dialogue data from the user A (S701). Subsequently, the user terminal 20A transmits the received dialogue data to the generation AI server 40 (S702).

[0322] The generation AI server 40 executes translation processing using the received dialogue data (S703), and then transmits the translated data to the server device 10 (S704).

[0323] The server device 10, which has received the translation data, extracts dialogue history data that satisfies predetermined conditions (S705). Next, the server device 10 generates learning data using the extracted dialogue history data (S706). Then, the server device 10 performs learning using the generated learning data (S707, S708).

[0324] If the learning termination condition is not met (No in S709), the server device 10 continues learning. On the other hand, if the learning termination condition is met (Yes in S709), the server device 10 ends learning.

[0325] The above-mentioned learning termination conditions are not particularly limited, and for example, the server device 10 may determine that the learning conditions are met when the number of learning times for the target generative model reaches a predetermined number of learning times or when the model accuracy of the generative model exceeds a predetermined threshold.

[0326] (effect) Next, the effects of the processing system 100 according to the fifth embodiment will be described. The input control unit 13 of the processing system 100 inputs dialogue data received from a user terminal to a generative model for translation, which is set with a prompt instructing the system to translate the input dialogue data into a specified language with a natural context. The training data generation unit 66 of the processing system 100 generates training data for training the generative model for translation, using translation data that is translated dialogue data output from the generative model for translation in response to input by the input control unit 13. The training unit 67 of the processing system 100 then trains the generative model for translation using the training data generated by the training data generation unit 66. Therefore, the processing system 100 according to the fifth embodiment has the effect of enabling the accuracy of interpretation content to be improved.

[0327] The storage unit 63 receives utterances associated with words designated by the user for the first translation data output from the generative model for translation, and stores the words and utterances in association with each other in the memory unit 70. The training data generation unit 66 then generates, as training data, association data between the words and quoted replies stored by the storage unit 63.

[0328] The storage unit 63 associates an inquiry about a wording included in the dialogue data by the first user with a response to the inquiry by the second user, and stores the associated information in the memory unit 70. The learning data generation unit 66 then generates data in which the inquiry and the response to the inquiry stored by the storage unit 63 are associated with each other, as learning data.

[0329] Therefore, the processing system 100 has the effect of making it easy to train the generative model by automatically generating training data from business dialogue data to be used for additional training to optimize the generative model for each organization and improve translation accuracy.

[0330] The learning unit 67 uses the learning data generated by the learning data generating unit 66 to perform learning so as to improve the accuracy of translation by the generative model for translation.

[0331] Therefore, the processing system 100 can easily perform a learning process to optimize the generative model for each organization by using learning data automatically generated from business dialogue data. As a result, the processing system 100 has the effect of optimizing the generative model for each organization, thereby improving the accuracy of the interpretation content.

[0332] <Modification> Modifications realized by the processing system 100 according to this embodiment will be described below.

[0333] (Data, etc.) The simultaneous interpretation service, retranslation (process / function), quote reply (process / function), translation alert (function), information regarding the agreed matters, agreement document, decisions, Action Items, Q&A list, automatic response to inquiries, learning data, names of functional parts of the processing system 100, steps, processes, names of steps or processes, etc. used in the description of the above embodiments are merely examples and can be changed as desired.

[0334] (About sequence diagrams) The processing steps explained using the sequence diagrams may be interchanged within the scope of not being inconsistent, and some steps may not be implemented. In addition, conjunctions such as "next," "continue," "further," "at this time," and "in this instance" used in the explanation of the sequence diagrams do not limit the order or timing of the processing in the sequence diagrams.

[0335] As described above, the server device 10 can use, as a translation generative model, a generative model in which a prompt is set to instruct the server device 10 to perform at least one of the following operations for input dialogue data: correction of homonyms, correction of grammatical errors, correction of kanji conversion errors, correction of speech recognition errors, correction of non-unique expressions, and correction of expressions that are rude to the dialogue partner. However, this is not limited to this. For example, in addition to the translation generative model, the server device 10 can use a correction generative model that performs the above-mentioned correction process. In this case, the server device 10 can perform the correction process using a correction generative model in which a prompt is set to perform the above-mentioned correction of homonyms, correction of grammatical errors, correction of kanji conversion errors, correction of speech recognition errors, correction of non-unique expressions, and correction of expressions that are rude to the dialogue partner.

[0336] [System Configuration of This Embodiment] The server device 10 is a functional concept and does not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of the functions of the server device 10 is not limited to that shown in the figure, and all or part of it can be functionally or physically distributed or integrated in any unit depending on various loads, usage conditions, etc.

[0337] Furthermore, all or any part of the processes performed by the server device 10 may be realized by a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and a program analyzed and executed by the CPU and the GPU. Furthermore, each process performed by the server device 10 may be realized as hardware using wired logic.

[0338] Furthermore, among the processes described in this embodiment, all or part of the processes described as being performed automatically can be performed manually. Alternatively, all or part of the processes described as being performed manually can be performed automatically using a known method. In addition, the processing procedures, control procedures, specific names, and information including various data and parameters described above and illustrated can be changed as appropriate unless otherwise specified.

[0339] [program] 48 is a diagram showing an example of a computer in which a program is executed to realize the server device 10. The computer 1000 has, for example, a memory 1010 and a CPU 1020. The computer 1000 also has a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.

[0340] The memory 1010 includes a ROM 1011 and a RAM 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1090. The disk drive interface 1040 is connected to a disk drive 1100. A removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to a mouse 1110 and a keyboard 1120, for example. The video adapter 1060 is connected to a display 1130, for example.

[0341] The hard disk drive 1090 stores, for example, an OS (Operating System) 1091, an application program 1092, a program module 1093, and program data 1094. That is, a program that defines each process of the server device 10 is implemented as a program module 1093 in which code executable by the computer 1000 is written. The program module 1093 is stored, for example, in the hard disk drive 1090. For example, a program module 1093 for executing the same process as the functional configuration of the server device 10 is stored in the hard disk drive 1090. The hard disk drive 1090 may be replaced with an SSD (Solid State Drive).

[0342] Furthermore, the setting data used in the processing of the present embodiment described above is stored as program data 1094, for example, in the memory 1010 or the hard disk drive 1090. Then, the CPU 1020 reads the program module 1093 or the program data 1094 stored in the memory 1010 or the hard disk drive 1090 into the RAM 1012 as necessary and executes them.

[0343] The program module 1093 and program data 1094 are not limited to being stored in the hard disk drive 1090, but may also be stored in, for example, a removable storage medium and read by the CPU 1020 via the disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (such as a local area network (LAN) or a wide area network (WAN)). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via the network interface 1070.

[0344] Although the present embodiment applying the invention made by the inventor has been described above, the present invention is not limited by the description and drawings that form part of the disclosure of the present invention according to the present embodiment. In other words, other embodiments, examples, operational techniques, etc. made by those skilled in the art based on the present embodiment are all included in the scope of the present invention. [Explanation of symbols]

[0345] 10 Server device 11 Pre-meeting information acquisition section 12 Prompt Creation Section 13 Input control section 14 Output control section 15 Summary Creation Department 20, 20A~20D User terminal 30A~30D Earphone Microphone 40,50 Generation AI Server 61 Translation Department 62 Consensus Building Department 63 Storage area 64 Display section 65 Inquiry Reception Department 66 Learning data generation unit 67 Learning Department 70 Memory section

Claims

1. an input control unit that inputs voice data or text data received from a user terminal into a generative model for translation, the generative model having a prompt configured to instruct the translation of the input voice data or text data into a specified language in a natural context; a translation unit that inputs first translation data, which is the translated speech data or the text data output from the generative model for translation in response to an input by the input control unit, into the generative model for translation and translates it into second translation data in a language different from the language of the first translation data; an output control unit that outputs the second translation data translated by the translation unit to the user terminal; A processing device comprising:

2. The translation unit When a translation command is received to translate the language of the first translation data translated into a first language into a second language that is a different language, the first translation data is input into the generative model for translation; The output control unit displaying the second translation data translated into the second language, which is output from the generative model for translation in response to the input by the translation unit, alongside the first translation data; 2. The processing device according to claim 1.

3. The output control unit receiving an utterance associated with a word designated by a user for the first translation data output from the generative model for translation; outputting the word and the utterance in association with each other; 2. The processing device according to claim 1.

4. The generative model for translation is a generative model in which a prompt is set to instruct the execution of at least one of the following for the input speech data or text data: correction of homonyms, correction of grammatical errors, correction of kanji conversion errors, correction of speech recognition errors, correction of non-unique expressions, and correction of rude expressions to a conversation partner; 3. The processing apparatus according to claim 1 or 2.

5. determining whether the conference is a specific field or a precision-oriented conference based on pre-conference information regarding the conference in which the user will participate; Based on the determined content, selecting either a first generative model, which is a natural language processing model fine-tuned to a specific field, or a second generative model, which is a large-scale natural language processing model; a generating unit that generates the prompt instructing the selected generative model to translate the input speech data or text data into the language used by each user, 4. The processing apparatus according to claim 1, wherein the processing apparatus is a processing apparatus for processing a substrate.

6. The processing device includes: The user terminal and each server device equipped with the generative model communicate with each other via a communication network related to an Innovative Optical and Wireless Network (IOWN).

4. The processing apparatus according to claim 1, wherein the processing apparatus is a processing apparatus for processing a substrate.

7. A processing method to be executed by a processing device, an input control step of inputting voice data or text data received from a user terminal into a generative model for translation, the generative model having a prompt set therein for instructing the translation of the input voice data or text data into a designated language in a natural context; a translation step of inputting first translation data, which is the translated speech data or the text data output from the generative model for translation in response to input in the input control step, into the generative model for translation and translating it into second translation data in a language different from the language of the first translation data; an output control step of outputting the second translation data translated by the translation step to the user terminal; A processing method comprising:

8. an input control step of inputting voice data or text data received from a user terminal into a generative model for translation, the generative model having a prompt set therein for instructing the translation of the input voice data or text data into a specified language in a natural context; a translation step of inputting first translation data, which is the translated speech data or the text data output from the generative model for translation in response to input in the input control step, into the generative model for translation and translating it into second translation data in a language different from the language of the first translation data; an output control step of outputting the second translation data translated by the translation step to the user terminal; A processing program that causes a computer to execute the above.

Citation Information

Patent Citations

  • Text translation method, device and equipment and readable storage medium

    CN116933807A

  • Text translation method and device, electronic equipment and storage medium

    CN118036619A

  • Voice processing system and terminal device

    JP2014016475A

  • Voice translating device, voice translating method, and voice translating program

    JP2017182394A

  • Web page processing device, web page processing method, and program

    JP2022017985A