Processing device, processing method, and processing program

The processing device adapts generative models for real-time, situation-specific translation in conferences using pre-meeting information, enhancing translation efficiency and reducing costs.

JP2025116742AActive Publication Date: 2025-08-08NTT DOCOMO BUSINESS INC

Patent Information

Application Number
JP2024011350
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-29
Publication Date
2025-08-08
Estimated Expiration
2044-01-29

AI Technical Summary

Technical Problem

Conventional interpretation systems lack the ability to provide optimal translation according to specific situations, failing to adapt to varying needs in conferences or meetings.

Method used

A processing device that utilizes a pre-meeting information acquisition unit to select appropriate generative models like Tsuzumi or ChatGPT based on meeting details, creating prompts for natural context translation, and outputs translated data via a low-latency IOWN network.

Benefits of technology

Enables real-time, situation-specific translation services, improving efficiency and reducing costs by providing optimal translation for multiple users in conferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025116742000001_ABST
    Figure 2025116742000001_ABST
Patent Text Reader

Abstract

To provide optimal translation according to the situation.SOLUTION: A server device 10 has a pre-conference information acquisition unit 11 configured to acquire pre-conference information related to a conference in which a user participates, a prompt generation unit 12 configured to select one of a plurality of generation models on the basis of the pre-conference information, and to generate a prompt that instructs the selected generation model to translate input voice data or text data into a specified language in a natural context, the prompt being added with content based on the pre-conference information, and an input control unit 13 configured to provide the generated prompt to the selected generation model and input voice data or text data received from a user terminal, and an output control unit 14 configured to output voice data or text data generated by the selected generation model to a predetermined terminal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a processing device, a processing method, and a processing program. [Background technology]

[0002] Conventionally, manual interpretation services have involved requesting interpretation from specialized companies for each type of interpretation service, and having interpreters under contract with those companies provide the interpretation to the client.

[0003] In recent years, various systems using IT technology have been provided as interpretation systems, such as automatic translation systems between languages and systems that convert voice and text using voice recognition technology and voice synthesis technology. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Publication No. 2022-003441 [Patent Document 2] Japanese Patent Application Publication No. 2019-139663 Summary of the Invention [Problem to be solved by the invention]

[0005] However, conventional interpretation systems only automate the translation process and are not optimal for each situation.

[0006] The present invention has been made in view of the above, and has as its object to provide a processing device, a processing method, and a processing program that can provide an optimal translation according to the situation. [Means for solving the problem]

[0007] In order to solve the above-mentioned problems and achieve the object, the processing device of the present invention is characterized by having an acquisition unit that acquires pre-meeting information regarding a conference in which a user will participate; a creation unit that selects one of a plurality of generative models based on the pre-meeting information and creates a prompt that instructs the selected generative model to translate input voice data or text data into a specified language with a natural context, the prompt including content based on the pre-meeting information; an input control unit that applies the prompt created by the creation unit to the generative model selected by the creation unit and inputs voice data or text data input from a user terminal; and an output control unit that outputs the voice data or text data output from the generative model selected by the creation unit to a specified terminal. [Effects of the Invention]

[0008] According to the present invention, it is possible to provide an optimum translation depending on the situation. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of a processing system according to an embodiment. [Figure 2] FIG. 2 is a diagram showing an overview of the IOWN technology. [Figure 3] FIG. 3 is a diagram illustrating an example of use of the processing system according to the embodiment. [Figure 4] FIG. 4 is a diagram illustrating the flow of processing in the processing system. [Figure 5] FIG. 5 is a diagram illustrating an outline of the processing of the processing system. [Figure 6] FIG. 6 is a diagram illustrating an example of pre-meeting information. [Figure 7] FIG. 7 is a diagram illustrating step S11 shown in FIG. [Figure 8] FIG. 8 is a diagram illustrating step S12 shown in FIG. [Figure 9] FIG. 9 is a diagram illustrating step S13 shown in FIG. [Figure 10] FIG. 10 is a diagram illustrating step S14 shown in FIG. [Figure 11] FIG. 11 is a diagram illustrating step S15 shown in FIG. [Figure 12] FIG. 12 is a diagram illustrating step S16 shown in FIG. [Figure 13] FIG. 13 is a diagram illustrating step S17 shown in FIG. [Figure 14] FIG. 14 is a diagram illustrating the processing of the processing system. [Figure 15] FIG. 15 is a diagram illustrating another example of the prompt generation process. [Figure 16] FIG. 16 is a diagram illustrating another example of the prompt generation process. [Figure 17] FIG. 17 is a diagram illustrating another example of the prompt generation process. [Figure 18] FIG. 18 is a diagram illustrating another example of use of the processing system. [Figure 19] FIG. 19 is a diagram illustrating another example of use of the processing system. [Figure 20] FIG. 20 is a diagram for explaining an outline of the task management support process of the processing system. [Figure 21] FIG. 21 is an example of a sequence diagram illustrating a processing procedure of the processing method according to the embodiment. [Figure 22] FIG. 22 is a sequence diagram illustrating an example of a processing procedure of another processing method according to the embodiment. [Figure 23] FIG. 23 is a diagram comparing a conventional translation service with a simultaneous interpretation service provided by the processing system according to the embodiment. [Figure 24] FIG. 24 is a diagram comparing a conventional translation service with a simultaneous interpretation service provided by the processing system according to the embodiment. [Figure 25] FIG. 25 is a diagram illustrating an example of a computer that implements a server device by executing a program. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited to this embodiment. In addition, in the description of the drawings, the same parts are designated by the same reference numerals.

[0011] [Embodiment Mode] [Processing System] The following describes the configuration of a processing system according to embodiment 1. The processing system according to the embodiment provides a simultaneous interpretation service in a conference between multiple users, using a generative AI (artificial intelligence) (generative model) to translate the voice data or text data of other users into a language designated by each user, and outputting the translated voice data or text data to a predetermined terminal.

[0012] FIG. 1 is a diagram showing an example of the configuration of a processing system according to an embodiment. As shown in FIG. 1, the processing system 100 according to the embodiment includes user terminals 20A to 20D used by users A to D, who are users of the simultaneous interpretation service, earphone microphones 30A to 30D worn by users A to D and which communicate with the user terminals 20A to 20D, and a cloud server device 10. The server device 10 communicates with a generation AI server 40 and a generation AI server 50. The user terminals 20A to 20D are collectively referred to as user terminals 20. The number of user terminals 20 may be two or more and is not limited to four.

[0013] The generative AI server 40 is equipped with Tsuzumi (registered trademark) 41 (first generative model), a generative AI. Tsuzumi 41 is a natural language processing model fine-tuned to specific fields, such as medicine, semiconductors, IT (Information Technology), academia, factories (plants), law, and office services. Tsuzumi 41 was built with an emphasis on low power consumption and has a faster processing speed than ChatGPT 51 (described below).

[0014] The generation AI server 50 is equipped with ChatGPT (registered trademark) 51 (second generation model), which is a generation AI. ChatGPT 51 is a large-scale natural language processing model that is slower than Tsuzumi 41 but more accurate. Tsuzumi 41 and ChatGPT 51 perform natural language processing on input voice data or text data according to set prompts, and generate and output voice data or text data. Note that the above generation AI is just one example, and additional servers equipped with multiple other generation AIs may be provided.

[0015] The user terminals 20A to 20D are terminal devices that can input and output voice data and text data, and communicate with the server device 10. The user terminals 20A to 20D may input and output voice data via earphone microphones 30A to 30D worn by each of the users A to D.

[0016] The user terminals 20A to 20D are, for example, personal computers (PCs), notebook PCs, tablet terminals, smartphones, etc. The user terminals 20A to 20D start a simultaneous interpretation application to communicate with the server device 10 regarding the simultaneous interpretation process. At this time, the user terminals 20A to 20D accept input of pre-conference information regarding the conference in which the user will participate, through an operation by the user before the conference, and transmit the pre-conference information to the server device 10.

[0017] The server device 10 is realized by loading a predetermined program into a computer or the like including, for example, a ROM (Read Only Memory), a RAM (Random Access Memory), a CPU (Central Processing Unit), etc., and having the CPU execute the predetermined program. The server device 10 also has a communication interface for transmitting and receiving various information to and from other devices (for example, user terminals 20A to 20D, generation AI servers 40, 50) connected via a network or the like.

[0018] In a conference between multiple users, the server device 10 uses Tsuzumi41 or ChatGPT51 to provide a simultaneous interpretation service in which the speech or text of other users is translated into a language specified by each user, and the translated speech or text data is output to a specified terminal.

[0019] Specifically, the server device 10 selects either Tsuzumi41 or ChatGPT51 based on pre-meeting information about the meeting the user will be attending. Then, the server device 10 adds content based on the pre-meeting information to the selected generation AI and sets a prompt to instruct the AI to translate input voice data or text data into a specified language with a natural context.

[0020] This allows the server device 10 to obtain the most suitable translation for the situation from the generating AI and provide it to the user.

[0021] Furthermore, the server device 10 communicates with the user terminals 20A to 20D and the generation AI servers 40 and 50 via a low-latency communication network related to the Innovative Optical and Wireless Network (IOWN) (hereinafter referred to as an IOWN network 60).

[0022] [IOWN Technology Overview] Here, we will explain the IOWN technology. Figure 2 is a diagram showing an overview of the IOWN technology. As shown in Figure 2, the IOWN technology consists of three main technology areas: "All-Photonics Network (APN)," "Digital Twin Computing (DTC)," and "Cognitive Foundation (CF) (registered trademark)."

[0023] [All Photonics Network] The APN related to IOWN technology is a technology that enables the construction of high-speed networks by processing all network transfer functions in the optical domain. Specifically, the APN related to IOWN technology is a technology that realizes low-power, high-quality, large-capacity, and low-latency communications based on optical-based (photonics-based) technologies such as "photonics-electronic convergence technology," "large-capacity optical transmission system and device technology," "optical Ising machine," and "optical lattice clock network."

[0024] [Digital Twin Computing] DTC, which is related to IOWN technology, is a technology that maps individual objects in the real world onto a virtual space using the vast amount of data collected by devices connected to the APN described above.

[0025] Conventional digital twin frameworks are used by mapping individual objects, such as automobiles and robots, into a virtual space, performing analysis and predictions on them, and then mapping the results of the analysis and predictions back onto the real world.

[0026] On the other hand, DTC related to IOWN technology expands on the conventional concept of digital twins, freely combining digital twins of various industries, objects, and people to perform calculations, thereby reproducing with high accuracy the combination of multiple objects, such as people and automobiles in a city. Furthermore, DTC related to IOWN technology enables not only the expression of a person's external appearance, but also the digital expression of their internal state, such as consciousness and thoughts, by combining technologies that enable "speech recognition," "speech synthesis," "understanding of emotions and intentions," etc. to collect information and build a digital twin environment.

[0027] In this way, DTC related to IOWN technology is a technology that enables the creation of digital twins that do not exist in the real world by combining multiple entities that are single in the real world and replicating them as digital twins in a virtual space, or by exchanging or merging some of the components between multiple digital twins.

[0028] [Cognitive Foundation] CF related to IOWN technology is a technology that centrally performs the deployment, configuration, linkage, management, and operation of ICT (Information and Communication Technology) resources at different layers, from the cloud to edge computers, network services, user equipment, etc. Specifically, CF related to IOWN technology treats various targets as a group of virtualized ICT resources, and optimally integrates multiple resources at different layers using multi-orchestration functions as a hub.

[0029] Furthermore, as shown in Figure 2, IOWN technology provides high-value-added services by linking the above-mentioned APN, DTC, and network services provided by operators.

[0030] For example, as shown in Figure 2 (1), IOWN technology provides a technology for transmitting information collected via APN to other terminal devices at high speed and with low latency. Also, as shown in Figure 2 (2), IOWN technology provides a technology for collecting large amounts of information from terminal devices and outputting information such as analysis results from the service provided by the operator at high speed and with low latency in services such as information analysis. Also, as shown in Figure 2 (3), IOWN technology provides a technology for transmitting large amounts of information at high speed and with low latency, using information obtained from surveillance cameras, automobile sensors, etc. to build a digital twin environment, make future predictions, and output the prediction results to the user.

[0031] It is believed that the high-capacity, high-speed, low-latency information transmission infrastructure based on the IOWN technology described above will advance the construction of digital twin environments and the linkage between different digital twin environments.

[0032] The processing system 100 communicates via a large-capacity, high-speed, low-latency information transmission infrastructure based on the above-mentioned IOWN technology. For example, when linking with the processing system 100, a low-latency simultaneous interpretation service is realized using an APN. Specifically, the processing system 100 can provide a simultaneous interpretation service that outputs translated speech data or text data in real time, even when speech data or text data to be translated is input, even when there are three or more users.

[0033] [Server device] 1, the server device 10 will be described. The server device 10 includes a pre-conference information acquisition unit 11 (acquisition unit), a prompt creation unit 12 (creation unit), an input control unit 13, an output control unit 14, and a summary creation unit 15.

[0034] The pre-conference information acquisition unit 11 acquires pre-conference information related to the conference in which the user will participate by communicating with the user terminals 20A to 20D.

[0035] The prompt creation unit 12 selects one of a plurality of generation AIs based on the pre-meeting information. The prompt creation unit 12 selects one of Tsuzumi41 and ChatGPT51 based on the pre-meeting information. The prompt creation unit 12 creates a prompt (command) that instructs the selected generation AI to translate input voice data or text data into a specified language with a natural context, and adds content based on the pre-meeting information.

[0036] The prompt creation unit 12 uses a generation AI (Tsuzumi41 or ChatGPT51) to determine, based on pre-meeting information, for example, a specified language (e.g., the user's language), the accuracy of the meeting, the response speed, the situation of the meeting, the industry related to the meeting, and / or whether the meeting is a specific field or an accuracy-oriented meeting. The prompt creation unit 12 may change the content of the determination depending on the industry, field, members, and situation, without being limited to the above. Furthermore, the server device 10 may determine, according to predetermined rules, the language used by user A, the accuracy of the meeting, the response speed, the situation of the meeting, the industry related to the meeting, and / or whether the meeting is a specific field or an accuracy-oriented meeting.

[0037] The prompt creation unit 12 selects either Tsuzumi41 or ChatGPT51 based on the content determined using the generation AI. For example, if the conference is in a specific field (e.g., medicine), the response speed is set to a relatively fast level, and speed is emphasized, the prompt creation unit 12 selects Tsuzumi41. On the other hand, if accuracy is emphasized, the prompt creation unit 12 selects ChatGPT51.

[0038] The prompt creation unit 12 then adds the content determined using the generation AI to the prompt. Based on the determination, the prompt creation unit 12 selects a generation AI suitable for the conference, adjusts the prompt to be used to instruct the generation AI, and then issues a translation instruction to the selected generation AI. The prompt creation unit 12 outputs the created prompt to the input control unit 13.

[0039] The input control unit 13 gives the prompt created by the prompt creating unit 12 to the generated AI selected by the prompt creating unit 12, and inputs the voice data or text data input from the user terminals 20A to 20D.

[0040] The output control unit 14 outputs the voice data or text data output from the generated AI selected by the prompt creating unit 12 to a predetermined terminal (user terminals 20A to 20D).

[0041] When the conference ends, the summary creation unit 15 causes the generation AI selected by the prompt creation unit 12 to create a summary of the conference contents in the language specified by each of the users A to D. The summary creation unit 15 transmits each summary created by this generation AI to the user terminals 20A to 20D of each user. The transmitted summaries are used by the users A to D for reviewing the conference, task management, etc. By using the generation AI, in addition to creating an automatic summary of the conference contents, the generation AI can also generate automatic Q&A and store it in a database, and it is also possible to provide ChatBot functionality.

[0042] [Application example] 3 is a diagram illustrating an example of use of the processing system 100 according to the embodiment. In FIG. 3, a conference is shown as an example between a user A who speaks Japanese, a user B who speaks English, a user C who speaks Chinese, and a user D who speaks French.

[0043] In this case, the prompt creation unit 12 creates a prompt that instructs the selected generation AI to translate the input voice data or text data into the language used by each user. Specifically, the voice data uttered by other users B, C, and D and to be output to user A is translated into Japanese. The voice data uttered by other users A, C, and D and to be output to user B is translated into English. The voice data uttered by other users A, B, and D and to be output to user C is translated into Chinese. The voice data uttered by other users A, B, and C and to be output to user D is translated into French.

[0044] As a result, the generation AI translates the speech of each of users A, B, C, and D into the language used by each of users A, B, C, and D. The output control unit 14 transmits the speech data or text data translated into the language used by each of users A to D to each of user terminals 20A to 20D used by each user.

[0045] As a result, the translated voice data is output via the user terminals 20A to 20D of the respective users, and may also be output from the earphone microphones 30A to 30D worn by the respective users A, B, C, and D. For example, the speech of users B, C, and D is translated into Japanese and output from the earphone microphone 30A of user A.

[0046] [Processing flow] The process flow of the processing system 100 in the case of Fig. 3 will be described. Fig. 4 is a diagram illustrating the process flow of the processing system 100. Here, the case where speech data of user A is translated for other users B to D will be described as an example.

[0047] User A starts an application (represented as an app in the figure) on user terminal 20A and inputs a summary of the meeting as pre-meeting information before the meeting (step S1). From the pre-meeting information, the generation AI (Tsuzumi41 or ChatGPT51) determines the specified language (language used by user A), the industry related to the meeting, the situation of the meeting, the accuracy of the meeting, the response time (degree of response speed), and / or whether the meeting is a specific field or an accuracy-oriented meeting.

[0048] The server device 10 uses the generation AI (Tsuzumi41 or ChatGPT51) to summarize user A's pre-meeting information and inputs the summarized information into a prompt (instruction) for user A to the generation AI (Tsuzumi41 or ChatGPT51). The server device 10 selects a generation AI appropriate for the meeting based on the language used by user A, the accuracy of the meeting, the degree of response speed, the industry related to the meeting, the situation of the meeting, and / or whether the meeting is in a specific field or an accuracy-oriented meeting, which have been determined by the generation AI using the pre-meeting information, and adjusts the prompt when instructing the generation AI, and then issues a translation instruction to the selected generation AI (step S2). In addition, the server device 10 may determine, according to predetermined rules, for example, the language used by user A, the accuracy of the meeting, the degree of response speed, the situation of the meeting, the industry related to the meeting, and / or whether the meeting is in a specific field or a meeting that emphasizes accuracy, and based on the determination, select a generation AI that is suitable for the meeting, adjust the prompts when instructing the generation AI, and then issue translation instructions to the selected generation AI.

[0049] The server device 10 sets a prompt for user A via an API (Application Programming Interface) and inputs the voice data (which may be text data) sent from the user terminal 20A into the selected generation AI (step S4).

[0050] From the pre-conference information provided by User A (the conversation organizer or administrator), the languages used by Users B to D are determined and input into the prompts for each of Users B to D sent to the generation AI (Tsuzumi41 or ChatGPT51).

[0051] The generation AI (Tsuzumi41 or ChatGPT51) translates the input voice data of user A (step S3) into the languages used by users B, C, and D according to each user's prompts (step S5).

[0052] Specifically, the generation AI translates user A's Japanese speech data into English for user B. The generation AI translates user A's Japanese speech data into Chinese for user C. The generation AI translates user A's Japanese speech data into French for user D.

[0053] The generation AI returns each piece of translated voice data (which may be text data) to the server device 10 (step S6). For example, the generation AI vocalizes the translated data using a synthesized voice.

[0054] Then, the server device 10 transmits voice data obtained by translating the Japanese voice data of user A into English to the user terminal 20B of user B, and causes it to be output (step S7-1). The server device 10 transmits voice data obtained by translating the Japanese voice data of user A into Chinese to the user terminal 20C of user C, and causes it to be output (step S7-2). The server device 10 transmits voice data obtained by translating the Japanese voice data of user A into French to the user terminal 20D of user D, and causes it to be output (step S7-3). Each piece of translation data may be text data.

[0055] Then, when voice data (which may be text data) is input from users B to D, the server device 10 similarly uses the generation AI to translate the voice data into the language used by each user, and outputs the translated voice data (which may be text data) from the corresponding user terminal 20.

[0056] The server device 10 can speed up each of the above steps by performing communication via the low-delay IOWN network 60, and can realize simultaneous interpretation even among three or more users.

[0057] [Processing Overview] An outline of the processing performed by the processing system 100 will now be described. FIG.

[0058] As shown in FIG. 5, the user starts an application on the user terminal 20 and selects a conference room ((1) in FIG. 5).

[0059] Next, before the start of the meeting, the user inputs pre-meeting information by inputting his / her name and a meeting summary ((2), (2-1) in FIG. 5).

[0060] The server device 10 acquires pre-meeting information and, using a generation AI, determines from the pre-meeting information the specified language (e.g., the language used by the user), the accuracy of the meeting, the degree of response speed, the situation of the meeting, the industry related to the meeting, and / or whether the meeting is a specific field or an accuracy-oriented meeting.

[0061] To create the prompt, the server device 10 causes the generation AI to determine the language used by the user from the pre-meeting information (step S11) and create a summary of the meeting information (step S12).Then, the server device 10 causes the generation AI to determine the accuracy of the meeting, the degree of response speed, the situation of the meeting, the industry related to the meeting, and / or whether the meeting is a specific field or an accuracy-oriented meeting based on the created summary (steps S13 to S17).

[0062] The server device 10 selects either Tsuzumi41 or ChatGPT51 based on the content determined using the generation AI. Furthermore, the server device 10 adds this determined content to the prompt and sets the created prompt to the selected generation AI ((3) in FIG. 5). For example, the instruction to Tsuzumi41 is, "Upcoming conversation, ID number, interpret Japanese, casual office setting."

[0063] Once the prompt settings are complete, the conversation begins ((4) in Figure 5). Tsuzumi 41 or ChatGPT 51 translates according to the set prompts (steps S21 and S22), and simultaneous interpretation services are provided to users A, B, C, and D ((5) in Figure 5). Note that in step S22, user X is a collective term for users B, C, and D.

[0064] [Create Prompt] Next, the prompt creation process will be described. Fig. 6 is a diagram showing an example of pre-meeting information. As shown in Fig. 6, for example, the user inputs pre-meeting information C11 such as "Today I will be meeting with Mr. A to discuss whether to continue the service in the future. This is an important meeting."

[0065] 7 to 13 are diagrams for explaining steps S11 to S17 shown in Fig. 5. Fig. 14 is a diagram for explaining the processing of the processing system 100.

[0066] When the server device 10 receives the pre-meeting information C11 from the user terminal 20, the server device 10 inputs the pre-meeting information C11 to the generation AI and determines the user's preferred language (step S11), as shown in Fig. 7. The generation AI returns that the user's preferred language is Japanese (answer A11).

[0067] As shown in FIG. 8, the server device 10 causes the generation AI to create a summary A12 of the pre-meeting information C11 (step S12).

[0068] 9, the server device 10 causes the generation AI to determine whether the accuracy of the meeting is between 60% and 90% from the summary A12 (step S13). The generation AI returns that the accuracy of the meeting is 90% (answer A13).

[0069] 10, the server device 10 causes the generation AI to determine whether the response speed is fast, normal, or slow from the summary A12 (step S14). The generation AI returns that the response speed is normal (answer A14).

[0070] 11, the server device 10 causes the generation AI to determine whether the situation of the meeting is casual, formal, or emergency from the summary A12 (step S15). The generation AI returns that the situation of the meeting is formal (answer A15).

[0071] 12, the server device 10 causes the generation AI to determine from the summary A12 whether the industry related to the conference is medical, semiconductor, service, academic, or factory (step S16). Note that the industry determined by the generation AI is not limited to the above. The generation AI returns that the industry related to the conference is service (answer A16).

[0072] As shown in FIG. 13, the server device 10 causes the generation AI to determine, from the summary A12 and the answers A13 to A16, whether the generation AI suitable for translation is Tsuzumi41, which is in a specific field and prioritizes speed, or ChatGPT51, which prioritizes accuracy (step S17). In addition to automatic selection by the generation AI, the user can also select ChatGPT51 or Tsuzumi41 by setting it themselves. The generation AI returns that the generation AI suitable for translation is ChatGPT51 (answer A17). In response to this, the server device 10 selects ChatGPT51 as the generation AI suitable for the conference.

[0073] 14, the server device 10 sets the prompt P2 including the answers A11, A13 to A16 in the ChatGPT 51, and the conversation between the users starts ((4) in FIG. 14). Once the users start talking, the server device 10 starts a simultaneous interpretation service using the generation AI ((5) in FIG. 14).

[0074] [Other examples of creating prompts] 15 to 17 are diagrams showing another example of the prompt creation process.

[0075] For example, consider the case of pre-meeting information C2 in Figure 15, which states, "I will now talk to the patient. I will explain the details of the illness and the treatment." In this case, the generation AI determines that the language used is Japanese, and based on the created summary, returns answer A2, which indicates that the meeting belongs to the medical field, the accuracy of the meeting is 90%, the response speed is slow, and the situation of the meeting is formal, and the generation AI suitable for translation is Tsuzumi41. A prompt including this answer A2 is set to Tsuzumi41.

[0076] Next, let us consider the case of pre-meeting information C3 in Figure 16, which states, "Today I will be meeting with Mr. A to discuss whether to continue the IT service in the future. This is an important meeting." In this case, the generation AI determines that the language used is Japanese, and based on the created summary, returns answer 3, which indicates that the meeting belongs to the IT service category, the accuracy of the meeting is 85%, the response speed is standard, and the situation of the meeting is formal, and the generation AI suitable for translation is Tsuzumi41. A prompt including this answer A3 is set to Tsuzumi41.

[0077] Next, let us consider the case of pre-meeting information C4 in Figure 17, which is "Meeting with a client. Very important." In this case, the generation AI determines that the language used is Japanese, and based on the created summary, returns answer A4, which indicates that the meeting belongs to office services, the accuracy of the meeting is 90%, the response speed is standard, the situation of the meeting is formal, and the generation AI suitable for translation is Tsuzumi41. A prompt including this answer A4 is set to Tsuzumi41.

[0078] 18 and 19 are diagrams illustrating other examples of use of the processing system 100. Fig. 18 shows an example in which voice data or text data of users A, B, and C is translated into English, which is a common language of users A, B, and C, and the translated voice data or text data is output to terminal 20E.

[0079] In this case, the prompt creation unit 12 creates a prompt that instructs the selected generation AI to translate the input voice data or text data into the common language of each user. The output control unit 14 transmits the voice data or text data corresponding to the common language (English) to a predetermined terminal (e.g., terminal 20E). Alternatively, the output control unit 14 may transmit the voice data or text data corresponding to the common language (English) to the user terminal used by each user. At this time, the server device 10 may automatically recognize the user using the AI and start interpretation using the generation AI.

[0080] FIG. 19 illustrates an example in which voice data of Japanese-speaking user A is translated into English, which is the language of user B, and the translated voice data (text data) is output from user terminal 20B in speaker mode. In this case, prompt creation unit 12 creates a prompt that instructs the selected generation AI to translate the input Japanese voice data or text data of user A (first user) into the language (e.g., English) used by user B (second user). Output control unit 14 transmits the voice data or text data translated into user B's language (English) to user terminal 20B. This is applicable to conversations at, for example, police stations, pharmacists, front desks, banks, ticket counters, kiosks, etc.

[0081] [Task management support] Next, a description will be given of the task management support process performed by the server device 10. FIG.

[0082] After the conference is over, the server device 10 causes the generation AI selected for the conference to automatically create a summary of the contents of the conference in the language designated by each of the users A to D ((6-1) in FIG. 20).

[0083] The server device 10 requests the generation AI selected for the meeting to create a summary of the meeting (step S26), create an Action Item (A / I) (step S27), and compile the summary on one slide (step S28). Note that this slide may contain not only a simple text summary but also an image. The server device 10 then saves the content summary, A / I, and slide created by the generation model AI in the cloud and in an application ((6-2) in Figure 20). By checking these, the user can review the meeting, manage tasks, etc. ((6) in Figure 20).

[0084] [Simultaneous interpretation processing] Next, the processing procedure of the processing method according to the embodiment will be described. Fig. 21 is an example of a sequence diagram showing the processing procedure of the processing method according to the embodiment. Fig. 21 shows the processing procedure of simultaneous interpretation processing in which voice data or text data from user A is translated into the language used by user B and output from user terminal 20B of user B, and also translated into the language used by user C and output from user terminal 20C of user C.

[0085] 21, for example, when an application is started on user terminal 20A (step S31), information about conference rooms in which user A can participate is output to user terminal 20A through communication between server device 10 and user terminal 20A (step S32). User A operates user terminal 20A to select a conference room to participate in (steps S33 and S34), and then inputs a conference outline as pre-conference information before the conference (step S35). User terminal 20A transmits the input pre-conference information to server device 10 (step S36).

[0086] The server device 10 uses the generation AI to determine the language used by user A from the pre-meeting information (step S37). The server device 10 uses the generation AI to create a summary of the pre-meeting information (step S38). The server device 10 uses the generation AI to determine, based on the summary, the accuracy of the meeting, the degree of response speed, the situation of the meeting, the industry related to the meeting, and / or whether the meeting is a specific field or an accuracy-oriented meeting (step S39).

[0087] Based on the determination made in step S39, the server device 10 selects either Tsuzumi41 or ChatGPT51, adds the determination to the prompt, and sets the created prompt to the selected generation AI (steps S40 to S42). The example in Fig. 21 shows a case where Tsuzumi41 is selected and a prompt is set. When ChatGPT51 is selected, the prompt is set to ChatGPT51, and ChatGPT51 is caused to perform the translation process (step S48) described below.

[0088] When the server device 10 receives the prompt setting notification (step S43), it transmits to the user terminal 20A a message indicating that conversation can be started (step S44).

[0089] Then, when the voice data or text data input by user A is sent from user terminal 20A (steps S45, S46), server device 10 inputs the voice data or text data input by user A to Tsuzumi41 (step S47) and has it translated into the languages used by users B and C, respectively (step S48).

[0090] The server device 10 transmits the voice data or text output from Tsuzumi 41 translated into the language used by user B (step S49) to user terminal 20B (step S50) and causes it to be output (step S53). The server device 10 transmits the voice data or text output from Tsuzumi 41 translated into the language used by user C (step S51) to user terminal 20C (step S52) and causes it to be output (step S54).

[0091] The same applies not only to user A, but also to cases where voice data or text data from users A, B, and C is input to server device 10. Server device 10 inputs the voice data or text data from users A, B, and C to Tsuzumi 41, and then outputs the voice data or text data translated into the languages used by users A, B, and C output from Tsuzumi 41 to user terminals 20A to 20C of users A to C.

[0092] [Task management support processing] Next, the task management support process will be described. Fig. 22 is a sequence diagram showing an example of a processing procedure of another processing method according to the embodiment.

[0093] 22, the server device 10 determines whether the conference of users A to C has ended (step S61). If the conference has not ended (step S61: No), the process returns to step S61. If the conference has ended (step S61: Yes), the server device 10 causes the generation AI selected for the conference to create a summary of the contents of the conference in the language used by each of users A to D (steps S62 and S63).

[0094] The server device 10 stores each summary output from the generation AI (steps S64, S65), and transmits the summary in the corresponding language to each of the user terminals 20A to 20C (steps S66 to S68).

[0095] [Effects of the embodiment] 23 and 24 are diagrams comparing a conventional translation service with the simultaneous interpretation service provided by the processing system according to the embodiment.

[0096] As shown in Figure 23, in conventional translation services that translate a user's Japanese voice data and output it as text data, it takes time for the translation to be output. In contrast, in the simultaneous interpretation service of the embodiment, each user uses their own terminal, which results in extra communication between the terminal and the server. However, since the IOWN network 60, which has high capacity, high speed, and low latency, is used, this delay can be almost ignored, resulting in faster dialogue. Note that the network used is not limited to the IOWN network 60, and other networks may also be used.

[0097] Therefore, for example, in the simultaneous interpretation service of the embodiment, the second dialogue can be completed by the time t1 when the first dialogue ends in the conventional translation service.

[0098] As described above, according to the embodiment, there is no need to wait for translation, which relieves stress during conversation. As a result, according to the embodiment, business efficiency can be improved and costs can be reduced. Furthermore, according to the embodiment, translated voice data is output, so the user only needs to listen to the voice data and does not need to read the text, which is highly convenient. Of course, the input and output may be either voice data or text data.

[0099] Furthermore, the processing system 100 communicates via the large-capacity, high-speed, low-latency IOWN network 60. Therefore, the processing system 100 can provide a simultaneous interpretation service that outputs translated audio data or text data in real time, even between three or more users.

[0100] In the processing system 100, even if there are, for example, four people in a conference, the voice data of one user is translated into each user's language and output from each user's user terminal 20. The simultaneous interpretation service in the embodiment can realize simultaneous interpretation even among three or more multiple users.

[0101] Furthermore, in the embodiment, simultaneous interpretation is performed by selecting either Tsuzumi41 or ChatGPT51 based on pre-conference information about the conference in which the user is participating, so that appropriate translation can be performed according to the conference.

[0102] In addition, in the embodiment, the server device 10 adds content based on the pre-conference information to the selected generation AI, and sets a prompt to instruct the selected generation AI to translate the input voice data or text data into a specified language with a natural context. Therefore, according to the embodiment, the optimal translation according to the situation can be obtained and provided to the user.

[0103] In this embodiment, Tsuzumi41 and ChatGPT51 have been used as examples of the generation AIs to be used, but other generation AIs may also be used, and the number of generation AIs is not limited to two, but may be any one of three or more generation AIs.

[0104] [System configuration of the embodiment] The server device 10 is a functional concept and does not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of the functions of the server device 10 is not limited to that shown in the figure, and all or part of it can be functionally or physically distributed or integrated in any unit depending on various loads, usage conditions, etc.

[0105] Furthermore, all or any part of the processes performed by the server device 10 may be realized by a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and a program analyzed and executed by the CPU and the GPU (Graphics Processing Unit). Furthermore, each process performed by the server device 10 may be realized as hardware using wired logic.

[0106] Furthermore, among the processes described in the embodiments, all or part of the processes described as being performed automatically can be performed manually. Alternatively, all or part of the processes described as being performed manually can be performed automatically using a known method. In addition, the processing procedures, control procedures, specific names, and information including various data and parameters described above and illustrated can be changed as appropriate unless otherwise specified.

[0107] [program] 25 is a diagram showing an example of a computer that executes a program to implement the server device 10. The computer 1000 includes, for example, a memory 1010 and a CPU 1020. The computer 1000 also includes a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.

[0108] The memory 1010 includes a ROM 1011 and a RAM 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1090. The disk drive interface 1040 is connected to a disk drive 1100. A removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to a mouse 1110 and a keyboard 1120, for example. The video adapter 1060 is connected to a display 1130, for example.

[0109] The hard disk drive 1090 stores, for example, an OS (Operating System) 1091, an application program 1092, a program module 1093, and program data 1094. That is, a program that defines each process of the server device 10 is implemented as a program module 1093 in which code executable by the computer 1000 is written. The program module 1093 is stored, for example, in the hard disk drive 1090. For example, a program module 1093 for executing the same process as the functional configuration of the server device 10 is stored in the hard disk drive 1090. The hard disk drive 1090 may be replaced with an SSD (Solid State Drive).

[0110] Furthermore, setting data used in the processing of the above-described embodiment is stored as program data 1094, for example, in memory 1010 or hard disk drive 1090. Then, CPU 1020 reads program module 1093 and program data 1094 stored in memory 1010 or hard disk drive 1090 into RAM 1012 as necessary and executes them.

[0111] The program module 1093 and program data 1094 are not limited to being stored in the hard disk drive 1090, but may also be stored in, for example, a removable storage medium and read by the CPU 1020 via the disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (such as a local area network (LAN) or a wide area network (WAN)). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via the network interface 1070.

[0112] Although the present invention has been described above as an embodiment, the present invention is not limited to the descriptions and drawings that form part of the disclosure of the present invention. In other words, other embodiments, examples, and operational techniques that can be made by those skilled in the art based on the present invention are all included in the scope of the present invention. [Explanation of symbols]

[0113] 10 Server device 11 Pre-meeting information acquisition section 12 Prompt Creation Section 13 Input control section 14 Output control section 15 Summary Creation Department 20, 20A~20D User terminal 30A~30D Earphone Microphone 40,50 Generation AI Server

Claims

1. an acquisition unit that acquires pre-conference information regarding a conference in which the user will participate; a generating unit that generates a prompt that selects one of a plurality of generative models based on the pre-meeting information or a predetermined rule, and instructs the selected generative model to translate input speech data or text data into a specified language with a natural context, the prompt including content based on the pre-meeting information; and an input control unit that provides the prompt created by the creation unit to the generative model selected by the creation unit and inputs voice data or text data input from a user terminal; an output control unit that outputs the voice data or text data output from the generative model selected by the creation unit to a predetermined terminal; A processing device comprising:

2. The processing device according to claim 1, characterized in that the creation unit determines, based on the pre-meeting information, the specified language, the accuracy of the meeting, the degree of response speed, the status of the meeting, the industry related to the meeting, and / or whether the meeting is in a specific field or a meeting that emphasizes accuracy, and, based on the determined content, selects either a first generative model which is a natural language processing model fine-tuned to a specific field, or a second generative model which is a large-scale natural language processing model.

3. The processing device according to claim 1, characterized in that the creation unit determines, based on the pre-meeting information, the specified language, the accuracy of the meeting, the degree of response speed, the status of the meeting, the industry related to the meeting, and / or whether the meeting is in a specific field or a meeting that emphasizes accuracy, and adds the determined content to the prompt.

4. the creation unit creates the prompt instructing the selected generative model to translate the input speech data or text data into the language used by each user, 2. The processing device according to claim 1, wherein the output control unit transmits, to each terminal used by each user, voice data or text data translated into the language used by each user.

5. the creation unit creates the prompt instructing the selected generative model to translate the input speech data or text data into a common language of each user; 2. The processing device according to claim 1, wherein the output control unit transmits the translated voice data or text data in the common language to each terminal used by each user or to a predetermined terminal.

6. the creation unit creates the prompt instructing the selected generative model to translate the input voice data or text data of the first user into a language used by the second user; The processing device according to claim 1 , wherein the output control unit transmits the translated voice data or text data in the language used by the second user to a terminal used by the second user.

7. a summary creation unit that, when the conference ends, causes the generative model selected by the creation unit to create a summary of the contents of the conference in the language specified by each user, and transmits each summary created by the generative model to each user; 2. The processing apparatus of claim 1, further comprising:

8. The processing device according to claim 1, characterized in that the processing device communicates with the user terminal and each server device equipped with the plurality of generative models via a communication network related to an Innovative Optical and Wireless Network (IOWN).

9. A processing method executed by a processing device, obtaining pre-conference information regarding a conference in which the user will be participating; a step of selecting one of a plurality of generative models based on the pre-meeting information or a predetermined rule, and creating a prompt that instructs the selected generative model to translate the input speech data or text data into a specified language in a natural context, the prompt including content based on the pre-meeting information; a step of providing the prompt created in the creating step to the generative model selected in the creating step, and inputting voice data or text data input from a user terminal; a step of outputting the speech data or text data output from the generative model selected in the creating step to a predetermined terminal; A processing method comprising:

10. obtaining pre-conference information regarding a conference in which the user will participate; creating a prompt that selects one of a plurality of generative models based on the pre-meeting information or a predetermined rule, and instructs the selected generative model to translate the input speech data or text data into a specified language in a natural context, the prompt including content based on the pre-meeting information; a step of providing the prompt created in the creating step to the generative model selected in the creating step, and inputting voice data or text data input from a user terminal; a step of outputting the speech data or text data output from the generative model selected in the creating step to a predetermined terminal; A processing program that causes a computer to execute the above.

Citation Information

Patent Citations

  • Conference control method and device, online conference system, equipment and medium

    CN116546159A

  • Audio and video translator

    US20230088322A1

  • Systems, methods, and apparatus for virtual meetings

    US20230254412A1

  • Recommendation device, recommendation method and recommendation program

    JP2019139663A

  • Multilingual asynchronous translation system, multilingual asynchronous translation method and program

    JP2022003441A

Cited By

  • Treatment support system

    JP7835948B1