Processing device, processing method, and processing program

The processing system addresses language barriers in emergencies by using AI translation to connect users to emergency services and provide real-time interpretation, ensuring effective communication.

JP2025116534AActive Publication Date: 2025-08-08NTT DOCOMO BUSINESS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024011017
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-29
Publication Date
2025-08-08
Estimated Expiration
2044-01-29

AI Technical Summary

Technical Problem

In emergency situations, foreign users often struggle to communicate their situation effectively to emergency services due to language barriers, even with conventional translation services, as these services are limited in their applicability and do not support emergency calls.

Method used

A processing system that includes a server device capable of determining the user's language and emergency destination, using generative AI models like Tsuzumi and ChatGPT to translate voice or text data in real-time, enabling simultaneous interpretation between the user and emergency responders via a high-capacity, low-latency network.

Benefits of technology

Facilitates immediate and appropriate communication with emergency services by automatically connecting users to the right destination and providing real-time interpretation, ensuring effective communication of the emergency situation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025116534000001_ABST
    Figure 2025116534000001_ABST
Patent Text Reader

Abstract

To dial an emergency contact number corresponding to the emergency situation in the event of an emergency for a user, and to enable the provision of appropriate interpretation services according to the emergency situation.SOLUTION: A server device 10 receives an emergency call request from a user terminal, determines a user's language and a call destination of the user terminal on the basis of acquired emergency situation description information, and connects the user terminal and the call destination for communication. The server device 10 sets a prompt to instruct a generative AI to translate input voice data or text data into a natural context interpretation language. The server device 10 instructs the generative AI to translate voice data input from a user terminal 20 into the language used at the call destination, and outputs the voice data output from the generative AI to the call destination. The server device 10 instructs the generative AI to translate the voice data input from the call destination into the user's language and outputs the voice data output from the generative AI to the user terminal 20.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a processing device, a processing method, and a processing program. [Background technology]

[0002] Conventionally, manual interpretation services have involved requesting interpretation from specialized companies for each type of interpretation service, and having interpreters under contract with those companies provide the interpretation to the client.

[0003] In recent years, various types of network-based systems have been provided as interpretation systems, such as automatic translation systems between languages and systems that convert speech to text using speech recognition technology. For example, a translation service has been provided that translates a user's Japanese speech data and outputs it as text data. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Publication No. 2022-003441 [Patent Document 2] Japanese Patent Application Publication No. 2019-139663 Summary of the Invention [Problem to be solved by the invention]

[0005] In the event of an emergency, a foreign user may not know who to call if they want to call an emergency contact in the country they are visiting. Even if they know the contact information, if the user's language differs from the language of the contact, they often cannot properly communicate the emergency situation, even with the use of conventional translation services.

[0006] As such, conventional translation services were limited in the situations in which they could be provided and did not support emergency calls.

[0007] The present invention has been made in consideration of the above, and aims to provide a processing device, processing method, and processing program that, in an emergency, makes a call to an emergency contact number that corresponds to the emergency situation, and enables the provision of an appropriate interpretation according to the emergency situation. [Means for solving the problem]

[0008] In order to solve the above-mentioned problems and achieve the object, the processing device of the present invention is characterized by having an acquisition unit that receives a request for an emergency call from a user terminal and acquires explanatory information explaining an emergency situation from the user terminal; a determination unit that determines the language used by the user based on the explanatory information; a determination unit that determines a call destination for the user terminal based on the explanatory information and connects the user terminal and the call destination so that a call can be made; a setting unit that sets a prompt that instructs a generative model to translate input voice data or text data into an interpreter language with a natural context; and an input / output control unit that causes the generative model to translate the voice data or text data input from the user terminal into the language used at the call destination and outputs the voice data or text data output from the generative model to the call destination, and causes the generative model to translate the voice data or text data input from the call destination into the language used by the user and outputs the voice data or text data output from the generative model to the user terminal. [Effects of the Invention]

[0009] According to the present invention, in an emergency, a call is made to an emergency contact number that can handle the emergency situation, and appropriate interpretation can be provided according to the emergency situation. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of a processing system according to an embodiment. [Figure 2] FIG. 2 is a diagram showing an overview of the IOWN technology. [Figure 3]FIG. 3 is a diagram illustrating an outline of the processing of the processing system. [Figure 4] FIG. 4 is a diagram illustrating the flow of processing in the processing system. [Figure 5] FIG. 5 is an example of a sequence diagram showing a processing procedure of the processing method according to the embodiment. [Figure 6] FIG. 6 is a diagram illustrating an emergency call service provided by the processing system according to the embodiment. [Figure 7] FIG. 7 is a diagram illustrating an example of a computer that implements a server device by executing a program. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited to this embodiment. In addition, in the description of the drawings, the same parts are designated by the same reference numerals.

[0012] [Embodiment Mode] [Processing System] The configuration of a processing system according to an embodiment will be described. The processing system according to the embodiment provides an emergency call service for when a user, in a country where a language other than the user's native language is used, wishes to call an emergency contact number corresponding to the emergency situation in the event of an emergency and when the user wishes to explain the emergency situation over the call. The emergency call service enables the user's terminal to make a call to the emergency contact number corresponding to the emergency situation in the event of an emergency, and provides appropriate mutual interpretation according to the emergency situation.

[0013] 1 is a diagram illustrating an example of the configuration of a processing system according to an embodiment. As shown in FIG. 1, a processing system 100 according to the embodiment includes a user terminal 20 used by a user who is a user of an emergency call service, and a cloud server device 10.

[0014] In the event of a user emergency, the server device 10 provides an emergency call service that determines the destination of the user's call and interprets the conversation between the user and the recipient of the emergency call. The server device 10 is realized, for example, by loading a predetermined program into a computer or the like including a ROM (Read Only Memory), RAM (Random Access Memory), CPU (Central Processing Unit), etc., and having the CPU execute the predetermined program. The server device 10 also has a communication interface for transmitting and receiving various information to and from other devices (e.g., a user terminal 20, generation AI servers 40, 50) connected via a network or the like.

[0015] The generative AI server 40 is equipped with Tsuzumi (registered trademark) 41 (first generative model), a generative AI (Artificial Intelligence) (generative model). Tsuzumi 41 is a natural language processing model fine-tuned to specific fields, such as medicine, semiconductors, IT (Information Technology), academia, factories (plants), law, and office services. Tsuzumi 41 was built with an emphasis on low power consumption and has a faster processing speed than ChatGPT 51 (described below).

[0016] The generation AI server 50 is equipped with ChatGPT (registered trademark) 51 (second generation model), which is a generation AI. ChatGPT 51 is a large-scale natural language processing model that is slower than Tsuzumi 41 but has higher accuracy. Tsuzumi 41 and ChatGPT 51 perform natural language processing on input voice data according to set prompts, generate voice data, and output it. The input and output of Tsuzumi 41 and ChatGPT 51 may be text data. Note that the above generation AI is an example, and additional servers equipped with multiple other generation AIs may be provided.

[0017] The user terminal 20 is a terminal device that can input voice data and text data and output voice data and text data, and communicates with the server device 10. The user terminal 20 has a call function. The user terminal 20 is, for example, a smartphone.

[0018] The user receives an emergency call service by activating an emergency call application on the user terminal 20. The user terminal 20 receives input of explanatory information (emergency situation explanatory information) that explains the user's emergency situation through speech from the user. The user terminal 20 transmits the emergency situation explanatory information to the server device 10. Thereafter, under the control of the server device 10, a call can be made between the user terminal 20 and the emergency destination terminal 30 according to the emergency situation. Here, the user can receive a mutual interpretation service in the conversation with the recipient of the emergency destination.

[0019] The server device 10 determines the call destination of the user in an emergency situation based on the emergency situation explanation information. The server device 10 connects the user terminal 20 and emergency destination terminals 30A and 30B so that they can communicate with each other via a Voice over Internet Protocol (VoIP) telephone line and a Public Switched Telephone Network (PSTN). Emergency destinations include, for example, a hospital (for requesting medical attention in the event of injury or sudden illness), a fire station (for requesting an ambulance in the event of injury or sudden illness), a police station (for the event of being involved in an incident such as robbery), and a government office (for applying for residence). The emergency destination terminals 30A and 30B are collectively referred to as emergency destination terminal 30. The number of emergency destinations is not limited to two.

[0020] Furthermore, the server device 10 provides the user and the recipient of the emergency call with an emergency call service that simultaneously translates the conversation between the user and the recipient of the emergency call. The server device 10 uses a generation AI to determine emergency contacts and provide simultaneous translation of the conversation between the user and the recipient of the emergency call. The server device 10 uses Tsuzumi41 or ChatGPT51 to provide the emergency call service.

[0021] In this way, in the event of a user emergency, the server device 10 makes a call to an emergency contact person corresponding to the emergency situation, and provides appropriate real-time mutual interpretation according to the emergency situation.

[0022] In addition, the server device 10 communicates with the user terminal 20, the emergency call destination terminal 30, and the generation AI servers 40 and 50 via a low-latency communication network related to IOWN (Innovative Optical and Wireless Network) (hereinafter, IOWN network 60).

[0023] [IOWN Technology Overview] Here, we will explain the IOWN technology. Figure 2 is a diagram showing an overview of the IOWN technology. As shown in Figure 2, the IOWN technology consists of three main technology areas: "All-Photonics Network (APN)," "Digital Twin Computing (DTC)," and "Cognitive Foundation (CF) (registered trademark)."

[0024] [All Photonics Network] The APN related to IOWN technology is a technology that enables the construction of high-speed networks by processing all network transfer functions in the optical domain. Specifically, the APN related to IOWN technology is a technology that realizes low-power, high-quality, large-capacity, and low-latency communications based on optical-based (photonics-based) technologies such as "photonics-electronic convergence technology," "large-capacity optical transmission system and device technology," "optical Ising machine," and "optical lattice clock network."

[0025] [Digital Twin Computing] DTC, which is related to IOWN technology, is a technology that maps individual objects in the real world onto a virtual space using the vast amount of data collected by devices connected to the APN mentioned above.

[0026] Conventional digital twin frameworks are used by mapping individual objects, such as automobiles and robots, into a virtual space, performing analysis and predictions on them, and then mapping the results of the analysis and predictions back onto the real world.

[0027] On the other hand, DTC related to IOWN technology expands on the conventional concept of digital twins, freely combining digital twins of various industries, objects, and people to perform calculations, thereby reproducing with high accuracy the combination of multiple objects, such as people and automobiles in a city. Furthermore, DTC related to IOWN technology enables not only the expression of a person's external appearance, but also the digital expression of their internal state, such as consciousness and thoughts, by combining technologies that enable "speech recognition," "speech synthesis," "understanding of emotions and intentions," etc. to collect information and build a digital twin environment.

[0028] In this way, DTC related to IOWN technology is a technology that enables the creation of digital twins that do not exist in the real world by combining multiple entities that are single in the real world and replicating them as digital twins in a virtual space, or by exchanging or merging some of the components between multiple digital twins.

[0029] [Cognitive Foundation] CF related to IOWN technology is a technology that centrally performs the deployment, configuration, linkage, management, and operation of ICT (Information and Communication Technology) resources at different layers, from the cloud to edge computers, network services, user equipment, etc. Specifically, CF related to IOWN technology treats various targets as a group of virtualized ICT resources, and optimally integrates multiple resources at different layers using multi-orchestration functions as a hub.

[0030] Furthermore, as shown in Figure 2, IOWN technology provides high-value-added services by linking the above-mentioned APN, DTC, and network services provided by operators.

[0031] For example, as shown in Figure 2 (1), IOWN technology provides a technology for transmitting information collected via APN to other terminal devices at high speed and with low latency. Also, as shown in Figure 2 (2), IOWN technology provides a technology for collecting large amounts of information from terminal devices and outputting information such as analysis results from the service provided by the operator at high speed and with low latency in services such as information analysis. Also, as shown in Figure 2 (3), IOWN technology provides a technology for transmitting large amounts of information at high speed and with low latency, using information obtained from surveillance cameras, automobile sensors, etc. to build a digital twin environment, make future predictions, and output the prediction results to the user.

[0032] It is believed that the high-capacity, high-speed, low-latency information transmission infrastructure based on the IOWN technology described above will advance the construction of digital twin environments and the linkage between different digital twin environments.

[0033] The processing system 100 communicates via a large-capacity, high-speed, low-latency information transmission infrastructure based on the above-mentioned IOWN technology. For example, when linking with the processing system 100, an APN is used to realize a low-latency emergency call service. In other words, the processing system 100 can provide an emergency call service that outputs interpreted speech in real time, even when speech to be interpreted is input.

[0034] [Server device] 1, the server device 10 will be described. The server device 10 includes an emergency situation acquisition unit 11 (acquisition unit), a language used determination unit 12 (determination unit), a generation AI selection unit 13 (selection unit), a destination determination unit (determination unit) 14, a prompt setting unit 15 (setting unit), and an input / output control unit 16.

[0035] The emergency situation acquisition unit 11 acquires emergency situation explanation information from the user terminal 20 through communication with the user terminal 20. The emergency situation explanation information is acquired from the speech of a user who uses the emergency call service.

[0036] The language used determination unit 12 determines the language used by the user based on the emergency situation explanation information. For example, the language used determination unit 12 determines the language used by the user using a generation AI (for example, ChatGPT51).

[0037] The generation AI selection unit 13 selects one of a plurality of generation AIs, which are natural language processing models, based on the emergency situation description information. For example, the generation AI selection unit 13 determines, based on the emergency situation description information, whether the user's emergency situation is a situation where speed is important and the situation is in a specific field, or a situation where accuracy is important, and selects one of a plurality of generation AIs based on the determination. The generation AI selection unit 13 may change the determination content of the generation AI depending on the time of year, the time zone, and the user's situation, without being limited to the above. The generation AI selection unit 13 uses a generation AI (e.g., ChatGPT51) to determine whether the user's emergency situation is a situation where speed is important and the situation is in a specific field, or a situation where accuracy is important. Furthermore, the generation AI selection unit 13 may determine, according to a predetermined rule, whether the user's emergency situation is a situation where speed is important and the situation is in a specific field, or a situation where accuracy is important, and select one of a plurality of generation AIs based on the determination content.

[0038] The generation AI selection unit 13 selects either Tsuzumi41 or ChatGPT51 based on the emergency situation description information. If the user's emergency situation is one that prioritizes speed and is in a specific field, the generation AI selection unit 13 selects Tsuzumi41. For example, if the user's emergency situation is an illness and is highly urgent, the generation AI selection unit 13 selects Tsuzumi41, which specializes in the medical field. Furthermore, if the user's emergency situation is one that prioritizes accuracy, the generation AI selection unit 13 selects ChatGPT51.

[0039] The destination decision unit 14 decides the destination of the user terminal 20 based on the emergency situation explanation information, and connects the user terminal 20 and the emergency destination terminal 30 so that they can communicate with each other.

[0040] The prompt setting unit 15 sets a prompt that instructs the generation AI to translate input voice data into an interpreter language with a natural context. The prompt setting unit 15 sets a prompt to the generation AI selected by the generation AI selection unit 13. In this case, the prompt setting unit 15 creates a prompt that instructs the user to translate the user's voice data from the user's language into the language (e.g., Japanese) of the recipient of the emergency call destination, and to translate the voice data input from the emergency call destination terminal 30 from the language of the recipient of the emergency call destination into the user's language. In this way, the prompt setting unit 15 also adjusts the prompt when instructing the selected generation AI, and then sets the prompt to the generation AI selected by the generation AI selection unit 13.

[0041] In other words, if the input voice data is voice data input from the user terminal 20, the prompt commands that the voice data be translated from the language used by the user into the language used by the recipient of the emergency destination (the user of the emergency destination terminal 30). Also, if the input voice data is voice data input from the emergency destination terminal 30, the prompt commands that the voice data be translated from the language used by the recipient of the emergency destination into the language used by the user.

[0042] The input / output control unit 16 causes the generation AI to translate voice data input from the user terminal 20 into the language used at the destination, and outputs the voice data output from the generation model to the destination emergency destination terminal 30. The input / output control unit 16 causes the generation AI to translate voice data input from the destination emergency destination terminal 30 into the language used by the user, and outputs the voice data output from the generation AI to the user terminal 20. The generation AI is the generation AI (Tsuzumi41 or ChatGPT51) selected by the selection unit. Furthermore, the input / output data of the generation AI is not limited to voice data, and may be text data.

[0043] [Processing Overview] An overview of the processing of the processing system 100 will be described below. Fig. 3 is a diagram illustrating an overview of the processing of the processing system 100. For example, an example will be described in which Japanese is used at hospitals, fire stations, and police stations that may be emergency call sources, and the user uses a language other than Japanese (a foreign language).

[0044] When an emergency situation occurs, the user starts an emergency call application (referred to as an application in the drawing) from the user terminal 20 and connects to the server device 10 (step S1).

[0045] The server device 10 receives a request for an emergency call from the user terminal 20, receives emergency situation explanation information from the user terminal 20 in response to the user's speech, and performs the following processes (step S2).

[0046] The server device 10 uses a generating AI (for example, ChatGPT51) to determine the language used by the user from the emergency situation explanation information ((1) in FIG. 3).

[0047] The server device 10 determines whether the situation is one in which speed is important and the situation is in a specific field, or one in which accuracy is important, based on the emergency situation explanation information ((2) in FIG. 3). If the situation is one in which speed is important and the situation is in a specific field, the server device 10 selects Tsuzumi 41. If the situation is one in which accuracy is important, the server device 10 selects ChatGPT 51.

[0048] Based on the emergency situation explanation information, the server device 10 determines a call destination (call destination) corresponding to the user's emergency situation ((3) in FIG. 3). For example, if the user is injured or suddenly ill and wishes to be examined, a nearby hospital that has a medical department that corresponds to the user's symptoms and is available to receive treatment at that time is determined as the emergency call destination. If the user is involved in an incident such as a robbery, the nearest police station is determined as the emergency call destination. If the user discovers a fire, the nearest fire station is determined as the emergency call destination. The server device 10 connects the user terminal 20 and the determined emergency call destination terminal 30 so that communication can be carried out.

[0049] The server device 10 corrects and adjusts the prompt for the selected generation AI ((4) in FIG. 3). At this time, the server device 10 creates a prompt instructing the translation of the user's voice data from the user's language into the language (e.g., Japanese) of the recipient of the emergency call, and the translation of the voice data input from the emergency call destination terminal 30 from the language of the recipient of the emergency call into the user's language.

[0050] The server device 10 sets the created prompt in the generation AI and inputs the user's voice data to the selected generation AI (step S3). When the server device 10 receives the translated voice data (Japanese) output from the generation AI (step S4), it outputs it to the emergency destination terminal 30 via the VoIP telephone line and PSTN (steps S5 and S6).

[0051] [Processing flow] Next, a description will be given of the processing flow of the processing system 100. Fig. 4 is a diagram illustrating the processing flow of the processing system 100. Fig. 4 is a diagram showing an example.

[0052] When an emergency situation arises, the user launches the emergency call application (referred to as an app in the figure) on the user terminal 20 ((1) in Figure 4), clicks button B11 displayed on the screen of the user terminal 20, and explains the emergency situation ((2) in Figure 4).

[0053] The cloud server device 10 receives the emergency situation explanation information ((3) in FIG. 4). Then, the server device 10 inputs the emergency situation explanation information into a generation AI (e.g., ChatGPT51) and causes it to determine the language used by the user (step S11). Based on the input emergency situation explanation information, the server device 10 causes ChatGPT51 to determine whether the user's emergency situation is one in which speed is important and in a specific field, or one in which accuracy is important. Based on this determination, the server device 10 selects either Tsuzumi41 or ChatGPT51 (step S12).

[0054] Next, the server device 10 causes the generation AI to create a summary of the emergency situation explanation information (step S13), and determines the optimal emergency destination based on the summary (step S14).The server device 10 also creates and sets a prompt for Tsuzumi41 or ChatGPT51 selected in step S12 based on the summary of the emergency situation explanation information.

[0055] The server device 10 connects the user terminal 20 and the emergency destination terminal 30 so that they can make a call. That is, the server device 10 controls the user terminal 20 to make a call to the emergency call source ((4) in FIG. 4).

[0056] Then, the server device 10 executes two-way interpretation using the generated AI selected in step S12 ((5) in FIG. 4).

[0057] Specifically, the server device 10 instructs the selected generation AI (Tsuzumi41 or ChatGPT51) to translate the voice data output from the user terminal 20 into Japanese voice data and output it as voice data (step S15), and to translate the voice data output from the emergency destination terminal 30 into the language used by the user and output it as voice data (step S16). In this example, the language of the emergency destination recipient is set to Japanese, but the language of the emergency destination recipient can be changed depending on the country in which the emergency call service is actually used.

[0058] This allows the cloud to use the generated AI to perform the interpretation via this emergency call service, and the call is conducted over a VoIP telephone line with two-way interpreted audio.

[0059] Specifically, the user's voice data is output from the user terminal 20 (step S15), and the cloud server device 10 inputs the data into the generation AI and acquires the user's voice data translated into Japanese (step S16).The server device 10 outputs the user's voice data translated into Japanese to the emergency call destination terminal 30 (step S17).

[0060] Next, the emergency call destination terminal 30 outputs the recipient's voice data (step S18), and the cloud server device 10 inputs the data into the generation AI and acquires the emergency call destination recipient's voice data translated into the user's language (step S19). The server device 10 outputs the emergency call destination recipient's voice data translated into the user's language to the user terminal 20 (step S20). The processing system 100 enables two-way simultaneous interpretation by repeating steps S15 to S20.

[0061] [Processing method] Next, a processing procedure of the processing method according to the embodiment will be described below. Fig. 5 is a sequence diagram showing an example of the processing procedure of the processing method according to the embodiment.

[0062] As shown in FIG. 5, for example, when an application is started in the user terminal 20 (step S31), communication between the server device 10 and the user terminal 20 starts (step S32).

[0063] The user terminal 20 receives emergency situation explanation information explaining the user's emergency situation through speech from the user (step S33), and transmits the received emergency situation explanation information to the server device 10 (step S34).

[0064] The server device 10 inputs emergency situation explanation information to the ChatGPT 51 and sets a determination command (step S35). The determination command instructs the server device 10 to determine the language used by the user, whether the user's emergency situation is one in which speed is important and in a specific field, or one in which accuracy is important, and to determine the destination of the user terminal 20.

[0065] The server device 10 determines the language used by the user based on the output from ChatGPT 51 (step S36). In addition, the server device 10 receives a determination result from ChatGPT 51 as to whether the user's emergency situation is one in which speed is important and in a specific field, or one in which accuracy is important, and selects either Tsuzumi41 or ChatGPT 51 (step S37). In FIG. 5, a case in which Tsuzumi41 is selected will be described as an example.

[0066] The server device 10 determines the destination of the user terminal 20 based on the output from the ChatGPT 51 (step S38). In Fig. 5, an example will be described in which the emergency destination terminal 30A is determined as the destination of the user terminal 20. The server device 10 connects the user terminal 20 and the emergency destination terminal 30A so that they can communicate with each other via the VoIP telephone line and the PSTN (step S39).

[0067] The server device 10 sets a prompt to Tsuzumi41 to instruct it to translate the user's voice data from the user's language to the language (e.g., Japanese) of the recipient of the emergency call, and to translate the voice data input from the emergency call destination terminal 30A from the language (Japanese) of the recipient of the emergency call to the user's language (steps S40 to S42).

[0068] When the server device 10 receives the prompt setting notification (step S43), it starts a conversation between the user terminal 20 and the emergency destination terminal 30A (step S44).

[0069] For example, when voice data in a foreign language input by the user is transmitted from the user terminal 20 (steps S45 and S46), the server device 10 inputs the user's voice data to Tsuzumi 41 (step S47) and causes it to be translated into Japanese (step S48). The server device 10 transmits the voice data translated into Japanese output from Tsuzumi 41 (step S49) to the emergency destination terminal 30A (step S50) and causes it to be output (step S51).

[0070] Then, when Japanese voice data input by the recipient of the emergency destination is transmitted from the emergency destination terminal 30A (steps S52, S53), the server device 10 inputs the Japanese voice data of the recipient of the emergency destination to Tsuzumi 41 (step S54) and has it translated into the language used by the user (foreign language) (step S55). The server device 10 transmits the voice data translated into the language used by the user output from Tsuzumi 41 (step S56) to the user terminal 20 (step S57) and has it output (step S58). The processing system 100 repeats the processes of steps S45 to S58 to perform simultaneous interpretation between the user and the recipient of the emergency destination.

[0071] [Effects of the embodiment] FIG. 6 is a diagram illustrating an emergency call service provided by the processing system according to the embodiment.

[0072] As shown in Figure 6, by the processing of the processing system 100 described above, in the event of an emergency, a user can simply click on the user terminal 20, and the call will be automatically connected to the optimal destination, and the conversation between the user and the recipient of the emergency call will be interpreted. Here, in the processing system 100, each user will have to perform additional communication between their own terminal and the server, but since the IOWN network 60, which has high capacity, high speed, and low latency, is used, the delay can be almost ignored, resulting in faster conversations. Note that the network used is not limited to the IOWN network 60, and other networks may also be used.

[0073] Therefore, even in an emergency, the user can automatically contact the most appropriate destination without having to go through the tedious process of searching for contact information, according to the processing system 100. Furthermore, the processing system 100 simultaneously translates the conversation between the user and the recipient of the emergency call, allowing the emergency situation to be properly communicated.

[0074] In this way, the processing system 100 makes it possible to call an emergency contact who can handle the emergency situation in the event of an emergency, and to provide an appropriate interpretation according to the emergency situation.

[0075] In this embodiment, Tsuzumi41 and ChatGPT51 have been used as examples of the generation AIs to be used, but other generation AIs may also be used, and the number of generation AIs is not limited to two, but may be any one of three or more generation AIs.

[0076] [System configuration of the embodiment] The server device 10 is a functional concept and does not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of the functions of the server device 10 is not limited to that shown in the figure, and all or part of it can be functionally or physically distributed or integrated in any unit depending on various loads, usage conditions, etc.

[0077] Furthermore, all or any part of the processes performed by the server device 10 may be realized by a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and a program analyzed and executed by the CPU and the GPU. Furthermore, each process performed by the server device 10 may be realized as hardware using wired logic.

[0078] Furthermore, among the processes described in the embodiments, all or part of the processes described as being performed automatically can be performed manually. Alternatively, all or part of the processes described as being performed manually can be performed automatically using a known method. In addition, the processing procedures, control procedures, specific names, and information including various data and parameters described above and illustrated can be changed as appropriate unless otherwise specified.

[0079] [program] 7 is a diagram showing an example of a computer that executes a program to implement the server device 10. The computer 1000 includes, for example, a memory 1010 and a CPU 1020. The computer 1000 also includes a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.

[0080] The memory 1010 includes a ROM 1011 and a RAM 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1090. The disk drive interface 1040 is connected to a disk drive 1100. A removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to a mouse 1110 and a keyboard 1120, for example. The video adapter 1060 is connected to a display 1130, for example.

[0081] The hard disk drive 1090 stores, for example, an OS (Operating System) 1091, an application program 1092, a program module 1093, and program data 1094. That is, a program that defines each process of the server device 10 is implemented as a program module 1093 in which code executable by the computer 1000 is written. The program module 1093 is stored, for example, in the hard disk drive 1090. For example, a program module 1093 for executing the same process as the functional configuration of the server device 10 is stored in the hard disk drive 1090. The hard disk drive 1090 may be replaced with an SSD (Solid State Drive).

[0082] Furthermore, setting data used in the processing of the above-described embodiment is stored as program data 1094, for example, in memory 1010 or hard disk drive 1090. Then, CPU 1020 reads program module 1093 and program data 1094 stored in memory 1010 or hard disk drive 1090 into RAM 1012 as necessary and executes them.

[0083] The program module 1093 and program data 1094 are not limited to being stored in the hard disk drive 1090, but may also be stored in, for example, a removable storage medium and read by the CPU 1020 via the disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (such as a local area network (LAN) or a wide area network (WAN)). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via the network interface 1070.

[0084] Although the present invention has been described above as an embodiment, the present invention is not limited to the descriptions and drawings that form part of the disclosure of the present invention. In other words, other embodiments, examples, and operational techniques that can be made by those skilled in the art based on the present invention are all included in the scope of the present invention. [Explanation of symbols]

[0085] 10 Server device 11 Emergency Situation Acquisition Department 12 Language determination unit 13 Generation AI selection section 14 Destination determination unit 15 Prompt Settings 16 Input / Output Control Unit 20 User terminal 30, 30A~30B Emergency call destination terminal 40,50 Generation AI Server

Claims

1. an acquisition unit that receives an emergency call request from a user terminal and acquires explanation information that explains an emergency situation from the user terminal; a determination unit that determines the language used by the user based on the explanation information; a determination unit that determines a destination of the user terminal based on the explanation information and connects the user terminal and the destination so that they can communicate with each other; a setting unit for setting a prompt that instructs the generative model to translate input speech data or text data into an interpreted language in a natural context; an input / output control unit that causes the generative model to translate voice data or text data input from the user terminal into a language used at the destination, and outputs the voice data or text data output from the generative model to the destination, and also causes the generative model to translate voice data or text data input from the destination into a language used by the user, and outputs the voice data or text data output from the generative model to the user terminal; A processing device comprising:

2. a selection unit that selects one of the plurality of generative models based on the explanation information; The processing device according to claim 1, characterized in that the input / output control unit uses the generative model selected by the selection unit to perform interpretation of the voice data or text data input from the user terminal and the voice data or text data input from the destination.

3. The processing device according to claim 2, characterized in that the selection unit determines, based on the explanatory information, whether the emergency situation is a situation in which speed is important and in a specific field, or a situation in which accuracy is important, and selects one of the plurality of generative models based on the determination result.

4. The processing device according to claim 3, characterized in that the selection unit selects a first generative model which is a natural language processing model fine-tuned to a specific field when the emergency situation is the speed-oriented situation and a situation in a specific field, and selects a second generative model which is a large-scale natural language processing model when the emergency situation is the accuracy-oriented situation.

5. The processing device according to claim 1, characterized in that the processing device communicates with the user terminal, the destination, and the server device equipped with the generative model via a communication network related to an Innovative Optical and Wireless Network (IOWN).

6. A processing method executed by a processing device, receiving a request for an emergency call from a user terminal and obtaining explanation information from the user terminal that explains the emergency situation; determining a language used by the user based on the description information; determining a destination of the user terminal based on the explanation information and connecting the user terminal and the destination so that communication can be performed; setting prompts that instruct the generative model to translate input speech or text data into an interpreted language in a natural context; a step of causing the generative model to translate the voice data or text data input from the user terminal into a language used at the destination, and outputting the voice data or text data output from the generative model to the destination; a step of causing the generative model to translate the voice data or text data input from the destination into the language used by the user, and outputting the voice data or text data output from the generative model to the user terminal; A processing method comprising:

7. receiving a request for an emergency call from a user terminal and obtaining explanation information explaining an emergency situation from the user terminal; determining a language used by the user based on the description information; determining a destination of the user terminal based on the explanation information and connecting the user terminal and the destination so that communication can be performed; setting a prompt that instructs the generative model to translate the input speech or text data into an interpreted language in a natural context; a step of causing the generative model to translate the voice data or text data input from the user terminal into a language used at the destination, and outputting the voice data or text data output from the generative model to the destination; a step of causing the generative model to translate the voice data or text data input from the destination into the language used by the user, and outputting the voice data or text data output from the generative model to the user terminal; A processing program that causes a computer to execute the above.

Citation Information

Patent Citations

  • Support device and method for emergency reporting, and storage medium

    JP2002099979A

  • Multilingual emergency report system

    JP2015061246A

  • Apparatus and method for informing emergency situation

    JP2017200159A

  • Recommendation device, recommendation method and recommendation program

    JP2019139663A

  • Multilingual asynchronous translation system, multilingual asynchronous translation method and program

    JP2022003441A