Call support device, call support method, and call support program
The call support system addresses the challenge of communication barriers between text and voice users by converting text to voice and voice to text using a large-scale language model, ensuring smooth and natural dialogue across different communication devices.
Patent Information
- Application Number
- JP2024034487
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-07
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2044-03-07
AI Technical Summary
Existing technologies fail to facilitate suitable call communication between text callers using text communication devices and voice callers using voice communication devices, as direct conversion between text and voice often results in inadequate communication due to differences in style and expression.
A call support system utilizing a first and second modification unit to convert text information into audio information optimized for voice calls and vice versa, incorporating a large-scale language model and natural language processing to ensure smooth communication between users with different communication methods, including apps like LINE and chat-based interfaces.
Enables seamless communication by converting text to voice and voice to text, accommodating diverse communication styles and preferences, ensuring real-time and natural dialogue between users with different devices and communication methods.
Smart Images

Figure 2025136205000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a communication support device, a communication support method, and a communication support program. [Background technology]
[0002] Patent Document 1 describes an information processing device that includes a calling unit, a storage unit, and an output unit. The calling unit automatically calls destination telephone terminals with phone numbers pre-stored in the storage unit based on a voice call command, broadcasts voice distribution information pre-stored in the storage unit by voice, and receives voice responses to the distribution information from the destination telephone terminals that respond. The storage unit stores the voice responses. The output unit outputs a response result based on the voice responses. The calling unit selects whether to distribute the distribution information as real voice or pseudo voice for each destination telephone number. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 7419472 Summary of the Invention [Problem to be solved by the invention]
[0004] However, the conventional technology including the information processing device of Patent Document 1 has room for improvement in terms of achieving suitable call communication between a text caller using a text communication device and a voice caller using a voice communication device.
[0005] The present invention has been made based on the above-mentioned awareness of the problem, and aims to provide a call support device, a call support method, and a call support program that enable suitable call communication between a text caller using a text call device and a voice caller using a voice call device. [Means for solving the problem]
[0006] The call support device of this embodiment is characterized by having a first modification unit that modifies text information input by a text call user to a text call device into audio information optimized for voice calls, a second modification unit that modifies audio information input by a voice call user to a voice call device into text information optimized for voice calls, a first output unit that outputs the audio information modified by the first modification unit to the voice call user via the voice call device, and a second output unit that outputs text information modified by the second modification unit to the text call user via the text call device. [Effects of the Invention]
[0007] According to the present invention, it is possible to provide a call support device, a call support method, and a call support program that enable suitable call communication between a text caller using a text call device and a voice caller using a voice call device. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a diagram illustrating an example of a schematic configuration of a communication support device according to an embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing an example of an internal configuration of a call assistance server according to the present embodiment. [Figure 3] 10 is a conceptual diagram showing an example of a process of modifying text information into audio information by a first modifying unit. FIG. [Figure 4] 10 is a conceptual diagram showing an example of a process of modifying audio information into text information by a second modifying unit. FIG. [Figure 5] FIG. 2 is a conceptual diagram showing an example of a first modification mode candidate group and a second modification mode candidate group held in a modification mode candidate group holding table. [Figure 6] FIG. 10 is a diagram showing an example of modification modes adopted by the first modification section and the second modification section. [Figure 7] FIG. 10 is a diagram illustrating an example of the operation of a text communication device and a voice communication device via a communication support server. [Figure 8] FIG. 2 is a diagram illustrating an example of functional configurations of a text communication device, a voice communication device, and a communication support server. [Figure 9] FIG. 2 is a diagram illustrating an example of the hardware configuration of a text communication device, a voice communication device, and a communication support server. DETAILED DESCRIPTION OF THE INVENTION
[0009] The following describes a call assistance device, a call assistance method, and a call assistance program according to this embodiment.
[0010] The call support device, call support method, and call system equipped with the call support program of this embodiment realize "Shin Autocall" as a new autocall system that is an "automated telephone for the Reiwa era." In particular, the Shin Autocall system realizes "Shin Autocall Kakehashi," equipped with a function as a "bridge" that connects people with different apps, devices, and communication methods by phone.
[0011] More specifically, Shin Auto Call Kakehashi has an intermediary function optimized for optimal telephone communication between text-based users (e.g., those who have difficulty or are not comfortable speaking by voice) and voice-based users (e.g., those who prefer to speak by voice or primarily use the phone). For example, it can convert messages from apps such as LINE ("LINE" is a registered trademark of LINE Corporation, Japan) or chat-based interfaces into telephone voice, and convert telephone voice replies into messages from apps such as LINE ("LINE" is a registered trademark of LINE Corporation, Japan) or chat-based interfaces. Furthermore, by utilizing generative AI (artificial intelligence) and large-scale language models, it can realize natural dialogue even between people with different apps, devices, and communication methods. Furthermore, it can combine speech and character recognition with content summarization using large-scale language models to enable real-time conversations between text input into apps such as LINE ("LINE" is a registered trademark of LINE Corporation, Japan) or chat-based interfaces and voice calls.
[0012] The "shin" in Shin Autocall or Shin Autocall Kakehashi can mean at least one of the following: "new era," "friendly," "aiming to truly solve problems," "synchronizing with customers," and "in-house efforts with conviction."
[0013] 1 is a diagram showing an example of a schematic configuration of a call support device 1 of this embodiment. The call support device 1 may be read as a "call support system" (the two may be synonymous). The call support method and call support program of this embodiment are realized by a computer (CPU: Central Processing Unit) included in the call support device 1 executing (causing) various processing steps.
[0014] As shown in FIG. 1, the communication assistance device 1 includes a text communication device 10, a voice communication device 20, a communication assistance server 30, and a large-scale language model 40. The text communication device 10, the voice communication device 20, the communication assistance server 30, and the large-scale language model 40 are communicably connected to each other via a network NW. In FIG. 1, the text communication device 10, the voice communication device 20, the communication assistance server 30, and the large-scale language model 40 are each depicted as a simplified single block for convenience of illustration. In reality, the text communication device 10, the voice communication device 20, the communication assistance server 30, and the large-scale language model 40 may have a more complex hardware and software configuration than that depicted in FIG. 1. For example, the communication assistance server 30 may include multiple servers, databases, and the like. The large-scale language model 40 may also constitute a part of the communication assistance server 30.
[0015] The communication support device (communication support system) 1 may be understood as a configuration including a text communication device 10, a voice communication device 20, a communication support server 30, and a large-scale language model 40, or as a configuration including only the communication support server 30 without including the text communication device 10, the voice communication device 20, and the large-scale language model 40. In other words, in the communication support device (communication support system) 1, the communication support server 30 is an essential component, but the text communication device 10, the voice communication device 20, and the large-scale language model 40 are not essential components.
[0016] The text communication device 10 is composed of a personal computer, telephone, smartphone, tablet terminal, or the like equipped with a function for a text communication user to execute a text communication. The text communication device 10 may be either a fixed type or a mobile (portable) type. The text communication user inputs messages, text, emoticons, stamps, etc. via an operation unit (dial keys, buttons, touch screen, etc.) of the text communication device 10, and can also check messages, text, emoticons, stamps, etc. output from the other party via a display unit (display) of the text communication device 10 (including voice information input by a voice communication user to the voice communication device 20 that has been converted into text information, as will be described in detail later).
[0017] The voice communication device 20 is composed of a personal computer, a telephone, a smartphone, a tablet terminal, or the like equipped with a function for a voice caller to make a voice call. The voice communication device 20 may be either a fixed type or a mobile (portable) type. A voice caller makes a voice call by exchanging voice information with a call partner via a microphone and a speaker of the voice communication device 20. As will be described in detail later, the voice information output from the speaker of the voice communication device 20 includes text information input by a text caller to the text communication device 10 that has been converted into voice information.
[0018] Here, the text communication device 10 is not limited to a device having only a text communication function and not a voice communication function, and conversely, the voice communication device 20 is not limited to a device having only a voice communication function and not a text communication function. The text communication device 10 and the voice communication device 20 may be the same (common) device having both a text communication function and a voice communication function, as long as the text communication function is performed when used by a text caller and the voice communication function is performed when used by a voice caller. In addition, the text communication device 10 and the voice communication device 20 may be devices that can connect text communication and voice communication with high convenience and in real time. Here, the "text communication user" may be a person who has difficulty or is not good at speaking by voice (for example, a person with a voice disorder, hearing impairment, or stage fright), and the "voice communication user" may be a person who prefers to speak by voice or a person who mainly uses the telephone (for example, a person with a visual impairment or difficulty inputting text information).
[0019] The text communication device 10 and the voice communication device 20 are communication devices capable of performing various types of communication with the communication assistance server 30. The text communication device 10 and the voice communication device 20 may communicate with a network (such as the Internet) via, for example, a wired and / or wireless connection (e.g., Long Term Evolution (LTE), 5th Generation New Radio (5G NR), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.). The text communication device 10 and the voice communication device 20 may execute a browser and various other applications in response to a user operation.
[0020] In this embodiment, the text communication device 10 and the voice communication device 20 can connect a text call made by a text caller using the text communication device 10 with a voice call made by a voice caller using the voice communication device 20 with high convenience and in real time when receiving a call support service (thin autocall call bridge) from the call support server 30. In other words, it becomes possible to have favorable call communication between a text caller using the text communication device 10 and a voice caller using the voice communication device 20. An application for receiving the call support service (thin autocall call bridge) provided by the call support server 30 may be downloaded (installed) to the text communication device 10 and the voice communication device 20. Hereinafter, this application may be referred to as the call support application "call bridge."
[0021] The call assistance server 30 may be any communication device capable of performing various communications between the text communication device 10 and the voice communication device 20, and may be interpreted as a device, a circuit, or the like. The call assistance server 30 controls and executes various processes to provide users (text callers and voice callers) with a suitable service that connects text calls and voice calls with high convenience in real time via communications between the text communication device 10 and the voice communication device 20.
[0022] 2 is a block diagram showing an example of the internal configuration of the call support server 30 of this embodiment. The call support server 30 has a first modification unit 31, a second modification unit 32, a first output unit 33, a second output unit 34, a first acquisition unit 35, a second acquisition unit 36, and a modification mode candidate group holding table 37.
[0023] The first modifying unit 31 modifies the text information input by a text caller to the text communication device 10 into voice information optimized for voice communication (hereinafter, this may be referred to as modified voice information).
[0024] The first modification unit 31 adds to the modified voice information voice information obtained by performing summarization, correction, and deletion processes on the text information input by the text caller to the text call device 10, voice information indicating the operation status and cooperation status of the text call device 10 and the voice call device 20, voice information indicating the recommended operation of the voice call device 20 by the voice caller, and voice information indicating the ring tone, call start tone, and call end tone (these are combined to form the modified voice information).
[0025] The summarization process refers to, for example, a process of briefly summarizing the gist and main points of the text information input by the text caller to the text call device 10. The correction process refers to, for example, a process of correcting obvious typographical errors in the text information input by the text caller to the text call device 10 to convert it into voice information, or a process of overwriting (updating) content written earlier in a sentence when the content is rewritten later in the sentence to convert it into voice information. The deletion process refers to, for example, a process of ignoring content that has become unnecessary as a result of the summarization process or correction process among the text information input by the text caller to the text call device 10, so as not to include it in the voice information.
[0026] The audio information indicating the operation status of the text communication device 10 and the voice communication device 20 means, for example, audio information indicating that an input operation of text information or preparation for the same is being performed on the text communication device 10, and audio information indicating that an input operation of voice information or preparation for the same is being performed on the voice communication device 20. The audio information indicating the cooperation status of the text communication device 10 and the voice communication device 20 means, for example, audio information indicating whether the text communication device 10 and the voice communication device 20 can be connected using the call assistance application "Kakehashi", audio information indicating that one of the text communication device 10 and the voice communication device 20 is calling the other, and audio information indicating whether the text communication device 10 and the voice communication device 20 are in a call using the call assistance application "Kakehashi".
[0027] The voice information indicating the recommended operation of the voice call device 20 by the voice caller means, for example, voice information for prompting the voice caller to perform a predetermined operation (pushing a specific dial number) to start a call using the call support application "Kakehashi", and voice information for prompting the voice caller to input voice information using the voice call device 20.
[0028] The audio information indicating the ring tone, call start tone, and call end tone means, for example, audio information such as "Ring..." or "The call is now ringing" when the call is received, and audio information such as "Click," "The call has started," or "The call has ended" when the call starts and ends.
[0029] In addition, the first modification unit 31 may perform modifications in the modified audio information, such as conversion of written / spoken language from text information to audio information, or language conversion from text information to audio information (for example, translation between Japanese and English), or dialect conversion from text information to audio information (for example, conversion from Tokyo dialect to Osaka dialect), or honorific expression conversion from text information to audio information (for example, mutual conversion between honorific language, humble language, and polite language).
[0030] 3A and 3B are conceptual diagrams showing an example of the process of modifying text information into voice information by the first modifying unit 31. Fig. 3A shows text information input by a text caller to the text communication device 10, and Fig. 3B shows modified voice information obtained by converting the text information into voice information by the first modifying unit 31. As shown in Figs. 3A and 3B, it can be seen that the text information input by the text caller to the text communication device 10 has been modified in the modified voice information to content optimized for voice communication.
[0031] The second modifying unit 32 modifies the voice information input by the voice caller to the voice communication device 20 into character information optimized for text communication (hereinafter, this may be referred to as modified character information).
[0032] The second modification unit 32 adds to the modified character information character information obtained by performing summarization, correction, and deletion processes on the voice information input by the voice caller to the voice call device 20, character information indicating the operation status and cooperation status of the voice call device 20 and the text call device 10, character information indicating the recommended operation of the text call device 10 by the text caller, and character information indicating the ring tone, call start tone, and call end tone (these are combined to form the modified character information).
[0033] The summarization process refers to, for example, a process of concisely summarizing the gist and main points of the voice information input by the voice caller to the voice communication device 20. The correction process refers to, for example, a process of correcting obvious mistakes in the voice information input by the voice caller to the voice communication device 20 and converting the voice information into text information, or a process of overwriting (updating) the earlier content of a recorded message when the latter content is restated later. The deletion process refers to, for example, a process of ignoring unnecessary content from the voice information input by the voice caller to the voice communication device 20 due to the summarization process or correction process, and not including it in the text information. The deletion process also refers to, for example, a process of ignoring unconscious utterances such as "uh," "ah," and "so, you see," and not including them in the text information.
[0034] The text information indicating the operation status of the voice communication device 20 and the text communication device 10 means, for example, text information indicating that an operation to input voice information or preparation for the same is being performed on the voice communication device 20, and text information indicating that an operation to input text information or preparation for the same is being performed on the text communication device 10. The text information indicating the cooperation status of the voice communication device 20 and the text communication device 10 means, for example, text information indicating whether the voice communication device 20 and the text communication device 10 can connect using the call assistance application "Kakehashi", text information indicating that one of the voice communication device 20 and the text communication device 10 is calling the other, and text information indicating whether the voice communication device 20 and the text communication device 10 are in a call using the call assistance application "Kakehashi".
[0035] The text information indicating the recommended operation of the text communication device 10 by the text communication user means, for example, text information for prompting the text communication user to perform a predetermined operation (pushing a specific dial number) to start a call using the call support application "Kakehashi", or text information for prompting the text communication user to input text information using the text communication device 10.
[0036] The text information indicating the ring tone, call start tone, and call end tone means, for example, text information (including pseudo-voice information) such as "Ring..." or "Calling now" when ringing, and text information (including pseudo-voice information) such as "Click," "Call has started," or "Call has ended" when starting and ending a call.
[0037] In addition, the second modification unit 32 may perform modifications in the modified text information, such as converting spoken / written language from audio information to text information, or may perform modifications related to language conversion from audio information to text information (for example, translating between Japanese and Chinese), or may perform modifications related to dialect conversion from audio information to text information (for example, converting from Tokyo dialect to Hakata dialect), or may perform modifications related to honorific expression conversion from audio information to text information (for example, mutual conversion between honorific language, humble language, and polite language).
[0038] 4A and 4B are conceptual diagrams showing an example of the process of modifying voice information into text information by the second modifying unit 32. Fig. 4A shows voice information input by the voice caller to the voice communication device 20, and Fig. 4B shows modified text information obtained by converting voice information into text information by the second modifying unit 32. As shown in Figs. 4A and 4B, it can be seen from the modified text information that the voice information input by the voice caller to the voice communication device 20 has been modified to content optimized for text communication.
[0039] The first output unit 33 outputs the modified voice information by the first modification unit 31 to the voice caller via the voice communication device 20 (for example, by outputting voice from a speaker) (see FIGS. 3A and 3B). The second output unit 34 outputs the modified text information by the second modification unit 32 to the text caller via the text communication device 10 (for example, by outputting text from a display) (see FIGS. 4A and 4B). In the above-mentioned FIGS. 3A, 3B, 4A, and 4B, only the beginning (starting point) of the conversation between the text caller and the voice caller is illustrated as an example, but the text caller and the voice caller can exchange conversation (play catch) by re-inputting the text information and the voice information while mutually checking the modified text information output from the text communication device 10 and the modified voice information output from the voice communication device 20.
[0040] With respect for diversity, there is a demand for technology that facilitates smooth communication between users who have difficulty or are not good at communicating by voice, such as by telephone, and users who wish to communicate by voice. One example of such technology under consideration is a technology that generates synthetic speech from input text and transmits it to the other party's phone, then responds to the caller with the synthetic speech, converting the user's voice into text and outputting it. However, in actual communication, there are differences in style and expression between reading and writing text and listening and speaking, so simply converting voice to text and vice versa does not ensure smooth communication. Therefore, there is a need for technology that can bridge this gap.
[0041] In this embodiment, the above-mentioned problem is considered to be an important technical challenge, and smooth communication between reading and writing using text and listening and speaking using voice is realized by connecting apps such as LINE ("LINE" is a registered trademark of LINE Corporation, Japan) and chat applications that communicate using text with telephones and other devices that communicate using voice using a natural language recognition function, a generative AI (large-scale language model) function, and an automatic telephone call function. To achieve smooth communication between the different styles of text and voice, interfaces such as apps and chat applications that use text, such as LINE ("LINE" is a registered trademark of LINE Corporation, Japan), may not only accept phone numbers and messages, but may also respond to and report user input using a generative AI (e.g., "Please wait" or "The call has started"). Furthermore, in voice-based telephones and other devices, in addition to reading and accepting messages, they may also summarize the received voice message, confirm it by voice, and communicate the status, such as "You are entering a message" while the other party is entering text.
[0042] Furthermore, in order to create a natural conversation (chat) in application software (LINE ("LINE" is a registered trademark of LINE Corporation of Japan)) on a smartphone or the like, the generation AA may be used to output text such as "I understand," "Please wait," onomatopoeia (pururu-ruru, click), and "The call has started." This may be achieved by instructing the generation AI to behave "as if accepting a telegram."
[0043] Also, to allow the telephone user to have a natural conversation, it may emit voices such as "I'm currently typing a message. Please wait a moment," "Your message has been received," "I'll read it out," "That's all," "Please enter your reply," "If you're happy with it, please say 'Yes,' or if you want to try again, please say 'No.'" This can be achieved by instructing the generation AI to behave "like a telephone operator."
[0044] Additionally, the AI can delete unnecessary parts of messages entered by text or voice (such as "uh," "I see," etc.) and summarize them, and then confirm the results with the user who entered the message. In this case, priority can be given to response speed over recognition accuracy to improve the real-time nature of the conversation. Furthermore, the conversation can be recorded based on the smartphone application ID and phone number.
[0045] The first acquisition unit 35 acquires identification information of the text communication device 10 and attribute information of the text caller linked thereto. The identification information of the text communication device 10 may include, for example, a telephone number, an International Mobile Equipment Identifier (IMEI), and other identification IDs. The attribute information of the text caller may include, for example, the age, sex, nationality, occupation, etc. of the text caller, as well as registration information and call history using the call support application "Kakehashi."
[0046] The second acquisition unit 36 acquires identification information of the voice communication device 20 and attribute information of the voice caller linked thereto. The identification information of the voice communication device 20 may include, for example, a telephone number, an International Mobile Equipment Identifier (IMEI), and other identification IDs. The attribute information of the voice caller may include, for example, the age, sex, nationality, occupation, etc. of the voice caller, as well as registration information and call history using the call support application "Kakehashi."
[0047] The first modification unit 31 may estimate the relationship between the text caller and the voice caller, etc., depending on a combination of the identification information of the text call device 10 and the attribute information of the text caller, and the identification information of the voice call device 20 and the attribute information of the voice caller, and may modify the manner of modifying text information to voice information accordingly. For example, if it is estimated from the identification information of the text call device 10 and the attribute information of the text caller that the text caller's native language is Japanese, and if it is estimated from the identification information of the voice call device 20 and the attribute information of the voice caller that the voice caller's native language is English, the first modification unit 31 may modify text information input by the text caller in Japanese to voice information in English. Also, if it is estimated from the identification information of the text call device 10 and the attribute information of the text caller that the text caller mainly speaks Tokyo dialect, and if it is estimated from the identification information of the voice call device 20 and the attribute information of the voice caller that the voice caller mainly speaks Osaka dialect, the first modification unit 31 may modify text information input by the text caller in Tokyo dialect to voice information in Osaka dialect. Furthermore, when it is estimated from the identification information of the text communication device 10 and the attribute information of the text communication person that the text communication person is a subordinate of the company, and when it is estimated from the identification information of the voice communication device 20 and the attribute information of the voice communication person that the voice communication person is a superior of the company, the first modification unit 31 may modify the voice information in a way that corrects the incorrect honorific expression entered by the text communication person.
[0048] The second modifying unit 32 may vary the manner of modifying the voice information to text information by estimating the relationship between the voice caller and the text caller, etc., depending on a combination of the identification information of the voice call device 20 and the attribute information of the voice caller, and the identification information of the text call device 10 and the attribute information of the text caller. For example, if it is estimated from the identification information of the voice call device 20 and the attribute information of the voice caller that the voice caller's native language is Chinese, and it is estimated from the identification information of the text call device 10 and the attribute information of the text caller that the text caller's native language is English, the second modifying unit 32 may modify the voice information input in Chinese by the voice caller into English text information. Furthermore, when it is estimated from the identification information of the voice communication device 20 and the attribute information of the voice communication person that the voice communication person mainly uses Hakata dialect, and when it is estimated from the identification information of the text communication device 10 and the attribute information of the text communication person that the text communication person mainly uses Aomori dialect, the second modifying unit 32 may modify the voice information input by the voice communication person in Hakata dialect into text information in Aomori dialect. Furthermore, when it is estimated from the identification information of the voice communication device 20 and the attribute information of the voice communication person that the voice communication person is a superior of the company, and when it is estimated from the identification information of the text communication device 10 and the attribute information of the text communication person that the text communication person is a subordinate of the company, the second modifying unit 32 may modify the voice information input by the voice communication person that is speaking in a somewhat rough manner into text information in appropriate polite language.
[0049] The modification mode candidate group holding table 37 holds a first modification mode candidate group for modification from character information to audio information, and a second modification mode candidate group for modification from audio information to character information. In the following, for the sake of convenience in explaining the invention, the first modification mode candidate group and the second modification mode candidate group held by the modification mode candidate group holding table 37 will be described as being the same (common), but the first modification mode candidate group and the second modification mode candidate group may be different from each other (the first modification from character information to audio information and the second modification from audio information to character information may be upgraded according to their respective characteristics).
[0050] The first modification mode candidate group and the second modification mode candidate group held in the modification mode candidate group holding table 37 may each include the following <Modification Mode A> to <Modification Mode E>. <Modification A> Modifications relating to summarization, correction, and deletion processes from one of text information and audio information to the other. <Modification B> The modification of written / spoken language from one of text information and audio information to the other. <Modification C> Modifications relating to the conversion of text and audio information into another language (for example, Japanese to English). <Modification D> Modifications related to dialect conversion from one of text information and audio information to the other (for example, conversion between Tokyo dialect and Osaka dialect). <Modification E> Modifications involved in converting honorific expressions from one of text information and audio information to the other (for example, mutual conversion between honorific language, humble language, and polite language).
[0051] 5 is a conceptual diagram showing an example of the first modification mode candidate group and the second modification mode candidate group held in the modification mode candidate group holding table 37. The above-mentioned <Modification mode A> to <Modification mode E> are held as the first modification mode candidate group and the second modification mode candidate group. The first modification unit 31 can extract at least one modification mode to be actually adopted from the first modification mode candidate group held in the modification mode candidate group holding table 37. Similarly, the second modification unit 32 can extract at least one modification mode to be actually adopted from the second modification mode candidate group held in the modification mode candidate group holding table 37.
[0052] The first modification unit 31 refers to the modification mode candidate group holding table 37 and extracts, from the first modification mode candidate group, modification modes corresponding to combinations of the identification information of the text call device 10 and the attribute information of the text caller, and the identification information of the voice call device 20 and the attribute information of the voice caller. In this case, the first modification unit 31 extracts modification modes A and B as "mandatory modification modes" and extracts modification modes C, D, and E as "selective modification modes." For example, the first modification unit 31 may estimate the relationship between the text caller and the voice caller in consideration of the ages, genders, nationalities, occupations, etc. of the text caller and the voice caller, and automatically select or determine some or all of modification modes C, D, and E based on the registration information and call history, etc., used in the call support application "Kakehashi."
[0053] The second modification unit 32 refers to the modification mode candidate group holding table 37 and extracts, from the second modification mode candidate group, modification modes corresponding to the combination of the identification information of the voice call device 20 and the attribute information of the voice caller, and the identification information of the text call device 10 and the attribute information of the text caller. In this case, the second modification unit 32 extracts modification modes A and B as "mandatory modification modes" and extracts modification modes C, D, and E as "selective modification modes." For example, the second modification unit 32 may automatically select or determine some or all of modification modes C, D, and E based on the registration information and call history using the call support application "Kakehashi" while estimating the relationship between the voice caller and the text caller in consideration of the age, sex, nationality, occupation, etc. of the voice caller and the text caller.
[0054] The modification modes adopted by the first modification unit 31 and the second modification unit 32 may be flexibly set by the text caller and the voice caller through manual settings using the call support application "Kakehashi."
[0055] In this way, by flexibly setting the modification modes adopted by the first modification unit 31 and the second modification unit 32 based on the relationship between the voice caller and the text caller, it is possible, for example, to deliberately not perform dialect conversion (deliberately not adopting modification mode D) or deliberately not perform honorific expression conversion (deliberately not adopting modification mode E).
[0056] 6A and 6B are diagrams showing examples of modification modes adopted by the first modification unit 31 and the second modification unit 32. In the example of Fig. 6A, the first modification unit 31 adopts modification modes A and B as required modification modes, and also adopts modification modes C and D as optional modification modes (modification mode E is not adopted). In the example of Fig. 6B, the second modification unit 32 adopts modification modes A and B as required modification modes, and also adopts modification mode E as optional modification mode (modification modes C and D are not adopted).
[0057] The call support server 30 may operate by appropriately referring to the large-scale language model 40. The large-scale language model 40 may be interpreted as a generative AI (Artificial Intelligence), and may cooperate with the call support server 30 to fulfill the functions of each block of the call support server 30 (first modification unit 31, second modification unit 32, first output unit 33, second output unit 34, first acquisition unit 35, second acquisition unit 36, and modification mode candidate group retention table 37).
[0058] The large-scale language model 40 stores, for example, specific information from the individual domain databases (e.g., modification information based on selective modification patterns) and common information from the common domain database (e.g., modification information based on required modification patterns) normalized into a unified data format so that they can be adapted to the training of language models for natural language processing. Furthermore, by combining one or more types of natural language processing based on the individual domain databases and the common domain database, a language model specialized for creating modified speech information and modified text information is trained and adjusted, and support information for assisting in the creation of modified speech information and modified text information can be generated by using the language model according to the content of the prompt. The language model may be stored on a cloud computer, which enables training of large datasets and training using high-performance computing resources.
[0059] Here, "natural language processing" refers to a process that enables a computer to understand text and voice data written in natural language and execute processing according to the purpose. Specifically, examples include morphological analysis, which breaks natural language down into "morphemes," the smallest units that make up the language, and assigns information such as parts of speech; syntactic analysis, which analyzes the grammatical structure of natural language to clarify the structure and meaning of a sentence; semantic analysis, which analyzes the meaning of natural language to understand the meaning of words and sentences and make logical judgments and inferences; contextual analysis, which understands natural language while taking into account the context before and after a sentence; and intention analysis, which extracts the intention of a speaker or writer from a conversation or sentence using natural language. Thus, "natural language processing" processes natural language by combining processes such as morphological analysis, syntactic analysis, semantic analysis, contextual analysis, and intention analysis, enabling the call support processing of the present embodiment (e.g., the creation of modified voice information by the first modification unit 31 and the creation of modified character information by the second modification unit 32).
[0060] A "language model" is a type of probabilistic model used in natural language processing, and is a model that probabilistically predicts how likely a given word or sentence is to occur in natural language. Specifically, a language model calculates the occurrence probability of a given word sequence or sentence, or compares the occurrence probabilities of multiple word sequences or sentences, making it possible to automatically generate the most likely word or sentence based on the context when predicting the next word or sentence.
[0061] In the field of natural language processing, the term "language model" refers to a mathematical model that learns language patterns, grammatical rules, etc., and generates and understands natural language. For example, the language model used in the creation of modified voice information by the first modification unit 31 and the creation of modified character information by the second modification unit 32 is learned specifically for the creation of modified voice information by the first modification unit 31 and the creation of modified character information by the second modification unit 32, and is used to generate information to support the creation of modified voice information by the first modification unit 31 and the creation of modified character information by the second modification unit 32, depending on the content of the prompt.
[0062] "Normalized into a unified data format" means that when handling large amounts of data such as individual domain databases and common domain databases, the data is converted into a common format and organized in a unified manner. When machine learning is used to create modified voice information by the first modification unit 31 and modified character information by the second modification unit 32, the efficiency and accuracy of the processing can be improved by using normalized unique information, common information, system overview, and other information as input. "Normalization" is a processing method for organizing and structuring data in database tables according to certain rules, thereby eliminating data redundancy and maintaining data consistency and integrity.
[0063] "Prompt" is information that displays to the user what to input, how to operate, etc., and allows the user to give instructions. "Information for supporting the creation of modified voice information by the first modifying unit 31 and the creation of modified character information by the second modifying unit 32" is information that makes the creation of modified voice information and modified character information more efficient, enabling more effective call support to be implemented.
[0064] In addition, the communication support server 30 may be configured to hold and refer to a dictionary database and a vocabulary database in addition to / instead of the large-scale language model 40 to perform the first conversion from text information to audio information and the second conversion from audio information to text information.
[0065] In addition to the components illustrated in the functional block diagram of FIG. 4 , the call assistance server 30 may additionally include various components that are expected to be included in a typical telephone terminal. For example, the call assistance server 30 may include a reception unit, a storage unit, a transmission unit, and an output unit. The reception unit may have a destination number reception function for receiving a destination telephone number and a distribution information reception function for receiving distribution information for the destination. The storage unit may have a function for storing the destination telephone number and distribution information received by the reception unit, and a function for storing a response to the distribution information received from the telephone terminal by the transmission unit. The transmission unit may have an automatic transmission function for automatically transmitting to the telephone terminal, a simultaneous distribution function for simultaneously distributing the distribution information to multiple telephone terminals in the form of voice (e.g., simulated voice, real voice) or text, a voice recognition function for voice recognition of a voice response to the voice distribution information, and a voice conversion function for converting the voice response to text. The output unit may have a function for outputting a response result based on the response received by the transmission unit from the telephone terminal. For example, the output unit may play back a voice response from a telephone terminal on the terminal device, or may output a response result based on a text response obtained by converting the voice response into text to a monitor of the terminal device.The output unit may also display a response result based on the text response from the telephone terminal on a monitor of the terminal device.
[0066] Fig. 7 is a diagram showing an example of the operation of the text communication device 10 and the voice communication device 20 via the communication support server 30. In Fig. 7, the communication support server 30 supports communication between Mr. A, a text communication participant using the text communication device 10, and Mr. B, a voice communication participant using the voice communication device 20.
[0067] First, person A operates the text communication device 10 and inputs "Call person B." Then, the display of the text communication device 10 displays the messages "Got it. I'm calling 03-xxx-xxx. Please wait." and "...ring-ring. If there's no answer after a while, please call again.", and the voice communication device 20 is called. When person B responds to the call on the voice communication device 20 (answers the phone), the display of the text communication device 10 displays "Click...call has started.", and person A inputs "Tell them that I was very helpful yesterday." At this time, the voice communication device 20 may issue a voice message to person B saying, "I'm currently inputting a text message, so please wait a moment." After person A has completed inputting the message, the message "Got it. I'll pass it on to them. Please wait a moment for my reply." is displayed, and the voice communication device 20 also displays a message to person B saying, "This is a call from someone else. You have a message from person A saying, 'Thank you very much for your help yesterday.'" "Please reply." is issued. In response, Person B records a voice message saying, "Yes, well, please tell them that I am very grateful for their help," and the display of the text communication device 10 displays, "The following reply was received over the phone: Thank you for your help. Thank you very much." Here, it is possible to extract and highlight (for example, by changing the color, highlighting with a highlighter, changing the font, or changing the size) only the reply portion of Person B, the voice caller, "Thank you for your help. Thank you very much," from the text message displayed on the display of the text communication device 10. Finally, when Person A inputs "end call" into the text communication device 10, the display of the text communication device 10 displays, "I understand. I will end the call. Thank you for using the service." and the call between the text communication device 10 and the voice communication device 20 is terminated.
[0068] In the example of Figure 7, Persons A and B simply exchange messages such as "Thank you for your help yesterday," but in practice, the messages may be more complex, as shown in Figures 3A, 3B, 4A, and 4B, and the first and second modifications may be made to a greater extent. In this case, the text communication party and the voice communication party can exchange conversations (play catch) by re-inputting the text information and voice information while mutually checking the modified text information output from the text communication device 10 and the modified voice information output from the voice communication device 20.
[0069] Fig. 8 is a diagram showing an example of the functional configuration of the text communication device 10, the voice communication device 20, and the call assistance server 30. As shown in Fig. 8, the text communication device 10, the voice communication device 20, and the call assistance server 30 each include a control unit 110, a storage unit 120, a communication unit 130, an input unit 140, and an output unit 150. Note that Fig. 8 shows an example of functional blocks of the text communication device 10, the voice communication device 20, and the call assistance server 30, and they may include other functional blocks not shown in Fig. 8. Furthermore, they may be configured not to include some of the functional blocks.
[0070] The control unit 110 controls the text communication device 10, the voice communication device 20, and the communication assistance server 30. The control unit 110 can be configured by a controller, a control circuit, or a control device that is described based on common understanding in the technical field related to the present invention.
[0071] The storage unit 120 stores information used by the text communication device 10, the voice communication device 20, and the communication assistance server 30. The storage unit 120 can be configured, for example, by a memory, a storage, a storage device, or the like that is described based on common understanding in the technical field related to the present invention.
[0072] The communication unit 130 performs communication (e.g., mutual communication) between the text communication device 10, the voice communication device 20, and the communication assistance server 30. The communication unit 130 can be configured by a transmitter / receiver, a transmission / reception circuit, or a transmission / reception device that is described based on common understanding in the technical field related to the present invention. Note that the communication unit 130 may be configured by a transmission unit and a reception unit.
[0073] The input unit 140 accepts inputs from the text communication device 10, the voice communication device 20, and the communication assistance server 30. The input unit 140 may also be connected to a predetermined device, a storage medium, or the like, and accept data input. The input unit 140 may output the input result to, for example, the control unit 110. The input unit 140 may be configured with input devices such as a keyboard, a mouse, and a button, an input / output terminal, an input / output circuit, and the like, which are explained based on common understanding in the technical field related to the present invention. The input unit 140 may also be configured as an integrated unit with a display unit (for example, a touch panel).
[0074] The output unit 150 performs output in the text communication device 10, the voice communication device 20, and the communication assistance server 30. For example, the output unit 150 may be configured to include a display unit that displays images, an audio output unit that outputs audio, etc. The display unit may be configured, for example, by a display device such as a monitor or the like that is described based on common understanding in the technical field related to the present invention. Furthermore, the audio output unit may be configured by an output device such as a speaker that is described based on common understanding in the technical field related to the present invention.
[0075] The block diagrams used to explain the above embodiments show functional blocks. These functional blocks (components) are realized by any combination of hardware and / or software. Furthermore, the means for realizing each functional block is not particularly limited. That is, each functional block may be realized by a single physically coupled device, or may be realized by two or more physically separate devices connected by wire or wirelessly.
[0076] For example, the devices (the text communication device 10, the voice communication device 20, and the communication assistance server 30) in one embodiment of the present disclosure may function as a computer that performs the communication assistance processing of the present disclosure. FIG. 9 is a diagram showing an example of the hardware configuration of the text communication device 10, the voice communication device 20, and the communication assistance server 30. The above-described text communication device 10, the voice communication device 20, and the communication assistance server 30 may be physically configured as a computer device including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, etc.
[0077] In the present disclosure, the terms apparatus, circuit, device, unit, server, etc. may be interchangeable. The hardware configurations of the text communication apparatus 10, the voice communication apparatus 20, and the communication assistance server 30 may be configured to include one or more of the devices shown in the drawings, or may be configured to exclude some of the devices.
[0078] For example, although only one processor 1001 is shown, there may be multiple processors. Furthermore, processing may be performed by one processor, or processing may be performed by two or more processors simultaneously, serially, or in other ways. Furthermore, processor 1001 may be implemented by one or more chips.
[0079] Each function of the text communication device 10, the voice communication device 20, and the communication assistance server 30 is realized by loading predetermined software (programs) onto hardware such as a processor 1001, a memory 1002, etc., so that the processor 1001 performs calculations and controls communication by a communication device 1004, reading and / or writing of data in the memory 1002 and the storage 1003, etc.
[0080] The processor 1001 controls the entire computer by running, for example, an operating system. The processor 1001 may be configured as a central processing unit (CPU) including an interface with peripheral devices, a control device, an arithmetic unit, a register, etc. Note that each unit such as the control unit 110 described above may be realized by the processor 1001.
[0081] The processor 1001 also reads programs (program codes), software modules, data, etc. from at least one of the storage 1003 and the communication device 1004 into the memory 1002, and executes various processes in accordance with these. The programs used are those that cause a computer to execute at least some of the operations described in the above-described embodiments. For example, the control unit 110 may be realized by a control program stored in the memory 1002 and running on the processor 1001, and similar implementations may be made for other functional blocks.
[0082] The memory 1002 is a computer-readable recording medium and may be configured by at least one of, for example, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an electrically EEPROM (EEPROM), a random access memory (RAM), or other suitable storage medium. The memory 1002 may also be referred to as a register, a cache, a main memory, or the like. The memory 1002 may store executable programs (program codes), software modules, and the like for implementing a method according to one embodiment.
[0083] Storage 1003 is a computer-readable recording medium, and may be configured by at least one of, for example, a flexible disk, a floppy disk, a magneto-optical disk (e.g., a compact disk (CD-ROM (Compact Disc ROM)), a digital versatile disk, a Blu-ray (registered trademark) disk), a removable disk, a hard disk drive, a smart card, a flash memory device (e.g., a card, a stick, a key drive), a magnetic stripe, a database, a server, or other suitable storage medium. Storage 1003 may also be called an auxiliary storage device. Note that the above-mentioned storage unit 120 may be realized by memory 1002 and / or storage 1003.
[0084] The communication device 1004 is hardware (transmission / reception device) for communicating between computers via at least one of a wired network and a wireless network, and is also referred to as a network device, a network controller, a network card, a communication module, etc. The communication device 1004 may include a SIM card. Note that the above-mentioned communication unit 130 may be realized by the communication device 1004.
[0085] The input device 1005 is an input device (for example, a keyboard, a mouse, etc.) that receives input from the outside. The output device 1006 is an output device (for example, a display, a speaker, etc.) that performs output to the outside. The input device 1005 and the output device 1006 may be integrated into one unit (for example, a touch panel). The above-mentioned input unit 140 and output unit 150 may be realized by the input device 1005 and the output device 1006, respectively.
[0086] Furthermore, each device such as the processor 1001 and the memory 1002 is connected by a bus 1008 for communicating information. The bus 1008 may be configured as a single bus, or may be configured as different buses between the devices.
[0087] The text communication device 10, the voice communication device 20, and the communication assistance server 30 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 1001 may be implemented by at least one of these pieces of hardware.
[0088] (Variation) In addition, terms explained in this disclosure and / or terms necessary for understanding this disclosure may be replaced with terms having the same or similar meanings.
[0089] The information, parameters, etc. described in this disclosure may be expressed using absolute values, relative values from a predetermined value, or other corresponding information. Furthermore, the names used for parameters, etc. in this disclosure are not limiting in any way.
[0090] The information, signals, etc. described in this disclosure may be represented using any of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.
[0091] Information, signals, etc. may be input / output via multiple network nodes. The input / output information, signals, etc. may be stored in a specific location (e.g., memory) or may be managed using a table. The input / output information, signals, etc. may be overwritten, updated, or added. Output information, signals, etc. may be deleted. Input information, signals, etc. may be transmitted to another device.
[0092] Furthermore, notification of specified information (e.g., notification that "it is X") is not limited to explicit notification, but may be made implicitly (e.g., by not notifying the specified information or by notifying other information).
[0093] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.
[0094] Software, instructions, information, etc. may also be transmitted and received via a transmission medium and / or signal waveform. For example, if software is transmitted from a website, server, or other remote source using wired technologies (such as coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL)), and / or wireless technologies (such as infrared, microwave), then these wired and / or wireless technologies are included within the definition of transmission media.
[0095] As used in this disclosure, the terms "system" and "network" may be used interchangeably.
[0096] Each aspect / embodiment described in this disclosure may be used alone, in combination, or switched depending on the implementation. Furthermore, the order of the processing procedures, sequences, flowcharts, etc. of each aspect / embodiment described in this disclosure may be changed unless inconsistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the specific order presented.
[0097] As used in this disclosure, the phrase "based on" does not mean "based only on," unless expressly stated otherwise. In other words, the phrase "based on" means both "based only on" and "based at least on."
[0098] As used in this disclosure, any reference to an element using a designation such as "first," "second," etc. does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed or that the first element must in some way precede the second element.
[0099] When used in this disclosure, the terms "include," "including," and variations thereof are intended to be inclusive, similar to the term "comprising." Furthermore, when used in this disclosure, the term "or" is not intended to be an exclusive or.
[0100] In this disclosure, where articles are added by translation, such as a, an, and the in English, the disclosure may include that the nouns following these articles are in the plural form.
[0101] Although the invention according to the present disclosure has been described in detail above, it is clear to those skilled in the art that the invention according to the present disclosure is not limited to the embodiments described in the present disclosure. The invention according to the present disclosure can be implemented in modified and altered forms without departing from the spirit and scope of the invention as defined by the description of the claims. Therefore, the description of the present disclosure is intended to be illustrative and does not impose any limiting meaning on the invention according to the present disclosure. [Explanation of symbols]
[0102] 1. Call support device (call support system) 10 character intercom 20 Voice communication device 30 Call support server 31 First Modification 32 Second Modification 33 First output section 34 Second output section 35 First Acquisition Section 36 Second Acquisition Section 37 Modification mode candidate group holding table 40 Large-scale language models NW Network
Claims
1. a first modifying unit that modifies text information input by a text call user into a text call device into voice information optimized for voice calls; a second modifying unit that modifies voice information input by a voice caller into a voice communication device into text information optimized for text communication; a first output unit that outputs the voice information modified by the first modification unit to the voice caller via the voice call device; a second output unit that outputs the modified character information by the second modification unit to the text caller via the text call device; A communication support device comprising:
2. the first modification unit adds to the modified voice information voice information obtained by performing summarization processing, correction processing, and deletion processing on the text information input by the text call user to the text call device, voice information indicating the operation status and cooperation status of the text call device and the voice call device, voice information indicating the recommended operation of the voice call device by the voice call user, and voice information indicating a ring tone, a call start tone, and a call end tone; 2. The communication support device according to claim 1.
3. the second modification unit adds, to the modified character information, character information obtained by performing summarization, correction, and deletion processes on the voice information input by the voice call user to the voice call device, character information indicating an operation status and a cooperation status of the voice call device and the character call device, character information indicating a recommended operation of the character call device by the character call user, and character information indicating a ring tone, a call start tone, and a call end tone.
2. The communication support device according to claim 1.
4. a first acquisition unit that acquires identification information of the text call device and attribute information of the text caller associated therewith; a second acquisition unit that acquires identification information of the voice communication device and attribute information of the voice communication party associated therewith; and the first modification unit varies a manner of modification from the text information to the voice information depending on a combination of identification information of the text communication device and attribute information of the text caller, and a combination of identification information of the voice communication device and attribute information of the voice caller; the second modification unit varies a manner of modification from the voice information to the text information depending on a combination of identification information of the voice communication device and attribute information of the voice caller, and identification information of the text communication device and attribute information of the text caller.
2. The communication support device according to claim 1.
5. The information processing device further includes an alteration pattern candidate group holding table that holds a first alteration pattern candidate group for altering the character information to the audio information and a second alteration pattern candidate group for altering the audio information to the character information, the first modification unit refers to the modification mode candidate group holding table, and extracts, from the first modification mode candidate group, a modification mode corresponding to a combination of the identification information of the text communication device and the attribute information of the text caller, and the identification information of the voice communication device and the attribute information of the voice caller; the second modification unit refers to the modification mode candidate group holding table and extracts, from the second modification mode candidate group, a modification mode corresponding to a combination of the identification information of the voice communication device and the attribute information of the voice caller, and the identification information of the text communication device and the attribute information of the text caller; 5. The communication support device according to claim 4.
6. The first modification mode candidate group and the second modification mode candidate group held in the modification mode candidate group holding table respectively include a modification mode A relating to summarization processing, correction processing, and deletion processing from one of the character information and the audio information to the other, a modification mode B relating to written / spoken conversion from one of the character information and the audio information to the other, a modification mode C relating to language conversion from one of the character information and the audio information to the other, a modification mode D relating to dialect conversion from one of the character information and the audio information to the other, and a modification mode E relating to honorific expression conversion from one of the character information and the audio information to the other, The first modification part and the second modification part extract the modification mode A and the modification mode B as essential modification modes, and extract the modification mode C, the modification mode D, and the modification mode E as selective modification modes.
6. The communication support device according to claim 5.
7. a first modification step of modifying text information input by a text call user into a text call device into voice information optimized for voice calls; a second modifying step of modifying voice information input by a voice caller into text information optimized for text communication; a first output step of outputting the modified voice information obtained by the first modification step to the voice caller via the voice call device; a second output step of outputting the modified character information obtained by the second modification step to the character caller via the character call device; A call support method characterized by being executed by a computer.
8. a first modification step of modifying text information input by a text call user into a text call device into voice information optimized for voice calls; a second modifying step of modifying voice information input by a voice caller into text information optimized for text communication; a first output step of outputting the modified voice information obtained by the first modification step to the voice caller via the voice call device; a second output step of outputting the modified character information obtained by the second modification step to the character caller via the character call device; A call support program that causes a computer to execute the above.
Citation Information
Patent Citations
Portable translation machine
JP1999110389A
Phone relay service method
JP2003092628A
Originating terminal, terminating terminal, mediate server and message output method
JP2003110732A
Information processing device, information processing method, and information processing program
JP7419472B1