Communication terminal, communication system, information processing method, and program
The communication system addresses network and device load issues by transmitting voice messages in text format with metadata for synthesized voice generation, effectively reducing traffic and load while maintaining voice authenticity.
Patent Information
- Application Number
- JP2026086200
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2026-05-22
- Publication Date
- 2026-08-26
- Estimated Expiration
- 2046-05-22
AI Technical Summary
The transmission of voice data imposes a significant load on networks and devices due to its large data volume, increasing network traffic and device load.
A communication system and method that transmits and receives voice messages in text format with accompanying metadata, allowing for synthesized voice generation at the receiving end based on user-specific metadata, reducing network traffic and device load.
Reduces network traffic and device load by transmitting small-volume text and metadata, enabling synthesized voice playback that mimics the original speaker's voice characteristics.
Smart Images

Figure 0007911655000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a communication terminal, a communication system, an information processing method, and a program.
Background Art
[0002] With the spread of communication terminals such as smartphones, many communication systems for transmitting and receiving messages between these communication terminals are known. In recent years, social network services capable of transmitting and receiving not only text messages but also voice data have been developed (see Patent Document 1). When transmitting and receiving voice data, the voice is digitized and transmitted and stored. However, the data volume of digitized voice is very large compared to text data, and the transmission of such digital data increases network traffic and the load on devices related to data transmission and reception and storage.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In view of such problems of the prior art, an object of the present invention is to provide a communication terminal, a communication system, an information processing method, and a program capable of transmitting and receiving voice messages without imposing a load on the network.
Means for Solving the Problems
[0005] According to a first aspect of the present invention, a communication terminal is provided that can send and receive voice messages without burdening a network. The communication terminal includes a sound output unit capable of outputting sound, a message data acquisition unit that acquires first message data which describes a first message sent by a first user in text format, a metadata generation unit that generates first metadata which describes information related to the first user in text format, and a communication unit that can communicate with another terminal owned by a second user. The communication unit includes a transmission unit that can transmit the first message data and the first metadata to the other terminal, and a reception unit that can receive second message data which describes a second message sent by the second user in text format and second metadata which describes information related to the second user in text format from the other terminal. The communication terminal includes a synthesized voice data generation unit that generates synthesized voice data corresponding to the second message data using the second metadata, and a voice playback unit that plays the synthesized voice data using the sound output unit. a text translation unit that converts input text data into text data in another language. To further prepare. The metadata generation unit is configured to include first language information representing the first language used by the first user in the first metadata. The second metadata includes second language information representing the second language used by the second user. The text translation unit translates the second message data from the second language to the first language to generate translated message data when the second language represented by the second language information in the second metadata is different from the first language. The synthesized speech data generation unit is configured to generate synthesized speech data from the second message data when the second language represented by the second language information in the second metadata is the same as the first language, and to generate synthesized speech data from the translated message data when the second language represented by the second language information in the second metadata is different from the first language. Alternatively, the text translation unit translates the first message data from the first language to a reference language to generate first reference language message data. The transmission unit is configured to transmit the first reference language message data in addition to the first message data to the other terminal. The receiving unit is configured to receive a second reference language message data from the other terminal in addition to the second message data. The text translation unit translates the second reference language message data from the reference language to the first language to generate translated message data when the second language represented by the second language information in the second metadata is different from the first language. The synthesized speech data generation unit is configured to generate synthesized speech data from the second message data when the second language represented by the second language information in the second metadata is the same as the first language, and to generate synthesized speech data from the translated message data when the second language represented by the second language information in the second metadata is different from the first language.
[0006] A second aspect of the present invention provides a communication system that can send and receive voice messages without burdening a network. The communication system comprises a first communication terminal owned by the first user, which is composed of the above-described communication terminals; a second communication terminal owned by the second user, which is composed of the above-described communication terminals; and a server device connected to each of the first and second communication terminals via a communication network, which is configured to mediate the sending and receiving of the first message data and the first metadata, and the second message data and the second metadata, between the first and second communication terminals.
[0007] A third aspect of the present invention provides an information processing method that enables sending and receiving voice messages without burdening a network. The information processing method includes acquiring first message data input from a first user, generating first metadata related to the user's voice, transmitting the first message data and the first metadata to another terminal owned by a second user, receiving second message data and second metadata related to the second user's voice input from the other terminal, generating synthesized voice data corresponding to the second message data using the second metadata, and playing the synthesized voice data via the sound output unit of the terminal. The first metadata includes first language information representing the first language used by the first user, and the second metadata includes second language information representing the second language used by the second user. The information processing method generates translated message data by translating the second message data from the second language to the first language when the second language represented by the second language information in the second metadata is different from the first language. The generation of synthesized speech data includes generating synthesized speech data from the second message data when the second language represented by the second language information in the second metadata is the same as the first language, and generating synthesized speech data from the translated message data when the second language represented by the second language information in the second metadata is different from the first language. Alternatively, the above information processing method translates the first message data from the first language to a reference language to generate first reference language message data, transmits the first reference language message data in addition to the first message data to the other terminal, and receives the second reference language message data in addition to the second message data from the other terminal. The above information processing method generates translated message data by translating the second reference language message data from the reference language to the first language when the second language represented by the second language information in the second metadata is different from the first language. The generation of the synthesized speech data includes generating the synthesized speech data from the second message data when the second language represented by the second language information in the second metadata is the same as the first language, and generating the synthesized speech data from the translated message data when the second language represented by the second language information in the second metadata is different from the first language.
[0008] According to a fourth aspect of the present invention, a program is provided for causing a computer to execute the above-described information processing method. [Brief explanation of the drawing]
[0009] [Figure 1] Figure 1 is a schematic conceptual diagram illustrating a communication system in a first embodiment of the present invention. [Figure 2] Figure 2 is a conceptual diagram showing the functional components of a terminal within the communication system shown in Figure 1. [Figure 3A] Figure 3A is a flowchart illustrating the process when a terminal in the communication system shown in Figure 1 sends a message. [Figure 3B] Figure 3B is a flowchart showing the processing performed by the server device that receives the data transmitted through the process shown in Figure 3A. [Figure 3C] Figure 3C is a flowchart showing the processing at the terminal that receives the data transmitted through the process shown in Figure 3B. [Figure 4A] Figure 4A is a schematic diagram showing an example of a screen display on a certain terminal. [Figure 4B] Figure 4B is a schematic diagram showing an example of the screen display of the terminal shown in Figure 4A. [Figure 4C]FIG. 4C is a schematic diagram showing an example of the screen display of the terminal in FIG. 4A. [Figure 4D] FIG. 4D is a schematic diagram showing an example of the screen display of the terminal in FIG. 4A. [Figure 5A] FIG. 5A is a schematic diagram showing an example of the screen display of a certain terminal. [Figure 5B] FIG. 5B is a schematic diagram showing an example of the screen display of the terminal in FIG. 5A. [Figure 6A] FIG. 6A is a flowchart showing the process when the terminal returns a message after the process shown in FIG. 3C. [Figure 6B] FIG. 6B is a flowchart showing the process in the server device that has received the data transmitted by the process shown in FIG. 6A. [Figure 6C] FIG. 6C is a flowchart showing the process in the terminal that has received the data transmitted by the process shown in FIG. 6B. [Figure 7A] FIG. 7A is a schematic diagram showing an example of the screen display of a certain terminal. [Figure 7B] FIG. 7B is a schematic diagram showing an example of the screen display of the terminal in FIG. 7A. [Figure 8] FIG. 8 is a conceptual diagram showing the functional components of the terminal in the communication system according to the second embodiment of the present invention. [Figure 9A] FIG. 9A is a flowchart showing the process when the terminal having the components shown in FIG. 8 transmits a message. <s [Figure 9B] FIG. <s is a flowchart showing the process in the server device that has received the data transmitted by the process shown in FIG. 9A. [Figure 9C] FIG. 9C is a flowchart showing the process in the terminal that has received the data transmitted by the process shown in FIG. 9B. [Figure 10A] FIG. 10A is a flowchart showing the process when the terminal returns a message after the process shown in FIG. 9C. [[ID=]] [Figure 10B] FIG. 10B is a flowchart showing the process in the server device that has received the data transmitted by the process shown in FIG. 10A. [Figure 10C] FIG. 10C is a flowchart showing processing in a terminal that has received data transmitted by the processing shown in FIG. 10B. Embodiments for Carrying Out the Invention
[0010] Hereinafter, embodiments of a communication terminal, a communication system, an information processing method, and a program according to the present invention will be described in detail with reference to FIGS. 1 to 10C. In FIGS. 1 to 10C, the same or corresponding components are denoted by the same reference numerals and redundant descriptions are omitted. Also, in FIGS. 1 to 10C, there are cases where the scales and dimensions of each component are exaggerated and cases where some components are omitted. In the following description and claims, unless otherwise specified, terms such as "first" and "second" are used only to distinguish components from each other and do not represent a specific rank or order.
[0011] FIG. 1 is a conceptual diagram schematically showing a communication system 1 in the first embodiment of the present invention. This communication system 1 includes a plurality of communication terminals (hereinafter simply referred to as "terminals") 10 such as smartphones and a server device 20. The terminals 10 and the server device 20 are connected to each other by a communication network 100 constructed by a mobile communication network (such as LTE, 5G), the Internet, a local area network (LAN), a wide area network (WAN), a short-range communication network, an intranet, or a combination thereof. Although various forms of the plurality of terminals 10 are conceivable as described below, since the basic configuration of any of the terminals 10 is the same, only the configuration of one terminal 10 will be described below.
[0012] Terminal 10 may consist of a smartphone, tablet computer, laptop computer, computing device, dedicated terminal with built-in speaker, browser extension, AR (Augmented Reality) glasses, IoT edge device, AR spatial projection device, earphones, headphones, smart speaker, etc., connected to the communication network 100 via Ethernet, USB, Bluetooth (registered trademark), WiFi, etc., whether wired or wireless. In this embodiment, terminal 10 will be described as a smartphone, but it is not limited to this.
[0013] As shown in Figure 1, each terminal 10 includes a display unit 11 such as a liquid crystal display or an organic EL (Electro-Luminescence) display, a touch sensor 12 (operation input unit) that receives user input, a microphone 13 (sound input unit) that can input sound, a speaker 14 (sound output unit) that can output sound, a communication unit 15 for connecting to a communication network 100 and performing data communication, a storage unit 16 including ROM, RAM, flash memory, etc., and a control unit 17 that controls the operation of each component. The storage unit 16 stores the OS (Operating System), programs for controlling the terminal 10, and various data. The control unit 17 includes a processor (CPU), ROM, RAM, etc., and realizes various functions by loading the programs stored in the storage unit 16 into RAM and executing them with the processor.
[0014] In the following description, when referring to the "screen" of terminal 10, it means the touch panel in which the display unit 11 and the touch sensor 12 are integrated. The user can perform the desired input by tapping the input interface displayed on the screen of terminal 10.
[0015] The server device 20 may consist of, for example, a server computer, a general-purpose computer, a dedicated computer, or a mobile terminal, and may be composed of multiple such devices combined. Furthermore, the server device 20 may share hardware with other devices. The server device 20 includes a storage unit 21 containing ROM, RAM, flash memory, etc., a communication unit 22 for connecting to a communication network 100 and performing data communication, and a control unit 23 for controlling the operation of each component. The storage unit 21 stores the OS, programs for controlling the server device 20, a database, and various data. The control unit 23 includes a processor (CPU), ROM, RAM, etc., and realizes various functions by loading programs stored in the storage unit 21 into RAM and executing them with the processor.
[0016] Terminal 10 and server device 20 are configured to send and receive data via a communication network 100 using general communication technology. In this embodiment, server device 20 has the role of mediating data sent and received between multiple terminals 10. The communication unit 15 of terminal 10 includes a transmission unit 18 that specifies the recipient terminal 10 of the message and sends data to server device 20, and a reception unit 19 that receives data from server device 20, for example, via push notification. For example, when the transmission unit 18 of terminal 10A shown in Figure 1 specifies terminal 10B as the destination and sends data to server device 20, and server device 20 receives the data via communication unit 22, server device 20 sends the data to terminal 10B, which was specified as the destination of the data. The reception unit 19 of terminal 10B receives the data from server device 20.
[0017] The memory unit 16 of terminal 10 stores a program (hereinafter referred to as the "message program") that enables the sending and receiving of messages between multiple terminals 10. Figure 2 is a conceptual diagram showing the functional components of terminal 10 realized when this message program is executed by the control unit 17. As shown in Figure 2, terminal 10 includes a message data acquisition unit 31, a feature information generation unit 34, an ambient sound information generation unit 35, a metadata acquisition unit 36, a metadata generation unit 38, a text translation unit 39, a synthesized speech data generation unit 40, and a speech playback unit 41.
[0018] The memory unit 16 also stores user identification information M1 that identifies the user using the terminal 10, language information M2 that represents the language used by the user, gender information M3 that represents the gender of the user, and voice type information M4 that represents the type of voice associated with the user. Predetermined initial values may be used for this information, or the user can arbitrarily set this information. For example, user identification information M1 includes an identifier that identifies the user and the user's name. Language information M2 includes an identifier that represents the language used by the user (e.g., "ja" or "en"). Gender information M3 includes an identifier that represents the gender of the user (e.g., "male" or "female"). Voice type information M4 includes an identifier that represents the voice type associated with the user (e.g., "child", "young", "adult", "elder").
[0019] Furthermore, the memory unit 16 stores a contact user list 60 that enumerates users who are contacts. This contact user list 60 may be imported from other applications or generated by a message program.
[0020] The message data acquisition unit 31 is configured to acquire message data T1 from the voice input by the user. This message data T1 is a text-based description of a message to be sent to another user. In this embodiment, the message data acquisition unit 31 includes an input voice data generation unit 32 that generates voice data (input voice data S1) from the user's voice, and a message data generation unit 33 that generates message data T1 from the input voice data S1.
[0021] The input audio data generation unit 32 is configured to acquire audio (transmitted message) input via the microphone 13 and generate input audio data S1 by digitally converting it at a predetermined sampling rate (e.g., 16kHz, 44.1kHz, etc.). The input audio data generation unit 32 stores the generated input audio data S1 in the storage unit 16. The input audio data S1 can be generated as an audio file in formats such as PCM, AAC, or MP3, but is not limited to these formats.
[0022] The message data generation unit 33 is configured to generate message data T1 by converting the speech contained in the input speech data S1 into text data through speech recognition processing. The message data generation unit 33 is configured to perform speech recognition processing using, for example, a speech recognition engine inside the terminal 10 or a speech recognition engine on the communication network 100. At this time, speech recognition is performed in the language specified by the language information M2. The speech recognition engine may, for example, use a machine learning model using a neural network. The message data generation unit 33 stores the generated text data as message data T1 in the storage unit 16.
[0023] The feature information generation unit 34 is configured to analyze the input audio data S1 using a predetermined algorithm to calculate feature quantities that represent the characteristics of user A's voice, and to generate feature information M5 that represents the characteristics of the voice from the calculated feature quantities. For example, the feature information generation unit 34 analyzes the waveform of the input audio data S1 in milliseconds and calculates feature quantities such as the peak volume and frequency characteristics of the voice. Based on the calculated feature quantities, the feature information generation unit 34 generates feature information M5 that represents, for example, user A's emotions. As feature information that represents emotions, for example, types of emotions such as "normal," "joy," "anger," and "sadness" can be used, and identifiers that represent such types of emotions can be used as feature information M5. The feature information generation unit 34 stores the generated feature information M5 in the storage unit 16.
[0024] The ambient sound information generation unit 35 is configured to analyze the input audio data S1 using a predetermined algorithm to calculate feature quantities representing ambient sounds, and to generate ambient sound information M6 representing ambient sounds from the calculated feature quantities. For example, the ambient sound information generation unit 35 analyzes the waveform pattern of ambient sounds filtered from user A's voice in the input audio data S1 and calculates feature quantities such as peak volume and frequency characteristics. Based on the calculated feature quantities, the ambient sound information generation unit 35 generates ambient sound information M6 representing ambient sounds such as "construction site noise," "train noise," and "wind noise." Identifiers representing the type of ambient sound can be used as ambient sound information M6. The ambient sound information generation unit 35 stores the generated ambient sound information M6 in the storage unit 16.
[0025] The metadata acquisition unit 36 is configured to acquire metadata related to the user's voice from the storage unit 16 or the like. For example, the metadata acquisition unit 36 can acquire user identification information M1, language information M2, gender information M3, voice type information M4, characteristic information M5, and ambient sound information M6 as metadata from the storage unit 16. User identification information M1 and language information M2 may be acquired by reading those stored in the storage unit 16 by the message program as described above, or they may be acquired from the OS setting information. The metadata acquisition unit 36 is also configured to acquire date and time information M7, which represents the date and time related to the message data T1, as metadata. The date and time related to the message data T1 may be the creation date and time of the input voice data S1 or the creation date and time of the message data T1, and the date and time information M7 includes, for example, a string that describes the date and time according to a predetermined format.
[0026] The metadata generation unit 38 is configured to use the metadata acquisition unit 36 described above to generate metadata 50 that describes user-related information, particularly information related to the user's voice (voice-related information), in text format. For example, the metadata generation unit 38 formats the information acquired by the metadata acquisition unit 36 into a predetermined format such as JSON format to generate text-format metadata 50. The voice-related information may include one or more of the following: user identification information M1, language information M2, gender information M3, voice type information M4, feature information M5, ambient sound information M6, and date and time information M7.
[0027] The text translation unit 39 is configured to convert input text data into text data in another language. As will be described later, when the message data T2 received by the receiving unit 19 needs to be translated, this message data T2 is input to the text translation unit 39. The text translation unit 39 is configured to perform conversion processing (translation processing) using, for example, a translation engine inside the terminal 10 or a translation engine on the communication network 100. The translation engine may, for example, use a machine learning model using a neural network. The text translation unit 39 stores the text data obtained by conversion as translated message data T3 in the storage unit 16.
[0028] The synthesized speech data generation unit 40 is configured to generate synthesized speech data S2 by analyzing input text data using a predetermined algorithm and performing speech synthesis processing based on the analysis results. The synthesized speech data generation unit 40 is configured to perform speech synthesis processing using, for example, a speech synthesis engine inside the terminal 10 or a speech synthesis engine on the communication network 100. The speech synthesis engine may, for example, use a machine learning model using a neural network. The synthesized speech data generation unit 40 stores the generated synthesized speech data S2 in the storage unit 16.
[0029] The audio playback unit 41 is configured to read the synthesized speech data S2 generated by the synthesized speech data generation unit 40 from the storage unit 16 and play it back as speech via the speaker 14. The audio playback unit 41 may, for example, start playback of the synthesized speech data S2 in response to a playback operation by the user, or it may start playback automatically when the generation of the synthesized speech data S2 is complete.
[0030] Here, we will explain an example of sending a message from terminal 10A owned by user A to terminal 10B owned by another user B. Figures 3A to 3C are flowcharts showing the process of sending such a message. In the following explanation, when referring specifically to the components of terminal 10A and information and data related to user A, we will add "A" after the code representing the component, information, and data. When referring specifically to the components of terminal 10B and information and data related to user B, we will add "B" after the code representing the component, information, and data. For example, the control unit 17 of terminal 10A is indicated by "17A", and the control unit 17 of terminal 10B is indicated by "17B".
[0031] First, we will explain the case where both User A and User B use Japanese. The messaging program is launched when User A taps an icon (not shown) corresponding to the messaging program on the screen of terminal 10A (step S1). When the messaging program is launched, the control unit 17A reads the contact user list 60A from the storage unit 16A and displays the contact list screen D1 shown in Figure 4A on the screen (step S2). This contact list screen D1 includes the name of each user and a message button B1. The contact list screen D1 also includes an add contact user button B2 for adding a new contact user.
[0032] When user A taps, for example, user B's message button B1 on the contact list screen D1, the control unit 17A displays a conversation screen D2 on the screen as shown in Figure 4B (step S3). This conversation screen D2 includes the name of the user having the conversation (user B), a gender selection pull-down menu B3 for selecting one's own gender, a voice type selection pull-down menu B4 for selecting one's own voice type, a conversation area B5 where past conversations are displayed, and a microphone button B6 for inputting user A's voice.
[0033] Prior to the display of the conversation screen D2, the metadata acquisition unit 36A acquires gender information M3A from the storage unit 16A, and the gender selection pull-down menu B3 displays the gender corresponding to the gender information M3A acquired by the metadata acquisition unit 36A. When user A taps the gender selection pull-down menu B3 on the conversation screen D2, a gender list such as "Male" and "Female" is displayed, and user A can select one of them. When user A selects a gender, the control unit 17A stores the information representing the selected gender as gender information M3A in the storage unit 16A.
[0034] Similarly, prior to the display of the conversation screen D2, the metadata acquisition unit 36A acquires voice type information M4A from the storage unit 16A, and the voice type selection pulldown menu B4 displays the voice type corresponding to the voice type information M4A acquired by the metadata acquisition unit 36A. When user A taps the voice type selection pulldown menu B4 on the conversation screen D2, a list of voice types such as "child," "young person," "adult," and "elderly" is displayed, and user A can select one of them. When user A selects a voice type, the control unit 17A stores the information representing the selected voice type as voice type information M4A in the storage unit 16A.
[0035] When user A sends a message to user B, user A taps the microphone button B6 on the conversation screen D2 and enters the message (message to send) (step S4). At this time, the input voice data generation unit 32A acquires the voice input by the microphone 13A, converts it digitally, and generates input voice data S1A (step S5). For example, voice input may be accepted between tapping the microphone button B6 and tapping it again, or voice input may be accepted while the microphone button B6 is being held down, or voice input may be accepted between tapping the microphone button B6 and a predetermined time has elapsed.
[0036] Once the generation of the input voice data S1A is complete, the message data generation unit 33A generates message data T1A corresponding to the voice of user A contained in the input voice data S1A (step S6). For example, if user A taps the microphone button B6 and says "Where are you now?", the message data generation unit 33A generates message data T1A that says "Where are you now?". As shown in Figure 4C, the control unit 17A displays the message data T1A generated by the message data generation unit 33A as user A's message C1 in the conversation area B5 of the conversation screen D2.
[0037] Next, the feature information generation unit 34A analyzes the input voice data S1A and generates feature information M5A that represents the characteristics of user A's voice (step S7). Also, the ambient sound information generation unit 35A analyzes the input voice data S1A and generates ambient sound information M6A that represents the surrounding ambient sounds (step S8).
[0038] Next, the metadata generation unit 38A uses the metadata acquisition unit 36A to acquire metadata related to user A's voice (user identification information M1A, language information M2A, gender information M3A, voice type information M4A, feature information M5A, ambient sound information M6A, and date and time information M7A), and formats the acquired metadata into a predetermined format to create text-format metadata 50A (step S9).
[0039] Then, the transmitting unit 18A specifies the terminal 10B owned by user B and sends message data T1A ("Where are you now?") and metadata 50A to the server device 20 (step S10). The message data T1A and metadata 50A are received by the server device 20 (step S11), and the server device 20 sends the message data T1A and metadata 50A to the terminal 10B, for example, by push notification (step S12).
[0040] When the receiving unit 19B of terminal 10B receives the message data T1A and metadata 50A sent by push notification (step S13), the control unit 17B starts the message program on terminal 10B if the message program is not already running.
[0041] The control unit 17B acquires language information M2B representing the language used by user B via the metadata acquisition unit 36B, and determines whether the language represented by this language information M2B is the same as the language represented by the language information M2A in the metadata 50A transmitted from terminal 10A (step S14).
[0042] In this example, since the language used by both User A and User B is Japanese and identical, the process proceeds to step S15, and the control unit 17B displays a conversation screen E1 on the screen as shown in Figure 5A (step S15). Similar to the conversation screen D2 of terminal 10A, this conversation screen E1 includes the name of the user having the conversation (User A), a gender selection pull-down menu F1 for selecting one's own gender, a voice type selection pull-down menu F2 for selecting one's own voice type, a conversation area F3 where past conversations are displayed, and a microphone button F4 for inputting User B's voice. Similar to the conversation screen D2 of terminal 10A, the gender selection pull-down menu F1 displays the gender corresponding to the gender information M3B acquired by the metadata acquisition unit 36B, and the voice type selection pull-down menu F2 displays the voice type corresponding to the voice type information M4B acquired by the metadata acquisition unit 36B. In addition, the received message data T1A ("Where are you now?") is displayed in the conversation area F3 as a message G1 from User A.
[0043] Then, the synthesized speech data generation unit 40B generates synthesized speech data S2A from the received message data T1A (step S16). At this time, the synthesized speech data generation unit 40B generates synthesized speech data S2A using the received metadata 50A. More specifically, the synthesized speech data generation unit 40B performs speech synthesis so that the synthesized speech data S2A becomes a voice of the gender represented by the gender information M3A contained in the metadata 50A, and a voice of the voice type represented by the voice type information M4A contained in the metadata 50A. For example, in terminal 10A shown in Figure 4C, the gender is set to "male" and the voice type to "young man," so the gender represented by the gender information M3A contained in the metadata 50A is "male," and the voice type represented by the voice type information M4A is "young man," and the synthesized speech data generation unit 40B generates synthesized speech data S2A to be a young man's voice. The synthesized speech data generation unit 40B also performs speech synthesis so as to reflect the features represented by the feature information M5A contained in the metadata 50A. For example, if the voice characteristic represented by feature information M5A is "joy," the synthesized voice data generation unit 40B generates synthesized voice data S2A such that the synthesized voice is a joyful voice. Furthermore, the synthesized voice data generation unit 40B adds the ambient sound represented by ambient sound information M6A included in the metadata 50 to the synthesized voice data S2A. For example, if the ambient sound represented by ambient sound information M6A is "train sound," the synthesized voice data generation unit 40B adds a pre-prepared train sound to the synthesized voice data S2A.
[0044] When the synthesized speech data generation unit 40B generates synthesized speech data S2A, the speech playback unit 41B outputs and plays the synthesized speech data S2A from the speaker 14B (step S17). This allows user B to hear the message sent by user A as speech. The playback of this synthesized speech data S2A may be performed automatically when message data T1A and metadata 50A are received, or it may be performed, for example, when user B taps the playback button F5 displayed together with message G1 in the conversation area F3 of Figure 5A.
[0045] The above process completes the transmission of a message from user A's terminal 10A to user B's terminal 10B (step S18). Thus, according to this embodiment, by transmitting the text-format message data T1A and metadata 50A input by user A at user A's terminal 10A from terminal 10A to terminal 10B, terminal 10B can play synthesized speech based on user A's voice, and user B can hear speech that is similar to user A's voice. At this time, since the data transmitted from terminal 10A to terminal 10B is text-format message data T1A and metadata 50A, which have a small data volume, an increase in traffic on the communication network 100 can be suppressed, and the load on equipment related to data transmission, reception and storage can also be reduced.
[0046] Next, we will explain an example of sending a reply from terminal 10B to terminal 10A in response to a message from user A. Figures 6A to 6C are flowcharts showing the process of sending such a message. When user B replies to a message from user A, they tap the microphone button F4 on the conversation screen E1 in Figure 5A to input a message (step S30). At this time, the input voice data generation unit 32B acquires the voice input by the microphone 13B and converts it digitally to generate input voice data S1B (step S31). Next, the message data generation unit 33B generates message data T1B corresponding to the voice of user B contained in the input voice data S1B (step S32). For example, if user B taps the microphone button F4 and says "Shibuya ni imasu", the message data generation unit 33B generates message data T1B that says "Shibuya ni imasu." The control unit 17B displays the message data T1B generated by the message data generation unit 33B as user B's message G2 in the conversation area F3 of the conversation screen E1, as shown in Figure 5B.
[0047] Next, the feature information generation unit 34B analyzes the input voice data S1B to generate feature information M5B representing the characteristics of user B's voice (step S33). The ambient sound information generation unit 35B also analyzes the input voice data S1B to generate ambient sound information M6B representing the surrounding ambient sounds (step S34). Then, the metadata generation unit 38B uses the metadata acquisition unit 36B to acquire metadata related to user B's voice (user identification information M1B, language information M2B, gender information M3B, voice type information M4B, feature information M5B, ambient sound information M6B, and date and time information M7B), and formats the acquired metadata into a predetermined format to create text-format metadata 50B (step S35).
[0048] Then, the transmitting unit 18B specifies the terminal 10A owned by user A and sends message data T1B ("I am in Shibuya.") and metadata 50B to the server device 20 (step S36). The message data T1B and metadata 50B are received by the server device 20 (step S37), and the server device 20 sends the message data T1B and metadata 50B to the terminal 10A, for example, by push notification (step S38).
[0049] When the receiving unit 19A of terminal 10A receives the message data T1B and metadata 50B sent by push notification (step S39), the control unit 17A obtains language information M2A representing the language used by user A via the metadata acquisition unit 36A, and determines whether the language represented by this language information M2A is the same as the language represented by the language information M2B in the metadata 50B sent from terminal 10B (step S40).
[0050] In this example, since both User A and User B use Japanese, the process proceeds to step S41, where the control unit 17A displays the received message data T1B ("I'm in Shibuya.") as a message C2 from User B in the conversation area B5 of the conversation screen D2, as shown in Figure 4D. The synthesized speech data generation unit 40A generates synthesized speech data S2B from the received message data T1B (step S42). At this time, the synthesized speech data generation unit 40A generates synthesized speech data S2B using the received metadata 50B. In this example, since the terminal 10B shown in Figure 5B is set to "female" and the voice type to "young person", the gender information M3B included in the metadata 50B represents "female" and the voice type information M4B represents "young person", and synthesized speech data S2B is generated in the voice of a young woman. Subsequently, the voice playback unit 41A outputs and plays the synthesized speech data S2B from the speaker 14A (step S43). This allows user A to hear the message sent by user B as audio. Playback of this synthesized speech data S2B may be performed automatically upon receiving the message data T1B and metadata 50B, or it may be performed, for example, when user A taps the play button B8 displayed in the conversation area B5 of Figure 4D along with the message C2.
[0051] The above process completes the transmission of a message from user B's terminal 10B to user A's terminal 10A (step S44). Thus, according to this embodiment, by transmitting the text-format message data T1B and metadata 50B input by user B from terminal 10B to terminal 10A, terminal 10A can play synthesized speech based on user B's voice, and user A can hear speech that is similar to user B's voice. At this time, since the data transmitted from terminal 10B to terminal 10A is text-format message data T1B and metadata 50B, which have a small amount of data, it is possible to suppress an increase in traffic on the communication network 100 and also reduce the load on equipment related to data transmission, reception and storage.
[0052] As is clear from the above description, the server device 20 in the communication system 1 of this embodiment only mediates the exchange of text-formatted data between terminal 10A and terminal 10B, and neither speech recognition nor speech synthesis is performed in the server device 20. The communication system 1 of this embodiment is an edge-to-edge P2P communication system that can control the audio output at the receiving terminal 10 by metadata 50 generated at the transmitting terminal 10.
[0053] Next, we will explain the case where the language used by user B is different from the language used by user A (Japanese), for example, English. When terminal 10A sends a message to terminal 10B, in step S13, when the receiving unit 19B of terminal 10B receives the message data T1A and metadata 50A, the control unit 17B obtains language information M2B representing the language used by user B via the metadata acquisition unit 36B, and determines whether the language represented by this language information M2B is the same as the language represented by the language information M2A in the metadata 50A sent from terminal 10A (step S14).
[0054] Here, the language identified by the language information M2A in the metadata 50A is Japanese, and the language identified by the language information M2B is English, so the process proceeds to step S19, where the text translation unit 39B converts the message data T1A sent from terminal 10A into the language represented by the language information M2B (i.e., English) to generate translated message data T3A. In this example, the text translation unit 39B translates the message data T1A ("Where are you now?") into the language identified by the language information M2B, i.e., English, to generate translated message data T3A ("Where are you now?").
[0055] Then, as shown in Figure 7A, the control unit 17B displays a conversation screen E2 on the screen that includes translated message data T3A ("Where are you now?") as message G3 from user A (step S20). In this conversation screen E2, the user's name, gender selection pull-down menu F1, and voice type selection pull-down menu F2 from the conversation screen E1 described above are displayed in English.
[0056] Then, the synthesized speech data generation unit 40B generates synthesized speech data S2A from the translated message data T3A (step S21). At this time, the synthesized speech data generation unit 40B performs speech synthesis in the language specified by the language information M2B, i.e., English. Therefore, the speech included in the generated synthesized speech data S2A will be in English. Also, similar to the example above, the synthesized speech data generation unit 40B generates synthesized speech data S2A using the gender information M3A, voice type information M4A, feature information M5A, and ambient sound information M6A in the metadata 50A transmitted from the terminal 10A. In this example, since the terminal 10A shown in Figure 4C is set to "male" and the voice type to "young man", the gender represented by the gender information M3A included in the metadata 50A is "male", and the voice type represented by the voice type information M4A is "young man", so synthesized speech data S2A is generated in the voice of a young man.
[0057] When the synthesized speech data generation unit 40B generates synthesized speech data S2A, the speech playback unit 41B outputs and plays the synthesized speech data S2A from the speaker 14B (step S17). This allows user B to hear the message sent by user A in the voice of their preferred language. The playback of this synthesized speech data S2A may be performed automatically when message data T1A and metadata 50A are received, or it may be performed, for example, when user B taps the playback button F7 displayed together with the message G3 in the conversation area F3 of Figure 7A.
[0058] The above process completes the transmission of a message from user A's terminal 10A to user B's terminal 10B (step S18). Thus, according to this embodiment, by transmitting the text-format message data T1A and metadata 50A input by user A from terminal 10A to terminal 10B, terminal 10B can play back synthesized speech translated into user B's language based on user A's voice, allowing user B to hear speech similar to user A's voice in their own language. At this time, since the data transmitted from terminal 10A to terminal 10B is text-format message data T1A and metadata 50A, which have a small data size, an increase in traffic on the communication network 100 can be suppressed, and the load on equipment related to data transmission, reception, and storage can also be reduced.
[0059] Furthermore, when sending a reply from terminal 10B to terminal 10A to a message from user A, the user taps the microphone button F4 on the conversation screen E2 in Figure 7A to input a message (step S30). The input voice data generation unit 32B acquires the voice input by microphone 13B, digitally converts it, and generates input voice data S1B (step S31). Next, the message data generation unit 33B generates message data T1B corresponding to the voice of user B contained in the input voice data S1B (step S32). For example, if user B taps the microphone button F4 and says "I'm in Shibuya", the message data generation unit 33B generates message data T1B "I'm in Shibuya." The control unit 17B displays the message data T1B generated by the message data generation unit 33B as user B's message G4 in the conversation area F3 of the conversation screen E2, as shown in Figure 7B.
[0060] Next, the feature information generation unit 34B analyzes the input voice data S1B to generate feature information M5B representing the characteristics of user B's voice (step S33). The ambient sound information generation unit 35B also analyzes the input voice data S1B to generate ambient sound information M6B representing the surrounding ambient sounds (step S34). Then, the metadata generation unit 38B uses the metadata acquisition unit 36B to acquire metadata related to user B's voice (user identification information M1B, language information M2B, gender information M3B, voice type information M4B, feature information M5B, ambient sound information M6B, and date and time information M7B), and formats the acquired metadata into a predetermined format to create text-format metadata 50B (step S35).
[0061] Then, the transmitting unit 18B specifies the terminal 10A owned by user A and transmits message data T1B ("I'm in Shibuya.") and metadata 50B to the server device 20 (step S36). The message data T1B and metadata 50B are received by the server device 20 (step S37), and the server device 20 transmits the message data T1B and metadata 50B to the terminal 10A, for example, by push notification (step S38).
[0062] When the receiving unit 19A of terminal 10A receives the message data T1B and metadata 50B sent by push notification (step S39), the control unit 17A obtains language information M2A representing the language used by user A via the metadata acquisition unit 36A, and determines whether the language represented by this language information M2A is the same as the language represented by the language information M2B included in the metadata 50B sent from terminal 10B (step S40).
[0063] Here, the language identified by the language information M2B in the metadata 50B is English, and the language identified by the language information M2A is Japanese, so the process proceeds to step S45, where the text translation unit 39A converts the message data T1B sent from terminal 10A into the language represented by the language information M2A (i.e., Japanese) to generate translated message data T3B. In this example, the text translation unit 39A translates the message data T1B ("I'm in Shibuya.") into Japanese to generate translated message data T3B ("I'm in Shibuya.").
[0064] Then, as shown in Figure 4D, the control unit 17A displays a conversation screen D2 on the screen containing translated message data T3B ("I'm in Shibuya.") as message C2 from user B (step S46). The synthesized speech data generation unit 40A generates synthesized speech data S2B from the received message data T1B (step S47). At this time, the synthesized speech data generation unit 40A performs speech synthesis in the language specified by the language information M2A, i.e., Japanese. Therefore, the speech included in the generated synthesized speech data S2B will be in Japanese. Also, similar to the example described above, the synthesized speech data generation unit 40A generates synthesized speech data S2B using gender information M3B, voice type information M4B, feature information M5B, and ambient sound information M6B from the metadata 50B transmitted from terminal 10B. In this example, terminal 10B, shown in Figure 5B, is set to "female" and voice type to "young person." Therefore, the gender information M3B included in metadata 50B represents "female," and the voice type information M4B represents "young person," resulting in the generation of synthesized speech data S2B in the voice of a young woman. Subsequently, the speech playback unit 41A outputs and plays the synthesized speech data S2B from speaker 14A (step S43). This allows user A to hear the message sent by user B as audio.
[0065] The above process completes the transmission of a message from user B's terminal 10B to user A's terminal 10A (step S44). Thus, according to this embodiment, by transmitting the text-format message data T1B and metadata 50B input by user B from terminal 10B to terminal 10A, terminal 10A can play back synthesized speech translated into user A's language based on user B's voice, and user A can hear speech similar to user B's voice in their own language. At this time, since the data transmitted from terminal 10B to terminal 10A is text-format message data T1B and metadata 50B, which have a small data volume, an increase in traffic on the communication network 100 can be suppressed, and the load on equipment related to data transmission, reception and storage can also be reduced.
[0066] In the modern construction, manufacturing, and logistics industries, reliance on foreign workers is rapidly increasing, but miscommunication due to language barriers is a direct cause of serious industrial accidents. According to the communication system 1 of this embodiment, users can communicate in their own language even if they speak different languages, thus suppressing such miscommunication.
[0067] Next, a communication system according to a second embodiment of the present invention will be described. This embodiment differs from the first embodiment described above only in the functional components of the terminal 10 and the content of the information and data transmitted and received. Therefore, components, information, and data common to the first embodiment will be described using the same reference numerals as in the first embodiment.
[0068] Figure 8 is a conceptual diagram showing the functional components of a terminal 10 in a communication system according to a second embodiment of the present invention. In this embodiment, the text translation unit 139 is configured to convert message data T1 generated by the message data generation unit 33 into text data in a predetermined reference language (e.g., English). The text translation unit 139 stores the text data obtained by the conversion as reference language message data T4 in the storage unit 16. In this embodiment, the communication unit 15 is configured to send and receive this reference language message data T4 in addition to the message data T1 and metadata 50. Furthermore, the text translation unit 139 is configured to convert the reference language message data T4 received by the receiving unit 19 into text data in the language used by the user when translation is required. The text translation unit 139 stores the text data obtained by the conversion as translated message data T3 in the storage unit 16.
[0069] Next, in this embodiment, we will describe an example of sending a message from terminal 10A owned by user A to terminal 10B owned by another user B. Figures 9A to 9C are flowcharts showing the process of sending such a message.
[0070] The process from starting the message program (step S1) to generating message data T1A corresponding to user A's voice (step S5) is the same as in the first embodiment. Once the generation of message data T1A is complete, the text translation unit 139A converts the generated message data T1A into a predetermined reference language (e.g., English) to generate reference language message data T4A (step S50).
[0071] After the reference language message data T4A is generated, the input voice data S1A is analyzed in the same manner as in the first embodiment to generate feature information M5A representing the characteristics of user A's voice and ambient sound information M6A representing the surrounding ambient sounds (steps S7 and S8), and metadata related to user A's voice is formatted into a predetermined format to create text-based metadata 50A (step S9).
[0072] Next, the transmitting unit 18A specifies the terminal 10B owned by user B and transmits message data T1A, reference language message data T4A, and metadata 50A to the server device 20 (step S51). The message data T1A, reference language message data T4A, and metadata 50A are received by the server device 20 (step S52), and the server device 20 transmits the message data T1A, reference language message data T4A, and metadata 50A to the terminal 10B, for example, by push notification (step S53).
[0073] The receiving unit 19B of terminal 10B receives the message data T1A, the reference language message data T4A, and the metadata 50A transmitted by push notification (step S54), and determines whether the language represented by the language information M2B acquired by the metadata acquisition unit 36B is the same as the language represented by the language information M2A in the metadata 50A transmitted from terminal 10A (step S55). If the language represented by the language information M2B and the language represented by the language information M2A are the same, the same processing as in the first embodiment is performed, and the synthesized speech data S2A corresponding to the message data T1A is output from the speaker 14B and played back, ending the process (steps S15, S16, S17, and S18).
[0074] On the other hand, if the language represented by language information M2B is different from the language represented by language information M2A, the text translation unit 139B converts the reference language message data T4A transmitted from terminal 10A into the language represented by language information M2B (for example, Thai) to generate translated message data T3A (step S56). Subsequently, a conversation screen including the translated message data T3A is displayed on the screen as a message from user A (step S20), and the synthesized speech data generation unit 40B generates synthesized speech data S2A from the translated message data T3A (step S21). Once the synthesized speech data generation unit 40B has generated the synthesized speech data S2A, the speech playback unit 41B outputs the synthesized speech data S2A from speaker 14B and plays it back (step S17), and the process ends (step S18).
[0075] Thus, in this embodiment, the reference language message data T4A, obtained by converting message data T1A into text data of the reference language, is transmitted from terminal 10A to terminal 10B. If the language used by user B is different from the language represented by the language information M2A in metadata 50A, the reference language message data T4A is translated into the language used by user B to generate translated message data T3A, and synthesized speech data S2A is generated from this translated message data T3A. Therefore, regardless of the language used by user A, terminal 10B can generate synthesized speech data S2A as long as it can convert from the reference language to the language used by user B. In other words, terminal 10B does not need to prepare translation engines for a wide variety of languages; it only needs to have a translation engine for the reference language to convert to the language used by user A. Furthermore, in terminal 10A, when generating the reference language message data T4A, only a translation engine from the language used by user A to the reference language is required, and translation engines for other languages are not necessary.
[0076] Next, we will describe an example in which a reply to a message from user A is sent from terminal 10B to terminal 10A. Figures 10A to 10C are flowcharts showing the process of sending such a message. From the input of the voice message (step S30) to the generation of message data T1B corresponding to user B's voice (step S32), the process is the same as in the first embodiment. Once the generation of message data T1B is complete, the text translation unit 139B converts the generated message data T1B into a reference language to generate reference language message data T4B (step S60).
[0077] After the reference language message data T4B is generated, the input voice data S1B is analyzed in the same manner as in the first embodiment to generate feature information M5B representing the characteristics of user B's voice and ambient sound information M6B representing the surrounding ambient sounds (steps S33 and S34), and metadata related to user B's voice is formatted into a predetermined format to create text-based metadata 50B (step S35).
[0078] Next, the transmitting unit 18B specifies terminal 10A and transmits message data T1B, reference language message data T4B, and metadata 50B to the server device 20 (step S61). The message data T1B, reference language message data T4B, and metadata 50B are received by the server device 20 (step S62), and the server device 20 transmits the message data T1B, reference language message data T4B, and metadata 50B to terminal 10A, for example, by push notification (step S63).
[0079] The receiving unit 19A of terminal 10A receives the message data T1B, the reference language message data T4B, and the metadata 50B transmitted by push notification (step S64), and determines whether the language represented by the language information M2A acquired by the metadata acquisition unit 36A is the same as the language represented by the language information M2B in the metadata 50B transmitted from terminal 10B (step S65). If the language represented by the language information M2A and the language represented by the language information M2B are the same, the same processing as in the first embodiment is performed, and the synthesized speech data S2B corresponding to the message data T1B is output from the speaker 14A and played back, ending the process (steps S41, S42, S43, S44).
[0080] On the other hand, if the language represented by language information M2A is different from the language represented by language information M2B, the text translation unit 139A converts the reference language message data T4B transmitted from terminal 10B into the language represented by language information M2A (for example, Japanese) to generate translated message data T3B (step S66). Subsequently, a conversation screen including the translated message data T3B as a message from user B is displayed on the screen (step S46), and the synthesized speech data generation unit 40A generates synthesized speech data S2B from the translated message data T3B (step S47). Once the synthesized speech data S2B is generated by the synthesized speech data generation unit 40A, the speech playback unit 41A outputs the synthesized speech data S2B from speaker 14A and plays it back (step S43), and the process ends (step S44).
[0081] Thus, in this embodiment, the reference language message data T4B, obtained by converting message data T1B into text data of the reference language, is transmitted from terminal 10B to terminal 10A. If the language used by user A is different from the language represented by the language information M2B in metadata 50B, the reference language message data T4B is translated into the language used by user A to generate translated message data T3B, and synthesized speech data S2B is generated from this translated message data T3B. Therefore, regardless of the language used by user B, terminal 10A can generate synthesized speech data S2B as long as it can convert from the reference language to the language used by user B. In other words, terminal 10A does not need to prepare translation engines for a wide variety of languages; it only needs to have a translation engine for the reference language to be able to convert to the language used by user B. Furthermore, in terminal 10B, when generating the reference language message data T4B, only a translation engine from the language used by user B to the reference language is required, and translation engines for other languages are not necessary.
[0082] As described above, in this embodiment, both terminal 10A and terminal 10B only need to be equipped with a translation engine that performs translation between the language used by each user and the reference language, and translation engines for a wide variety of languages are not required. Therefore, the memory capacity required by terminals 10A and 10B can be reduced.
[0083] In the above embodiment, an example of sending and receiving messages between two terminals 10A and 10B was described, but the communication system 1 can also be configured to send and receive messages with three or more terminals participating. It is also possible to configure the system in such a way that, after messages have been sent and received between specific terminals 10, another terminal 10 can join the conversation.
[0084] In the embodiment described above, the storage unit 16A of terminal 10A may store message data T1A generated by the message data generation unit 33A, metadata 50A generated by the metadata generation unit 38A, and message data T1B or translated message data T3B and metadata 50B received from terminal 10B. In this case, when user A taps the play button B7 (see Figure 4D) on the conversation screen D2, the synthesized voice data generation unit 40A may generate synthesized voice data S2A from the message data T1A and metadata 50A stored in the storage unit 16A, and the voice playback unit 41A may play the synthesized voice data S2A. Alternatively, when user A taps the play button B8 (see Figure 4D) on the conversation screen D2, the synthesized voice data generation unit 40A may generate synthesized voice data S2B from the message data T1B or translated message data T3B and metadata 50B stored in the storage unit 16A, and the voice playback unit 41A may play the synthesized voice data S2A.
[0085] Similarly, the storage unit 16B of terminal 10B may store message data T1B generated by the message data generation unit 33B, metadata 50B generated by the metadata generation unit 38B, and message data T1A or translated message data T3A and metadata 50A received from terminal 10A. In this case, when user B taps the play button F5 (see Figure 5B) on the conversation screen E1, the synthesized voice data generation unit 40B may generate synthesized voice data S2A from the message data T1A and metadata 50A stored in the storage unit 16B, and the voice playback unit 41B may play the synthesized voice data S2A. Alternatively, when user B taps the play button F6 (see Figure 5B) on the conversation screen E1 or the play button F8 (see Figure 7B) on the conversation screen E2, the synthesized voice data generation unit 40B may generate synthesized voice data S2B from the message data T1B and metadata 50B stored in the storage unit 16B, and the voice playback unit 41B may play the synthesized voice data S2B. Alternatively, when user B taps the play button F7 (see Figure 7B) on the conversation screen E2, the synthesized speech data generation unit 40B generates synthesized speech data S2A from the translated message data T3A and metadata 50A stored in the storage unit 16B, and the speech playback unit 41B plays the synthesized speech data S2A.
[0086] Furthermore, the message data T1 and metadata 50 transmitted from the terminal 10 may be stored in the storage unit 21 of the server device 20. In this case, as described above, if another terminal 10 joins the conversation after messages have been sent and received between specific terminals 10, the past message data T1 and metadata 50 may be sent to the newly joined terminal 10 in chronological order, the terminal 10 may generate synthesized speech data S2, and the speech playback unit 41 may play back the synthesized speech data S2.
[0087] Alternatively, the server device 20 may be equipped with a user authentication function to authenticate users on terminal 10. In this case, by assigning specific attributes (for example, voice types or voice characteristics that evoke distrust or discomfort) to the voice type information M4 or feature information M5 in the metadata 50 sent from the terminal 10 of an unauthenticated user and sending it to other terminals 10, users on other terminals 10 can know that the message is from an unauthenticated user, thereby promoting the health of the community built by the communication system 1.
[0088] In the embodiment described above, the message data acquisition unit 31 of the terminal 10 includes an input voice data generation unit 32 that generates input voice data S1 from the user's voice input via the microphone 13, and a message data generation unit 33 that generates message data T1 by converting the voice contained in the input voice data S1 into text data through voice recognition processing. However, the message data acquisition unit 31 may also directly input text-format message data T1 by inputting text using, for example, an input interface displayed on the screen.
[0089] The metadata 50 described above is merely an example and is not limited to it. For example, GPS information of terminal 10 may be included in metadata 50. Also, avatar information representing an avatar associated with the user (such as a 2D character, 3D character, or user's face photo) may be included in metadata 50. In this case, when the receiving terminal 10 displays a message from the sender, it may display the avatar corresponding to the avatar information in the received metadata 50. Furthermore, an animation of the avatar can be generated based on the feature information M5 in the received metadata 50 and displayed together with the message from the sender. This avatar animation may include, for example, the avatar moving its hands and feet, moving its mouth, or changing its facial expressions.
[0090] Furthermore, the voice types that users can select (voice type information M4) may be restricted in advance, and in order to select a restricted voice type, the user may need to purchase the voice type via the server device 20 or by other means. Also, if avatar information is included in metadata 50, the avatars that users can select (avatar information) may be restricted in advance, and in order to select a restricted avatar, the user may need to purchase the voice type via the server device 20 or by other means.
[0091] Alternatively, each terminal 10 may be made capable of purchasing tokens with monetary value, and when the feature information generation unit 34 generates feature information M5, it may send an amount of tokens corresponding to the calculated feature quantity to the conversation partner's terminal 10. In this way, it becomes possible to assign monetary value to the conversation partner (e.g., the broadcaster) in response to emotions such as laughter, exclamation, or excitement.
[0092] In the embodiment described above, data transmission and reception between terminals 10 is mediated by the server device 20. However, if the terminals 10 can be configured to directly transmit and receive data to each other via a communication network such as Bluetooth® or WiFi, the server device 20 is not necessary. Alternatively, for example, the server device 20 could function as a signaling server, exchange network address information between terminals 10, and then switch to the UDP direct communication protocol without going through the server device 20 to directly transmit and receive data between terminals 10. Furthermore, it is possible to make the terminal 10 a transceiver-like device by omitting the display unit 11 and touch sensor 12 from the terminal 10.
[0093] In this specification, even if a device or structure is described as comprising certain components, those components may be optional. Furthermore, the device or structure described above may include additional components not mentioned herein. Similarly, even if a method is described as including certain steps or operations, those steps or operations may be optional. Furthermore, the method described above may include additional steps or operations not mentioned herein. The steps or operations described in the embodiments described above may be performed in any order, unless it is impossible to perform them, and may be performed simultaneously if possible.
[0094] The various effects described in the embodiments above are not always obtainable by the present invention. On the other hand, effects not described in the embodiments above may also be obtainable by the present invention.
[0095] As described above, the communication terminal according to the present invention can adopt the following configuration. [Configuration 1] A sound output section capable of outputting sound, A message data acquisition unit acquires first message data, which is a text-formatted description of the first message sent by the first user. A metadata generation unit generates first metadata that describes information related to the above-mentioned first user in text format, A communication unit capable of communicating with another terminal owned by a second user, A transmission unit capable of transmitting the above-mentioned first message data and the above-mentioned first metadata to the above-mentioned other terminal, A receiving unit capable of receiving from another terminal the second message data, which is a text-formatted description of the second message sent by the second user, and the second metadata, which is a text-formatted description of information related to the second user. The communications department, including, A synthesized speech data generation unit generates synthesized speech data corresponding to the second message data using the second metadata described above, The above synthesized speech data is played back by the above sound output unit and A communication terminal equipped with the following features.
[0096] [Configuration 2] It also features an audio input section that allows for sound input, The above message data acquisition unit is: An input audio data generation unit generates input audio data from the audio of the first message transmitted by the first user, which is input via the above audio input unit, A message data generation unit generates text data corresponding to the first transmission message from the input voice data using speech recognition processing, and Includes, The above-mentioned first metadata describes information related to the audio of the above-mentioned first message sent by the above-mentioned first user, The above second message data is text data corresponding to the above second message, generated from the audio of the above second message sent by the above second user. The above second metadata describes information related to the audio of the above second user's second sent message. The communication terminal described in Configuration 1.
[0097] [Configuration 3] The system further includes a feature information generation unit that analyzes the above input audio data to generate first feature information representing the characteristics of the first user's voice, The metadata generation unit described above is configured to include the first feature information generated by the feature information generation unit described above in the first metadata, The above second metadata includes second feature information representing the voice characteristics of the above second user, The above-mentioned synthesized speech data generation unit is configured to generate the synthesized speech data so as to reflect the features represented by the second feature information in the second metadata. The communication terminal described in Configuration 2.
[0098] [Structure 4] The system further includes an ambient sound information generation unit that analyzes the above input audio data to generate first ambient sound information representing the surrounding ambient sounds. The metadata generation unit described above is configured to include the first ambient sound information generated by the ambient sound information generation unit described above in the first metadata, The above second metadata includes second ambient sound information representing the ambient sound of the other terminal, The above-mentioned synthesized speech data generation unit is configured to add the ambient sound represented by the second ambient sound information in the second metadata to the synthesized speech data. A communication terminal as described in configuration 2 or 3.
[0099] [Composition 5] It further includes a text translation unit that converts input text data into text data in another language. The metadata generation unit described above is configured to include first language information representing the first language used by the first user in the first metadata. The above-mentioned second metadata includes second language information representing the second language used by the above-mentioned second user, The above text translation unit, when the second language represented by the second language information in the second metadata is different from the first language, translates the second message data from the second language to the first language to generate translated message data. The above synthesized speech data generation unit is: If the second language represented by the second language information in the second metadata is the same as the first language, the synthesized speech data is generated from the second message data. If the second language represented by the second language information in the second metadata is different from the first language, the synthesized speech data is generated from the translated message data. It is configured in such a way. A communication terminal as described in any of configurations 1 to 4.
[0100] [Composition 6] It further includes a text translation unit that converts input text data into text data in another language. The metadata generation unit described above is configured to include first language information representing the first language used by the first user in the first metadata. The above-mentioned second metadata includes second language information representing the second language used by the above-mentioned second user, The above text translation unit translates the above first message data from the above first language to the base language to generate first base language message data. The above-mentioned transmission unit is configured to transmit the first standard language message data in addition to the first message data to the above-mentioned other terminal. The receiving unit described above is configured to receive a second reference language message data from the other terminal in addition to the second message data described above. The above text translation unit, when the second language represented by the second language information in the second metadata is different from the first language, translates the second reference language message data from the reference language to the first language to generate translated message data. The above synthesized speech data generation unit is: If the second language represented by the second language information in the second metadata is the same as the first language, the synthesized speech data is generated from the second message data. If the second language represented by the second language information in the second metadata is different from the first language, the synthesized speech data is generated from the translated message data. It is configured in such a way. A communication terminal as described in any of configurations 1 to 4.
[0101] [Composition 7] The metadata generation unit described above is configured to include first gender information representing the gender of the first user in the first metadata. The above second metadata includes second gender information representing the gender of the above second user, The above-mentioned synthesized voice data generation unit is configured to generate the synthesized voice data using a voice of the gender represented by the second gender information in the second metadata. A communication terminal as described in any of configurations 1 to 6.
[0102] [Structure 8] The metadata generation unit described above is configured to include first voice type information representing the voice type associated with the first user in the first metadata. The above second metadata includes second voice type information representing the voice type associated with the above second user, The above-mentioned synthesized speech data generation unit is configured to generate the synthesized speech data using the voice of the voice type represented by the second voice type information in the second metadata. A communication terminal as described in any of configurations 1 through 7.
[0103] The communication system according to the present invention can employ the following configuration. [Composition 9] It consists of a communication terminal as described in any of configurations 1 to 8, and the first communication terminal owned by the first user, It consists of a communication terminal as described in any of Configurations 1 to 8, and the second communication terminal owned by the second user, A server device is connected to each of the first and second communication terminals via a communication network and configured to mediate the transmission and reception of the first message data and the first metadata, and the second message data and the second metadata, between the first and second communication terminals. A communication system equipped with these features.
[0104] The information processing method according to the present invention can employ the following configuration. [Configuration 10] To obtain the first message data entered by the first user, To generate first metadata related to the voice of the first user mentioned above, The above-mentioned first message data and the above-mentioned first metadata are to be transmitted to another terminal owned by the second user, The second message data input by the second user and the second metadata related to the voice of the second user are received from the other terminal. Using the above second metadata, synthesized speech data corresponding to the above second message data is generated, The above synthesized speech data is played back via the sound output section of the above terminal. Information processing methods, including those mentioned above.
[0105] [Composition 11] The acquisition of the above first message data is as follows: The process involves generating input audio data from the voice of the first user, which is input via the audio input section of the above-mentioned terminal, The process involves generating text data corresponding to the voice of the first user from the input voice data using speech recognition processing, and generating the first message data from the input voice data. The information processing method described in configuration 10, including the above.
[0106] [Composition 12] The above input audio data is analyzed to generate first feature information representing the characteristics of the first user's voice. The above-mentioned first metadata includes the above-mentioned first feature information, The above second metadata includes second feature information representing the voice characteristics of the above second user, The generation of the synthesized speech data described above includes generating the synthesized speech data so as to reflect the features represented by the second feature information in the second metadata described above. The information processing method described in configuration 11.
[0107] [Composition 13] The above input audio data is analyzed to generate first ambient sound information representing the surrounding ambient sounds. The above-mentioned first metadata includes the above-mentioned first ambient sound information, The above second metadata includes second ambient sound information representing the ambient sound of the other terminal, The generation of the above-mentioned synthesized speech data includes adding the ambient sound represented by the second ambient sound information in the second metadata to the above-mentioned synthesized speech data. The information processing method described in configuration 11 or 12.
[0108] [Composition 14] The above-mentioned first metadata includes first language information representing the first language used by the above-mentioned first user, The above-mentioned second metadata includes second language information representing the second language used by the above-mentioned second user, If the second language represented by the second language information in the second metadata is different from the first language, the second message data is translated from the second language to the first language to generate translated message data. The above synthesized speech data is generated by When the second language represented by the second language information in the second metadata is the same as the first language, the synthesized speech data is generated from the second message data. When the second language represented by the second language information in the second metadata is different from the first language, the synthesized speech data is generated from the translated message data. including, An information processing method described in any of configurations 10 to 13.
[0109] [Composition 15] The above-mentioned first metadata includes first language information representing the first language used by the above-mentioned first user, The above-mentioned second metadata includes second language information representing the second language used by the above-mentioned second user, The above first message data is translated from the above first language to the base language to generate the first base language message data. In addition to the first message data mentioned above, the first standard language message data is sent to the other terminal mentioned above. In addition to the second message data mentioned above, the second standard language message data is received from the other terminal mentioned above. If the second language represented by the second language information in the second metadata is different from the first language, the second reference language message data is translated from the reference language to the first language to generate translated message data. The above synthesized speech data is generated by When the second language represented by the second language information in the second metadata is the same as the first language, the synthesized speech data is generated from the second message data. When the second language represented by the second language information in the second metadata is different from the first language, the synthesized speech data is generated from the translated message data. including, An information processing method described in any of configurations 10 to 13.
[0110] [Composition 16] The above-mentioned first metadata includes first gender information representing the gender of the above-mentioned first user, The above second metadata includes second gender information representing the gender of the above second user, The generation of the above-mentioned synthesized voice data includes generating the above-mentioned synthesized voice data in a voice of the gender represented by the second gender information in the second metadata. An information processing method described in any of configurations 10 to 15.
[0111] [Composition 17] The above-mentioned first metadata includes first voice type information representing the voice type associated with the above-mentioned first user, The above second metadata includes second voice type information representing the voice type associated with the above second user, The generation of the above-mentioned synthesized speech data includes generating the above-mentioned synthesized speech data using the voice of the voice type represented by the second voice type information in the second metadata. An information processing method described in any of configurations 10 to 16.
[0112] The program according to the present invention can employ the following configuration. [Composition 18] A program for causing a computer to execute one of the information processing methods described in configuration 10 to 17.
[0113] Although preferred embodiments of the present invention have been described above, it goes without saying that the present invention is not limited to the embodiments described above and may be implemented in various different forms within the scope of its technical concept. [Explanation of Symbols]
[0114] 1. Communication System 10, 10A, 10B terminals 11 Display section 12 touch sensors 13, 13A, 13B Microphone (sound input section) 14, 14A, 14B Speaker (sound output section) 15 Communications Department 16,16A,16B Storage section 17, 17A, 17B Control Unit 18, 18A, 18B Transmitter 19, 19A, 19B Receiver 20 Server Devices 21 Memory section 22 Communications Department 23 Control Unit 31 Message data acquisition unit 32, 32A, 32B Input audio data generation unit 33, 33A, 33B Message data generation unit 34, 34A, 34B Feature Information Generation Unit 35,35A,35B Environmental sound information generation section 36, 36A, 36B Metadata acquisition unit 38, 38A, 38B Metadata generation unit 39, 39A, 39B Text Translation Department 40, 40A, 40B Synthesized Speech Data Generation Unit 41, 41A, 41B Audio playback section 50, 50A, 50B metadata 60,60A Contact User List 100 Communication Networks
Claims
1. A sound output section capable of outputting sound, A message data acquisition unit acquires first message data, which is a text-formatted description of the first message sent by the first user. A metadata generation unit generates first metadata that describes information related to the first user in text format, A communication unit capable of communicating with another terminal owned by a second user, A transmission unit capable of transmitting the first message data and the first metadata to the other terminal, A receiving unit capable of receiving from another terminal a second message data, which is a text-formatted description of the second message sent by the second user, and a second metadata, which is a text-formatted description of information related to the second user. The communications department, including, A synthesized speech data generation unit generates synthesized speech data corresponding to the second message data using the second metadata, A sound playback unit that plays back the synthesized speech data using the sound output unit, A text translation unit that converts input text data into text data in another language. Equipped with, The metadata generation unit is configured to include first language information representing the first language used by the first user in the first metadata, The second metadata includes second language information representing the second language used by the second user, The text translation unit, when the second language represented by the second language information in the second metadata is different from the first language, translates the second message data from the second language to the first language to generate translated message data. The synthesized speech data generation unit, If the second language represented by the second language information in the second metadata is the same as the first language, the synthesized speech data is generated from the second message data. If the second language represented by the second language information in the second metadata is different from the first language, the synthesized speech data is generated from the translated message data. It is configured in such a way. Communication terminal.
2. It also features an audio input section that allows for sound input, The message data acquisition unit, An input audio data generation unit generates input audio data from the audio of the first message transmitted by the first user, which is input via the audio input unit, A message data generation unit generates text data corresponding to the first transmission message from the input voice data using speech recognition processing, and Includes, The first metadata describes information related to the audio of the first message sent by the first user, The second message data is text data corresponding to the second message, generated from the audio of the second message sent by the second user. The second metadata describes information related to the audio of the second message sent by the second user. The communication terminal according to claim 1.
3. The system further includes a feature information generation unit that analyzes the input audio data to generate first feature information representing the characteristics of the first user's voice, The metadata generation unit is configured to include the first feature information generated by the feature information generation unit in the first metadata, The second metadata includes second feature information representing the voice characteristics of the second user, The synthesized speech data generation unit is configured to generate the synthesized speech data so as to reflect the features represented by the second feature information in the second metadata. The communication terminal according to claim 2.
4. The system further includes an ambient sound information generation unit that analyzes the input audio data to generate first ambient sound information representing ambient sounds, The metadata generation unit is configured to include the first ambient sound information generated by the ambient sound information generation unit in the first metadata. The second metadata includes second ambient sound information representing the ambient sound of the other terminal, The synthesized speech data generation unit is configured to add the ambient sound represented by the second ambient sound information in the second metadata to the synthesized speech data. The communication terminal according to claim 2.
5. The metadata generation unit is configured to include first gender information representing the gender of the first user in the first metadata, The second metadata includes second gender information representing the gender of the second user, The synthesized voice data generation unit is configured to generate the synthesized voice data in a voice of the gender represented by the second gender information in the second metadata. The communication terminal according to claim 1.
6. The metadata generation unit is configured to include first voice type information representing the voice type associated with the first user in the first metadata, The second metadata includes second voice type information representing the voice type associated with the second user, The synthesized speech data generation unit is configured to generate the synthesized speech data using the voice of the voice type represented by the second voice type information in the second metadata. The communication terminal according to claim 1.
7. A sound output section capable of outputting sound, A message data acquisition unit acquires first message data, which is a text-formatted description of the first message sent by the first user. A metadata generation unit generates first metadata that describes information related to the first user in text format, A communication unit capable of communicating with another terminal owned by a second user, A transmission unit capable of transmitting the first message data and the first metadata to the other terminal, A receiving unit capable of receiving from another terminal a second message data, which is a text-formatted description of the second message sent by the second user, and a second metadata, which is a text-formatted description of information related to the second user. The communications department, including, A synthesized speech data generation unit generates synthesized speech data corresponding to the second message data using the second metadata, A sound playback unit that plays back the synthesized speech data using the sound output unit, A text translation unit that converts input text data into text data in another language. Equipped with, The metadata generation unit is configured to include first language information representing the first language used by the first user in the first metadata, The second metadata includes second language information representing the second language used by the second user, The text translation unit translates the first message data from the first language to the reference language to generate first reference language message data. The transmission unit is configured to transmit first reference language message data in addition to the first message data to the other terminal. The receiving unit is configured to receive a second reference language message data from the other terminal in addition to the second message data. The text translation unit, when the second language represented by the second language information in the second metadata is different from the first language, translates the second reference language message data from the reference language to the first language to generate translated message data. The synthesized speech data generation unit, If the second language represented by the second language information in the second metadata is the same as the first language, the synthesized speech data is generated from the second message data. If the second language represented by the second language information in the second metadata is different from the first language, the synthesized speech data is generated from the translated message data. A communication terminal configured in such a way.
8. It also features an audio input section that allows for sound input, The message data acquisition unit, An input audio data generation unit generates input audio data from the audio of the first message transmitted by the first user, which is input via the audio input unit, A message data generation unit generates text data corresponding to the first transmission message from the input voice data using speech recognition processing, and Includes, The first metadata describes information related to the audio of the first message sent by the first user, The second message data is text data corresponding to the second message, generated from the audio of the second message sent by the second user. The second metadata describes information related to the audio of the second message sent by the second user. The communication terminal according to claim 7.
9. The system further includes a feature information generation unit that analyzes the input audio data to generate first feature information representing the characteristics of the first user's voice, The metadata generation unit is configured to include the first feature information generated by the feature information generation unit in the first metadata, The second metadata includes second feature information representing the voice characteristics of the second user, The synthesized speech data generation unit is configured to generate the synthesized speech data so as to reflect the features represented by the second feature information in the second metadata. The communication terminal according to claim 8.
10. The system further includes an ambient sound information generation unit that analyzes the input audio data to generate first ambient sound information representing ambient sounds, The metadata generation unit is configured to include the first ambient sound information generated by the ambient sound information generation unit in the first metadata. The second metadata includes second ambient sound information representing the ambient sound of the other terminal, The synthesized speech data generation unit is configured to add the ambient sound represented by the second ambient sound information in the second metadata to the synthesized speech data. The communication terminal according to claim 8.
11. The metadata generation unit is configured to include first gender information representing the gender of the first user in the first metadata, The second metadata includes second gender information representing the gender of the second user, The synthesized voice data generation unit is configured to generate the synthesized voice data in a voice of the gender represented by the second gender information in the second metadata. The communication terminal according to claim 7.
12. The metadata generation unit is configured to include first voice type information representing the voice type associated with the first user in the first metadata, The second metadata includes second voice type information representing the voice type associated with the second user, The synthesized speech data generation unit is configured to generate the synthesized speech data using the voice of the voice type represented by the second voice type information in the second metadata. The communication terminal according to claim 7.
13. The first communication terminal is comprised of the communication terminal described in any one of claims 1 to 12, and the first communication terminal is owned by the first user, A second communication terminal owned by the second user, comprising a communication terminal according to any one of claims 1 to 12, A server device connected to each of the first and second communication terminals via a communication network, and configured to mediate the transmission and reception of the first message data and the first metadata, and the second message data and the second metadata, between the first and second communication terminals. A communication system equipped with these features.
14. Obtain the first message data, which is a text-formatted description of the first message sent by the first user. To generate first metadata that describes information related to the first user in text format, Transmitting the first message data and the first metadata to another terminal owned by the second user, Receiving from the other terminal a second message data in text format, which is a second message sent by the second user, and a second metadata in text format, which is information related to the second user. Using the second metadata, generate synthesized speech data corresponding to the second message data, The synthesized speech data is played back via the sound output unit of the aforementioned terminal. Includes, The first metadata includes first language information representing the first language used by the first user, The second metadata includes second language information representing the second language used by the second user, If the second language represented by the second language information in the second metadata is different from the first language, the second message data is translated from the second language to the first language to generate translated message data. The generation of the aforementioned synthesized speech data is performed by When the second language represented by the second language information in the second metadata is the same as the first language, the synthesized speech data is generated from the second message data, When the second language represented by the second language information in the second metadata is different from the first language, the synthesized speech data is generated from the translated message data. Information processing methods, including those mentioned above.
15. The acquisition of the first message data is as follows: The process involves generating input audio data from the audio of the first message transmitted by the first user, which is input via the audio input section of the terminal, The process involves generating text data corresponding to the first transmission message from the input voice data using speech recognition processing, and generating the first message data. The information processing method according to claim 14, including the method described in claim 14.
16. The input audio data is analyzed to generate first feature information representing the characteristics of the first user's voice. The first metadata includes the first feature information, The second metadata includes second feature information representing the voice characteristics of the second user, The generation of the synthesized speech data includes generating the synthesized speech data so as to reflect the features represented by the second feature information in the second metadata. The information processing method according to claim 15.
17. The input audio data is analyzed to generate first ambient sound information representing the surrounding ambient sounds. The first metadata includes the first ambient sound information, The second metadata includes second ambient sound information representing the ambient sound of the other terminal, The generation of the synthesized speech data includes adding the ambient sound represented by the second ambient sound information in the second metadata to the synthesized speech data. The information processing method according to claim 15.
18. The first metadata includes first gender information representing the gender of the first user, The second metadata includes second gender information representing the gender of the second user, The generation of the synthesized voice data includes generating the synthesized voice data in a voice of the gender represented by the second gender information in the second metadata. The information processing method according to claim 14.
19. The first metadata includes first voice type information representing the voice type associated with the first user, The second metadata includes second voice type information representing the voice type associated with the second user, The generation of the synthesized speech data includes generating the synthesized speech data using the voice of the voice type represented by the second voice type information in the second metadata. The information processing method according to claim 14.
20. Obtain the first message data, which is a text-formatted description of the first message sent by the first user. To generate first metadata that describes information related to the first user in text format, Transmitting the first message data and the first metadata to another terminal owned by the second user, Receiving from the other terminal a second message data in text format, which is a second message sent by the second user, and a second metadata in text format, which is information related to the second user. Using the second metadata, generate synthesized speech data corresponding to the second message data, The synthesized speech data is played back via the sound output unit of the aforementioned terminal. Includes, The first metadata includes first language information representing the first language used by the first user, The second metadata includes second language information representing the second language used by the second user, The first message data is translated from the first language to the reference language to generate the first reference language message data. In addition to the first message data, the first standard language message data is transmitted to the other terminal. In addition to the second message data, a second standard language message data is received from the other terminal. If the second language represented by the second language information in the second metadata is different from the first language, the second reference language message data is translated from the reference language to the first language to generate translated message data. The generation of the aforementioned synthesized speech data is performed by When the second language represented by the second language information in the second metadata is the same as the first language, the synthesized speech data is generated from the second message data, When the second language represented by the second language information in the second metadata is different from the first language, the synthesized speech data is generated from the translated message data. Information processing methods, including those mentioned above.
21. The acquisition of the first message data is as follows: The process involves generating input audio data from the audio of the first message transmitted by the first user, which is input via the audio input section of the terminal, The process involves generating text data corresponding to the first transmission message from the input voice data using speech recognition processing, and generating the first message data. The information processing method according to claim 20, including the method described in claim 20.
22. The input audio data is analyzed to generate first feature information representing the characteristics of the first user's voice. The first metadata includes the first feature information, The second metadata includes second feature information representing the voice characteristics of the second user, The generation of the synthesized speech data includes generating the synthesized speech data so as to reflect the features represented by the second feature information in the second metadata. The information processing method according to claim 21.
23. The input audio data is analyzed to generate first ambient sound information representing the surrounding ambient sounds. The first metadata includes the first ambient sound information, The second metadata includes second ambient sound information representing the ambient sound of the other terminal, The generation of the synthesized speech data includes adding the ambient sound represented by the second ambient sound information in the second metadata to the synthesized speech data. The information processing method according to claim 21.
24. The first metadata includes first gender information representing the gender of the first user, The second metadata includes second gender information representing the gender of the second user, The generation of the synthesized voice data includes generating the synthesized voice data in a voice of the gender represented by the second gender information in the second metadata. The information processing method according to claim 20.
25. The first metadata includes first voice type information representing the voice type associated with the first user, The second metadata includes second voice type information representing the voice type associated with the second user, The generation of the synthesized speech data includes generating the synthesized speech data using the voice of the voice type represented by the second voice type information in the second metadata. The information processing method according to claim 20.
26. A program for causing a computer to execute the information processing method described in any one of claims 14 to 25.
Citation Information
Patent Citations
Device and method for transmitting audio signal
JP2000284799A
Communication equipment, telephone set and recording medium with recorded communication processing program
JP2001127900A
Information processor, information processing method, information transmission system, medium for making information processor run information processing program, and information processing program
JP2002244688A
Voice translation device and voice translation system
JP2017215555A
system
JP2026068480A