Communication system and vehicle

The communication system addresses the challenge of safe and efficient voice-based content exchange for users with mobility issues by utilizing voice input and non-overlapping audio playback, ensuring safe and effective operation in vehicles.

JP7802417B2Active Publication Date: 2026-01-20CASE CHARTER CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025069689
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2012-03-28
Filing Date
2025-04-21
Publication Date
2026-01-20
Estimated Expiration
2032-12-14

AI Technical Summary

Technical Problem

Users with difficulty operating or viewing terminal devices, such as mobile phones, face challenges in safely sending and receiving content like tweets, especially while driving a vehicle.

Method used

A communication system with terminal devices equipped with voice input conversion, transmission, and reproduction capabilities, allowing users to input and receive content via voice, with features like non-overlapping audio playback, motion detection for safe operation, and integration with vehicle systems to enable hands-free input.

Benefits of technology

Enables safe and efficient voice-based communication for users with mobility issues, enhancing safety and usability for vehicle occupants by allowing hands-free input and clear audio playback without overlap.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007802417000001
    Figure 0007802417000001
  • Figure 0007802417000002
    Figure 0007802417000002
  • Figure 0007802417000003
    Figure 0007802417000003
Patent Text Reader

Abstract

To provide a communication system and a terminal device that allow even users who have difficulty in operating or viewing the terminal device to safely send and receive outgoing contents such as contents they want to murmur.SOLUTION: A communication system 1 includes a plurality of terminal devices 31, 32, and 33 that can communicate with each other. Each terminal device includes audio input conversion means, audio transmission means, audio reception means, and audio reproduction means. When there are a plurality of audio signals the reproduction of which has not been completed, the audio reproduction means arranges pieces of audio corresponding to the respective audio signals so as not to overlap with each other and then reproduces them.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This international application claims priority based on Japanese Patent Application Nos. 2011-273578 and 2011-273579, filed with the Japan Patent Office on December 14, 2011, and Japanese Patent Application Nos. 2012-074627 and 2012-074628, filed with the Japan Patent Office on March 28, 2012, and the entire contents of Japanese Patent Application Nos. 2011-273578, 2011-273579, 2012-074627, and 2012-074628 are incorporated herein by reference. [Technical Field]

[0002] The present invention relates to a communication system including a plurality of terminal devices, and to the terminal devices. [Background technology]

[0003] A system is known in which a large number of users use terminal devices such as mobile phones to input text content, such as what they want to tweet, and a server distributes the content in order to each terminal device (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2011-215861 Summary of the Invention [Problem to be solved by the invention]

[0005] It is preferable that even a user who has difficulty operating or viewing a terminal device can safely send and receive content such as tweets. [Means for solving the problem]

[0006] One aspect of the present invention is a method for producing a medicament for the treatment of a cancer, comprising: A communication system including a plurality of terminal devices capable of communicating with each other, Each terminal device a voice input conversion means for converting voice input by a user into a voice signal; a voice transmitting means for transmitting the voice signal to a device including another terminal device; an audio receiving means for receiving an audio signal transmitted by another terminal device; an audio reproducing means for reproducing the received audio signal; Equipped with When there are a plurality of audio signals whose reproduction has not been completed, the audio reproduction means aligns the audio signals so that the audio corresponding to each audio signal does not overlap, and then reproduces the audio signals.

[0007] Another aspect of the present invention is A communication system including a plurality of terminal devices and a server device capable of communicating with each of the plurality of terminal devices, Each terminal device a voice input conversion means for converting voice input by a user into a voice signal; an audio transmitting means for transmitting the audio signal to the server device; an audio receiving means for receiving an audio signal transmitted by the server device; an audio reproducing means for reproducing the received audio signal; Equipped with The server device The system includes a distribution means for receiving the audio signals transmitted from the plurality of terminal devices, aligning the audio signals so that the audio corresponding to each audio signal does not overlap, and then distributing the audio signals to the plurality of terminal devices.

[0008] With these communication systems, it is possible to input the content of a message by voice and play back the content of a message from another communication device by voice, so even users who have difficulty operating or seeing a terminal device can safely send and receive content such as tweets.

[0009] Furthermore, in these communication systems, when audio is played back, it is possible to play back audio without overlapping, making it easier to hear the audio. In any of the above communication systems, the audio reproducing means may be configured to re-play an audio signal that has already been reproduced, in particular, the audio signal that has just been reproduced, upon receiving a specific command from the user.

[0010] According to such a communication system, an audio signal that cannot be heard due to noise or the like can be reproduced again at the user's command. Furthermore, in any of the above communication systems, the terminal device may be mounted on a vehicle, and the terminal device may include a motion detection means for detecting the movement of the user's hands on the steering wheel of the vehicle, and an operation control means for causing the voice input conversion means to start operation when the motion detection means detects a specific start motion that starts voice input, and for causing the voice input conversion means to end operation when the motion detection means detects a specific end motion that ends voice input.

[0011] According to this communication system, the user can input voice signals from the steering wheel, which allows the user to input voice signals more safely while driving the vehicle. In any of the above communication systems, the specific start motion and the specific end motion may be a hand movement in a direction different from the operating direction when operating the horn.

[0012] According to such a communication system, it is possible to prevent the horn from being sounded when a specific start or end action is performed. Next, the terminal device may be configured as a terminal device that constitutes any one of the above communication systems.

[0013] Such a terminal device can provide the same effects as any of the above-mentioned communication systems. In addition, in the above communication system (terminal device), the voice transmitting means may generate a new voice signal by adding, to the voice signal input by the user, identification information for identifying the user in the form of voice, and transmit the generated voice signal.

[0014] In addition, in a communication system having a plurality of terminal devices and a server device capable of communicating with each of the plurality of terminal devices, the distribution means may generate a new audio signal by adding, in audio form, identification information for identifying the user to the audio signal input by the user, and distribute the generated audio signal.

[0015] With these communication systems, it is possible to input the content of a message by voice and play back the content of a message from another communication device by voice, so even users who have difficulty operating or seeing a terminal device can safely send and receive content such as tweets.

[0016] In addition, in such communication systems, identification information for identifying a user can be added by voice, so that a user can transmit his / her own identification information to other users by voice without speaking it. The source user can be notified.

[0017] Furthermore, in any of the above communication systems, the added identification information may be the user's real name or a nickname other than the real name. Furthermore, the content added as identification information (e.g., real name, first nickname, second nickname, etc.) may be changed depending on the communication mode. For example, each party may exchange information for identifying itself in advance, and the added identification information may be changed depending on whether the communication partner is a previously registered partner.

[0018] Furthermore, any of the above communication systems may be provided with a keyword extraction means for extracting a predetermined keyword contained in the audio signal from the audio signal before playing the audio signal, and a voice replacement means for replacing the keyword contained in the audio signal with the voice of another word or a predetermined sound.

[0019] With such a communication system, words that are undesirable for distribution (such as obscene words, words indicating people's names, or vulgar words) can be registered as keywords and then replaced with words or sounds that are acceptable for distribution before distribution.

[0020] The voice replacement means may or may not be activated depending on conditions such as the communication partner. Furthermore, meaningless words typical of spoken language, such as "um" and "errr," may be omitted from the generated voice signal. In this case, the playback time of the voice signal can be shortened by the amount of omission.

[0021] In addition, in any of the above communication systems, each terminal device may be equipped with a location information transmitting means for acquiring its own location information and transmitting the location information to other terminal devices in association with an audio signal transmitted from its own terminal device, a location information acquiring means for acquiring the location information transmitted by other terminal devices, and a volume control means for controlling the volume of the audio output when the audio is played back by the audio playback means, depending on the positional relationship between the position of the terminal device corresponding to the audio to be played back and the position of its own terminal device.

[0022] According to such a communication system, the volume is controlled according to the positional relationship between the position of the terminal device corresponding to the audio being played and the position of the user's own terminal device, so that the user listening to the audio can intuitively grasp the positional relationship between other terminal devices and their own terminal device.

[0023] The volume control means may control the volume so that the volume decreases as the distance to the other terminal device increases, or, when audio is output from multiple speakers, may control the volume so that the volume increases from the direction in which the other terminal device is located as seen from the user, depending on the orientation in which the other terminal device is located relative to the position of the user's own terminal device. Also, such controls may be combined.

[0024] Furthermore, the above communication system may comprise a character conversion means for converting a voice signal into characters, and a character output means for outputting the converted characters to a display device. Such a communication system allows users to check information they missed in text form.

[0025] In addition, in any of the above communication systems, information indicating the vehicle's status and the user's driving operations (operation status of lights, wipers, television, radio, etc., driving status such as driving speed and direction of travel, detection values ​​from vehicle sensors, control status such as whether or not there is a malfunction, etc.) may be transmitted along with the voice.

[0026] In this case, this information may be output by voice or text. For example, various information such as traffic congestion information may be generated based on information obtained from other terminal devices, and the generated information may be output.

[0027] In any of the above communication systems, the configuration in which each terminal device exchanges voice signals directly with each other and the configuration in which each terminal device exchanges voice signals with each other via a server device may be switched depending on pre-set conditions (for example, the distance to the base station connected to the server device, the communication status between the terminal devices, or user settings).

[0028] Furthermore, advertisements may be played by audio at regular intervals, at regular intervals of a certain number of audio signal playbacks, or according to the location of the terminal device. The content of the advertisement may be transmitted from the advertiser's terminal device or a server device to the terminal device that plays the audio signal.

[0029] At this time, the user may input a predetermined operation into the terminal device to enable a call (communication) with the advertiser, or to guide the user of the terminal device to the advertiser's store.

[0030] Furthermore, in any of the above communication systems, the communication partners may be an unspecified number of people, or may be limited. For example, in the case of a configuration in which position information is exchanged, the communication partner may be a terminal device within a certain distance, a specific user (such as a preceding vehicle or an oncoming vehicle (identified by capturing a license plate image using a camera or by GPS)), or a user within a preset group. Also, the communication partner may be changed by setting (the communication mode may be changed).

[0031] Also, if the terminal device is mounted on a vehicle, the communication partner may be a terminal device traveling in the same direction on the same road, or a terminal device with matching vehicle behavior (for example, a terminal device that is removed from the communication partner if the vehicle deviates onto a side road).

[0032] In addition, it may be configured so that other users can be specified and registered as favorites or excluded in advance, so that a user can select and communicate with a preferred partner. In this case, communication partner information specifying the communication partner may be transmitted when transmitting voice, or, if a server device is present, communication partner information may be registered in the server device, so that voice signals can be transmitted only to specified partners.

[0033] Alternatively, a plurality of switches such as buttons may be associated with respective directions, and when the user uses the switch to specify the direction in which the communication partner is located, only users in that direction may be set as communication partners.

[0034] Furthermore, the emotion and politeness of the user's speech when inputting the voice may be determined by detecting the amplitude of the input level and polite expressions contained in the voice from the voice signal, and voices from users who are emotional or do not speak politely may not be distributed.

[0035] Furthermore, a user who speaks ill of others, or who is emotional or does not speak politely may be notified of a gaffe, etc. In this case, the decision may be made based on whether or not a word registered as a keyword for abusive language is included in the voice signal.

[0036] Furthermore, in any of the above communication systems, audio signals are exchanged, but it is also possible to acquire information consisting of text, convert this information into audio, and play it back. When transmitting voice, the voice may be converted into text before transmission, and then restored to voice before playback. This configuration can reduce the amount of data to be transmitted via communication.

[0037] Furthermore, in a configuration in which data is transmitted by characters in this way, the data may be transmitted after being translated into the language that produces the smallest amount of data from among a plurality of preset languages.

[0038] In recent years, however, there have been many cases where users have posted information on social media platforms such as Twitter (registered trademark) and Facebook (registered trademark). The number of users of services that manage information sent by users in a manner that allows other users to view the information is rapidly increasing. In these services, terminal devices such as personal computers and mobile phones (including so-called smartphones) are used to send and view information (Japanese Patent Application Laid-Open No. 2011-215861).

[0039] On the other hand, a user (e.g., a driver of a vehicle) who is in a vehicle (e.g., an automobile) is often limited in his or her actions compared to when not in the vehicle, and therefore has a lot of free time. For this reason, using the above-mentioned services while in the vehicle allows the user to make effective use of their time. However, it is difficult to say that conventional services have fully taken into consideration the usage patterns of users in the vehicle.

[0040] Therefore, there is a need to provide technology for building useful services for users riding in vehicles. As a technology for this purpose, the terminal device of the first reference invention comprises an input means for inputting voice, an acquisition means for acquiring location information, a transmission means for transmitting voice information representing the voice input by the input means to a server shared by a plurality of the terminal devices in a form in which the utterance position, which is the location information acquired by the acquisition means at the time of input of the voice, can be identified, and a playback means for receiving the voice information transmitted from another of the terminal devices to the server from the server and playing back the voice represented by the voice information, and the playback means plays back the voice represented by the voice information transmitted from the other of the terminal devices to the server, whose utterance position is located within a surrounding area set based on the location information acquired by the acquisition means.

[0041] Voices emitted by a user at a certain location are often useful to other users who are present at (or heading to) that location. In particular, the location of a user in a vehicle can change significantly over time (as the vehicle travels) compared to a user who is not in a vehicle. Therefore, for a user in a vehicle, voices emitted by other users within the surrounding area based on the current location can be useful information.

[0042] In this regard, with the above-described configuration, a user of a certain terminal device A can hear, among the voices uttered by a user of another terminal device B (another vehicle), voices uttered within a surrounding area set based on the current location of the terminal device A. Therefore, with such a terminal device, it is possible to provide useful information to a user in a vehicle.

[0043] In the terminal device described above, the surrounding area may be an area biased toward the direction of travel relative to the location information acquired by the acquisition means. With this configuration, sounds emitted in places already passed can be reduced from the playback targets, and sounds emitted in places to be headed can be increased. Therefore, with such a terminal device, the usefulness of information can be increased.

[0044] In the terminal device described above, the surrounding area may be an area along the planned travel route. According to this configuration, the voice uttered in an area other than the area along the planned travel route is not detected. Therefore, with such a terminal device, the usefulness of information can be increased.

[0045] In addition, the terminal device of the second reference invention includes a detection means for detecting the occurrence of a specific event related to a vehicle, a transmission means for transmitting, when the detection means detects that the specific event has occurred, audio information representing audio corresponding to the specific event to another terminal device or a server shared by multiple terminal devices, and a playback means for receiving the audio information from the other terminal device or the server and playing back the audio represented by the audio information.

[0046] According to this configuration, a user of a certain terminal device A can hear the audio corresponding to a specific event related to the vehicle transmitted from another terminal device B (another vehicle). Therefore, the user of the terminal device can grasp the status of the other vehicle while driving.

[0047] In the above-described terminal device, the detection means may detect, as the specific event, a specific driving operation performed by a user of the vehicle, and the transmission means may transmit, when the detection means detects the specific driving operation, the audio information representing a voice notifying the user that the specific driving operation has been performed. With this configuration, the user of the terminal device can know, while driving, that a specific driving operation has been performed in another vehicle.

[0048] In the above-described terminal device, the detection means may detect a sudden braking operation, and the transmission means may transmit the audio information representing a sound notifying the user of the sudden braking operation when the detection means detects the sudden braking operation. With this configuration, the user of the terminal device can know that another vehicle has performed a sudden braking operation while driving. Therefore, safer driving can be achieved compared to checking the status of other vehicles only by visual inspection. [Brief explanation of the drawings]

[0049] [Figure 1] 1 is a block diagram showing a schematic configuration of a communication system; [Figure 2] 1 is a block diagram showing a schematic configuration of a device mounted on a vehicle. [Figure 3A-3C] FIG. 2 is an explanatory diagram illustrating a configuration of an input unit. [Figure 4] FIG. 2 is a block diagram showing a schematic configuration of a server. [Figure 5A] 10 is a flowchart showing a part of a voice transmission process. [Figure 5B] 10 is a flowchart showing the remaining part of the voice transmission process. [Figure 6] 10 is a flowchart showing a classification process. [Figure 7] 10 is a flowchart showing an operation voice determination process. [Figure 8] 10 is a flowchart showing a state sound determination process. [Figure 9] 10 is a flowchart showing an authentication process. [Figure 10A] 10 is a flowchart showing a part of an audio distribution process. [Figure 10B] 10 is a flowchart showing another part of the audio distribution process. [Figure 10C] 10 is a flowchart showing the remaining part of the audio distribution process. [Figure 11A] 10 is a flowchart showing a part of an audio reproduction process. [Figure 11B]10 is a flowchart showing the remaining part of the audio playback process. [Figure 12] 10 is a flowchart showing a speaker setting process. [Figure 13] 10 is a flowchart showing a mode transition process. [Figure 14] 10 is a flowchart showing a replay process. [Figure 15] 10 is a flowchart showing a storage process. [Figure 16] 10 is a flowchart showing a request process. [Figure 17A] 10 is a flowchart showing a part of a response process. [Figure 17B] 10 is a flowchart showing the remaining part of the response process. [Figure 18] FIG. 2 is a diagram showing a surrounding area set within a predetermined distance from the vehicle position. [Figure 19] FIG. 10 is a diagram showing a surrounding area set within a predetermined distance from a reference position. [Figure 20] FIG. 2 is a diagram showing surrounding areas set along a planned travel route. [Figure 21A] 10 is a flowchart illustrating a part of a response process according to a modified example. [Figure 21B] 10 is a flowchart showing the remaining part of the response process of the modified example. [Figure 22A] 10 is a flowchart showing a part of a modified example of an audio reproduction process. [Figure 22B] 10 is a flowchart showing another part of the audio reproduction process according to the modified example. [Figure 22C] 10 is a flowchart showing yet another part of the audio reproduction process according to the modified example. [Figure 22D] 10 is a flowchart showing the remaining part of the audio reproduction process according to the modified example. [Figure 23] 10 is a flowchart showing a drowsy driving determination process. [Figure 24] 10 is a flowchart showing additional information display processing. DETAILED DESCRIPTION OF THE INVENTION

[0050] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS An embodiment of the present invention will be described below with reference to the drawings. [Configuration of this embodiment] Fig. 1 is a block diagram showing a schematic configuration of a communication system 1 to which the present invention is applied, Fig. 2 is a block diagram showing a schematic configuration of a device mounted on a vehicle, Figs. 3A-3C are explanatory diagrams showing the configuration of an input unit 45, and Fig. 4 is a block diagram showing a schematic configuration of a server 10.

[0051] The communication system 1 is a system having a function that enables users of terminal devices 31 to 33 to communicate with each other by voice with simple operations. The users here refer to vehicle occupants (for example, the driver). As shown in FIG. 1, the communication system 1 includes a plurality of terminal devices 31 to 33 that can communicate with each other, and a server 10 that can communicate with each of these terminal devices 31 to 33 via base stations 20 and 21. In other words, the server 10 is shared by the plurality of terminal devices 31 to 33. Note that in this embodiment, for convenience of explanation, three terminal devices 31 to 33 are shown, but more terminal devices may be used.

[0052] The terminal devices 31 to 33 are configured as on-board devices mounted on vehicles such as passenger cars and trucks, and are configured as well-known microcomputers including a CPU 41, a ROM 42, a RAM 43, etc., as shown in Fig. 2. The CPU 41 performs various processes, such as a voice transmission process and a voice playback process, which will be described later, based on programs stored in the ROM 42, etc.

[0053] Furthermore, the terminal devices 31 to 33 are also equipped with a communication unit 44 and an input unit 45. The communication unit 44 has a function of communicating with the base stations 20, 21 configured as wireless base stations for mobile phones, for example (a communication unit for communicating with the server 10 and the other terminal devices 31 to 33 via a mobile phone network), and a function of directly communicating with the other terminal devices 31 to 33 located within its own line-of-sight range (terminal devices 31 to 33 mounted on other vehicles, such as a preceding vehicle, a following vehicle, or an oncoming vehicle) (a communication unit for directly communicating with the other terminal devices 31 to 33).

[0054] The input unit 45 is configured as buttons, switches, etc. for inputting commands by the users of the terminal devices 31 to 33. The input unit 45 also has a function for reading the biometric characteristics of the users.

[0055] 3A, the input unit 45 is also configured as a plurality of touch panels 61, 62 on the user-facing surface of the vehicle steering wheel 60. These touch panels 61, 62 are arranged around the horn switch 63 and in positions that are not touched when the steering wheel is operated to turn the vehicle. These touch panels 61, 62 detect the movement of the user's hand (fingers) on the vehicle steering wheel 60.

[0056] As shown in FIG. 2, the terminal devices 31 to 33 are connected to a navigation device 48, an audio device 50, an ABSECU 51, a microphone 52, a speaker 53, a brake sensor 54, a passing sensor 55, a turn signal sensor 56, an underside sensor 57, a crosswind sensor 58, a collision / rollover sensor 59, and other multiple well-known sensors provided in the vehicle via communication lines (e.g., an in-vehicle LAN).

[0057] Like well-known navigation devices, the navigation device 48 includes a current location detection unit that detects the current location of the vehicle and a display that displays images. The navigation device 48 executes well-known navigation processing based on information such as the vehicle's position coordinates (current location) detected by a GPS receiver 49. When the navigation device 48 receives a request for the current location from the terminal devices 31 to 33, it returns the latest information on the current location to the terminal devices 31 to 33. When the navigation device 48 receives a command from the terminal devices 31 to 33 to display text information, it displays text corresponding to the text information on the display.

[0058] The audio device 50 is a well-known sound reproducing device that reproduces music, etc. The music, etc. reproduced by the audio device 50 is output from a speaker 53. The ABSECU 51 is an electronic control unit (ECU) that controls the antilock braking system (ABS). The ABSECU 51 controls the braking force (brake hydraulic pressure) so that the wheel slip ratio falls within a preset range (a slip ratio that allows the vehicle to be braked safely and quickly), thereby preventing the brakes from locking. In other words, the ABSECU 51 functions as a brake lock prevention device. Furthermore, the ABSECU 51 outputs a notification signal to the terminal devices 31 to 33 when the ABS is activated (when control of the braking force is started).

[0059] The microphone 52 inputs the voice uttered by the user. The speaker 53 is configured as a 5.1ch surround system including, for example, five speakers and a subwoofer, and these speakers are arranged so as to surround the user.

[0060] The brake sensor 54 detects a sudden braking operation when the speed at which the user presses the brake pedal is equal to or greater than a reference value. However, the determination of a sudden braking operation is not limited to this, and may be made based on, for example, the amount of depression of the brake pedal or the acceleration (deceleration) of the vehicle.

[0061] The passing sensor 55 detects a passing operation by the user (an operation to temporarily change the headlights from low beam to high beam). The turn signal sensor 56 detects a turn signal operation (an operation to flash either the left or right turn signal) by the user.

[0062] The underside sensor 57 detects a state in which the underside of the vehicle is in contact with a road step or the like, or a state in which there is a high risk of contact. In this embodiment, a detection member that is easily deformed or displaced when it comes into contact with an external object (such as a road step) and returns to its original (pre-deformation or pre-displacement) state when released from contact is provided on the underside of the front and rear bumpers (the lowest part of the underside of the vehicle excluding the wheel area). The underside sensor 57 detects the deformation or displacement of this detection member, thereby detecting a state in which the underside of the vehicle is in contact with a road step or the like, or a state in which there is a high risk of contact. Note that a sensor that detects a state in which the underside of the vehicle is in contact with a road step or the like, or a state in which there is a high risk of contact, without coming into contact with an external object, may also be used.

[0063] The crosswind sensor 58 detects a strong crosswind against the vehicle (for example, a crosswind whose wind pressure is equal to or greater than a reference value). The crosswind sensor 58 may be a sensor that detects wind pressure, or may be a sensor that detects wind volume.

[0064] The collision / rollover sensor 59 detects a state in which the vehicle is highly likely to have caused an accident, such as a state in which the vehicle has collided with an external object (including another vehicle) or a state in which the vehicle has rolled over. The collision / rollover sensor 59 can be configured as a sensor that detects a strong impact (acceleration) on the vehicle. Note that the collision state and the rollover state may be detected independently. For example, the rollover state may be determined based on the degree of inclination of the vehicle (vehicle body). Other states in which the vehicle is highly likely to have caused an accident (for example, an airbag deployment state) may also be included.

[0065] 4, the server 10 has a hardware configuration as a well-known server including a CPU 11, a ROM 12, a RAM 13, a database (storage device) 14, a communication unit 15, etc. The server 10 performs various processes such as an audio distribution process described later.

[0066] The database 14 of the server 10 includes a name / nickname DB that associates the names of users of the terminal devices 31 to 33 with the nicknames of these users, a communication partner DB that describes the IDs of the terminal devices 31 to 33 and communication partners, a prohibited word DB that registers prohibited words, an advertisement DB that describes the IDs of advertisers' terminal devices and the range of advertising, etc. Users can use the services of this system by registering their names, nicknames, communication partners, etc. in the server 10.

[0067] [Processing of this embodiment] 5A and 5B onward, the processing performed in such a communication system 1 will be described. Figures 5A and 5B are flowcharts showing the voice transmission processing executed by the CPU 41 of each of the terminal devices 31-33.

[0068] The voice transmission process is started when the terminal devices 31 to 33 are powered on, and is executed repeatedly thereafter. The voice transmission process is also a process of transmitting voice information representing voices uttered by the user to other devices (such as the server 10 or other terminal devices 31 to 33). In the following description, the voice information to be transmitted to other devices is also referred to as "transmitted voice information."

[0069] 5A and 5B, first, it is determined whether or not an input operation has been performed via the input unit 45 (S101). Here, the input operation includes a normal input operation and an emergency input operation, which are registered in a memory such as a ROM as different input operations.

[0070] As a normal input operation, for example, as shown in Fig. 3B, a hand movement sliding in the radial direction of the steering wheel 60 on the touch panels 61, 62 is registered, and as an emergency input operation, for example, as shown in Fig. 3C, a hand movement sliding in the circumferential direction of the steering wheel 60 on the touch panels 61, 62 is registered. In this process, it is determined whether any of these operations has been input via the touch panels 61, 62.

[0071] It is preferable to set the input operation to operate in a direction different from the operating direction (pushing direction) of the horn switch 63 in this way, as this can prevent the horn from being operated erroneously.

[0072] If there is no input operation (S101: NO), the process proceeds to S110, which will be described later. If there is an input operation (S101: YES), recording of the voice uttered by the user (voice to be used as transmission voice information) starts (S102).

[0073] Then, it is determined whether or not a termination operation has been performed via the input unit 45 (S103). The termination operation is registered in advance in a memory such as a ROM, similar to the input operation. This termination operation may be the same as the input operation, or may be a different operation.

[0074] If there is no termination operation (S103: NO), it is determined whether a preset reference time (e.g., 20 seconds) has elapsed since the input operation (S104). The reference time is set to prohibit long-term voice input, and when communicating with a specific communication partner, the same or a longer reference time is set than when communicating with an unspecified communication partner. Alternatively, the reference time may be set to an infinite time.

[0075] If the reference time has not elapsed (S104: NO), the process returns to S102. If the reference time has elapsed (S104: YES) and an end operation has been performed in the process of S103 (S103: YES), recording is ended, and it is determined whether the input operation is an emergency operation input (S105). If it is an emergency operation input (S105: YES), an emergency information flag indicating that it is emergency voice information is added (S106), and the process proceeds to S107.

[0076] If the operation input is not for emergency use (S105: NO), the process immediately proceeds to S107. Next, the recorded voice (natural voice) is converted into synthetic voice (mechanical sound) (S107). Specifically, first, the recorded voice is converted into text (character data) using well-known voice recognition technology. Next, the text obtained by the conversion is converted into synthetic voice using well-known voice synthesis technology. In this way, the recorded voice is converted into synthetic voice.

[0077] Next, the operation content information is set to be added to the transmitted voice information (S108). Here, the operation content information indicates information indicating the vehicle state and the user's driving operation (operational states such as blinker operation, passing operation, sudden braking, etc.; operating states of lights, wipers, television, radio, etc.; driving states such as driving speed, acceleration, and direction of travel; control states such as values ​​detected by various sensors and the presence or absence of malfunctions, etc.). This information is acquired via a communication line within the vehicle (for example, an in-vehicle LAN).

[0078] Next, a classification process is executed to classify the transmitted voice information into one of a plurality of predetermined types (genres) and set so that classification information indicating the classified type is added to the transmitted voice information (S109). Such classification is used, for example, in the server 10 or the terminal devices 31 to 33 when extracting (searching) necessary voice information from a large amount of voice information. The classification of the transmitted voice information may be performed manually by the user or automatically. For example, selection items (a plurality of predetermined types) may be displayed on a display so that the user can select one. Also, for example, the content of a speech may be analyzed and the speech may be classified into a type according to the content. Also, for example, the speech may be classified into a type according to the situation at the time of speech (such as the state of the vehicle or driving operation).

[0079] 6 is a flowchart showing an example of a classification process when classification of transmitted voice information is performed automatically. First, it is determined whether a specific driving operation was performed at the time of speech (S201). In this example, a passing operation detected by the passing sensor 55 and a sudden braking operation detected by the brake sensor 54 correspond to the specific driving operation. In this example, the voice information is classified into three types: "normal," "caution," and "warning."

[0080] If it is determined that a specific driving operation is not being performed at the time of the statement (S201: NO), the transmitted voice information is classified as "normal" and classification information representing "normal" is set to be added to the transmitted voice information (S202).

[0081] On the other hand, if it is determined that a passing operation was performed when the user spoke (S201: YES, S203: YES), the transmitted voice information is classified as "Caution" and the classification information representing "Caution" is set to be added to the transmitted voice information (S204). Also, if it is determined that an emergency braking operation was performed when the user spoke (S203: NO, S205: YES), the transmitted voice information is classified as "Warning" and the classification information representing "Warning" is set to be added to the transmitted voice information (S206).

[0082] After such classification processing, the process proceeds to S114 (FIG. 5B) described later. Note that if it is determined in S205 that sudden braking was not performed while speaking (S205: NO), the process returns to S101 (FIG. 5A).

[0083] Incidentally, in this example, the transmitted audio information is classified into one type, but this is not limited to this, and it may be possible to assign two or more types to one transmitted audio information, for example, like tag information.

[0084] 5A, if it is determined in S101 that no input operation has been performed (S101: NO), it is determined whether a specific driving operation has been performed (S110). The specific driving operation here refers to an operation that triggers the transmission of voice not uttered by the user as transmission voice information, and may be different from the specific driving operation determined in S201. If it is determined that a specific driving operation has been performed (S110: YES), an operation voice determination process is executed to determine transmission voice information corresponding to the driving operation (S111).

[0085] 7 is a flowchart showing an example of the operation sound determination process. In this example, the specific driving operations in S110 include a turn signal operation detected by the turn signal sensor 56, a passing operation detected by the passing sensor 55, and an emergency braking operation detected by the brake sensor 54.

[0086] First, it is determined whether a turn signal operation has been detected as a specific driving operation as defined in S110 (S301). If it is determined that a turn signal operation has been detected (S301: YES), a voice notifying that a turn signal operation has been performed (for example, a voice such as "Turn signal operation has been performed," "Moving left (right)," or "Please be careful") is set as transmission voice information (S302).

[0087] On the other hand, if it is determined that a turn signal operation has not been detected (S301: NO), it is determined (S303) whether a passing operation has been detected as a specific driving operation as defined in S110. If it is determined that a passing operation has been detected (S303: YES), a voice notifying that a passing operation has been performed (for example, a voice such as "You have performed a passing operation" or "Please be careful") is set as the transmission voice information (S304).

[0088] Then, in both cases where a turn signal operation is detected (S301: YES) and where a passing operation is detected (S303: YES), the transmitted voice information is classified as "Caution" and the classification information representing "Caution" is added to the transmitted voice information (S305). After that, the operation voice determination process is terminated, and the process proceeds to the process of S114 (FIG. 5B) described later.

[0089] On the other hand, if it is determined that a passing operation has not been detected (S303: NO), it is determined whether or not a sudden braking operation has been detected as the specific driving operation referred to in S110 (S306). If it is determined that a sudden braking operation has been detected (S306: YES), a voice notifying that a sudden braking operation has been performed (for example, a voice such as "A sudden braking operation has been performed") is set as the transmission voice information (S307). Then, the transmission voice information is set to "Warning" The operation voice determination process is then terminated, and the process proceeds to S114 (FIG. 5B) described later. Note that if it is determined that a sudden braking operation has not been detected as the specific driving operation referred to in S110 (S306: NO), the process returns to S101 (FIG. 5A).

[0090] Returning to FIG. 5A, if it is determined in S110 that a specific driving operation is not being performed (S110: NO), it is determined whether a specific state has been detected (S112). The specific state here refers to a state that triggers the transmission of voice not uttered by the user as transmission voice information. If it is determined that a specific state has not been detected (S112: NO), the voice transmission process is terminated. On the other hand, if it is determined that a specific state has been detected (S112: YES), a state voice determination process is executed to determine transmission voice information according to the state (S113).

[0091] 8 is a flowchart showing an example of the state sound determination process. In this example, the specific states referred to in S112 are a state in which the ABS is activated, a state in which the underside of the vehicle is in contact with a bump or the like on the road surface or there is a high risk of contact, a state in which a strong crosswind is detected against the vehicle, and a state in which the vehicle has collided or rolled over.

[0092] First, it is determined whether the ABS has been activated (S401) based on the notification signal output from the ABS ECU 51. If it is determined that the ABS has been activated (S401: YES), a voice notifying that the ABS has been activated (for example, "ABS has been activated" or "Be careful of slipping") is set as transmission voice information (S402).

[0093] On the other hand, if it is determined that the ABS is not operating (S401: NO), it is determined whether the underside sensor 57 has detected a state in which the underside of the vehicle is in contact with a bump or the like in the road surface, or a state in which there is a high risk of contact (S403).If it is determined that a contact state or a state in which there is a high risk of contact has been detected (S403: YES), a voice notifying that there is a risk of scraping the underside of the vehicle (for example, a voice such as "There is a risk of scraping the underside of the vehicle" or "A bump exists") is set as transmission voice information (S404).

[0094] On the other hand, if it is determined that a contact state or a state with a high risk of contact has not been detected (S403: NO), it is determined whether a strong crosswind has been detected against the vehicle by the crosswind sensor 58 (S405). If it is determined that a strong crosswind has been detected (S405: YES), a voice notifying that a strong crosswind has been detected (for example, a voice such as "Be careful of crosswinds") is set as the transmission voice information (S406).

[0095] Then, when the ABS is activated (S401: YES), when a state in which the underside of the vehicle is in contact with a bump or the like on the road surface or a state in which there is a high risk of contact is detected (S403: YES), or when a strong crosswind is detected (S405: YES), the transmitted voice information is classified as "Caution" and the classification information representing "Caution" is set to be added to the transmitted voice information (S407). Thereafter, the state voice determination process is ended, and the process proceeds to the process of S114 (FIG. 5B) described later.

[0096] On the other hand, if it is determined that a strong crosswind has not been detected (S405: NO), it is determined whether the collision / rollover sensor 59 has detected a state in which the vehicle has collided or rolled over (a state in which the vehicle is likely to have caused an accident) (S408). If it is determined that a state in which the vehicle has collided or rolled over has been detected (S408: YES), a voice notifying that the vehicle has collided or rolled over (for example, a voice such as "The vehicle has collided (or rolled over)" or "An accident has occurred") is set as the transmission voice information (S409). Then, the transmission voice information is classified as "warning," and classification information representing "warning" is set to be added to the transmission voice information (S410). Thereafter, the state sound determination process is terminated, and the process proceeds to S114 (FIG. 5B) described later. Note that if it is determined that the vehicle has not been detected as having collided or rolled over (S408: NO), the process returns to S101 (FIG. 5A).

[0097] Returning to FIG. 5B, in S114, location information is acquired by requesting location information from the navigation device 48, and this location information is set to be added to the transmitted voice information. The location information is information that represents the current location of the navigation device 48 at the time of acquisition (in other words, the current location of the vehicle, the current location of the terminal devices 31 to 33, or the current location of the user) as an absolute location (latitude, longitude, etc.). In particular, the location information acquired in S114 represents the location where the user uttered the voice (utterance location). Note that the utterance location may be anything that can identify the location where the user uttered the voice, and some variation may occur due to factors such as the time difference between the timing of inputting the voice and the timing of acquiring the location information (changes in position as the vehicle travels).

[0098] Next, a voice packet is generated as transmission data containing information including transmission voice information, emergency information, location information, operation content information, classification information, and the user's nickname set in the authentication process described below (S115).The generated voice packet is then wirelessly transmitted to the communication partner device (server 10 or another terminal device 31-33) (S116), and the voice transmission process is terminated.In other words, the transmission voice information representing the voice uttered by the user is transmitted in a form in which the emergency information, location information (utterance location), operation content information, classification information, and nickname are identifiable (associated).It is also possible to transmit all voice packets to server 10, and to transmit specific voice packets (for example, those containing transmission voice information with classification information such as "Caution" or "Warning") to other terminal devices 31-33 with which direct communication is possible.

[0099] Furthermore, when the terminal devices 31 to 33 are started up or when a specific operation is input by the user, an authentication process shown in Fig. 9 is carried out. The authentication process is a process for identifying the user who performs voice input.

[0100] 9, in the authentication process, first, the user's biometric characteristics (for example, any one of fingerprint, retina, palm, voiceprint, etc.) are acquired via the input unit 45 of the terminal device 31-33 (S501). Then, the user is identified by referring to a database that associates user names (IDs) and biometric characteristics registered in advance in a memory such as RAM, and a nickname set by the user is set to be inserted into the voice (S502). When this process is completed, the authentication process ends.

[0101] Next, the audio distribution process will be described with reference to the flowcharts shown in Figures 10A to 10C. The audio distribution process is a process in which the server 10 performs a predetermined operation (processing) on ​​audio packets transmitted from the terminal devices 31 to 33, and distributes the audio packets after the operation to the other terminal devices 31 to 33.

[0102] 10A-10C, first, it is determined whether or not a voice packet transmitted from the terminal device 31 to 33 has been received (S701). If a voice packet has not been received (S701: NO), the process proceeds to S725, which will be described later.

[0103] If a voice packet is received (S701: YES), the voice packet is acquired (S702) and a nickname is added to the voice packet by referencing the name / nickname DB (S703). In this process, a new voice packet is generated by adding a nickname, which is identification information for identifying the user, to the voice packet input by the user. At this time, the nickname is added before or after the voice input by the user to prevent the voice input and the nickname from overlapping.

[0104] Next, the voice representing the voice information contained in the new voice packet is converted into text (S704). This process utilizes a well-known process for converting voice into text that can be read by other users. At this time, location information, operation content information, etc. are also converted into text data in the same way. Note that when the voice-to-text conversion process is performed in the terminal device 31-33 that is the source of the voice packet, the conversion process in the server 10 may be omitted (or simplified) by acquiring the text data obtained by the conversion process from the terminal device 31-33.

[0105] Next, the name (person's name) is extracted from the voice (or converted character data) and it is determined whether or not a name is present (S705). If emergency information is included in the voice packet, the process from S705 to S720 (described later) is omitted and the process proceeds to S721 (described later).

[0106] If there is no name in the speech (S705: NO), the process proceeds to S709, which will be described later. If there is a name in the speech (S705: YES), it is determined whether this name exists in the name / nickname DB (S706). If the name in the speech exists in the name / nickname DB (S706: YES), this name is replaced with the corresponding nickname in the speech (in the character data, the name is replaced with the nickname in text) (S707), and the process proceeds to S709.

[0107] Also, if the name in the voice does not exist in the name / nickname DB (S706: NO), the name is replaced with a "beep" sound (in character data, the name is replaced with other characters such as "***") (S708), and the process proceeds to S709.

[0108] Next, predetermined prohibited words are extracted from the voice (or converted character data) (S709). Prohibited words are generally vulgar, obscene, or other words that may offend listeners and are registered in a prohibited word database. In this process, prohibited words are extracted by comparing the words in the voice (or converted character data) with the prohibited words in the prohibited word database.

[0109] Forbidden words are replaced with the sound "boo" (in character data, names are replaced with other characters such as "xxx") (S710). Next, intonation checks and politeness checks are carried out (S711, S712).

[0110] In these processes, the emotion and politeness of the user's speech when inputting voice are determined by detecting the amplitude (absolute value and rate of change) of the input level to the microphone 52 and the polite expressions contained in the voice from the voice information, and it is determined whether the user is emotional or not speaking politely.

[0111] If the check results show that the input voice has normal emotion and is spoken in a polite manner (S713: YES), the process immediately proceeds to step S715, which will be described later. If the check results show that the input voice is emotional or not spoken in a polite manner (S713: NO), the process sets a voice distribution prohibition flag (S714), which indicates that distribution of this voice packet is prohibited, and proceeds to step S715.

[0112] Next, the behavior of the terminal device 31-33 that sent the corresponding voice packet is identified by tracking the location information, and it is determined whether the behavior of the terminal device 31-33 deviates from the main road onto a side road or whether it has stopped for a certain period of time (for example, about 5 minutes) or more (S715).

[0113] If the behavior of the terminal device 31-33 is that it has deviated from the main road onto a side road or has stopped (S715: YES), it is determined that the terminal device 31-33 is taking a rest, and a rest flag is set (S716), and the process proceeds to S717. If it is not the case (S715: NO), the process immediately proceeds to S717.

[0114] Then, a pre-set terminal device (for example, all terminal devices that are powered on, some terminal devices extracted based on the setting, etc.) is set as the communication partner (S717), and the results of the gaffe (results of the presence or absence of a name or prohibited words, results of intonation check, polite expression check, etc.) are notified to the terminal device 31 to 33 that is the source of the voice input (S718).

[0115] Next, it is determined whether the position of the terminal device 31-33 (recipient vehicle) performing the distribution has passed a specific position (S719). Here, the server 10 has registered a plurality of advertisements in an advertisement DB, and in this advertisement DB, an area to which the advertisement is to be distributed is set for each advertisement. In this process, it is detected that the terminal device 31-33 has entered an area to which the advertisement (CM) is to be distributed.

[0116] If the destination vehicle has passed the specific position (S719: YES), an advertisement corresponding to the specific position is added to the voice packet to be transmitted to the terminal device 31-33 (S720). If the destination vehicle has not passed the specific position (S719: NO), the process proceeds to S721.

[0117] Next, it is determined whether there are multiple voice packets to be distributed (S721). If there are multiple voice packets (S721: YES), the voice packets are sorted so that they do not overlap (S722), and the process proceeds to S723. During this process, emergency information packets are sorted first in the playback order, and voice packets with the break flag set are sorted later in the order.

[0118] If there is only one voice packet (S721: NO), the process immediately proceeds to S723, where it is determined whether or not a voice packet is already being transmitted (S723). If a voice packet is currently being transmitted (S723: YES), the voice packet to be transmitted is set to be transmitted after the packet currently being transmitted (S724), and the process proceeds to S726. If a voice packet is not currently being transmitted (S723: NO), the process immediately proceeds to S726.

[0119] In the process of S725, it is determined whether there are any voice packets to be transmitted (S725). If there are any voice packets (S725: YES), the voice packets to be transmitted are transmitted (S726), and the voice distribution process ends. On the other hand, if there are no voice packets (S726: NO), the voice distribution process ends.

[0120] Next, the audio reproduction process in which the CPU 41 of the terminal devices 31 to 33 reproduces audio will be described with reference to the flowcharts shown in Figures 11A and 11B. In the audio reproduction process, as shown in Figures 11A and 11B, first, it is determined whether or not a packet has been received from another device such as another terminal device 31 to 33 or the server 10 (S801). If no packet has been received (S801: NO), the audio reproduction process ends.

[0121] If a packet is received (S801: YES), it is determined whether or not this packet is an audio packet (S802). If it is not an audio packet (S802: NO), the process proceeds to S812, which will be described later.

[0122] If the packet is an audio packet (S802: YES), the audio information currently being played is stopped and the audio information being played is erased (S803). Then, it is determined whether or not a distribution prohibition flag is set (S804). If the distribution prohibition flag is set (S804: YES), an audio message (stored in a memory such as a ROM) indicating that audio playback is prohibited is displayed. The voice message is generated (pre-recorded in the library) and a text message corresponding to the voice message is generated (S805).

[0123] Next, in order to make the sound easier to hear, the playback volume of the music or the like being played by the audio device 50 is turned off (S806). Note that the playback volume may be lowered, or the playback itself may be stopped.

[0124] Next, a speaker setting process is performed to set the volume of the sound to be output from the speaker 53 when the sound is played back, depending on the positional relationship between the location (utterance location) of the terminal device 31-33 corresponding to the source of the sound to be played back and the location (current location) of the terminal device 31-33 itself (S807), and the sound message generated in the process of S805 is played back (playback starts) at the set volume (S808). At this time, character data including the character message is output to the display of the navigation device 48 and displayed (S812).

[0125] The character data displayed in this process includes character data indicating the content of the voice represented by the voice information included in the voice packet, character data indicating the content of the operation by the user, and a voice message.

[0126] Furthermore, if the distribution prohibition flag is not set in the processing of S804 (S804: NO), the playback volume of the music or the like being played by the audio device 50 is turned off (S809), as in the above-described S806. Furthermore, speaker setting processing is performed (S810), and the voice represented by the voice information included in the voice packet is played (S811), and the processing proceeds to S812. Because the voice represented by the voice information is synthesized voice, the playback mode can be easily changed. For example, the playback speed can be adjusted, the emotional components included in the words can be suppressed, the intonation can be adjusted, the tone of the voice can be adjusted, and other changes can be made based on user operation or automatically.

[0127] In the processing of S811, the audio represented by the audio information contained in the aligned audio packets is played back in order, so it is assumed that the audio packets will not overlap. However, if there is a risk of audio packets overlapping, such as when audio packets are received from multiple servers 10, the processing of S722 described above (processing to align the audio packets) can be performed.

[0128] Next, it is determined whether or not a voice representing a command has been input by the user (S813). In this embodiment, if a voice representing a command to skip the currently played audio (e.g., the voice saying "cut") is input during audio playback (S813: YES), the currently played audio information is skipped in predetermined units, and the next audio information is played (S814). In other words, audio that you do not want to hear can be skipped and played back by voice operation. The predetermined unit here may be, for example, a statement (comment) unit, or a speaker unit. When audio playback of all received audio information has finished (S815: YES), the playback volume of music, etc. being played by the audio device 50 is turned back on (S816), and the audio playback process ends.

[0129] Next, details of the speaker setting process executed in S807 and S810 will be described with reference to Fig. 12. Fig. 12 is a flowchart showing the speaker setting process of the audio reproduction process. In the speaker setting process, as shown in Fig. 12, first, position information indicating the vehicle position (including past position information) is acquired from the navigation device 48 (S901).

[0130] Next, the position information (utterance position) included in the voice packet to be played back is extracted (S902), and the positions (distances) of the other terminal devices 31 to 33 relative to the movement vector of the own vehicle are detected (S903). The direction in which the terminal device that is the sender of the voice packet to be played back is located and the distance to this terminal device are detected as a reference.

[0131] Next, the volume ratio of the speakers is adjusted according to the direction of the terminal device that is the sender of the voice packet (S904). In this process, the volume ratio is set so that the volume of the speaker located in the direction of the sender terminal device is the loudest so that the voice can be heard from the direction of the sender terminal device.

[0132] Then, the maximum volume is set for the speaker with the largest volume ratio according to the distance to the terminal device that sent the voice packet (S905). This process also allows the volume for other speakers to be set based on the volume ratio. When this process is completed, the speaker setting process is terminated.

[0133] Next, the mode transition process will be described. Fig. 13 is a flowchart showing the mode transition process executed by the CPU 41 of the terminal devices 31 to 33. The mode transition process is a process for switching between a distribution mode in which communication partners are an unspecified number of people and a dialogue mode in which communication partners are specified people, and is started when the power of the terminal devices 31 to 33 is turned on, and is thereafter repeatedly executed.

[0134] 13, first, it is determined whether the current setting mode is the interactive mode or the distribution mode (S1001, S1004). If the setting mode is neither the interactive mode nor the distribution mode (S1001: NO, S1004: NO), the process returns to S1001.

[0135] If the setting mode is the interactive mode (S1001: YES), it is determined (S1002) whether a mode change operation has been performed via the input unit 45. If no mode change operation has been performed (S1002: NO), the mode change process ends.

[0136] If a mode change operation is performed (S1002: YES), the mode changes to the distribution mode (S1003). At this time, the server 10 is notified to set the communication partner to an unspecified number of people, and the server 10 receives this notification and changes the communication partner setting (description in the communication partner DB) for the terminal device 31-33.

[0137] If the set mode is the distribution mode (S1001: NO, S1004: YES), it is determined (S1005) whether a mode change operation has been performed via the input unit 45. If no mode change operation has been performed (S1005: NO), the mode change process ends.

[0138] If a mode transition operation is performed (S1005: YES), a dialogue request is sent to the other party who is playing back the voice packet (S1006). In detail, when a dialogue request is sent to the server 10 by specifying (the terminal device that sent the voice packet), the server 10 sends the dialogue request to the corresponding terminal device. The server 10 then receives a response to the dialogue request from this terminal device and sends the response to the terminal device 31 to 33 that originated the dialogue request.

[0139] If the terminal devices 31 to 33 respond that they accept the interaction request (S1007: YES), they transition to an interaction mode in which the terminal device that accepted the interaction request is used as the communication partner (S1008). At this time, they notify the server 10 that they should set a specific terminal device as the communication partner, and the server 10 changes the setting of the communication partner upon receiving this notification.

[0140] If there is no response to the effect that the interaction request is accepted (S1007: NO), the mode transition process ends. 14 is a flowchart showing the replay process executed by the CPU 41 of the terminal devices 31 to 33. The replay process is an interrupt process that starts when a replay operation for repeating playback is input during playback of the audio represented by the audio information included in the audio packet. This process is performed by interrupting the audio playback process.

[0141] In the replay process, first, a specified audio packet is selected (S1011) as shown in Fig. 14. In this process, the audio packet being played or the audio packet immediately after playback (within a reference time (e.g., 3 seconds) after playback ends) is selected according to the time since the start of playback of the currently played audio packet.

[0142] Next, the selected audio packet is played back from the beginning (S1012), and the replay process is completed. After this process is completed, the audio playback process is resumed. Meanwhile, the server 10 may receive voice packets transmitted from the terminal devices 31 to 33, store the voice information, and when a request is received from the terminal device 31 to 33, extract the voice information corresponding to the requesting terminal device 31 to 33 from the stored voice information, and transmit the extracted voice information.

[0143] 15 is a flowchart showing the accumulation process executed by the CPU 11 of the server 10. The accumulation process is a process for accumulating (storing) voice packets transmitted from the terminal devices 31 to 33 in the above-described voice transmission process (FIGS. 5A and 5B). Note that the accumulation process may be performed instead of the above-described voice distribution process (FIGS. 10A-10C), or may be performed together with the voice distribution process (for example, in parallel).

[0144] In the storage process, first, it is determined whether or not a voice packet transmitted from the terminal device 31 to 33 has been received (S1101). If a voice packet has not been received (S1101: NO), the storage process ends.

[0145] If a voice packet is received (S1101: YES), the received voice packet is stored in the database 14 (S1102). Specifically, the voice information is stored in a form that allows it to be searched based on various information contained in the voice packet (voice information, emergency information, location information, operation content information, classification information, user nickname, etc.). Thereafter, the storage process ends. Note that voice information that satisfies predetermined deletion conditions (for example, old voice information, voice information of canceled users, etc.) may be deleted from the database 14.

[0146] 16 is a flowchart showing the request processing executed by the CPU 41 of each of the terminal devices 31 to 33. The request processing is executed every predetermined time (for example, every several tens of seconds or several minutes). In the request processing, first, it is determined whether or not route guidance processing is being executed by the navigation device 48 (in other words, whether a guided route to the destination has been set) (S1201). If it is determined that route guidance processing is being executed (S1201: YES), location information indicating the vehicle's location (current location information) and information on the guided route (information on the destination and the route to the destination) are acquired as vehicle location information from the navigation device 48 (S1202). Thereafter, the process proceeds to S1206. The vehicle location information is information for causing the server 10 to set a surrounding area, which will be described later.

[0147] On the other hand, if it is determined that route guidance processing is not being performed (S1201: NO), it is determined whether the vehicle is moving (S1203). Specifically, if the vehicle's traveling speed exceeds a determination reference value, it is determined that the vehicle is moving. This determination reference value is set to a value higher than the speed at which a typical right or left turn occurs (a speed that is lower than that for traveling straight ahead). Therefore, if it is determined that the vehicle is moving, there is a high possibility that the direction of travel of the vehicle will not change significantly. In other words, in S1203, the traveling direction is predicted based on changes in the vehicle's position information. However, the present invention is not limited to this, and the reference value may be set to, for example, 0 or a value close to 0.

[0148] If it is determined that the vehicle is traveling (S1203: YES), the process acquires, as vehicle position information, position information indicating the vehicle's position (current position information) and one or more pieces of past position information from the navigation device 48 (S1204). Thereafter, the process proceeds to S1206. Note that when acquiring multiple pieces of past position information, position information at equal time intervals, such as position information from 5 seconds ago and position information from 10 seconds ago, may be acquired.

[0149] On the other hand, if it is determined that the vehicle is not moving (S1203: NO), the control unit 41 acquires, as vehicle position information, position information (current position information) indicating the vehicle's own position from the navigation device 48 (S1205). Then, the control unit 41 proceeds to S1206.

[0150] Next, favorite information is acquired (S1206). Favorite information is information that can be used to extract (search) audio information that the user needs from among a large amount of audio information, and is set manually or automatically by the user (for example, by analyzing preferences from the user's utterances, etc.). For example, when a specific user (for example, a user designated in advance) is set as favorite information, audio information transmitted by the specific user is preferentially extracted. Similarly, for example, when a specific keyword is set as favorite information, audio information containing the specific keyword is preferentially extracted. Note that in the process of S1206, favorite information that has already been set may be acquired, or the user may set it at this stage.

[0151] Next, the type (genre) of the requested audio information is acquired (S1207). The type here refers to the type classified in the process of S109 and the like (in this embodiment, "normal", "caution", and "warning").

[0152] Next, the terminal device 31-33 transmits an information request including the vehicle position information, the favorite information, the type of voice information, and its own (terminal device 31-33) ID (information for the server 10 to identify a reply destination) to the server 10 (S1208). Thereafter, the request process is terminated.

[0153] 17A and 17B are flowcharts showing the response processing executed by the CPU 11 of the server 10. The response processing is processing for responding to the information request transmitted from the terminal devices 31 to 33 in the request processing (FIG. 16) described above. Note that the response processing may be performed instead of the audio distribution processing (FIGS. 10A-10C) described above, or may be performed together with the audio distribution processing (for example, in parallel).

[0154] In the response process, first, it is determined whether or not an information request transmitted from the terminal device 31 to 33 has been received (S1301). If an information request has not been received (S1301: NO), the response process ends.

[0155] When an information request is received (S1301: YES), a surrounding area is set according to the vehicle position information included in the received information request (S1302 to S1306). In this embodiment, as described above (S1201 to S1205, S1208), any one of (P1) position information indicating the vehicle's position, (P2) position information indicating the vehicle's position and one or more pieces of past position information, and (P3) position information indicating the vehicle's position and information on a guided route is provided as vehicle position information.

[0156] When it is determined that the vehicle position information included in the received information request is (P1) position information indicating the vehicle's own position (S1302: YES), the vehicle position is calculated as shown in FIG. An area A1 within a predetermined distance L1 from (the vehicle's current location) C is set as the surrounding area (S1303). Then, the process proceeds to S1307. In Fig. 18, the vertical and horizontal lines indicate roads on which the vehicle can travel, and "x" indicates the position where the voice information stored in the server 10 is spoken.

[0157] On the other hand, if it is determined that the vehicle position information included in the received information request is (P2) position information indicating the vehicle's own position and one or more pieces of past position information (S1302: NO, S1304: YES), an area A2 within a predetermined distance L2 from a reference position Cb that is shifted in the traveling direction (right side in FIG. 19) from the vehicle's own position (current vehicle position) C is set as the surrounding area (S1305), as shown in FIG. 19. Then, the process proceeds to S1307.

[0158] Specifically, the vehicle's traveling direction is estimated based on the positional relationship between the current position information C and the past position information Cp. The distance from the current position C to the reference position Cb may be greater than the radius L2 of the area. In other words, an area A2 that does not include the current position C may be set as the surrounding area.

[0159] On the other hand, if it is determined that the vehicle position information included in the received information request is (P3) position information indicating the vehicle's own position and information on the guided route (S1304: NO), the area along the planned driving route is set as the surrounding area (S1306). Then, the process proceeds to S1307.

[0160] Specifically, as shown in FIG. 20, for example, of the planned driving route (thick line portion) from the departure point S to the destination G, an area A3 along the route from the current location C to the destination G (area determined to be on the route) is set as the surrounding area. Here, the area along the route may be set as a range within a predetermined distance L3 from the route, for example. In this case, the value of L3 may be changed according to road attributes (for example, road width, type, etc.). Also, instead of the range to the destination G, it may be a range within a predetermined distance L4 from the current location C. Also, although previously traveled routes (route from the departure point S to the current location C) are excluded here, this is not limited to this, and previously traveled routes may also be included.

[0161] Next, voice information whose utterance position is within the surrounding area is extracted (S1307) from the voice information stored in the database 14. In the examples of Figs. 18 to 20, voice information whose utterance position is within the surrounding areas (areas surrounded by dashed lines) A1, A2, and A3 is extracted.

[0162] Next, from the voice information extracted in S1307, voice information corresponding to the type of voice information included in the received information request is extracted (S1308).Furthermore, from the voice information extracted in S1308, voice information corresponding to the favorite information included in the received information request is extracted (S1309).

[0163] Through the above processing (S1302 to S1309), the voice information corresponding to the information request received from the terminal device 31 to 33 is extracted as transmission candidate voice information to be transmitted to the terminal device 31 to 33 that sent the information request.

[0164] Next, the content of the voice is analyzed for each of the voice information candidates for transmission (S1310). Specifically, the voice represented by the voice information is converted into text (character data), and keywords contained in the converted text are extracted. Note that the text (character data) representing the content of the voice information may be stored in the database 14 together with the voice information, and the stored information may be used.

[0165] Next, for each of the candidate voice information to be transmitted, it is checked whether the content overlaps with other voice information. It is determined whether or not the contents of the speech information A and B overlap (S1311). For example, if the keywords included in one piece of speech information A and the keywords included in another piece of speech information B match at a predetermined rate or more (for example, more than half), it is determined that the contents of the speech information A and B overlap. Note that the overlap of the contents may be determined based on the context, etc., rather than simply on the rate of keywords.

[0166] Then, when it is determined that the content of all the candidate voice information for transmission does not overlap with other voice information (S1311: NO), the process proceeds to S1313, where the candidate voice information for transmission is transmitted to the terminal device 31-33 that sent the information request (S1313), and the response process ends. Note that, for each destination terminal device 31-33, voice information that has already been transmitted may be stored, and voice information that has already been transmitted may not be transmitted.

[0167] On the other hand, if it is determined that the voice information candidates for transmission contain voice information whose content overlaps with other voice information (S1311: YES), the voice information candidates for transmission are selected and all but one voice information to be retained as the voice information candidates for transmission are removed from the voice information candidates for transmission to eliminate the content overlap (S1312). For example, the newest voice information of the voice information candidates for transmission may be retained, or conversely, the oldest voice information may be retained. Alternatively, the voice information containing the most keywords may be retained. Thereafter, the voice information candidates for transmission are transmitted to the terminal device 31 to 33 that sent the information request (S1313), and the response process is terminated.

[0168] The audio information transmitted from the server 10 to the terminal devices 31 to 33 is played back in the audio playback process (FIGS. 11A and 11B) described above. When the terminal devices 31 to 33 receive a new audio packet while playing back audio represented by an already received audio packet, they prioritize playback of the audio represented by the newly received audio packet. Furthermore, audio represented by older audio packets is deleted without being played back. As a result, audio represented by audio information according to the vehicle's position is played back with priority.

[0169] [Effects of this embodiment] As described above, according to this embodiment, the following effects can be obtained. (1) The terminal devices 31 to 33 input voice uttered by the user in the vehicle (S102) and acquire location information (current location) (S114). Then, the terminal devices 31 to 33 transmit voice information representing the input voice to the server 10 shared by the multiple terminal devices 31 to 33 in a form that enables identification of the location information (utterance location) acquired at the time of input of the voice (S116). Furthermore, the terminal devices 31 to 33 receive voice information transmitted from the other terminal devices 31 to 33 to the server 10, and play back the voice represented by the received voice information (S811). Specifically, the terminal devices 31 to 33 play back the voice represented by the voice information transmitted from the other terminal devices 31 to 33 to the server 10, where the utterance location is located within a surrounding area set based on the current location (S1303, S1305, S1306).

[0170] Voices emitted by a user at a certain location are often useful to other users who are present at (or heading to) that location. In particular, the location of a user in a vehicle can change significantly over time (as the vehicle travels) compared to a user who is not in a vehicle. Therefore, for a user in a vehicle, voices emitted by other users within the surrounding area based on the current location can be useful information.

[0171] In this regard, in this embodiment, a user of a certain terminal device 31 can hear, among the voices uttered by users of other terminal devices 32 and 33 (other vehicles), voices uttered within a surrounding area set based on the current location of the terminal device 31. Therefore, according to this embodiment, useful information can be provided to users in vehicles.

[0172] The audio generated within the surrounding area is not limited to real-time audio, but may also include audio from the past (e.g., audio generated several hours ago). This audio can provide more useful information than real-time audio generated outside the surrounding area (e.g., a distant location). For example, if lane restrictions are in place at a specific point on a road, a user can tweet this information, allowing other users heading to that point to choose to avoid the point. Also, if the view from a specific point on a road is spectacular, a user can tweet this information, allowing other users heading to that point to avoid missing the view.

[0173] (2) When the terminal devices 31 to 33 transmit, as vehicle position information, past position information or information on the planned driving route (planned travel route) in addition to the current position information of the vehicle to the server 10, the surrounding area is set to an area biased toward the direction of travel of the vehicle's current location (FIGS. 19 and 20). Therefore, it is possible to reduce the number of voices emitted from places that the vehicle has already passed through from the playback targets and increase the number of voices emitted from places the vehicle is heading to. As a result, it is possible to increase the usefulness of the information.

[0174] (3) If the surrounding area is set to an area along the planned driving route, voices emitted outside the area along the planned driving route can be excluded, thereby further improving the usefulness of the information.

[0175] (4) The terminal devices 31 to 33 automatically acquire vehicle position information at predetermined time intervals (S1202, S1204, S1205) and transmit it to the server 10 (S1208), thereby acquiring voice information corresponding to the current location of the vehicle. Therefore, even if the position changes as the vehicle travels, voice information corresponding to the position can be acquired. In particular, when the terminal devices 31 to 33 receive new voice information from the server 10, they erase the received voice information (S803) and play the new voice information (S811). In other words, the voice information to be played back is updated to voice information corresponding to the latest vehicle position information at predetermined time intervals. Therefore, the played back voice can be made to follow changes in the position as the vehicle travels.

[0176] (5) The terminal devices 31 to 33 transmit vehicle position information, etc., including the current location of the vehicle, to the server 10 (S1208), cause the server 10 to extract voice information corresponding to the current location, etc. (S1302 to S1312), and play back the voice represented by the voice information received from the server 10 (S811). Therefore, compared to a configuration in which the terminal devices 31 to 33 extract voice information, the amount of voice information received from the server 10 can be reduced.

[0177] (6) The terminal devices 31 to 33 receive input of the user's voice (S102) and convert the input voice into synthetic voice (S107). By converting the voice into synthetic voice in this way and then playing it back, it becomes easier to change the playback mode compared to playing back the voice as is. For example, it is easy to adjust the playback speed, suppress the emotional components contained in the words, adjust the tone, adjust the voice tone, and so on.

[0178] (7) When a voice representing a command is input during the playback of a voice, the terminal devices 31 to 33 change the playback mode of the voice in accordance with the command. Specifically, when a voice representing a command to skip the voice being played is input (S813: YES), the voice being played is skipped in predetermined units (S814). Therefore, voice that one does not want to hear can be skipped in predetermined units, such as by comment or by speaker, before being played.

[0179] (8) The server 10 determines whether or not the content of each of the candidate voice information for transmission overlaps with other voice information (S1311), and if it determines that there is overlap (S1311: YES), deletes the voice information with the overlapping content (S1312). For example, When multiple users make similar comments about the same event, such as a comment about an accident that occurred, the phenomenon of similar comments being played back repeatedly is likely to occur. In this regard, according to the present embodiment, the playback of audio with overlapping content is omitted, making it possible to make it less likely that the phenomenon of similar comments being played back repeatedly will occur.

[0180] (9) The server 10 extracts the audio information corresponding to the favorite information from the audio information of the transmission candidates (S1309), and transmits it to the terminal devices 31 to 33 (S1313). Therefore, it is possible to efficiently extract the audio information required by the user (satisfying the priority conditions according to the user) from the large amount of audio information transmitted by many other users, and play it preferentially.

[0181] (10) The terminal devices 31 to 33 are configured to be able to communicate with an audio device mounted on a vehicle, and the volume of the audio device during audio playback is lowered compared to the volume before audio playback (S806). This makes it easier to hear the audio.

[0182] (11) The terminal devices 31 to 33 detect that a specific event related to the vehicle has occurred in the vehicle (S110, S112). When the terminal devices 31 to 33 detect that a specific event has occurred (S110: YES or S112: YES), they transmit audio information representing audio corresponding to the specific event to the other terminal devices 31 to 33 or the server 10 (S116). Furthermore, the terminal devices 31 to 33 receive audio information from the other terminal devices 31 to 33 or the server 10, and play back the audio represented by the received audio information (S811).

[0183] According to this embodiment, a user of a certain terminal device 31 can hear audio corresponding to a specific event related to the vehicle, which is transmitted from other terminal devices 32 and 33 (other vehicles). Therefore, the users of the terminal devices 31 to 33 can grasp the status of the other vehicles while driving.

[0184] (12) The terminal devices 31 to 33 detect, as a specific event, a specific driving operation performed by a user of a vehicle (S301, S303, S306). Then, when the terminal devices 31 to 33 detect a specific driving operation (S301: YES, S303: YES, or S306: YES), they transmit audio information representing a voice notifying that the specific driving operation has been performed (S302, S304, S307, S116). Therefore, the users of the terminal devices 31 to 33 can know, while driving, that a specific driving operation has been performed in another vehicle.

[0185] (13) When the terminal devices 31 to 33 detect a sudden braking operation (S306: YES), they transmit audio information representing a voice notifying the driver that a sudden braking operation has been performed (S307). This allows the user of the terminal device 31 to 33 to know that a sudden braking operation has been performed in another vehicle while driving. This allows for safer driving compared to checking the status of other vehicles only visually.

[0186] (14) When the terminal device 31 to 33 detects a passing operation (S303: YES), the terminal device 31 to 33 transmits audio information representing a sound notifying the user that a passing operation has been performed (S304). This allows the user of the terminal device 31 to 33 to know that another vehicle has performed a passing operation while driving. This makes it easier to understand signals from other vehicles compared to checking the status of other vehicles only by visual inspection.

[0187] (15) When the terminal device 31 to 33 detects the blinker operation (S301: YES), the terminal device 31 to 33 transmits voice information representing a voice notifying that the blinker operation has been performed (S30 2) Therefore, the users of the terminal devices 31 to 33 can know that the blinkers of other vehicles have been operated while driving. Therefore, it is easier to predict the movements of other vehicles compared to when checking the status of other vehicles only by visual inspection.

[0188] (16) The terminal devices 31 to 33 detect a specific state related to the vehicle as a specific event (S401, S403, S405, S408). Then, when the terminal devices 31 to 33 detect the specific state (S401: YES, S403: YES, S405: YES, or S408: YES), they transmit audio information representing audio notifying that the specific state has been detected (S402, S404, S406, S409, S116). Therefore, the users of the terminal devices 31 to 33 can know, while driving, that a specific state has been detected in another vehicle.

[0189] (17) When the terminal device 31 to 33 detects that the ABS has been activated (S401: YES), the terminal device 31 to 33 transmits audio information representing audio notifying the user that the ABS has been activated (S402). This allows the user of the terminal device 31 to 33 to know that the ABS has been activated in another vehicle while driving. This makes it easier to predict situations where slippage is likely compared to checking the road surface conditions visually alone.

[0190] (18) When the terminal devices 31 to 33 detect a state in which the underside of the vehicle is in contact with a bump or the like on the road surface or a state in which there is a high risk of contact (S403: YES), they transmit audio information representing a voice notifying that there is a risk of the underside of the vehicle scraping (S404). This makes it easier to predict a situation in which there is a high risk of scraping the underside of the vehicle compared to when the road surface condition is checked only visually.

[0191] (19) When the terminal device 31 to 33 detects that the vehicle is experiencing a crosswind (S405: YES), the terminal device 31 to 33 transmits audio information representing a voice notifying the driver that the vehicle is experiencing a crosswind (S406). This makes it easier to predict situations where the vehicle is susceptible to the effects of a crosswind compared to when the driver only visually checks the situation outside the vehicle.

[0192] (20) When the terminal devices 31 to 33 detect that a vehicle has fallen or collided (S408: YES), they transmit audio information representing audio notifying the user that the vehicle has fallen or collided (S409). This allows the user of the terminal device 31 to 33 to know that another vehicle has fallen or collided while driving. This allows for safer driving compared to checking the status of other vehicles only by visual inspection.

[0193] (21) In addition, the following effects can be obtained: The communication system 1 described above in detail includes a plurality of terminal devices 31 to 33 that can communicate with each other, and also includes a server 10 that can communicate with each of the plurality of terminal devices 31 to 33. When a user inputs voice, the CPU 41 of each of the terminal devices 31 to 33 converts the voice into voice information and transmits the voice information to devices including the other terminal devices 31 to 33.

[0194] The CPU 41 then receives and plays back audio information transmitted from the other terminal devices 31 to 33. The server 10 also receives audio information transmitted from the plurality of terminal devices 31 to 33, aligns the audio information so that the audio corresponding to each piece of audio information does not overlap, and then distributes the audio information to the plurality of terminal devices 31 to 33.

[0195] According to such a communication system 1, it is possible to input the contents of a message by voice and to play back the contents of messages from other terminal devices 31 to 33 by voice, so that even a user who is driving can safely send and receive the contents of a message he or she wants to tweet. Also, in such a communication system 1, when playing back the voice, the voice can be played back so that the voice does not overlap, so that the voice can be played back without overlapping. It can make your voice easier to hear.

[0196] Furthermore, in the communication system 1, when a specific command is received from the user, the CPU 41 replays audio information that has already been played back. In particular, the audio information that was played back immediately before may be played back. According to such a communication system 1, audio information that was inaudible due to noise or the like can be replayed in response to a command from the user.

[0197] Furthermore, in the above communication system 1, the terminal devices 31 to 33 are mounted on vehicles, and the CPU 41 of the terminal devices 31 to 33 starts a process of generating a voice packet when it detects a specific start action to start voice input via the touch panels 61 and 62, which detect the movement of the user's hands on the steering wheel of the vehicle, and ends the process of generating the voice packet when it detects a specific end action to end voice input via the touch panels 61 and 62. According to this communication system 1, an operation for inputting voice can be performed on the steering wheel. Therefore, the user can input voice more safely while driving the vehicle.

[0198] Furthermore, in the communication system 1, the CPU 41 detects the start specific motion and the end specific motion input via the touch panels 61, 62 as hand movements in a direction different from the operation direction when operating the horn. With this communication system 1, it is possible to prevent the horn from being sounded when the start specific motion or the end specific motion is performed.

[0199] Furthermore, in the above-described communication system 1, the server 10 generates new voice information by adding, to the voice information input by the user, identification information for identifying the user, and distributes the generated voice information. In such communication system 1, since identification information for identifying a user can be added by voice, the user can notify other users of the user who sent the voice without speaking his / her own identification information.

[0200] Furthermore, in the above communication system 1, before playing back the audio information, the server 10 extracts from the audio information any preset keywords contained in the audio information and replaces the keywords contained in the audio information with the sound of another word or a preset sound. According to such communication system 1, words that are undesirable for distribution (obscene words, words indicating people's names, vulgar words, etc.) can be registered as keywords and then replaced with words or sounds that are acceptable for distribution before distribution.

[0201] Furthermore, in the communication system 1, the CPU 41 of each of the terminal devices 31 to 33 acquires its own location information, transmits the location information to the other terminal devices 31 to 33 in association with audio information transmitted from its own terminal device 31 to 33, and acquires the location information transmitted by the other terminal devices 31 to 33. Then, the volume of the audio to be output when the audio is played back is controlled according to the positional relationship between the position of the terminal device 31 to 33 corresponding to the audio to be played back and the position of its own terminal device 31 to 33.

[0202] In particular, the CPU 41 controls the volume so that it decreases as the distance to the other terminal devices 31 to 33 increases, and when outputting sound from multiple speakers, controls the volume so that the volume increases from the direction in which the other terminal devices 31 to 33 are located as seen from the user, depending on the orientation in which the other terminal devices 31 to 33 are located relative to the position of the user's own terminal device 31 to 33.

[0203] According to such a communication system 1, the volume is controlled in accordance with the positional relationship between the position of the terminal device 31 to 33 corresponding to the audio being played and the position of the user's own terminal device 31 to 33. Therefore, the user listening to the audio can sense the positional relationship between the other terminal devices 31 to 33 and the user's own terminal device 31 to 33. It can be intuitively grasped.

[0204] Furthermore, in the communication system 1, the server 10 converts the audio information into text and outputs the converted text to a display device. According to such a communication system 1, it is possible to check, in text, information that was missed in audio.

[0205] Furthermore, in the communication system 1, the CPU 41 of each of the terminal devices 31 to 33 transmits information indicating the vehicle state and the user's driving operation (operation status of lights, wipers, television, radio, etc., driving status such as driving speed and direction, detection values ​​by vehicle sensors, control status such as the presence or absence of malfunctions, etc.) along with the voice. In such a communication system 1, information indicating the user's driving operation can also be sent to the other terminal devices 31 to 33.

[0206] Furthermore, in the communication system 1, advertisements are played by voice in accordance with the locations of the terminal devices 31 to 33. The content of the advertisements is inserted by the server 10, and the server 10 transmits the advertisements to the terminal devices 31 to 33 that play back the audio information. Such a communication system 1 can provide a business model that can earn advertising revenue.

[0207] Furthermore, the communication system 1 is set to be switchable between a distribution mode in which communication partners are an unspecified number of people and a conversation mode in which communication partners are specific people. With this communication system 1, you can enjoy talking with specific people.

[0208] Furthermore, in the above-mentioned communication system 1, the server 10 determines the emotion and politeness of the user's speech when inputting the voice by detecting the amplitude of the input level and polite expressions contained in the voice from the voice information, and does not distribute voices from users who are emotional or do not speak politely. With this communication system 1, it is possible to eliminate comments from users that could cause arguments.

[0209] Furthermore, in the above-described communication system 1, the server 10 notifies a user who speaks ill of others, or who is emotional or does not speak politely of a gaffe or the like. In this case, the server 10 makes a determination based on whether or not a word registered as a keyword for a gaffe or the like is included in the voice information. With such a communication system 1, it is possible to urge a user who has made a gaffe to stop making such a gaffe.

[0210] Note that S102 corresponds to an example of processing as an input means, S114, S1202, S1204, and S1205 correspond to an example of processing as an acquisition means, S116 corresponds to an example of processing as a transmission means, and S811 corresponds to an example of processing as a playback means.

[0211] [Other embodiments] The present invention is not limited to the above-described embodiment, and various other forms may be adopted as long as they fall within the technical scope of the present invention.

[0212] (1) In the above embodiment, the server 10 executes the extraction process of extracting audio information corresponding to the terminal devices 31 to 33 that play the audio from among the multiple pieces of audio information stored (S1302 to S1312), but this is not limited to this. For example, the terminal devices 31 to 33 that play the audio may execute part or all of this extraction process. For example, the server 10 may transmit audio information to the terminal devices 31 to 33 in a form that may include unnecessary audio information that may not be played on the terminal devices 31 to 33, and the terminal devices 31 to 33 may extract audio information to be played from the received audio information (discard unnecessary audio information). Note that the transmission of audio information from the server 10 to the terminal devices 31 to 33 may be performed in response to a request from the terminal devices 31 to 33, or may be performed regardless of a request. You can do this.

[0213] (2) In the above embodiment, the server 10 determines whether each piece of candidate speech information for transmission overlaps with other pieces of speech information (S1311), and if it determines that there is overlap (S1311: YES), it eliminates the overlapping content (S1312). This can prevent the phenomenon of similar utterances being repeatedly played back, but from another perspective, the greater the number of pieces of speech information with overlapping content, the higher the credibility of that content. Therefore, a condition for playing back speech with that content may be that there is a certain number or more pieces of speech information with overlapping content.

[0214] 21A and 21B is obtained by replacing the processes of S1311 to S1312 in the response process of Figures 17A and 17B with the processes of S1401 to S1403. That is, in the response process shown in Figures 21A and 21B, it is determined whether or not each piece of voice information that is a candidate for transmission contains a specific keyword that has been set in advance, such as "accident" or "traffic regulation" (S1401).

[0215] Then, if it is determined that voice information containing a specific keyword exists among the voice information of transmission candidates (S1401: YES), it is determined whether or not there is less than a certain number (for example, three) of voice information with the same content (duplicate content) among the voice information containing the specific keyword (S1402). Here, if it is determined that there is less than a certain number of voice information with the same content (S1402: YES), the voice information with that content is excluded from the voice information of transmission candidates (S1403). Thereafter, the voice information of transmission candidates is transmitted to the terminal device 31 to 33 that sent the information request (S1313), and the response process is terminated.

[0216] On the other hand, if it is determined that there is no voice information containing a specific keyword among the voice information candidates for transmission (S1401: NO), or if it is determined that there is no voice information with the same content less than a certain number of voice information (S1402: NO), the process proceeds directly to S1313, and the voice information candidates for transmission are transmitted to the terminal device 31 to 33 that sent the information request (S1313), and the response process is terminated.

[0217] 21A and 21B, voice information containing a specific keyword is not transmitted to the terminal devices 31 to 33 until a certain number of voice information pieces with the same content are accumulated, and is transmitted to the terminal devices 31 to 33 on the condition that the certain number is reached. In this way, it is possible to prevent voices with low credibility, such as false information, from being played back on the terminal devices 31 to 33.

[0218] 21A and 21B, the condition for transmission is that a certain number of pieces of voice information with the same content exist for voice information containing a specific keyword, but this is not limiting. For example, the condition for transmission may be that a certain number of pieces of voice information with the same content exist for all of the voice information candidates for transmission, regardless of whether they contain a specific keyword. Furthermore, in determining whether a certain number of pieces of voice information with the same content exist, even if there are multiple pieces of voice information with the same sender, they may be counted as one (not counted as multiple). In this way, it is possible to prevent voices with low credibility from being played on the terminal devices 31 to 33 of other users, which may be caused by the same user uttering voices with the same content multiple times.

[0219] (3) In the above embodiment, the terminal devices 31 to 33 that transmit the voice information convert the recorded voice into synthetic voice (S107), but this is not limiting. For example, part or all of this conversion process may be performed by the server 10, or may be performed by the terminal devices 31 to 33 that play back the voice.

[0220] (4) In the above embodiment, the terminal devices 31 to 33 execute the request process (FIG. 16) at predetermined time intervals, but this is not limiting and the process may be executed at other predetermined timings. For example, the process may be executed every time a predetermined distance is traveled, when a user performs a request operation, at a specified time, when a specified point is passed, when the vehicle leaves the surrounding area, etc.

[0221] (5) In the above embodiment, audio information within the surrounding area is extracted and played back, but this is not limited to this. For example, an expanded area that includes the surrounding area (larger than the surrounding area) may be set, audio information within the expanded area may be extracted, and audio information within the surrounding area may be further extracted and played back from that expanded area. Specifically, in S1307 of the above embodiment, audio information whose utterance position is within the expanded area is extracted. Then, the terminal devices 31 to 33 store (cache) the audio information within the expanded area received from the server 10, and when playing back the audio represented by the audio information, they extract and play back the audio information whose utterance position is within the surrounding area. In this way, even if the surrounding area based on the current location changes as the vehicle travels, as long as the new surrounding area is within the expanded area at the time the audio information was received, the audio represented by the audio information corresponding to the surrounding area can be played back.

[0222] (6) In the above embodiment, the surrounding area is set based on the current location (actual position) of the vehicle. However, this is not limited to this. Location information other than the current location may be acquired, and the surrounding area may be set based on the acquired location information. Specifically, location information set (input) by the user may be acquired. In this way, for example, by setting the surrounding area based on the destination, audio information transmitted near the destination can be confirmed in advance from a location far from the destination.

[0223] (7) In the above embodiment, the terminal devices 31 to 33 configured as in-vehicle devices mounted on a vehicle are exemplified as terminal devices, but the present invention is not limited to this. For example, the terminal devices may be devices (such as mobile phones) that are portable by the user and can be used in a vehicle. Furthermore, the terminal devices may or may not share at least a part of their configuration (for example, a configuration for detecting a position, a configuration for inputting voice, a configuration for playing voice, etc.) with the devices mounted on the vehicle.

[0224] (8) In the above embodiment, an example was given of a configuration in which, among the voice information transmitted from other terminal devices 31 to 33 to the server 10, the voice representing the voice information whose utterance position is located within a surrounding area set based on the current location is played back. However, this is not limited to this, and the voice information may be extracted and played back regardless of the location information.

[0225] (9) Information input by the user to the terminal device is not limited to the user's voice, but may be, for example, text information input by the user. In other words, the terminal device may input text information from the user instead of (or in addition to) the user's voice. In this case, the terminal device may display the text information so that the user can see it, or may convert the text information into synthetic voice and play it back.

[0226] (10) The above-described embodiment is merely an example of an embodiment to which the present invention is applied. The present invention can be realized in various forms, such as a system, an apparatus, a method, a program, and a recording medium.

[0227] (11) In addition, the following may be done: For example, in the above communication system 1, the content (for example, real name, first nickname, second nickname, etc.) to be added as identification information may be changed depending on the communication mode. For example, information for identifying each party may be exchanged in advance, and when a communication partner is present, the content may be changed. The identification information may be changed depending on whether the other party is registered in advance or not.

[0228] Furthermore, in the above embodiment, the keyword is replaced with another word or the like in any case, but the process of replacing the keyword with another word or the like may not be performed depending on conditions such as the communication partner, communication mode, etc. This is because when communicating with a specific partner, unlike when distributing to the public, such a process does not need to be performed.

[0229] Furthermore, meaningless words that are typical of spoken language, such as "um" and "er," may be omitted from the generated audio information. In this case, the playback time of the audio information can be shortened by the amount of omission.

[0230] In the above embodiment, the CPU 41 of each terminal device 31 to 33 outputs information indicating the vehicle state and the user's driving operation in text form, but it may also output the information in audio form. Also, various information such as traffic congestion information may be generated based on information obtained from other terminal devices 31 to 33, and the generated information may be output.

[0231] Furthermore, in the above communication system 1, it is possible to switch between a configuration in which the terminal devices 31 to 33 exchange voice information directly with each other and a configuration in which the terminal devices 31 to 33 exchange voice information with each other via the server 10, depending on pre-set conditions (for example, the distance to the base station connected to the server 10, the communication status between the terminal devices 31 to 33, or user settings).

[0232] Furthermore, in the above embodiment, the advertisement is played by voice depending on the location of the terminal device 31-33, but the advertisement may be played by voice at fixed time intervals or after a fixed number of audio information playbacks. The content of the advertisement may be transmitted from the terminal device 31-33 of the advertiser to the other terminal devices 31-33. In this case, the user may input a predetermined operation into the terminal device 31-33 to make a call (communicate) with the advertiser, or the user of the terminal device 31-33 may be guided to the advertiser's store.

[0233] Furthermore, in the above embodiment, a mode for communicating with a plurality of specific communication partners may be set. For example, in the case of a configuration for exchanging position information, the communication partners may be the terminal devices 31 to 33 within a certain distance, or a specific user (such as a preceding vehicle or an oncoming vehicle (identified by capturing an image of the license plate using a camera or by GPS)) or a user within a preset group. Also, the communication partner may be changed by setting (the communication mode may be changed).

[0234] Furthermore, when the terminal devices 31 to 33 are mounted on vehicles, the terminal devices 31 to 33 traveling in the same direction on the same road may be set as communication partners, or the terminal devices 31 to 33 with matching vehicle behavior (for example, terminal devices 31 to 33 may be removed from communication partners if they veer off onto a side road) may be set as communication partners.

[0235] In addition, it may be configured so that other users can be specified and registered as favorites or excluded in advance, allowing communication with a selected partner. In this case, communication partner information specifying the communication partner may be transmitted when transmitting voice, or, if a server exists, communication partner information may be registered in the server, so that voice information can be transmitted only to specified partners.

[0236] Alternatively, a plurality of switches such as buttons may be associated with respective directions, and when the user uses the switch to specify the direction in which the communication partner is located, only users in that direction may be set as communication partners.

[0237] Furthermore, although the communication system 1 is configured to exchange voice information, it is also possible to acquire information consisting of text, convert this information into voice, and play it back. Also, when transmitting voice, it is also possible to convert the voice into text before transmitting, and then restore it to voice before playing it back. This configuration can reduce the amount of data to be transmitted by communication.

[0238] Furthermore, in a configuration in which data is transmitted by characters in this way, the data may be transmitted after being translated into the language that produces the smallest amount of data from among a plurality of languages ​​set in advance.

[0239] Furthermore, some of the functions of the server 10 in the above embodiment may be implemented by the terminal devices 31 to 33. For example, when there are multiple pieces of voice information that have not yet been played back in the CPU 41 of the terminal devices 31 to 33, the voices corresponding to the pieces of voice information may be aligned so as not to overlap, and then played back. The CPU 41 may generate new voice information by adding, to the voice information input by the user, voice-based identification information for identifying the user, and transmit the generated voice information.

[0240] Specifically, part of the audio distribution process may be performed in the audio transmission process or the audio playback process. For example, when part of the audio distribution process is performed in the audio playback process, the processes of S703 to S722 (excluding S717) of the audio distribution process may be performed between S802 and S804 of the audio playback process, as shown in Figures 22A-22D.

[0241] Even in such a communication system, when audio is played back, it is possible to play back audio without overlapping, making it easier to hear the audio. Also, even in such communication system 1, identification information for identifying a user can be added to the audio, so that a user can notify other users of the user who sent the audio without speaking their own identification information.

[0242] Although the above embodiment prohibits overlapping playback of audio, audio may be partially overlapped. In this case, only the last few seconds of the earlier audio and the first few seconds of the later audio may overlap. In this case, a fade-out effect, which gradually reduces the volume at the end of the earlier audio, and a fade-in effect, which gradually increases the volume at the beginning of the later audio, may be combined.

[0243] Furthermore, whether the driver is dozing off or not may be determined based on the presence or absence of a voice uttered by the user. To realize this configuration, for example, a dozing off driver determination process shown in FIG. The drowsiness determination process is a process performed in parallel with the voice transmission process in the terminal devices 31 to 33, and is a process for issuing an alarm or controlling the vehicle, such as stopping the vehicle, when the user does not speak for a certain period of time or when the user speaks but is estimated to have a low level of consciousness. In detail, as shown in Fig. 23, first, the time since the last user's utterance is counted (S1401).

[0244] Then, it is determined whether a preset first reference time (for example, about 10 minutes) has elapsed since the user's last utterance (S1402). If the first reference time has not elapsed (S1402: NO), the dozing determination process ends.

[0245] If the first reference time has elapsed (S1402: YES), a question is asked to the user (S1403). In this process, a question is asked to prompt the user to speak in order to confirm the user's level of consciousness. For this purpose, a direct question such as "Are you awake?" is used. Alternatively, an indirect question such as "What time is it now?" may be used.

[0246] Next, it is determined whether or not a response to such an inquiry has been obtained (S1404). If a response to the inquiry has been obtained (S1404: YES), the level of the response is determined (S1405).

[0247] Here, the response level refers to items such as the volume of the voice and clarity of speech, the level of which (volume of sound, accuracy of voice recognition) decreases when the person becomes drowsy, and this process makes a judgment by comparing one of these items with a preset threshold value.

[0248] If the response level is sufficiently high (S1405: YES), the dozing determination process is terminated. If the response level is insufficient (S1405: NO), or if no response to the question is obtained in the process of S1404 (S1404: NO), it is determined whether a second reference time (for example, about 10 minutes and 10 seconds) set to a value equal to or greater than the first reference time has elapsed since the user's last utterance (S1406).

[0249] If the second reference time has not elapsed (S1406: NO), the process returns to S1403. If the second reference time has elapsed (S1406: YES), an alarm is issued via the speaker 53 or the like (S1407).

[0250] Next, it is determined whether or not a response to the alarm (for example, a voice such as "Wow" or an operation via the input unit 45 to cancel the alarm) has been obtained (S1408). If a response to the alarm has been obtained (S1408: YES), the level of the response is determined (S1409) in the same manner as in the processing of S1405. Note that if an operation has been input via the input unit 45, it is determined that the level of the response is sufficient.

[0251] If the response level is sufficiently high (S1409: YES), the dozing determination process is terminated. If the response level is insufficient (S1409: NO), or if no response to the question is obtained in the process of S1408 (S1408: NO), it is determined whether a third reference time (e.g., about 10 minutes and 30 seconds), which is set to a value equal to or greater than the second reference time, has elapsed since the user's last utterance (S1410).

[0252] If the third reference time has not elapsed (S1410: NO), the process returns to S1407. If the third reference time has elapsed (S1410: YES), vehicle control such as stopping the vehicle is performed (S1411), and the drowsiness determination process ends.

[0253] By carrying out such processing, it is possible to prevent the user (driver) from falling asleep at the wheel. Furthermore, when an advertisement is inserted by voice (S720) or when voice is also displayed as text (S812), additional text information may be displayed on the display of the navigation device 48 or the like.

[0254] In order to realize such a configuration, for example, the additional information display process shown in Fig. 24 may be performed. The additional information display process is a process that is performed in the server 10 in parallel with the audio distribution process.

[0255] In the additional information display process, as shown in Fig. 24, first, it is determined whether an advertisement (CM) has been inserted in the audio distribution process (S1501) and whether text has been displayed (S1502). If an advertisement has been inserted (S1501: YES) or text has been displayed (S1502: YES), information related to the text to be displayed is obtained (S1503).

[0256] Here, the information related to the characters corresponds to information associated with the character string, such as advertisements or local information corresponding to the character string (keyword) to be displayed, or information explaining the information related to the character string, etc. This information may be obtained from within the server 10 or from outside the server 10 (such as from the Internet).

[0257] When the information related to the text is acquired, this information is displayed on the display as text or an image (S1504), and the additional information display process ends. It should be noted that if an advertisement is not inserted and no text is displayed in the processes of S1501 and S1502 (S1501: NO and S1502: NO), the additional information display process ends.

[0258] According to this configuration, not only can the voice be displayed as text, but also related information can be displayed. The display that displays the text information is not limited to the display of the navigation device 48, but may be any display such as a head-up display inside the vehicle. [Explanation of symbols]

[0259] 1...communication system, 10...server, 20...base station, 31-33...terminal device, 41...CPU, 42...ROM, 43...RAM, 44...communication unit, 45...input unit, 48...navigation device, 49...GPS receiver, 52...microphone, 53...speaker

Claims

1. A communication system including a terminal device and a server device capable of communicating with the terminal device, The terminal device When a user inputs voice, the voice is converted into a voice signal; transmitting the audio signal to the server device; acquiring its own location information, and transmitting the location information in association with the voice signal transmitted from its own terminal device; receiving the transmitted audio signal from the server device; The server device receiving a voice signal transmitted from the terminal device; generating a new voice signal by adding identification information for identifying the user who input the voice to the received voice signal; delivering the audio signal with the identification information attached to it to the terminal device; extracting audio information corresponding to the location information and delivering the extracted audio signal to the terminal device; Communication system.

2. 2. The communication system of claim 1, The server device receiving the audio signals transmitted from the terminal device, aligning the audio signals so that the audio corresponding to each audio signal does not overlap, and then distributing the audio signals to the terminal device; Communication system.

3. 3. A communication system according to claim 1 or 2, the server device generates a voice signal from the received voice signal in which meaningless words have been omitted, and distributes the voice signal in which meaningless words have been omitted. Communication system.

4. The communication system according to any one of claims 1 to 3, The server device converts the received voice signal into text. Communication system.

5. The communication system according to any one of claims 1 to 4, The terminal device is mounted on a vehicle, The terminal device transmits the voice signal to which information indicating the state of the vehicle or the driving operation of the user is added. Communication system.

6. The communication system according to any one of claims 1 to 5, The terminal device is mounted on a vehicle, The terminal device transmits text information indicating the state of the vehicle or the driving operation of the user to the server device. Communication system.

7. A communication system according to any one of claims 1 to 6, transmitting data by characters between the plurality of terminal devices and the server device; If the data is transmitted in text, the data is translated before being transmitted; Communication system.

8. The communication system according to any one of claims 1 to 7, A plurality of the terminal devices can communicate with each other. Communication system.

9. A communication system according to any one of claims 1 to 8, The terminal device reproduces the received audio signal; When a command is input by the user, the audio signal being reproduced is skipped by a predetermined unit. Communication system.

10. A communication system according to any one of claims 1 to 9, The terminal device reproduces the received audio signal; When a replay operation is input during playback of the audio signal, the terminal device plays back the audio packet being played back or the audio packet immediately after playback from the beginning. Communication system.

11. A communication system according to any one of claims 1 to 9, The terminal device is mounted on a vehicle, the terminal device acquires location information of the vehicle at predetermined time intervals and transmits the acquired location information of the vehicle to the server device; Communication system.

12. 12. The communication system of claim 11, when the vehicle passes a specific position, the server device adds an advertisement corresponding to the specific position to the audio signal to be transmitted to the terminal device mounted on the vehicle that has passed the specific position; Communication system.

13. A vehicle equipped with a terminal device described in any one of claims 1 to 12.

Citation Information

Patent Citations

  • Information terminal device and information server

    JP2003106850A

  • Communication method and system, and communication device constituting the communication system

    JP2004343784A

  • Radio terminal, management server and PTT system

    JP2006246201A

  • Radio communication system, and information deployment method

    JP2011114738A

  • Terminal, system and method for displaying information, and program

    JP2011215861A