Communication systems and vehicles
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-03-25
AI Technical Summary
Existing communication systems do not adequately address the needs of users who have difficulty operating or visually recognizing terminal devices, particularly in vehicles, for safe and efficient transmission and reception of text messages.
A communication system comprising multiple terminal devices and a server, equipped with voice input conversion, transmission, and playback capabilities, ensuring non-overlapping audio playback and integration with vehicle sensors for safe voice input during driving, along with features like motion detection and identification information addition.
Enables safe and efficient voice-based message transmission and reception, even for users with disabilities, while reducing distractions and enhancing safety during vehicle operation.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Cross - reference to related applications , , , ,
[0005] , , ,
[0003] , , , , , , , , ,
[0002] , ,
[0004] , , , , , , , , , ,
[0006] ,
[0001] This international application claims priority based on Japanese Patent Application No. 2011 - 273578 and Japanese Patent Application No. 2011 - 273579, which were filed with the Japan Patent Office on December 14, 2011, and Japanese Patent Application No. 2012 - 074627 and Japanese Patent Application No. 2012 - 074628, which were filed with the Japan Patent Office on March 28, 2012, and incorporates by reference the entire contents of Japanese Patent Application No. 2011 - 273578, No. 2011 - 273579, No. 2012 - 074627, and No. 2012 - 074628 into this international application.
Technical Field
[0002] The present invention relates to a communication system including a plurality of terminal devices and a terminal device.
Background Art
[0003] There is known a system in which when a large number of users use a terminal device such as a mobile phone to input text messages and other transmission contents, the server distributes the transmission contents arranged in order to each terminal device (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] It is preferable that even a user who has difficulty operating or visually recognizing a terminal device can safely transmit and receive transmission contents such as text messages.
Means for Solving the Problems
[0006] One aspect of the present invention is A communication system comprising multiple terminal devices capable of communicating with each other, Each terminal device, A voice input conversion means that converts the voice input by the user into an audio signal, A voice transmission means for transmitting the aforementioned voice signal to a device including other terminal devices, A voice receiving means for receiving voice signals transmitted from another terminal device, Audio playback means for playing back the received audio signal, Equipped with, If there are multiple audio signals that have not yet been played back, the audio playback means will arrange the audio corresponding to each audio signal so that there is no overlap before playing them back.
[0007] Another aspect of the present invention is, A communication system comprising a plurality of terminal devices and a server device capable of communicating with each of the plurality of terminal devices, Each terminal device, A voice input conversion means that converts the voice input by the user into an audio signal, A voice transmission means for transmitting the aforementioned voice signal to the server device, A voice receiving means for receiving voice signals transmitted by the server device, Audio playback means for playing back the received audio signal, Equipped with, The server device is The system includes distribution means that receives audio signals transmitted from the aforementioned multiple terminal devices, arranges the audio corresponding to each audio signal so that there is no overlap, and then distributes the audio signals to the aforementioned multiple terminal devices.
[0008] With these types of communication systems, users can input messages by voice and play back messages from other communication devices by voice, allowing even users who have difficulty operating or visually inspecting terminal devices to safely send and receive messages such as tweets.
[0009] Furthermore, these types of communication systems can play audio without overlapping, making it easier to hear. Furthermore, in any of the above communication systems, the audio playback means may, upon receiving a specific command from the user, replay an audio signal that has already been played. In particular, it may replay the audio signal that was played immediately before.
[0010] With such a communication system, audio signals that were not audible due to noise or other factors can be played back at the user's command. Furthermore, in any of the above communication systems, the terminal device may be mounted on a vehicle, and the terminal device may include motion detection means for detecting the movement of the user's hand on the steering wheel of the vehicle, and motion control means for causing the voice input conversion means to start operation when the motion detection means detects a start-specific operation to start voice input, and for causing the voice input conversion means to stop operation when the motion detection means detects a end-specific operation to end voice input.
[0011] With this type of communication system, voice input can be performed on the steering wheel. Therefore, users can input voice more safely while driving the vehicle. Furthermore, in any of the above communication systems, the start-specific operation and the end-specific operation may be hand movements in a direction different from the direction of operation when operating the horn.
[0012] Such a communication system can suppress the sounding of the horn during specific start or end operations. Next, the terminal device may be configured as a terminal device that constitutes any of the above-mentioned communication systems.
[0013] Such terminal devices can achieve the same effects as any of the communication systems described above. In the above communication system (terminal device), the voice transmission means may generate a new voice signal by adding, in voice, identification information for identifying the user to the voice signal input by the user, and transmit the generated voice signal.
[0014] Further, in a communication system including a server device capable of communicating with each of a plurality of terminal devices and the plurality of terminal devices, the distribution means may generate a new voice signal by adding, in voice, identification information for identifying the user to the voice signal input by the user, and distribute the generated voice signal.
[0015] According to such communication systems, since the transmission content can be input in voice and the transmission content from other communication devices can be reproduced in voice, even a user who has difficulty operating or visually recognizing the terminal device can safely transmit and receive the transmission content such as what they want to whisper.
[0016] Also, in such communication systems, since identification information for identifying the user can be added in voice, the user can notify other users of the voice transmission source without speaking their own identification information. the user of the transmission source can be notified.
[0017] Furthermore, in any of the above communication systems, the identification information to be added may be the user's real name or a nickname other than the real name. Also, the content to be added as identification information (for example, real name, first nickname, second nickname, etc.) may be changed according to the communication mode. For example, information for identifying each other may be exchanged in advance, and the identification information to be added may be changed according to whether the communication partner is a pre-registered partner or not.
[0018] Also, in any of the above communication systems, before reproducing the voice signal, it may be provided with keyword extraction means for extracting a preset keyword included in the voice signal from the voice signal, and voice replacement means for replacing the keyword included in the voice signal with the voice of another word or a preset sound.
[0019] According to such a communication system, it is possible to register words that are not preferable for distribution (such as obscene words, words indicating personal names, or barbaric words, etc.) as keywords, and perform distribution by replacing them with words or sounds that are acceptable for distribution.
[0020] Note that, for example, the voice replacement means may be activated or not activated depending on conditions such as the communication partner. Also, for words that have no specific meaning in spoken language, such as "um" or "uh," the voice signal may be generated with these words omitted. In this case, the playback time of the voice signal can be shortened by the omitted amount.
[0021] In any of the above communication systems, each terminal device may include position information transmission means for acquiring its own position information and transmitting it to another terminal device in association with the voice signal transmitted from its own terminal device, position information acquisition means for acquiring the position information transmitted by another terminal device, and volume control means for controlling the volume output when the voice is played back by the voice playback means according to the positional relationship between the position of the terminal device corresponding to the played-back voice and the position of its own terminal device.
[0022] According to such a communication system, since the volume is controlled according to the positional relationship between the position of the terminal device corresponding to the played-back voice and the position of its own terminal device, the user listening to the voice can intuitively grasp the positional relationship between the other terminal device and its own terminal device.
[0023] Note that the volume control means may control the volume to decrease as the distance to the other terminal device increases, or when outputting voice from a plurality of speakers, control the volume so that the volume from the direction where the other terminal device is located as viewed by the user increases according to the azimuth of the other terminal device relative to the position of its own terminal device. Also, such controls may be combined.
[0024] Furthermore, the above communication system may also include a character conversion means for converting audio signals into characters, and a character output means for outputting the converted characters to a display device. With this type of communication system, information that was missed during audio can be confirmed in text.
[0025] Furthermore, any of the above communication systems may transmit information indicating the vehicle's status and the user's driving operations (such as the operating status of lights, wipers, television, radio, etc., driving conditions such as speed and direction of travel, and control conditions such as detected values from the vehicle's sensors and whether or not there are malfunctions) along with the voice.
[0026] In this case, this information may be output as audio or as text. Also, for example... For example, the system may generate various types of information, such as traffic congestion information, based on information obtained from other terminal devices, and output the generated information.
[0027] Furthermore, in any of the above communication systems, the system may be configured to switch between a configuration in which each terminal device directly exchanges voice signals with each other and a configuration in which each terminal device exchanges voice signals via a server device, based on pre-set conditions (for example, the distance to the base station connected to the server device, the communication status between terminal devices, or user settings).
[0028] Furthermore, advertisements may be played via audio at regular intervals, after a certain number of audio signal playbacks, or depending on the location of the terminal device. The content of the advertisement can be transmitted from the advertiser's terminal device or server device to the terminal device that plays the audio signal.
[0029] In this case, the system may be configured to allow users to make calls (communicate) with advertisers by inputting predetermined operations into the terminal device, or to guide users of the terminal device to the advertiser's store.
[0030] Furthermore, in any of the above communication systems, the communication partners may be an unspecified number of people, or they may be limited to specific individuals. For example, in a configuration for exchanging location information, the communication partner may be a terminal device within a certain distance, a specific user (such as a preceding vehicle or an oncoming vehicle (identified by license plate imaging using a camera or GPS)), or a user within a pre-configured group. Furthermore, the communication partner may be changed through settings (the communication mode may be changed).
[0031] Furthermore, if the terminal device is mounted on a vehicle, it may be used as a communication partner if it is a terminal device moving on the same road in the same direction, or if it is a terminal device whose vehicle behavior matches (for example, if it deviates onto a side road, it may be excluded from the communication partner).
[0032] Alternatively, the system could be configured to allow users to identify other users and pre-set settings such as adding them to favorites or excluding them, enabling them to select their preferred partners for communication. In this case, when sending voice, the system could send communication partner information to identify the recipient, or, if a server device exists, register the communication partner information with the server device, thereby enabling the transmission of voice signals only to designated recipients.
[0033] Alternatively, multiple buttons or switches could be assigned to specific directions, allowing the user to specify the direction of their communication partner using these switches. In this case, only users in that direction could be designated as communication partners.
[0034] Furthermore, the system may determine whether a user is expressing emotion or speaking politely when inputting voice by detecting the amplitude of the input level and the polite expressions contained in the voice signal. The system may then avoid distributing voices from users who are emotionally agitated or not speaking politely.
[0035] Additionally, the system could be configured to notify users who speak ill of others, or who are emotionally agitated or speak in an impolite manner. In this case, the system could determine whether or not a word registered as a keyword for insults or other offensive language is included in the audio signal.
[0036] Furthermore, although the above communication systems are configured to exchange voice signals, they may also be configured to acquire information consisting of text, convert this information into voice, and then play it back. When transmitting audio, it is also possible to convert the audio to text before transmission, and then restore it to audio before playback. This configuration can reduce the amount of data that needs to be transmitted via communication.
[0037] Furthermore, in a configuration where data is transmitted using text, the data may be translated into the language that uses the least amount of data among several pre-configured languages before being transmitted.
[0038] By the way, in recent years, for example, Twitter (registered trademark) and Facebook (registered trademark), users have been The number of users of services that manage information they share in a way that other users can view is rapidly increasing. In such services, terminal devices such as personal computers and mobile phones (including so-called smartphones) are used for sending and viewing information (Japanese Patent Publication No. 2011-215861).
[0039] On the other hand, users (e.g., drivers) who are in a vehicle (e.g., a car) often have a lot of free time on their hands because their actions are more restricted compared to when they are not in a vehicle. Therefore, using the aforementioned services while in a vehicle allows them to make effective use of their time. However, conventional services have not adequately considered how users in vehicles use them.
[0040] Therefore, there is a need to provide technology that can build useful services for users riding in vehicles. As a technology for this purpose, the terminal device of the first reference invention comprises: an input means for inputting voice; an acquisition means for acquiring location information; a transmission means for transmitting voice information representing the voice input by the input means to a server shared by a plurality of terminal devices in a form that allows the voice position, which is location information acquired by the acquisition means at the time of voice input, to be identifiable; and a playback means for receiving the voice information transmitted from other terminal devices to the server from the server and playing back the voice represented by the voice information, wherein the playback means plays back the voice represented by the voice information transmitted from other terminal devices to the server, the voice position of which is located within a surrounding area set based on location information acquired by the acquisition means.
[0041] Voices emitted by a user in a particular location are often useful to other users present in (or heading towards) that location. In particular, the location of a user in a vehicle can change significantly over time (due to the vehicle's movement) compared to a user not in a vehicle. Therefore, for a user in a vehicle, voices emitted by other users within the surrounding area relative to their current location can be useful information.
[0042] In this regard, according to the configuration described above, a user of a terminal device A can hear voices emitted by a user of another terminal device B (another vehicle) that are emitted within a surrounding area set based on the current location of terminal device A. Therefore, such a terminal device can provide useful information to users riding in a vehicle.
[0043] Furthermore, in the aforementioned terminal device, the surrounding area may be an area that is biased toward the direction of travel with respect to the location information acquired by the acquisition means. With this configuration, it is possible to reduce the amount of audio emitted at places already passed from the playback target and increase the amount of audio emitted at places to be reached. Therefore, such a terminal device can increase the usefulness of the information.
[0044] Furthermore, in the aforementioned terminal device, the surrounding area may be an area along the planned route. With this configuration, audio emitted outside the area along the planned route can be detected. It can be excluded from playback. Therefore, such a terminal device can increase the usefulness of the information.
[0045] Furthermore, the terminal device of the second reference invention comprises: detection means for detecting that a specific event relating to a vehicle has occurred; transmission means for transmitting audio information representing the audio corresponding to the specific event to another terminal device or a server shared by multiple terminal devices when the detection means has detected that the specific event has occurred; and playback means for receiving the audio information from another terminal device or the server and playing the audio represented by the audio information.
[0046] In this configuration, a user of terminal device A can hear audio corresponding to specific events related to the vehicle, transmitted from another terminal device B (another vehicle). Therefore, the user of the terminal device can understand the status of the other vehicle while driving.
[0047] Furthermore, in the aforementioned terminal device, the detection means may detect that a specific driving operation has been performed by a vehicle user as a specific event, and the transmission means may transmit the voice information, which represents a voice message notifying that the specific driving operation has been performed, when the detection means has detected the specific driving operation. With this configuration, the user of the terminal device can be aware that a specific driving operation has been performed in another vehicle while driving.
[0048] Furthermore, in the aforementioned terminal device, the detection means may detect a sudden braking operation, and the transmission means may transmit the voice information, which represents a voice message notifying that a sudden braking operation has been performed, when the detection means detects a sudden braking operation. With this configuration, the user of the terminal device can be aware that a sudden braking operation has been performed in another vehicle while driving. Therefore, safer driving can be achieved compared to checking the status of other vehicles by visual inspection alone. [Brief explanation of the drawing]
[0049] [Figure 1] This is a block diagram illustrating the schematic configuration of a communication system. [Figure 2] This is a block diagram showing the schematic configuration of the equipment installed in the vehicle. [Figure 3A-3C] This is an explanatory diagram showing the configuration of the input section. [Figure 4] This is a block diagram showing the general configuration of the server. [Figure 5A] This is a flowchart showing part of the voice transmission process. [Figure 5B] This is a flowchart showing the remaining part of the voice transmission process. [Figure 6] This is a flowchart of the classification process. [Figure 7] This is a flowchart showing the process of determining operation by voice. [Figure 8] This is a flowchart showing the state voice determination process. [Figure 9] This is a flowchart of the authentication process. [Figure 10A] This is a flowchart showing part of the audio distribution process. [Figure 10B] This is a flowchart showing another part of the audio distribution processing. [Figure 10C] This is a flowchart showing the remaining part of the audio distribution process. [Figure 11A] This is a flowchart showing part of the audio playback process. [Figure 11B]This is a flowchart showing the remaining part of the audio playback process. [Figure 12] This is a flowchart showing the speaker setup process. [Figure 13] This is a flowchart showing the mode transition process. [Figure 14] This is a flowchart showing the re-playback process. [Figure 15] This is a flowchart showing the data accumulation process. [Figure 16] This is a flowchart showing the request processing. [Figure 17A] This flowchart shows a portion of the response processing. [Figure 17B] This flowchart shows the remaining part of the response processing. [Figure 18] This diagram shows the surrounding area, which is set to be within a predetermined distance from the vehicle's position. [Figure 19] This diagram shows the surrounding area set within a predetermined distance from the reference position. [Figure 20] This is a diagram showing the surrounding area along the planned route. [Figure 21A] This flowchart shows part of the response processing for a variation. [Figure 21B] This flowchart shows the remaining part of the response processing for the modified form. [Figure 22A] This is a flowchart showing part of the audio playback process for torture. [Figure 22B] This flowchart shows another part of the audio playback processing for torture. [Figure 22C] This flowchart shows yet another part of the audio playback processing for torture. [Figure 22D] This flowchart shows the remaining part of the audio playback process for the torture. [Figure 23] This is a flowchart showing the process for detecting drowsiness. [Figure 24] This is a flowchart showing the process for displaying additional information. [Modes for carrying out the invention]
[0050] Embodiments of the present invention will be described below with reference to the drawings. [Structure of this embodiment] Figure 1 is a block diagram showing the schematic configuration of a communication system 1 to which the present invention is applied, and Figure 2 is a block diagram showing the schematic configuration of a device mounted on a vehicle. Figures 3A-3C are explanatory diagrams showing the configuration of the input unit 45. Figure 4 is a block diagram showing the schematic configuration of the server 10.
[0051] Communication system 1 is a system that has the function of enabling users of terminal devices 31 to 33 to communicate with each other by voice with simple operation. Here, "user" refers to the occupants of a vehicle (for example, the driver). As shown in Figure 1, communication system 1 comprises a plurality of terminal devices 31 to 33 that can communicate with each other, and a server 10 that can communicate with each of these terminal devices 31 to 33 via base stations 20 and 21. In other words, server 10 is shared by the plurality of terminal devices 31 to 33. In this embodiment, for the sake of explanation, three terminal devices 31 to 33 are shown, but more terminal devices may be used.
[0052] Terminal devices 31 to 33 are configured as in-vehicle devices mounted on vehicles such as passenger cars and trucks, and as shown in Figure 2, they are configured as well-known microcontrollers equipped with a CPU 41, ROM 42, RAM 43, etc. The CPU 41 performs various processes, such as voice transmission processing and voice playback processing, which will be described later, based on the programs stored in the ROM 42, etc.
[0053] Furthermore, terminal devices 31-33 are also equipped with a communication unit 44 and an input unit 45. The communication unit 44 has a function for communicating with base stations 20, 21 configured as wireless base stations for mobile phones (a communication unit for communicating with the server 10 and other terminal devices 31-33 via the mobile phone network), and a function for directly communicating with other terminal devices 31-33 located within its line of sight (terminal devices 31-33 mounted on other vehicles such as preceding vehicles, following vehicles, and oncoming vehicles) (a communication unit for directly communicating with other terminal devices 31-33).
[0054] The input unit 45 is configured as buttons, switches, etc., for each user of terminal devices 31 to 33 to input commands. The input unit 45 also has a function for reading the user's biological characteristics.
[0055] Furthermore, as shown in Figure 3A, the input unit 45 is also configured as a plurality of touch panels 61, 62 on the user-facing side of the vehicle's steering wheel 60. These touch panels 61, 62 are positioned around the horn switch 63 and in a location where they are not touched when the steering wheel is operated during vehicle turns. These touch panels 61, 62 detect the movement of the user's hand (fingers) on the vehicle's steering wheel 60.
[0056] Furthermore, as shown in Figure 2, terminal devices 31 to 33 are connected via communication lines (e.g., in-vehicle LAN) to a navigation system 48, audio system 50, ABSECU 51, microphone 52, speaker 53, brake sensor 54, passing sensor 55, turn signal sensor 56, underbody sensor 57, crosswind sensor 58, collision / rollover sensor 59, and several other well-known sensors installed in the vehicle.
[0057] The navigation device 48, like other well-known navigation devices, is equipped with a current location detection unit for detecting the vehicle's current location and a display for displaying images. The navigation device 48 performs well-known navigation processing based on information such as the vehicle's position coordinates (current location) detected by the GPS receiver 49. When the navigation device 48 receives a request for its current location from terminal devices 31-33, it returns the latest current location information to terminal devices 31-33. Furthermore, when the navigation device 48 receives a command from terminal devices 31-33 to display text information, it displays the corresponding text information on the display.
[0058] The audio device 50 is a well-known sound reproduction device for playing music, etc. The music, etc., played by the audio device 50 is output from the speaker 53. ABSECU51 is an electronic control unit (ECU) that controls the anti-lock braking system (ABS). ABSECU51 suppresses brake lock by controlling the braking force (brake hydraulic pressure) so that the wheel slip ratio remains within a preset range (a slip ratio that allows for safe and rapid braking of the vehicle). In other words, ABSECU51 functions as a brake lock suppression device. Furthermore, ABSECU51 outputs a notification signal to terminal devices 31-33 when the ABS is activated (when braking force control begins).
[0059] Microphone 52 takes in the user's voice. Speaker 53 is configured as a 5.1ch surround sound system, for example, consisting of five speakers and a subwoofer. Each of these speakers is positioned to surround the user.
[0060] The brake sensor 54 detects an emergency braking operation when the speed at which the user depresses the brake pedal exceeds a certain threshold value. However, the determination of an emergency braking operation is not limited to this; for example, it may be determined based on the amount the brake pedal is depressed, or based on the vehicle's acceleration (deceleration).
[0061] The passing sensor 55 detects the user's passing operation (the operation of temporarily switching the headlights from low beam to high beam). The turn signal sensor 56 detects the user's turn signal operation (the operation of flashing either the left or right turn signal).
[0062] The underside sensor 57 detects when the underside of the vehicle is in contact with a bump or unevenness in the road surface, or when there is a high risk of such contact. In this embodiment, a detection member is provided on the underside of the front and rear bumpers (the lowest part of the underside of the vehicle excluding the wheel area) that easily deforms or displaces when it comes into contact with an external object (such as a bump or unevenness in the road surface) and returns to its original (pre-deformation or pre-displacement) state when released from contact. The underside sensor 57 detects when the underside of the vehicle is in contact with a bump or unevenness in the road surface, or when there is a high risk of such contact, by detecting the deformation or displacement of this detection member. Alternatively, a sensor that detects when the underside of the vehicle is in contact with a bump or unevenness in the road surface, or when there is a high risk of such contact, without contact with an external object may be used.
[0063] The crosswind sensor 58 detects strong crosswinds against the vehicle (for example, crosswinds where the wind pressure exceeds a certain threshold). Detect. The crosswind sensor 58 may also be used as a sensor to detect wind pressure, or as a sensor to detect airflow.
[0064] The collision / tip sensor 59 detects conditions in which a vehicle is likely to have been involved in an accident, such as when the vehicle has collided with an external object (including other vehicles) or when the vehicle has overturned. The collision / tip sensor 59 can be configured as a sensor that detects when a strong impact (acceleration) occurs to the vehicle. The collision condition and the overturning condition may be detected independently. For example, whether or not the vehicle is overturned may be determined based on the degree of tilt of the vehicle (body). In addition, other conditions in which the vehicle is likely to have been involved in an accident (for example, when the airbags have deployed) may also be included.
[0065] As shown in Figure 4, Server 10 has a well-known server hardware configuration, including a CPU 11, ROM 12, RAM 13, database (storage device) 14, communication unit 15, etc. This Server 10 performs various processes, such as voice distribution processing, which will be described later.
[0066] Furthermore, the server 10's database 14 includes a name / nickname database that associates the names of users on terminal devices 31-33 with their nicknames, a communication partner database that describes the IDs of terminal devices 31-33 and their communication partners, a prohibited word database that registers prohibited words, and an advertising database that describes the IDs of advertisers' terminal devices and the scope of advertising. Users can use the services of this system by registering their name, nickname, communication partners, etc., with the server 10.
[0067] [Processing in this embodiment] The processes performed in such a communication system 1 will be explained using the diagrams 5A and 5B and subsequent figures. Figures 5A and 5B are flowcharts showing the voice transmission process executed by the CPU 41 of terminal devices 31 to 33.
[0068] The voice transmission process is initiated when the power to terminal devices 31-33 is turned on and is subsequently executed repeatedly. The voice transmission process also involves transmitting voice information, such as voices spoken by the user, to other devices (such as server 10 or other terminal devices 31-33). In the following explanation, the voice information to be transmitted to other devices is also referred to as "transmitted voice information."
[0069] In detail, as shown in Figures 5A and 5B, first, it is determined whether or not an input operation has been performed via the input unit 45 (S101). Here, there are normal input operations and emergency input operations, and these are registered as different input operations in the device's own memory, such as its ROM.
[0070] For normal input operations, as shown in Figure 3B, for example, a hand movement sliding radially around the handle 60 is registered on the touch panels 61 and 62. For emergency input operations, as shown in Figure 3C, for example, a hand movement sliding circumferentially around the handle 60 is registered on the touch panels 61 and 62. This process determines whether or not one of these operations was input via the touch panels 61 and 62.
[0071] It is preferable to set the input operation to operate in a direction different from the operating direction (pressing direction) of the horn switch 63, as this can suppress accidental activation of the horn.
[0072] If no input is received (S101:NO), the process proceeds to S110, which will be described later. If an input is received (S101:YES), recording of the user's voice (voice to be used as transmitted voice information) begins (S102).
[0073] Then, it is determined whether or not a termination operation has been performed via the input unit 45 (S103). Termination operations, like input operations, are pre-registered in memory such as ROM. This termination operation may be the same as an input operation or may be a different operation.
[0074] If no termination operation is performed (S103: NO), it is determined whether a pre-set reference time (e.g., 20 seconds) has elapsed since the input operation (S104). Here, the reference time is set to prohibit long-duration voice input, and when communicating with a specific communication partner, the same or longer reference time is set compared to when communicating with an unspecified communication partner. The reference time may also be set to an infinite time.
[0075] If the reference time has not elapsed (S104: NO), the process returns to S102. If the reference time has elapsed (S104: YES) or if the termination operation has been performed in S103 (S103: YES), the recording is terminated and it is determined whether the input operation is an emergency operation input (S105). If it is an emergency operation input (S105: YES), an emergency information flag indicating that it is emergency voice information is added (S106), and the process proceeds to S107.
[0076] If the input is not for an emergency (S105: NO), the process immediately proceeds to S107. Next, the recorded voice (human voice) is converted into synthesized voice (machine voice) (S107). Specifically, first, the recorded voice is converted into text (character data) using well-known speech recognition technology. Next, the text obtained from the conversion is converted into synthesized voice using well-known speech synthesis technology. In this way, the recorded voice is converted into synthesized voice.
[0077] Next, the system is configured to add operation information to the transmitted voice information (S108). Here, operation information refers to information indicating the vehicle's status and the user's driving operations (such as turn signal operation, flashing headlights operation, sudden braking operation, etc.; operating status of lights, wipers, TV, radio, etc.; driving status such as speed, acceleration, direction of travel; and control status such as detected values from various sensors and whether or not there are malfunctions). This information is acquired via the vehicle's communication lines (e.g., in-vehicle LAN).
[0078] Next, a classification process is performed (S109) in which the transmitted voice information is classified into one of several predetermined types (genres), and classification information representing the classified type is added to the transmitted voice information. Such classification is used, for example, when the server 10 or terminal devices 31-33 extract (search) necessary voice information from a large amount of voice information. The classification of transmitted voice information may be performed manually by the user or automatically. For example, selection items (several predetermined types) may be displayed on the display and the user may make a selection. Alternatively, for example, the content of the utterance may be analyzed and classified into a type according to its content. Alternatively, for example, the utterance may be classified into a type according to the circumstances at the time of utterance (vehicle condition, driving operation, etc.).
[0079] Figure 6 is a flowchart showing an example of the classification process when transmitting voice information is classified automatically. First, it is determined whether a specific driving operation was performed at the time of speaking (S201). In this example, the passing operation detected by the passing sensor 55 and the sudden braking operation detected by the brake sensor 54 are considered the specific driving operations. In this example, the voice information is classified into three types: "normal," "caution," and "warning."
[0080] If it is determined that a specific driving operation was not performed at the time of speaking (S201: NO), the transmitted voice information is classified as "normal," and classification information indicating "normal" is added to the transmitted voice information (S202).
[0081] On the other hand, if it is determined that a passing operation was performed at the time of speaking (S201:YES, S203:YES), the transmitted voice information is classified as "Caution" and classification information indicating "Caution" is added to the transmitted voice information (S204). Also, if it is determined that an emergency braking operation was performed at the time of speaking (S203:NO, S205:YES), the transmitted voice information is classified as "Warning" and classification information indicating "Warning" is added to the transmitted voice information (S206).
[0082] After this classification process, the process proceeds to S114 (Figure 5B), which will be described later. If it is determined in S205 that an emergency braking operation was not performed at the time of the statement (S205: NO), the process returns to S101 (Figure 5A).
[0083] Incidentally, in this example, the transmitted audio information is classified into one type, but this is not the only option. For example, it may be possible to assign two or more types to a single transmitted audio information, such as tag information.
[0084] Returning to Figure 5A, if it is determined in S101 that there was no input operation (S101: NO), it is determined whether a specific driving operation was performed (S110). The specific driving operation referred to here is an operation that acts as a trigger to transmit voice information that was not spoken by the user, and may be different from the specific driving operation determined in S201 above. If it is determined that a specific driving operation was performed (S110: YES), an operation voice determination process is executed to determine the transmission voice information corresponding to the driving operation (S111).
[0085] Figure 7 is a flowchart showing an example of the operation voice determination process. In this example, the turn signal operation detected by the turn signal sensor 56, the passing operation detected by the passing sensor 55, and the sudden braking operation detected by the brake sensor 54 correspond to the specific driving operations referred to in S110.
[0086] First, it is determined whether or not a turn signal operation was detected as a specific driving operation as described in S110 (S301). If it is determined that a turn signal operation was detected (S301: YES), an audio message notifying that the turn signal has been activated (for example, "You have activated the turn signal," "Moving left (right)," "Please be careful," etc.) is set as the transmitted audio information (S302).
[0087] On the other hand, if it is determined that no turn signal operation has been detected (S301: NO), it is determined whether or not a passing operation has been detected as a specific driving operation as referred to in S110 (S303). If it is determined that a passing operation has been detected (S303: YES), an audio message notifying that a passing operation has been performed (for example, "You have performed a passing operation," "Please be careful," etc.) is set in the transmitted audio information (S304).
[0088] Then, in both cases where a turn signal operation is detected (S301:YES) and a passing operation is detected (S303:YES), the transmitted voice information is classified as "Caution," and classification information indicating "Caution" is added to the transmitted voice information (S305). After that, the operation voice determination process is terminated, and the process proceeds to S114 (Figure 5B), which will be described later.
[0089] On the other hand, if it is determined that no flashing operation has been detected (S303: NO), it is determined whether or not an emergency braking operation has been detected as a specific driving operation as referred to in S110 (S306). If it is determined that an emergency braking operation has been detected (S306: YES), an audio message notifying that an emergency braking operation has been performed (for example, "An emergency braking operation has been performed") is set as the transmitted audio information (S307). Then, the transmitted audio information is set as "Warning". The system is classified and configured to add classification information indicating "warning" to the transmitted voice information (S308). After that, the operation voice determination process is terminated and the system proceeds to the process in S114 (Figure 5B) described later. If it is determined that an emergency braking operation has not been detected as a specific driving operation as described in S110 (S306: NO), the system returns to S101 (Figure 5A).
[0090] Returning to Figure 5A, if it is determined in S110 that no specific driving operation has been performed (S110: NO), then it is determined whether or not a specific state has been detected (S112). The specific state referred to here is a state that acts as a trigger for transmitting voice information that has not been spoken by the user. If it is determined that no specific state has been detected (S112: NO), the voice transmission process is terminated. On the other hand, if it is determined that a specific state has been detected (S112: YES), then a state voice determination process is executed to determine the transmission voice information according to the state (S113).
[0091] Figure 8 is a flowchart showing an example of the state voice determination process. In this example, the specific states referred to in S112 are: when the ABS is activated, when the underside of the vehicle is in contact with a bump or other uneven surface in the road or when there is a high risk of contact, when a strong crosswind is detected against the vehicle, and when the vehicle has collided or overturned.
[0092] First, based on the notification signal output from ABSECU51, it is determined whether or not the ABS has activated (S401). If it is determined that the ABS has activated (S401: YES), an audio message notifying that the ABS has activated (for example, "ABS has activated," "Please be careful of slipping," etc.) is set in the transmitted audio information (S402).
[0093] On the other hand, if it is determined that the ABS is not operating (S401: NO), the underside sensor 57 determines whether a condition in which the underside of the vehicle is in contact with a bump or other unevenness in the road surface, or a condition in which there is a high risk of contact, has been detected (S403). If it is determined that a condition in which contact has occurred or a condition in which there is a high risk of contact has been detected (S403: YES), an audio message notifying that there is a risk of scraping the underside of the vehicle (for example, "There is a risk of scraping the underside of the vehicle," "There is a bump") is set in the transmitted audio information (S404).
[0094] On the other hand, if it is determined that contact has not occurred or that there is a high risk of contact (S403: NO), the crosswind sensor 58 determines whether or not a strong crosswind has been detected on the vehicle (S405). If it is determined that a strong crosswind has been detected (S405: YES), an audio message notifying that a strong crosswind has been detected (for example, "Please be careful of crosswinds") is set in the transmitted audio information (S406).
[0095] Then, in any of the following cases, the system is configured to classify the transmitted voice information as "Caution" and add classification information indicating "Caution" to the transmitted voice information (S407). After that, the status voice determination process is terminated and the system proceeds to the process described in S114 (Figure 5B) below.
[0096] On the other hand, if it is determined that no strong crosswinds have been detected (S405: NO), the collision / rollover sensor 59 determines whether or not a collision or rollover has been detected in the vehicle (a state in which there is a high probability that the vehicle has been involved in an accident) (S408). If it is determined that a collision or rollover has been detected in the vehicle (S408: YES), an audio message notifying that the vehicle has collided or rolled over (for example, "The vehicle has collided (or rolled over)," "An accident has occurred," etc.) is set as the transmitted audio information (S409). Then, the transmitted audio information is classified as a "warning," and classification information indicating "warning" is added to the transmitted audio information (S410). After this, the state voice determination process ends and the process proceeds to S114 (Figure 5B), which will be described later. If it is determined that no collision or overturning state of the vehicle has been detected (S408: NO), the process returns to S101 (Figure 5A).
[0097] Returning to Figure 5B, in S114, location information is obtained by requesting it from the navigation device 48, and this location information is set to be added to the transmitted voice information. The location information is information that represents the current location of the navigation device 48 at the time of acquisition (in other words, the current location of the vehicle, the current locations of terminal devices 31-33, and the current location of the user) in absolute position (latitude, longitude, etc.). In particular, the location information acquired in S114 represents the location where the user made the sound (speech position). Note that the speech position only needs to be able to identify the location where the user made the sound, and some variation is acceptable due to factors such as the time difference between the timing of voice input and the timing of location information acquisition (changes in position due to vehicle movement).
[0098] Next, a voice packet is generated using transmission data that includes transmission voice information, emergency information, location information, operation content information, classification information, and the user's nickname, which is set in the authentication process described later (S115). Then, the generated voice packet is wirelessly transmitted to the communication partner's device (server 10 or other terminal devices 31-33) (S116), and the voice transmission process ends. In other words, the transmission voice information representing the voice spoken by the user is transmitted in a form that allows identification of emergency information, location information (speaking location), operation content information, classification information, and nickname (associated state). Note that all voice packets may be sent to server 10, and certain voice packets (for example, those containing transmission voice information with "caution" or "warning" added as classification information) may also be sent to other terminal devices 31-33 with which direct communication is possible.
[0099] Furthermore, terminal devices 31-33 perform the authentication process shown in Figure 9 when they are started up or when a specific operation is entered by the user. The authentication process is for identifying the user who is performing the voice input.
[0100] In the authentication process, as shown in Figure 9, first, the user's biometric characteristics (for example, fingerprints, retina, palm, voiceprint, etc.) are acquired via the input unit 45 of terminal devices 31-33 (S501). Then, the user is identified by referring to a database that associates usernames (IDs) and biometric characteristics, which are pre-registered in memory such as RAM, and the system is configured to insert the nickname set by the user into the voice (S502). Once this process is complete, the authentication process is terminated.
[0101] Next, the audio distribution process will be explained using the flowchart shown in Figures 10A-10C. The audio distribution process is a process in which the server 10 performs predetermined operations (processing) on audio packets transmitted from terminal devices 31-33, and then distributes the processed audio packets to the other terminal devices 31-33.
[0102] In detail, as shown in Figures 10A-10C, first, it is determined whether or not a voice packet transmitted from terminal devices 31-33 has been received (S701). If no voice packet has been received (S701: NO), the process proceeds to S725, which will be described later.
[0103] Furthermore, if a voice packet has been received (S701:YES), this voice packet is acquired (S702), and a nickname is added to the voice packet by referring to the name / nickname DB (S703). In this process, a new voice packet is generated by adding a nickname, which is identification information for identifying the user, to the voice packet entered by the user. At this time, the nickname is added before or after the voice entered by the user so that the voice entered by the user does not overlap with the nickname.
[0104] Next, the audio information contained in the new audio packet is converted into text (S704). This process utilizes a well-known process for converting audio into text that other users can read. Location information, operation content information, etc., are also converted into text data in the same manner. If the terminal devices 31-33 that sent the audio packet perform the audio-to-text conversion process, the conversion process in server 10 may be omitted (or simplified) by obtaining the text data obtained from the terminal devices 31-33.
[0105] Next, names (personal names) are extracted from the audio (or converted text data), and it is determined whether or not a name is present (S705). If the audio packet contains emergency information, the processing from S705 to S720 (described later) is omitted, and the process proceeds to S721 (described later).
[0106] If the name is not present in the audio (S705: NO), the process proceeds to S709, which will be described later. If the name is present in the audio (S705: YES), it is determined whether or not this name exists in the Name / Nickname DB (S706). If the name present in the audio exists in the Name / Nickname DB (S706: YES), this name is replaced with the corresponding nickname in the audio (or the name is replaced with the nickname in text data) (S707), and the process proceeds to S709.
[0107] Furthermore, if the name in the voice does not exist in the Name / Nickname DB (S706: NO), the name is replaced with the sound "beep" (or, in the case of text data, the name is replaced with other characters such as "***") (S708), and the process proceeds to S709.
[0108] Next, pre-set prohibited words are extracted from the audio (or converted text data) (S709). Prohibited words generally include words that are vulgar, obscene, or otherwise likely to offend the listener, and these words are registered in the prohibited word database. In this process, prohibited words are extracted by comparing the words in the audio (or converted text data) with the prohibited words in the prohibited word database.
[0109] For prohibited words, they are replaced with the sound "boo" (or, in the case of text data, names are replaced with other characters such as "xxx") (S710). Next, intonation checks and polite expression checks are performed (S711, S712).
[0110] These processes determine whether the user is emotional or speaking politely when inputting voice by detecting the amplitude (absolute value and rate of change) of the input level to microphone 52 and polite expressions contained in the voice from the voice information, thereby detecting whether the user is emotionally agitated or not speaking politely.
[0111] If the check confirms that the input voice expresses normal emotions and is spoken politely (S713:YES), the process immediately proceeds to S715, which will be described later. If the check confirms that the input voice expresses heightened emotions or is not spoken politely (S713:NO), the voice distribution prohibition flag is set (S714) to indicate that the distribution of this voice packet is prohibited, and the process proceeds to S715.
[0112] Next, the behavior of terminal devices 31-33 that transmitted the relevant voice packet is identified by tracking their location information, and it is determined whether the behavior of terminal devices 31-33 deviates from the main road onto a side road, or whether they stopped for a certain period of time (for example, about 5 minutes) or longer (S715).
[0113] If terminal devices 31-33 deviate from the main road onto a side road or stop (S715: YES), the rest flag is set as if they are resting (S716), and the process proceeds to S717. If the action does not apply (S715:NO), proceed immediately to processing S717.
[0114] Then, the system sets pre-configured terminal devices (for example, all terminal devices that are powered on, some terminal devices selected based on the settings, etc.) as communication partners (S717), and also notifies the voice input terminal devices 31-33 of the results of the slip of the tongue (results of whether or not names or prohibited words were used, intonation checks, polite expression checks, etc.) (S718).
[0115] Next, it is determined whether the terminal devices 31-33 (destination vehicles) that will deliver the advertisements have passed a specific location (S719). Here, the server 10 has multiple advertisements registered in the advertisement database, and in this advertisement database, an area for delivering advertisements is set for each advertisement. This process detects when terminal devices 31-33 have entered the area for delivering advertisements (CMs).
[0116] If the destination vehicle is passing through a specific location (S719: YES), an advertisement corresponding to that location is added to the voice packet sent to this terminal device 31-33 (S720). If the destination vehicle is not passing through a specific location (S719: NO), the process proceeds to S721.
[0117] Next, it is determined whether there are multiple audio packets to be delivered (S721). If there are multiple audio packets (S721: YES), the audio packets are sorted so that they do not overlap (S722), and the process proceeds to S723. During this process, the packets are sorted with priority given to the playback order of emergency information, and audio packets with the rest flag set are sorted to appear later in the sequence.
[0118] Furthermore, if there is only one voice packet (S721: NO), the process immediately proceeds to S723. Next, it is determined whether or not a voice packet is already being transmitted (S723). If a voice packet is currently being transmitted (S723:YES), the system configures the voice packet to be sent to be transmitted after the currently transmitted packet (S724), and then proceeds to process S726. If a voice packet is not currently being transmitted (S723:NO), the system immediately proceeds to process S726.
[0119] By the way, in the S725 process, it is determined whether or not there are any voice packets to send (S725). If there are voice packets (S725: YES), the voice packets to be sent are sent (S726), and the voice distribution process is terminated. If there are no voice packets (S726: NO), the voice distribution process is terminated.
[0120] Next, the audio playback process, in which the CPU 41 of terminal devices 31-33 plays back audio, will be explained using the flowcharts shown in Figures 11A and 11B. In the audio playback process, as shown in Figures 11A and 11B, it is first determined whether or not a packet has been received from other terminal devices 31-33 or other devices such as the server 10 (S801). If no packet has been received (S801: NO), the audio playback process is terminated.
[0121] Furthermore, if a packet has been received (S801:YES), it is determined whether or not this packet is a voice packet (S802). If it is not a voice packet (S802:NO), the process proceeds to S812, which will be described later.
[0122] Furthermore, if it is an audio packet (S802:YES), playback of the currently playing audio information is stopped and the audio information that was being played is erased (S803). Then, it is determined whether or not the distribution prohibition flag is set (S804). If the distribution prohibition flag is set (S804:YES), an audio message (from ROM, etc.) indicating that audio playback is prohibited is sent. It generates a voice message (pre-recorded in the device) and also generates a corresponding text message (S805).
[0123] Next, in order to make the sound easier to hear, the playback volume of music or other audio being played by the audio device 50 is turned off (S806). Alternatively, the playback volume may be reduced, or playback may be stopped altogether.
[0124] Next, a speaker setting process is performed to set the volume output when the audio is played from speaker 53, according to the positional relationship between the positions of terminal devices 31-33 corresponding to the source of the audio to be played (speech position) and the position of the terminal device 31-33 itself (current location) (S807). The audio message generated in the process of S805 is then played (playback begins) at the set volume (S808). At this time, text data including the text message is output and displayed on the display of the navigation device 48 (S812).
[0125] The character data displayed during this process includes character data representing the content of the audio information contained in the audio packet, character data representing the user's actions, and audio messages.
[0126] Furthermore, if the distribution prohibition flag is not set in the S804 process (S804:NO), the playback volume of music etc. being played by the audio device 50 is turned off (S809), similar to the S806 process described above. In addition, speaker setting processing is performed (S810), the audio represented by the audio information contained in the audio packet is played (S811), and the process proceeds to S812. Since the audio represented by the audio information is synthesized speech, the playback mode can be easily changed. For example, the playback speed can be adjusted, emotional components contained in the words can be suppressed, the tone can be adjusted, and the voice color can be adjusted, either based on user operation or automatically.
[0127] Note that the S811 process plays the audio information contained in the aligned audio packets in order, so it is assumed that there will be no duplication of audio packets. However, if there is a risk of duplication of audio packets, such as when receiving audio packets from multiple servers 10, the S722 process (the process of aligning audio packets) described above should be performed.
[0128] Next, it is determined whether or not a voice command has been input by the user (S813). In this embodiment, if a voice command to skip the currently playing audio (for example, the voice "cut") is input during audio playback (S813:YES), the currently playing audio information is skipped in predetermined units and the next audio information is played (S814). In other words, audio that you do not want to hear can be skipped and played back by voice operation. The predetermined unit here may be, for example, a comment unit, or a speaker unit. When playback of all received audio information has finished (S815:YES), the playback volume of the music or other audio being played by the audio device 50 is turned back on (S816) and the audio playback process is terminated.
[0129] Next, the details of the speaker setting process performed in S807 and S810 will be explained using Figure 12. Figure 12 is a flowchart of the speaker setting process within the audio playback process. In the speaker setting process, as shown in Figure 12, first, location information indicating the vehicle's position (including past location information) is obtained from the navigation device 48 (S901).
[0130] Next, the position information (speech position) contained in the audio packet to be played back is extracted (S902), and the positions (distance) of the other terminal devices 31-33 relative to the vehicle's movement vector are detected (S903). In this process, the direction of travel of the terminal devices 31-33 (vehicle) is considered. The system detects the orientation of the terminal device that is the source of the audio packet to be played back, and the distance to this terminal device, using the reference point.
[0131] Next, the speaker volume ratio is adjusted according to the direction in which the terminal device that is the source of the voice packet is located (S904). In this process, the volume ratio is set so that the volume of the speaker located in the direction in which the source terminal device is located is the highest, so that the voice can be heard from the direction in which the source terminal device is located.
[0132] Then, the maximum volume for the speaker with the highest volume ratio is set according to the distance to the terminal device that sent the voice packet (S905). This process also allows the volume for other speakers to be set based on their volume ratios. Once this process is complete, the speaker setting process is terminated.
[0133] Next, the mode transition process will be explained. Figure 13 is a flowchart showing the mode transition process executed by the CPU 41 of terminal devices 31 to 33. The mode transition process is a process for switching between a distribution mode, which communicates with an unspecified number of communication partners, and a dialogue mode, which communicates with a specific partner. It starts when the power to terminal devices 31 to 33 is turned on and is then performed repeatedly.
[0134] In detail, as shown in Figure 13, first, it is determined whether the current setting mode is interactive mode or distribution mode (S1001, S1004). If the setting mode is neither interactive mode nor distribution mode (S1001: NO, S1004: NO), the process returns to S1001.
[0135] Furthermore, if the setting mode is interactive mode (S1001:YES), it is determined whether or not a mode transition operation has occurred via the input unit 45 (S1002). If no mode transition operation has occurred (S1002:NO), the mode transition process is terminated.
[0136] If a mode transition operation is performed (S1002:YES), the system switches to distribution mode (S1003). At this time, the server 10 is notified to set the communication partners to an unspecified number of people, and upon receiving this notification, the server 10 changes the communication partner settings (description in the communication partner DB) for these terminal devices 31-33.
[0137] Furthermore, if the setting mode is distribution mode (S1001: NO, S1004: YES), it is determined whether or not a mode transition operation has occurred via the input unit 45 (S1005). If no mode transition operation has occurred (S1005: NO), the mode transition process is terminated.
[0138] If a mode transition operation is performed (S1005: YES), a dialogue request is sent to the party currently playing the voice packet (S1006). Specifically, when a dialogue request is sent to the server 10 specifying the voice packet (the terminal device that sent it), the server 10 sends a dialogue request to the corresponding terminal device. The server 10 then receives a response to the dialogue request from this terminal device and sends the response to the terminal devices 31-33 that initiated the dialogue request.
[0139] If terminal devices 31-33 respond that they accept the dialogue request (S1007: YES), they switch to dialogue mode with the terminal device that accepted the dialogue request as their communication partner (S1008). At this time, they notify the server 10 to set the communication partner to a specific terminal device, and the server 10 receives this notification and changes the setting of the communication partner.
[0140] Furthermore, if there is no response indicating acceptance of the dialogue request (S1007: NO), the mode transition process is terminated. Next, Figure 14 is a flowchart showing the replay process executed by the CPU 41 of terminal devices 31-33. The replay process is an interrupt process that is started when a replay operation is input to repeat playback during the playback of the audio represented by the audio information contained in the audio packet. This process is performed by interrupting the audio playback process.
[0141] In the replay process, as shown in Figure 14, the first step is to select a specified audio packet (S1011). In this process, the audio packet currently being played or an audio packet immediately after playback (within a reference time (e.g., 3 seconds) after the end of playback) is selected according to the time elapsed since the start of playback of the currently playing audio packet.
[0142] Next, the selected audio packet is played back from the beginning (S1012), and the replay process is terminated. Once this process is complete, the audio playback process is resumed. Incidentally, the server 10 may receive voice packets transmitted from terminal devices 31 to 33, store voice information, and when it receives a request from terminal devices 31 to 33, it may extract voice information corresponding to the requesting terminal device 31 to 33 from the stored voice information and transmit the extracted voice information.
[0143] Figure 15 is a flowchart showing the storage process executed by the CPU 11 of the server 10. The storage process is for storing (remembering) the voice packets transmitted from terminal devices 31-33 in the voice transmission process described above (Figures 5A and 5B). Note that the storage process may be performed instead of the voice distribution process described above (Figures 10A-10C), or it may be performed together with the voice distribution process (for example, in parallel).
[0144] In the storage process, it is first determined whether or not a voice packet transmitted from terminal devices 31-33 has been received (S1101). If no voice packet has been received (S1101: NO), the storage process is terminated.
[0145] If a voice packet is received (S1101:YES), the received voice packet is stored in the database 14 (S1102). Specifically, the voice information is stored in a searchable format based on the various information contained in the voice packet (voice information, emergency information, location information, operation details, classification information, user nickname, etc.). After that, the storage process is terminated. Voice information that meets predetermined deletion conditions (for example, old voice information, voice information of users who have canceled their service, etc.) may be deleted from the database 14.
[0146] Figure 16 is a flowchart showing the request processing performed by the CPU 41 of terminal devices 31-33. Request processing is performed at predetermined intervals (e.g., tens of seconds or several minutes). In request processing, it is first determined whether the navigation device 48 is performing route guidance (in other words, whether a guidance route to the destination has been set) (S1201). If it is determined that route guidance is being performed (S1201: YES), the system obtains location information indicating the vehicle's current position (current location information) and guidance route information (destination and route information to the destination) from the navigation device 48 as vehicle location information (S1202). The system then proceeds to S1206. The vehicle location information is used to set the surrounding area on the server 10, as described later.
[0147] On the other hand, if it is determined that route guidance processing is not being performed (S1201: NO), it is determined whether the vehicle is in motion (S1203). Specifically, it is determined that the vehicle is in motion if its speed exceeds a certain threshold value. This threshold value is set to be higher than the speed at which a typical right or left turn occurs (a lower speed compared to straight-ahead driving). Therefore, if it is determined that the vehicle is in motion, there is a high probability that the direction of travel will not change significantly. In other words, in S1203, the direction of travel is predicted based on the change in the vehicle's position information. Determine whether the driving conditions are favorable or not. However, this is not the only option; for example, a value of 0 or close to 0 may be set as the criterion value.
[0148] If it is determined that the vehicle is in motion (S1203: YES), the navigation device 48 is used to obtain location information indicating the current vehicle's position (current location information) and one or more past location information as vehicle location information (S1204). The system then proceeds to S1206. When obtaining multiple past location information, it is also possible to obtain location information with the same time interval, for example, location information from 5 seconds ago and location information from 10 seconds ago.
[0149] On the other hand, if it is determined that the vehicle is not in motion (S1203: NO), the system obtains location information (current location information) indicating the vehicle's position from the navigation device 48 as vehicle location information (S1205). The system then proceeds to S1206.
[0150] Next, the favorites information is retrieved (S1206). Favorites information is information that can be used to extract (search) audio information that the user needs from a large amount of audio information, and it is set by the user manually or automatically (for example, by analyzing preferences from the user's statements). For example, if a specific user (for example, a user specified in advance) is set as favorites information, audio information transmitted by that specific user will be extracted preferentially. Similarly, for example, if a specific keyword is set as favorites information, audio information containing that specific keyword will be extracted preferentially. Note that in this S1206 process, already set favorites information may be retrieved, or the user may be allowed to set it at this stage.
[0151] Next, the type (genre) of the requested audio information is obtained (S1207). The type referred to here is the type classified in the process described above in S109, etc. (in this embodiment, "normal", "caution", and "warning").
[0152] Next, an information request is sent to the server 10 (S1208) that includes vehicle location information, favorite information, type of voice information, and the ID of the terminal device (terminal devices 31-33) (information that the server 10 uses to identify the reply destination). After that, the request processing is terminated.
[0153] Figures 17A and 17B are flowcharts showing the response processing performed by the CPU 11 of the server 10. Response processing is the process of responding to information requests sent from terminal devices 31-33 in the request processing (Figure 16) described above. Note that response processing may be performed instead of the voice distribution processing (Figures 10A-10C) described above, or it may be performed together with (for example, in parallel with) the voice distribution processing.
[0154] In the response processing, it is first determined whether or not an information request sent from terminal devices 31-33 has been received (S1301). If no information request has been received (S1301: NO), the response processing is terminated.
[0155] If an information request is received (S1301:YES), the surrounding area is set according to the vehicle location information included in the received information request (S1302~S1306). In this embodiment, as described above (S1201~S1205,S1208), the vehicle location information is provided as any of the following: (P1) location information indicating the current vehicle's position, (P2) location information indicating the current vehicle's position and one or more past location pieces, or (P3) location information indicating the current vehicle's position and information on the guidance route.
[0156] If it is determined that the vehicle location information included in the received information request is (P1) location information indicating the position of the vehicle itself (S1302:YES), then, as illustrated in Figure 18, the vehicle's position (Vehicle's current location) Area A1 within a predetermined distance L1 from C is set as the surrounding area (S1303). Then proceed to S1307. In Figure 18, the vertical and horizontal lines indicate roads that the vehicle can travel on, and the "×" indicates the location where the voice information stored in server 10 was spoken.
[0157] On the other hand, if it is determined that the vehicle location information included in the received information request is (P2) location information indicating the current vehicle's position and one or more past location information (S1302: NO, S1304: YES), then, as illustrated in Figure 19, an area A2 within a predetermined distance L2 from a reference position Cb that is shifted to the right side in the direction of travel (to the right in Figure 19) from the current vehicle's position C is set as the surrounding area (S1305). After that, the process proceeds to S1307.
[0158] Specifically, the vehicle's direction of travel is estimated based on the positional relationship between the current location information C and the past location information Cp. The distance from the current location C to the reference location Cb may be greater than the area radius L2. In other words, an area A2 that does not include the current location C may be set as the surrounding area.
[0159] On the other hand, if it is determined that the vehicle location information included in the received information request is (P3) location information indicating the vehicle's position and information on the guidance route (S1304: NO), the area along the planned travel route is set as the surrounding area (S1306). Then, the process proceeds to S1307.
[0160] Specifically, as shown in Figure 20, for example, the area A3 along the route from the current location C to destination G (the area determined to be on the route) within the planned route from departure point S to destination G (thick line portion) is set as the surrounding area. Here, the area along the route may be set as, for example, a range within a predetermined distance L3 from the route. In this case, the value of L3 may be changed according to road attributes (e.g., road width and type). Alternatively, instead of the range to destination G, it may be a range within a predetermined distance L4 from the current location C. Also, although previously traveled routes (the route from departure point S to current location C) are excluded here, this is not limited to this, and previously traveled routes may also be included.
[0161] Next, from the audio information stored in database 14, audio information where the voice is uttered within the surrounding area is extracted (S1307). In the examples in Figures 18 to 20, audio information where the voice is uttered within the surrounding area (area enclosed by dashed lines) A1, A2, and A3 is extracted.
[0162] Next, from the audio information extracted in S1307, audio information corresponding to the type of audio information included in the received information request is extracted (S1308). Furthermore, from the audio information extracted in S1308, audio information corresponding to the favorite information included in the received information request is extracted (S1309).
[0163] Through the above processing (S1302~S1309), the voice information corresponding to the information request received from terminal devices 31~33 is extracted as candidate voice information to be sent to the terminal devices 31~33 that originated the information request.
[0164] Next, the content of each of the candidate audio pieces is analyzed (S1310). Specifically, the audio represented by the audio piece is converted into text (character data), and keywords contained in the converted text are extracted. Alternatively, the text (character data) representing the content of the audio piece may be stored in the database 14 along with the audio piece, and the stored information may be used.
[0165] Next, for each of the candidate audio pieces to be sent, check whether the content overlaps with other audio pieces. The system determines whether the content of audio information A and B overlaps (S1311). For example, if a predetermined percentage or more (e.g., more than half) of the keywords contained in audio information A and B match those keywords in audio information B, the system determines that the content of audio information A and B overlaps. Note that the system may determine content overlap not only based on the percentage of keywords, but also based on context, etc.
[0166] Then, if it is determined that none of the candidate voice information do not overlap in content with other voice information (S1311: NO), the process proceeds to S1313, where the candidate voice information is sent to the terminal devices 31-33 that sent the information request (S1313), and the response processing ends. Alternatively, the system may store previously sent voice information for each of the destination terminal devices 31-33 and refrain from sending voice information that has already been sent.
[0167] On the other hand, if it is determined that there is audio information among the candidate audio information that overlaps in content with other audio information (S1311: YES), the content overlap is resolved by removing all but one of the audio information that is to be kept as the candidate audio information from the candidate audio information (S1312). For example, among the multiple audio information that overlaps in content, the most recent audio information may be kept, or conversely, the oldest audio information may be kept. Alternatively, for example, the audio information with the most keywords may be kept. After that, the candidate audio information is sent to the terminal devices 31-33 that sent the information request (S1313), and the response processing is terminated.
[0168] The audio information transmitted from server 10 to terminal devices 31-33 is played back using the audio playback process described above (Figures 11A and 11B). If terminal devices 31-33 receive a new audio packet while playing back audio from an already received audio packet, they prioritize playing back the audio from the newly received packet. They also delete the audio from the older audio packets without playing it back. As a result, the audio information corresponding to the vehicle's location is prioritized for playback.
[0169] [Effects of this embodiment] As described above, the following effects can be obtained according to this embodiment. (1) Terminal devices 31-33 input voice spoken by the user in the vehicle (S102) and acquire location information (current location) (S114). Then, terminal devices 31-33 transmit voice information representing the input voice to the server 10, which is shared by multiple terminal devices 31-33, in a form that allows the location information (speaking location) acquired at the time of voice input to be identified (S116). In addition, terminal devices 31-33 receive voice information transmitted to the server 10 from other terminal devices 31-33 and play back the voice represented by the received voice information (S811). Specifically, terminal devices 31-33 play back the voice information represented by voice information transmitted to the server 10 from other terminal devices 31-33, whose speaking location is within the surrounding area set based on the current location (S1303, S1305, S1306).
[0170] Voices emitted by a user in a particular location are often useful to other users present in (or heading towards) that location. In particular, the location of a user in a vehicle can change significantly over time (due to the vehicle's movement) compared to a user not in a vehicle. Therefore, for a user in a vehicle, voices emitted by other users within the surrounding area relative to their current location can be useful information.
[0171] In this embodiment, a user of a terminal device 31 can hear voices emitted by users of other terminal devices 32, 33 (other vehicles) that are emitted within a surrounding area set based on the current location of terminal device 31. Therefore, according to this embodiment, useful information can be provided to users riding in vehicles.
[0172] Audio generated within the surrounding area may include not only real-time audio but also audio from the past (for example, several hours ago). Such audio can be more useful than real-time audio generated outside the surrounding area (for example, in a distant location). For example, if there is a lane closure at a specific point on a road, a user tweeting this information allows another user heading to that point to choose to avoid it. Also, if the view from a specific point on a road is spectacular, a user tweeting this information allows another user heading to that point to avoid missing the view.
[0173] (2) When terminal devices 31-33 transmit vehicle location information to server 10, including the vehicle's current location, past location information, or planned route information, the surrounding area is set to an area that is biased toward the direction of travel relative to the vehicle's current location (Figures 19 and 20). This reduces the amount of audio emitted from locations already passed and increases the amount of audio emitted from locations the vehicle is heading towards. As a result, the usefulness of the information can be increased.
[0174] (3) If the surrounding area is set to an area along the planned driving route, voices emitted from areas outside the planned driving route can be excluded, thereby further improving the usefulness of the information.
[0175] (4) Terminal devices 31 to 33 automatically acquire vehicle location information at predetermined intervals (S1202, S1204, S1205) and transmit it to the server 10 (S1208) to acquire voice information corresponding to the vehicle's current location. Therefore, even if the location changes as the vehicle moves, voice information corresponding to that location can be acquired. In particular, when terminal devices 31 to 33 receive new voice information from the server 10, they erase the previously received voice information (S803) and play the new voice information (S811). In other words, at predetermined intervals, the voice information to be played is updated to voice information corresponding to the latest vehicle location information. Therefore, the played voice can be made to follow the changes in location as the vehicle moves.
[0176] (5) Terminal devices 31 to 33 transmit vehicle location information, including the vehicle's current location, to the server 10 (S1208), causing the server 10 to extract voice information corresponding to the current location (S1302 to S1312), and then play back the voice represented by the voice information received from the server 10 (S811). Therefore, compared to a configuration in which terminal devices 31 to 33 extract voice information, the amount of voice information received from the server 10 can be reduced.
[0177] (6) Terminal devices 31-33 input the voice spoken by the user (S102) and convert the input voice into synthesized voice (S107). By converting to synthesized voice in this way before playback, it is possible to change the playback manner compared to playing the voice as is. For example, it is possible to easily adjust the playback speed, suppress emotional components contained in the words, adjust the tone, and adjust the voice color.
[0178] (7) When a command is input while audio is being played back, terminal devices 31-33 change the audio playback mode according to the command. Specifically, when a command to skip the audio being played is input (S813:YES), the audio being played is skipped in predetermined units (S814). Therefore, audio that you do not want to hear can be skipped in predetermined units, such as by statement (comment) or by speaker.
[0179] (8) The server 10 determines whether each of the candidate audio information to be transmitted has overlapping content with other audio information (S1311), and if it determines that there is overlap (S1311: YES), it deletes the audio information with overlapping content (S1312). For example, if an event occurs in a certain location When multiple users make similar comments about the same event, such as comments about an accident, the phenomenon of repeated playback of the same comments is likely to occur. In this regard, according to this embodiment, since the playback of overlapping audio is omitted, the phenomenon of repeated playback of the same comments can be made less likely.
[0180] (9) The server 10 extracts audio information corresponding to the favorite information from the candidate audio information to be transmitted (S1309) and transmits it to terminal devices 31-33 (S1313). In this way, it is possible to efficiently extract and preferentially play audio information that the user needs (that satisfies the priority conditions according to the user) from the large amount of audio information transmitted by many other users.
[0181] (10) Terminal devices 31 to 33 are configured to communicate with the audio equipment installed in the vehicle, and the volume of the audio equipment during audio playback is lowered compared to the volume before audio playback (S806). This makes the audio easier to hear.
[0182] (11) Terminal devices 31-33 detect that a specific event related to the vehicle has occurred in the vehicle (S110, S112). If terminal devices 31-33 detect that a specific event has occurred (S110: YES or S112: YES), they transmit audio information representing the audio corresponding to the specific event to other terminal devices 31-33 or the server 10 (S116). Terminal devices 31-33 also receive audio information from other terminal devices 31-33 or the server 10 and play the audio represented by the received audio information (S811).
[0183] According to this embodiment, a user of a terminal device 31 can hear audio corresponding to specific events related to the vehicle, transmitted from other terminal devices 32, 33 (other vehicles). Therefore, users of terminal devices 31 to 33 can understand the status of other vehicles while driving.
[0184] (12) Terminal devices 31-33 detect when a specific driving operation is performed by a vehicle user as a specific event (S301, S303, S306). When terminal devices 31-33 detect a specific driving operation (S301: YES, S303: YES, or S306: YES), they transmit voice information representing a voice message notifying that a specific driving operation has been performed (S302, S304, S307, S116). As a result, users of terminal devices 31-33 can be aware that a specific driving operation has been performed in another vehicle while driving.
[0185] (13) When terminal devices 31 to 33 detect an emergency braking operation (S306: YES), they transmit voice information representing a voice message notifying that an emergency braking operation has been performed (S307). As a result, users of terminal devices 31 to 33 can be aware that an emergency braking operation has been performed in another vehicle while driving. Therefore, safer driving can be achieved compared to checking the status of other vehicles by visual inspection alone.
[0186] (14) When terminal devices 31 to 33 detect a passing operation (S303: YES), they transmit voice information representing a voice message notifying that a passing operation has been performed (S304). As a result, users of terminal devices 31 to 33 can be aware that another vehicle has performed a passing operation while driving. Therefore, it is possible to understand signals from other vehicles more easily compared to checking the status of other vehicles by visual inspection alone.
[0187] (15) When terminal devices 31-33 detect a turn signal operation (S301:YES), they transmit voice information representing a voice message notifying that a turn signal operation has been performed (S30 2) Therefore, users of terminal devices 31-33 can be aware that turn signals have been activated in other vehicles while driving. Thus, it becomes easier to predict the movements of other vehicles compared to checking the status of other vehicles by visual inspection alone.
[0188] (16) Terminal devices 31-33 detect specific conditions relating to the vehicle as specific events (S401, S403, S405, S408). When terminal devices 31-33 detect a specific condition (S401: YES, S403: YES, S405: YES, or S408: YES), they transmit voice information representing a voice message notifying that a specific condition has been detected (S402, S404, S406, S409, S116). As a result, users of terminal devices 31-33 can be aware that a specific condition has been detected in another vehicle while driving.
[0189] (17) When terminal devices 31-33 detect the activation of ABS (S401: YES), they transmit voice information representing a voice message notifying that ABS has been activated (S402). As a result, users of terminal devices 31-33 can be aware that ABS has been activated in other vehicles while driving. Therefore, it is possible to predict situations where slipping is likely to occur more easily compared to checking the road surface conditions visually.
[0190] (18) When terminal devices 31-33 detect a condition in which the underside of the vehicle is in contact with a bump or other unevenness in the road surface, or a condition in which there is a high risk of contact (S403: YES), they transmit voice information representing a message notifying the driver that there is a risk of scraping the underside of the vehicle (S404). Therefore, compared to checking the road surface condition by visual inspection alone, it is possible to predict situations in which there is a high risk of scraping the underside of the vehicle.
[0191] (19) When terminal devices 31 to 33 detect that the vehicle is being affected by a crosswind (S405: YES), they transmit voice information representing a message notifying the vehicle that it is being affected by a crosswind (S406). This makes it easier to predict situations that are susceptible to the effects of crosswinds compared to simply checking the conditions outside the vehicle visually.
[0192] (20) When terminal devices 31 to 33 detect that a vehicle has fallen or collided (S408: YES), they transmit voice information representing a voice message notifying the user that a vehicle has fallen or collided (S409). As a result, users of terminal devices 31 to 33 can be aware that another vehicle has fallen or collided while driving. Therefore, safer driving can be achieved compared to checking the status of other vehicles by visual inspection alone.
[0193] (21) In addition, the following effects can also be obtained. As described above, the communication system 1 comprises a plurality of terminal devices 31-33 that can communicate with each other, and a server 10 that can communicate with each of the plurality of terminal devices 31-33. The CPU 41 of each terminal device 31-33 converts the voice input from the user into voice information and transmits the voice information to the device including the other terminal devices 31-33.
[0194] The CPU 41 then receives audio information transmitted by other terminal devices 31-33 and plays back the received audio information. The server 10 also receives audio information transmitted from multiple terminal devices 31-33, arranges the audio so that the corresponding audio does not overlap, and then distributes the audio information to the multiple terminal devices 31-33.
[0195] With this communication system 1, it is possible to input the content to be transmitted by voice and to play back the content transmitted from other terminal devices 31-33 by voice, so even a user who is driving can safely send and receive content such as what they want to tweet. In addition, with this communication system 1, when playing back audio, it is possible to play back the audio so that the audio does not overlap, so the sound It can make voices easier to hear.
[0196] Furthermore, in the above communication system 1, when the CPU 41 receives a specific command from the user, it will replay the audio information that has already been played. In particular, it may replay the audio information that was played immediately before. With such a communication system 1, audio information that could not be heard due to noise or other reasons can be replayed at the user's command.
[0197] Furthermore, in the above communication system 1, terminal devices 31-33 are mounted in the vehicle, and the CPU 41 of terminal devices 31-33 detects the user's hand movements on the vehicle's steering wheel using touch panels 61 and 62. When it detects an initiation action to start voice input via the touch panels 61 and 62, it starts the process of generating voice packets, and when it detects an termination action to end voice input via the touch panels 61 and 62, it terminates the process of generating voice packets. With such a communication system 1, voice input operations can be performed on the steering wheel. Therefore, users can input voice more safely while driving the vehicle.
[0198] Furthermore, in the above communication system 1, the CPU 41 detects the start and end identification actions input via the touch panels 61 and 62 as hand movements in a direction different from the direction of operation when honking the horn. With such a communication system 1, it is possible to suppress the honking of the horn during the start or end identification action.
[0199] Furthermore, in the above communication system 1, the server 10 generates new voice information by adding identification information to the voice information input by the user, and distributes the generated voice information. In such a communication system 1, since identification information to identify the user can be added by voice, the user can notify other users of the user who initiated the voice message without having to speak their own identification information.
[0200] Furthermore, in the above communication system 1, before playing the audio information, the server 10 extracts pre-set keywords contained in the audio information and replaces the keywords contained in the audio information with the sound of another word or a pre-set sound. With such a communication system 1, words that are undesirable to distribute (obscene words, words that represent people's names, barbaric words, etc.) can be registered as keywords and replaced with words or sounds that are acceptable for distribution before distribution.
[0201] Furthermore, in the above communication system 1, the CPU 41 of each terminal device 31-33 acquires its own location information, associates this location information with the voice information transmitted from its own terminal device 31-33, transmits it to the other terminal devices 31-33, and acquires the location information transmitted by the other terminal devices 31-33. Then, it controls the volume output when playing the voice according to the positional relationship between the position of the terminal device 31-33 corresponding to the voice to be played and the position of its own terminal device 31-33.
[0202] In particular, the CPU 41 controls the volume so that it decreases as the distance to other terminal devices 31-33 increases, and when outputting sound from multiple speakers, it controls the volume so that the volume from the direction in which the other terminal devices 31-33 are located relative to the position of the CPU 41 is increased.
[0203] According to this communication system 1, the volume is controlled according to the positional relationship between the terminal devices 31-33 corresponding to the audio being played and the position of the user's own terminal device 31-33. Therefore, the user listening to the audio can perceive the positional relationship between other terminal devices 31-33 and their own terminal device 31-33. It can be grasped consciously.
[0204] Furthermore, in the above communication system 1, the server 10 converts the audio information into text and outputs the converted text to a display device. With such a communication system 1, information that was missed during the audio can be confirmed in text.
[0205] Furthermore, in the above-mentioned communication system 1, the CPU 41 of each terminal device 31 to 33 transmits information indicating the vehicle's status and the user's driving operations (operating status of lights, wipers, TV, radio, etc., driving status such as speed and direction of travel, and control status such as detected values from vehicle sensors and whether or not there are malfunctions) along with the voice. In such a communication system 1, information indicating the user's driving operations can also be sent to other terminal devices 31 to 33.
[0206] Furthermore, the above-described communication system 1 plays advertisements via audio depending on the location of the terminal devices 31 to 33. The content of the advertisements is inserted by the server 10, and the server 10 transmits the advertisements to the terminal devices 31 to 33 that play the audio information. Such a communication system 1 can provide a business model that can generate advertising revenue.
[0207] Furthermore, the above-mentioned communication system 1 is configured to allow switching between a distribution mode for communicating with an unspecified number of people and a dialogue mode for making calls with specific communication partners. With such a communication system 1, it is also possible to enjoy making calls with specific people.
[0208] Furthermore, in the above communication system 1, the server 10 determines whether the user is expressing emotion or speaking politely when inputting voice by detecting the amplitude of the input level and polite expressions contained in the voice from the voice information, and prevents the distribution of voices from users who are emotionally agitated or not speaking politely. With such a communication system 1, it is possible to eliminate statements from users that could cause arguments.
[0209] Furthermore, in the above communication system 1, server 10 notifies users who speak ill of others, or who are emotionally agitated or speak in an impolite manner. In this case, the determination is made based on whether or not words registered as keywords for slander, etc., are included in the voice information. With such a communication system 1, it is possible to encourage users who have made inappropriate remarks to refrain from doing so again.
[0210] Note that S102 corresponds to an example of processing as an input means, S114, S1202, S1204, and S1205 correspond to an example of processing as an acquisition means, S116 corresponds to an example of processing as a transmission means, and S811 corresponds to an example of processing as a playback means.
[0211] [Other embodiments] The embodiments of the present invention are not limited in any way to the embodiments described above, and various forms can be taken as long as they fall within the technical scope of the present invention.
[0212] (1) In the above embodiment, the server 10 performs an extraction process to extract audio information corresponding to the terminal devices 31-33 that play audio from a plurality of stored audio information (S1302-S1312), but it is not limited to this. For example, the terminal devices 31-33 that play audio may perform part or all of such an extraction process. For example, the server 10 may send audio information to the terminal devices 31-33 in a form that may include extra audio information that may not be played back by the terminal devices 31-33, and the terminal devices 31-33 may extract the audio information to be played back from the received audio information (discarding unnecessary audio information). The transmission of audio information from the server 10 to the terminal devices 31-33 may be performed in response to a request from the terminal devices 31-33, or it may be performed regardless of the request. You can do that.
[0213] (2) In the above embodiment, the server 10 determines whether each of the candidate audio information to be transmitted has overlapping content with other audio information (S1311), and if it determines that there is overlap (S1311: YES), it resolves the overlapping content (S1312). This makes it less likely for the same statement to be played back repeatedly, but from another perspective, the more audio information with overlapping content there is, the more credible the content is considered to be. Therefore, the condition for playing back the audio with overlapping content may be set as having a certain number of audio information with overlapping content or more.
[0214] For example, the response processing shown in Figures 21A and 21B replaces the processing S1311 to S1312 in the response processing shown in Figures 17A and 17B with the processing S1401 to S1403. That is, in the response processing shown in Figures 21A and 21B, it is determined whether each of the candidate voice information for transmission contains a pre-set specific keyword, such as "accident" or "traffic restriction" (S1401).
[0215] Then, if it is determined that there is audio information containing a specific keyword among the candidate audio information to be transmitted (S1401:YES), it is determined whether there are fewer than a certain number (for example, three) of audio information with the same content (duplicate content) among the audio information containing the specific keyword (S1402). If it is determined that there are fewer than a certain number of audio information with the same content (S1402:YES), the audio information with that content is excluded from the candidate audio information to be transmitted (S1403). After that, the candidate audio information to be transmitted is sent to the terminal devices 31-33 that sent the information request (S1313), and the response processing is terminated.
[0216] On the other hand, if it is determined that there is no audio information containing a specific keyword among the candidate audio information to be transmitted (S1401: NO), or if it is determined that there are fewer than a certain number of audio information with the same content (S1402: NO), the process proceeds directly to S1313, and the candidate audio information to be transmitted is sent to the terminal devices 31-33 that sent the information request (S1313), and the response processing is terminated.
[0217] In other words, according to the response processing illustrated in Figures 21A and 21B, audio information containing specific keywords is not transmitted to terminal devices 31-33 until a certain number of identical audio pieces are accumulated. Once this number is reached, the audio is transmitted to terminal devices 31-33. This prevents audio with unreliable content, such as false information, from being played back by terminal devices 31-33.
[0218] In the examples in Figures 21A and 21B, the condition for transmission was that a certain number of audio pieces containing the same content exist, targeting audio information containing a specific keyword. However, the system is not limited to this. For example, the condition for transmission could be that a certain number of audio pieces containing the same content exist for all candidate audio pieces, regardless of whether they contain a specific keyword or not. Furthermore, when determining whether a certain number of audio pieces with the same content exist, even if multiple audio pieces originate from the same source, they may be counted as one (not as multiple). This prevents the playback of unreliable audio on other users' terminal devices 31-33 due to the same user uttering the same audio multiple times.
[0219] (3) In the above embodiment, terminal devices 31 to 33 that transmit voice information perform processing to convert the recorded voice into synthesized voice (S107), but the embodiment is not limited to this. For example, part or all of such conversion processing may be performed on the server 10, or on the terminal devices 31 to 33 that play back the voice.
[0220] (4) In the above embodiment, terminal devices 31 to 33 are configured to execute request processing (Figure 16) at predetermined intervals, but the system is not limited to this and may be executed at other predetermined timings. For example, it may be executed each time a predetermined distance is traveled, when a request is made by the user, when a specified time arrives, when a specified point is passed, when leaving the surrounding area, etc.
[0221] (5) In the above embodiment, audio information within the surrounding area is extracted and played back, but the embodiment is not limited to this. For example, an expanded area encompassing the surrounding area (larger than the surrounding area) may be set, audio information within the expanded area may be extracted, and then audio information within the surrounding area may be further extracted and played back from there. Specifically, in S1307 of the above embodiment, audio information where the voice is uttered within the expanded area is extracted. Then, terminal devices 31 to 33 store (cache) the audio information within the expanded area received from the server 10, and at the stage of playing back the audio represented by the audio information, audio information where the voice is uttered within the surrounding area is extracted and played back. In this way, even if the surrounding area based on the current location changes as the vehicle moves, if the changed surrounding area is within the expanded area at the time the audio information was received, the audio represented by the audio information corresponding to that surrounding area can be played back.
[0222] (6) In the above embodiment, the surrounding area is set based on the vehicle's current location (actual position), but it is not limited to this, and location information other than the current location may be acquired, and the surrounding area may be set based on the acquired location information. Specifically, location information set (input) by the user may be acquired. In this way, for example, by setting the surrounding area based on the destination, voice information transmitted near the destination can be checked in advance at a location far from the destination.
[0223] (7) In the above embodiments, terminal devices 31 to 33 configured as in-vehicle devices mounted on a vehicle have been given as examples, but the invention is not limited thereto. For example, a device that is portable by the user and can be used in a vehicle (such as a mobile phone) may also be used. Furthermore, the terminal device may or may not share at least a part of its configuration (for example, a configuration for detecting location, a configuration for inputting voice, a configuration for playing back voice, etc.) with a device mounted on a vehicle.
[0224] (8) In the above embodiment, an example was given in which the audio information transmitted from other terminal devices 31 to 33 to the server 10 is played back if the location of the voice is within the surrounding area set based on the current location. However, the embodiment is not limited to this, and audio information may be extracted and played back regardless of location information.
[0225] (9) The information that the terminal device receives from the user is not limited to the voice spoken by the user, but may also be, for example, text information entered by the user. In other words, the terminal device may receive text information from the user in place of (or along with) the voice spoken by the user. In this case, the terminal device may display the text information in a way that is visible to the user, or it may convert the text information into synthesized speech and play it back.
[0226] (10) The embodiments described above are merely examples of embodiments to which the present invention may be applied. The present invention can be realized in various forms such as systems, apparatus, methods, programs, recorded recording media, etc.
[0227] (11) Alternatively, the following may also be done: For example, in the above communication system 1, the content added as identification information (e.g., real name, first nickname, second nickname, etc.) may be changed depending on the communication method. For example, information to identify each other may be exchanged in advance, and the communication partner The identification information should be changed depending on whether or not the person is registered in advance.
[0228] Furthermore, in the above embodiment, the system is configured to replace keywords with other words in all cases. However, depending on conditions such as the communication partner or communication mode, the process of replacing keywords with other words may be omitted. This is because, unlike when distributing to the public, such processing is not necessary when communicating with a specific party.
[0229] Furthermore, meaningless words specific to spoken language, such as "um" or "uh," may be omitted when generating audio information. In this case, the playback time of the audio information can be shortened by the amount of the omitted words.
[0230] Furthermore, in the above embodiment, the CPU 41 of each terminal device 31 to 33 output information indicating the vehicle status and the user's driving operations in text, but it may also output it in voice. In addition, various information such as traffic congestion information may be generated based on information obtained from other terminal devices 31 to 33, and the generated information may be output.
[0231] Furthermore, in the above-described communication system 1, a configuration in which each terminal device 31 to 33 directly exchanges voice information with each other and a configuration in which each terminal device 31 to 33 exchanges voice information with each other via the server 10 may be switched depending on pre-set conditions (for example, the distance to the base station connected to the server 10, the communication status between the terminal devices 31 to 33, or settings by the user).
[0232] Furthermore, in the above embodiment, the system is configured to play advertisements by voice according to the location of the terminal devices 31-33, but the advertisements may also be played by voice at regular intervals or after a certain number of audio information playbacks. In addition, the content of the advertisement may be transmitted from the advertiser's terminal devices 31-33 to other terminal devices 31-33. In this case, the system may be configured to allow users to make calls (communicate) with the advertiser by inputting predetermined operations to the terminal devices 31-33, or to guide users of terminal devices 31-33 to the advertiser's store.
[0233] Furthermore, in the above embodiment, a mode may be set for communicating with a specific number of communication partners. For example, in a configuration for exchanging location information, terminal devices 31 to 33 within a certain distance may be used as communication partners, or specific users (such as preceding vehicles or oncoming vehicles (identified by license plate imaging using a camera or GPS) or users within a pre-set group may be used as communication partners. In addition, the communication partners may be changed by setting (the communication mode may be changed).
[0234] Furthermore, if terminal devices 31-33 are mounted on a vehicle, the communication partners may be terminal devices 31-33 that are moving on the same road in the same direction, or terminal devices 31-33 whose vehicle behavior matches (for example, terminal devices that deviate from a side road may be excluded from the communication partners).
[0235] Alternatively, the system could be configured to allow users to identify other users and pre-configure them as favorites or excluded, enabling them to select their preferred partners for communication. In this case, when sending voice messages, the system could send information identifying the recipient, or, if a server exists, register the recipient information on the server, thereby ensuring that voice messages are sent only to designated recipients.
[0236] Alternatively, multiple buttons or switches could be assigned to specific directions, allowing the user to specify the direction of their communication partner using these switches. In this case, only users in that direction could be designated as communication partners.
[0237] Furthermore, although the above-described communication system 1 is configured to exchange voice information, it may also be configured to acquire information consisting of text, convert this information into voice, and then play it back. Alternatively, when transmitting voice, it may be configured to convert the voice into text before transmission, and then restore it to voice before playback. This configuration can reduce the amount of data that needs to be transmitted through communication.
[0238] Furthermore, in a configuration where data is transmitted using text, the data may be translated into the language that uses the least amount of data among several pre-configured languages before being transmitted.
[0239] Furthermore, terminal devices 31 to 33 may perform some of the functions of the server 10 in the above embodiment. For example, if there are multiple audio pieces of information that have not been played in the CPU 41 of terminal devices 31 to 33, the CPU 41 may arrange the audio pieces so that the corresponding audio pieces do not overlap before playing them. The CPU 41 may also generate new audio pieces by adding identification information to identify the user in audio to the audio piece input by the user, and then transmit the generated audio piece.
[0240] Specifically, a portion of the audio distribution process may be performed in the audio transmission process or the audio playback process. For example, if a portion of the audio distribution process is to be performed in the audio playback process, as shown in Figure 22A-22D, the processes S703 to S722 (excluding S717) of the audio distribution process may be performed between S802 and S804 of the audio playback process.
[0241] Even with such a communication system, when playing back audio, it is possible to play it back without overlapping audio, making it easier to hear. Furthermore, even with such a communication system 1, it is possible to add identification information to identify the user via audio, so that the user can notify other users of the source of the audio without having to speak their own identification information.
[0242] Furthermore, although the above embodiment prohibits the overlapping playback of audio, it is permissible to intentionally overlap parts of the audio during playback. In this case, only the last few seconds of the first audio and the first few seconds of the second audio may overlap. In this case, a fade-out effect that gradually decreases the volume at the end of the first audio and a fade-in effect that gradually increases the volume at the beginning of the second audio may be combined.
[0243] Furthermore, the presence or absence of voice utterances by the user may be used to determine whether or not the driver is drowsy. To realize this configuration, for example, the drowsiness detection process shown in Figure 23 can be implemented. The drowsiness detection process is performed in terminal devices 31-33 in parallel with the voice transmission process. If the user does not speak for a certain period of time, or if the user speaks but is estimated to be at a low level of consciousness, the process issues an alarm or takes vehicle control such as stopping. Specifically, as shown in Figure 23, the time since the last time the user spoke is first counted (S1401).
[0244] Then, it is determined whether a predetermined first reference time (for example, about 10 minutes) has elapsed since the user's last speech (S1402). If the first reference time has not elapsed (S1402: NO), the drowsiness detection process is terminated.
[0245] Furthermore, if the first reference time has elapsed (S1402: YES), the user is questioned (S1403). In this process, questions are asked to prompt the user to speak in order to confirm their level of consciousness. For this reason, direct questions such as "Are you awake?" are used. You can ask a direct question, or you can use an indirect question such as, "What is the current time?"
[0246] Next, it is determined whether or not a response to such a question has been obtained (S1404). If a response to the question has been obtained (S1404: YES), the level of the response is determined (S1405).
[0247] Here, the response level refers to items such as voice volume and articulation clarity, which decrease in level (sound volume, accuracy of speech recognition) when the user becomes sleepy. This process compares one of these items with a pre-set threshold to make a judgment.
[0248] If the response level is sufficiently high (S1405:YES), the dozing detection process is terminated. If the response level is insufficient (S1405:NO), or if no response to the question was obtained in the S1404 process (S1404:NO), it is determined whether a second reference time (for example, about 10 minutes and 10 seconds) set to be greater than or equal to the first reference time mentioned above has elapsed since the user's last utterance (S1406).
[0249] If the second reference time has not elapsed (S1406: NO), the process returns to the process of S1403. If the second reference time has elapsed (S1406: YES), an alarm is issued via the speaker 53 or the like (S1407).
[0250] Subsequently, it is determined whether a response to the alarm (for example, a voice such as "Wow" or an operation via the input unit 45 for解除 the alarm) has been obtained (S1408). If a response to the alarm has been obtained (S1408: YES), the level of the response is determined in the same manner as the process of S1405 (S1409). When an operation via the input unit 45 is input, it is determined that the level of the response is sufficient.
[0251] If the level of the response is sufficiently high (S1409: YES), the drowsiness determination process ends. If the level of the response is insufficient (S1409: NO), and if no response to the inquiry is obtained in the process of S1408 (S1408: NO), it is determined whether a third reference time (for example, about 10 minutes and 30 seconds) set to a value equal to or greater than the aforementioned second reference time has elapsed from the user's last utterance (S1410).
[0252] If the third reference time has not elapsed (S1410: NO), the process returns to the process of S1407. If the third reference time has elapsed (S1410: YES), vehicle control such as stopping the vehicle is performed (S1411), and the drowsiness determination process ends.
[0253] By performing such processing, it is possible to suppress drowsiness of the user (driver). Also, when inserting an advertisement by voice (S720) or when displaying the voice in characters (S812), additional character information may be further displayed on a display of the navigation device 48 or the like.
[0254] To achieve this configuration, for example, the additional information display process shown in Figure 24 can be implemented. Note that the additional information display process is performed in parallel with the audio distribution process on server 10.
[0255] In the additional information display process, as shown in Figure 24, it is first determined whether or not an advertisement (CM) was inserted in the audio distribution process (S1501) and whether or not text was displayed (S1502). If an advertisement was inserted (S1501: YES) or text was displayed (S1502: YES), information related to the text to be displayed is obtained (S1503).
[0256] Here, information related to the text includes, for example, advertisements or local information corresponding to the string (keyword) to be displayed, or information explaining the information related to the string. This information may be obtained from within server 10 or from outside server 10 (such as from the internet).
[0257] When information related to the characters is obtained, this information is displayed on the screen as characters or images (S1504), and the additional information display process is terminated. Furthermore, if no advertisement is inserted during processing S1501 and S1502, and no text is displayed (S1501: NO and S1502: NO), the additional information display process will be terminated.
[0258] With this configuration, not only can audio be displayed as text, but related information can also be displayed. The display used to show text information is not limited to the display of the navigation device 48; any other display, such as a head-up display in the vehicle, may be used. [Explanation of symbols]
[0259] 1...Communication system, 10...Server, 20...Base station, 31-33...Terminal device, 41...CPU, 42...ROM, 43...RAM, 44...Communication unit, 45...Input unit, 48...Navigation device, 49...GPS receiver, 52...Microphone, 53...Speaker
Claims
1. A communication system comprising a terminal device and a server device capable of communicating with the terminal device, The aforementioned terminal device is When a user inputs voice, the system generates text data by converting the voice into text. The character data is transmitted to the server device. The device acquires its own location information, associates the location information with the character data transmitted from its terminal device, and transmits it. The server device receives the character data transmitted, The server device is The terminal device receives the location information and the character data transmitted from the terminal device. New character data is generated that includes a response corresponding to the received location information. A communication system that transmits the aforementioned new character data to the terminal device.
2. The communication system according to Claim 1, The aforementioned terminal device is The system generates text data from the voice input by the user, omitting meaningless words, and transmits the text data with meaningless words omitted to the server device. Communication system.
3. A communication system according to Claim 1, The aforementioned terminal device is mounted on the vehicle, The terminal device transmits the character data to which information indicating the status of the vehicle or the user's driving operations has been added. Communication system.
4. The communication system according to claim 3, Multiple terminal devices can communicate with each other. Communication system.
5. A vehicle equipped with the terminal device described in Claim 4.
6. A vehicle equipped with a terminal device capable of communicating with a specific communication partner, The aforementioned terminal device is A dialogue request is sent to the aforementioned specific communication partner, and a dialogue mode is established with the aforementioned specific communication partner. Converts the voice input by the user into an audio signal, By obtaining location information to determine your current location, A vehicle that, in the aforementioned dialogue mode, associates the converted voice signal with the acquired location information and identification information for identifying the user, transmits it to the specific communication partner, and receives a voice signal transmitted from the specific communication partner.
7. The vehicle according to claim 6, The system detects the user's hand movement on the vehicle's steering wheel as a starting action to initiate voice input, as it is a movement in a direction different from the direction used when operating the horn. A vehicle that, upon detecting the aforementioned start-specific action, converts the voice input by the user into an audio signal.
8. The vehicle according to claim 7, A vehicle that, upon receiving an audio signal from the aforementioned specific communication partner, plays the received audio signal with the playback volume of the vehicle's audio system reduced.
9. The vehicle according to claim 6, A vehicle that reproduces the aforementioned audio signal in a way that prevents overlapping audio.
10. The vehicle according to claim 6, The aforementioned audio signal is converted into text, A vehicle that outputs the converted characters to a display device.
11. The vehicle according to claim 6, A vehicle that transmits the aforementioned audio signal along with operation information indicating at least one of the vehicle's status and the user's driving operation to the specified communication partner.
12. The vehicle according to claim 6, A vehicle that stores voice signals transmitted from the aforementioned specific communication partner in a pre-prepared database in a searchable format.
13. The vehicle according to claim 6, A sensor that detects when the vehicle tilts and accelerates, Based on the detection results from the aforementioned sensors, it is determined whether or not a collision or overturning has occurred to the vehicle. A vehicle that notifies the aforementioned vehicle of a collision or overturning when the aforementioned vehicle is involved in a collision or overturning.
14. The vehicle according to claim 6, A vehicle that associates the converted audio signal with the acquired location information and identification information for identifying the user, and transmits it to another terminal device.
15. The vehicle according to claim 6, A vehicle that determines whether the driver is drowsy based on voice commands made by the user.