Terminal device and communication system
The communication system addresses the challenge of users with mobility challenges by enabling voice-based content transmission and reception through non-overlapping audio playback, enhancing usability and safety during vehicle operation.
Patent Information
- Application Number
- JP2025069689
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2012-03-28
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-23
- Estimated Expiration
- 2032-12-14
AI Technical Summary
Existing communication systems do not adequately cater to users who have difficulty operating or visually recognizing terminal devices, making it challenging for them to safely transmit and receive transmission contents, especially in vehicles where user interactions are restricted.
A communication system comprising terminal devices equipped with voice input conversion, transmission, and reproduction means, which arranges and reproduces multiple voice signals without overlap, and includes a server for distributing voice signals to ensure non-overlapping playback, allowing users to input and receive content via voice commands, and integrates with vehicle operations for safe use during driving.
Enables safe and efficient voice-based content transmission and reception for users with mobility challenges, enhancing usability and safety during vehicle operation by allowing hands-free interaction and non-overlapping audio playback.
Smart Images

Figure 2025108675000001_ABST
Abstract
Description
Cross - reference to related applications
[0001] This international application claims priority based on Japanese Patent Application No. 2011 - 273578 and Japanese Patent Application No. 2011 - 273579, which were filed with the Japan Patent Office on December 14, 2011, and Japanese Patent Application No. 2012 - 074627 and Japanese Patent Application No. 2012 - 074628, which were filed with the Japan Patent Office on March 28, 2012, and incorporates by reference the entire contents of Japanese Patent Application No. 2011 - 273578, No. 2011 - 273579, No. 2012 - 074627, and No. 2012 - 074628 into this international application.
Technical Field
[0002] The present invention relates to a communication system including a plurality of terminal devices and a terminal device.
Background Art
[0003] There is known a system in which a large number of users each input transmission contents such as what they want to tweet in characters using a terminal device such as a mobile phone, and a server distributes the transmission contents arranged in order to each terminal device (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] It is preferable that even a user who has difficulty in operating or visually recognizing a terminal device can safely transmit and receive transmission contents such as what they want to tweet.
Means for Solving the Problems
[0006] One aspect of the present invention is A communication system including a plurality of terminal devices capable of communicating with each other, each terminal device includes: voice input conversion means for converting voice input by a user into a voice signal; voice transmission means for transmitting the voice signal to a device including other terminal devices; voice reception means for receiving a voice signal transmitted by another terminal device; voice reproduction means for reproducing the received voice signal; and is provided with: when there are a plurality of voice signals for which reproduction has not been completed, the voice reproduction means arranges them so that the voices corresponding to the respective voice signals do not overlap and then reproduces them.
[0007] Further, another aspect of the present invention is: a communication system including a plurality of terminal devices and a server device capable of communicating with each of the plurality of terminal devices, each terminal device includes: voice input conversion means for converting voice input by a user into a voice signal; voice transmission means for transmitting the voice signal to the server device; voice reception means for receiving a voice signal transmitted by the server device; voice reproduction means for reproducing the received voice signal; and is provided with: the server device includes: distribution means for receiving voice signals transmitted from the plurality of terminal devices, arranging them so that the voices corresponding to the respective voice signals do not overlap, and then distributing the voice signals to the plurality of terminal devices.
[0008] According to such a communication system, since the content of a transmission can be input by voice and the content of a transmission from another communication device can be reproduced by voice, even a user who has difficulty operating or visually recognizing a terminal device can safely transmit and receive the content of a transmission such as what they want to whisper.
[0009] Also, in such communication systems, when playing back audio, it is possible to play it back so that the audio does not overlap, making it easier to hear the audio. Also, in any of the above communication systems, the audio playback means may play back the already played audio signal again when receiving a specific command from the user. In particular, the audio signal played back immediately before may be played back.
[0010] According to such a communication system, an audio signal that could not be heard due to noise or the like can be played back again by the user's command. Furthermore, in any of the above communication systems, the terminal device may be mounted on a vehicle. The terminal device includes an operation detection means for detecting the movement of the user's hand on the vehicle's steering wheel, and when detecting a start specific operation for starting the input of audio by the operation detection means, the operation control means for starting the operation of the audio input conversion means, and when detecting an end specific operation for ending the input of audio by the operation detection means, the operation control means for ending the operation of the audio input conversion means.
[0011] According to such a communication system, an operation for inputting audio can be performed on the steering wheel. Therefore, a user during vehicle driving can input audio more safely. Also, in any of the above communication systems, the start specific operation and the end specific operation may be hand movements in a direction different from the operation direction when operating the horn.
[0012] According to such a communication system, it is possible to suppress the horn from sounding during the start specific operation or the end specific operation. Next, the terminal device may be configured as a terminal device constituting any of the above communication systems.
[0013] According to such a terminal device, the same effects as any of the above communication systems can be enjoyed. In the above communication system (terminal device), the voice transmission means may generate a new voice signal by adding, in voice, identification information for identifying the user to the voice signal input by the user, and transmit the generated voice signal.
[0014] Further, in a communication system including a plurality of terminal devices and a server device capable of communicating with each of the plurality of terminal devices, the distribution means may generate a new voice signal by adding, in voice, identification information for identifying the user to the voice signal input by the user, and distribute the generated voice signal.
[0015] According to such communication systems, since the transmission content can be input in voice and the transmission content from other communication devices can be reproduced in voice, even a user who has difficulty operating or visually recognizing the terminal device can safely transmit and receive the transmission content such as what they want to whisper.
[0016] Also, in such communication systems, since identification information for identifying the user can be added in voice, the user can notify other users of the voice transmission source without speaking their own identification information. The user of the transmission source can be notified.
[0017] Furthermore, in any of the above communication systems, the identification information to be added may be the user's real name or a nickname other than the real name. Also, the content to be added as identification information (for example, real name, first nickname, second nickname, etc.) may be changed according to the communication mode. For example, information for identifying each other may be exchanged in advance, and the identification information to be added may be changed according to whether the communication partner is a pre-registered partner or not.
[0018] Also, in any of the above communication systems, before reproducing the voice signal, it may be provided with keyword extraction means for extracting a preset keyword included in the voice signal from the voice signal, and voice replacement means for replacing the keyword included in the voice signal with the voice of another word or a preset sound.
[0019] According to such a communication system, words that are not preferable for distribution (such as obscene words, words indicating personal names, or barbarous words, etc.) can be registered as keywords and distributed after being replaced with words or sounds that can be distributed.
[0020] Note that, for example, the voice replacement means may or may not be activated according to conditions such as the communication partner. Also, for words that have no specific meaning in spoken language, such as "ano-" or "etto", the voice signal with these words omitted may be generated. In this case, the playback time of the voice signal can be shortened by the omitted part.
[0021] In any of the above communication systems, each terminal device may include a position information transmission means for acquiring its own position information and transmitting it to other terminal devices in association with the voice signal transmitted from its own terminal device, a position information acquisition means for acquiring the position information transmitted by other terminal devices, and a volume control means for controlling the volume output when the voice is played back by the voice playback means according to the positional relationship between the position of the terminal device corresponding to the played-back voice and the position of its own terminal device.
[0022] According to such a communication system, since the volume is controlled according to the positional relationship between the position of the terminal device corresponding to the played-back voice and the position of its own terminal device, the user who listens to the voice can intuitively grasp the positional relationship between the other terminal device and its own terminal device.
[0023] Note that the volume control means may control the volume to decrease as the distance to the other terminal device increases, or when outputting voice from a plurality of speakers, control the volume so that the volume from the direction where the other terminal device is located as seen by the user increases according to the azimuth of the other terminal device relative to the position of its own terminal device. Also, such controls may be combined.
[0024] Furthermore, the communication system may further include a character conversion means for converting an audio signal into characters, and a character output means for outputting the converted characters to a display device. According to such a communication system, information missed by listening can be confirmed in characters.
[0025] In any of the above communication systems, information indicating the state of the vehicle and the driving operation of the user (operating states of lights, wipers, televisions, radios, etc., driving states such as traveling speed and traveling direction, control states such as detection values by vehicle sensors and presence or absence of failures) may be transmitted together with the audio.
[0026] At this time, these pieces of information may be output in audio or in characters. Also, for example various information such as traffic jam information may be generated based on information obtained from other terminal devices, and the generated information may be output.
[0027] In any of the above communication systems, a configuration in which terminal devices directly exchange audio signals with each other and a configuration in which terminal devices exchange audio signals via a server device may be switched according to preset conditions (for example, the distance to a base station connected to the server device, the communication state between terminal devices, or settings by the user, etc.).
[0028] Also, advertisements may be played in audio at regular intervals, at regular numbers of audio signal reproductions, or according to the position of the terminal device. The content of the advertisement may be transmitted to the terminal device that plays the audio signal from the terminal device of the advertiser, the server device, etc.
[0029] At this time, a user may be able to make a call (communicate) with the advertiser by inputting a predetermined operation to the terminal device, or the user of the terminal device may be guided to the advertiser's store.
[0030] Furthermore, in any of the above communication systems, the communication partner may be an unspecified large number, or the communication partner may be limited. For example, in the case of a configuration for exchanging location information, terminal devices within a certain distance may be used as communication partners, or specific users (such as a preceding vehicle or an oncoming vehicle (identified by license plate imaging using a camera or by GPS)), users within a preset group may be used as communication partners. Also, it may be possible to change the communication partner (change the communication mode) according to the settings.
[0031] Also, when the terminal device is installed in a vehicle, terminal devices moving in the same direction on the same road may be used as communication partners, or terminal devices with matching vehicle behaviors (such as removing them from the communication partners when deviating onto a side road) may be used as communication partners.
[0032] Also, it may be configured so that other users can be identified in advance for favorite registration, exclusion registration, etc., and the user can select a preferred partner for communication. In this case, when transmitting voice, communication partner information for identifying the communication partner may be transmitted, or when there is a server device, the communication partner information may be registered with the server device so that the voice signal can be transmitted only to a predetermined partner.
[0033] Also, directions may be associated with switches such as a plurality of buttons respectively. When the user designates the direction in which the communication partner is located using this switch, only the users in that direction may be set as communication partners.
[0034] Furthermore, when the user inputs voice, it may be determined by detecting the amplitude of the input level and the polite expressions included in the voice from the voice signal whether the emotion or the way of speaking is polite, etc., and the voice by users with high emotions or impolite ways of speaking may not be distributed.
[0035] Also, users who speak ill of others, or users with high emotions or impolite ways of speaking may be notified of slips of the tongue, etc. In this case, the determination may be made based on whether words registered as keywords for insults, etc. are included in the voice signal.
[0036] Furthermore, in any of the above communication systems, although the configuration is such that voice signals are exchanged, it is also possible to obtain information consisting of characters, convert this information into voice, and play it back. Also, when transmitting voice, it is also possible to convert the voice into characters, transmit it, and then restore it to voice and play it back. With this configuration, the amount of data to be transmitted by communication can be reduced.
[0037] Furthermore, in a configuration where data is transmitted by characters in this way, it is also possible to translate the data into the language with the smallest amount of data among a plurality of preset languages and then transmit the data.
[0038] By the way, in recent years, the number of users of services that manage information sent by users, such as Twitter (registered trademark) and Facebook (registered trademark), in a state where other users can view it has been increasing rapidly. In such services, terminal devices such as personal computers and mobile phones (including so-called smartphones) are used for information transmission and viewing (Japanese Patent Application Laid-Open No. 2011-215861). On the other hand, a user (for example, a driver of a vehicle) who is in a state of riding in a vehicle (for example, an automobile) has restricted actions compared to the state of not riding in the vehicle, and thus often has spare time. Therefore, using the above-described services during vehicle rides can effectively utilize time. However, it is difficult to say that conventional services have fully considered the usage patterns of users riding in vehicles.
[0039] Therefore, there is a need to provide a technology for constructing a service useful for users riding in vehicles.
[0040] Therefore, there is a need to provide a technology for constructing a service useful for users riding in vehicles. As a technology for this, the terminal device of the first reference invention includes an input means for inputting voice, an acquisition means for acquiring position information, and transmits voice information representing the voice input by the input means to a server shared by a plurality of the terminal devices in a form in which the voice position, which is the position information acquired by the acquisition means at the time of input of the voice, can be specified, and receives the voice information transmitted from the other terminal device to the server from the server, and a playback means for playing back the voice represented by the voice information. The playback means plays back the voice represented by the voice information among the voice information transmitted from the other terminal device to the server, the voice position of which is located within a peripheral area set based on the position information acquired by the acquisition means.
[0041] Voice uttered by a user at a certain location is often useful to other users present at (or heading towards) that location. In particular, the position of a user riding in a vehicle can change significantly with the passage of time (travel of the vehicle) compared to a user not riding in the vehicle. Therefore, for a user riding in a vehicle, the voice of other users uttered within a peripheral area based on the current location can be useful information.
[0042] In this regard, according to the above-described configuration, a user of a certain terminal device A can listen to the voice uttered within a peripheral area set based on the current location of the terminal device A among the voices uttered by a user of another terminal device B (another vehicle). Therefore, such a terminal device can provide useful information to a user riding in a vehicle.
[0043] Also, in the above-described terminal device, the peripheral area may be an area biased towards the traveling direction side with respect to the position information acquired by the acquisition means. According to this configuration, the voice uttered at a location already passed can be reduced from the playback target, and the voice uttered at a location to come can be increased. Therefore, such a terminal device can increase the usefulness of the information.
[0044] Also, in the terminal device described above, the peripheral area may be an area along the planned route. According to this configuration, voices emitted outside the area along the planned route can be excluded from the reproduction targets. Therefore, according to such a terminal device, the usefulness of information can be enhanced.
[0045] The terminal device of the second reference invention further includes: detection means for detecting that a specific event related to the vehicle has occurred; and transmission means for transmitting, when the detection means detects that the specific event has occurred, voice information representing the voice corresponding to the specific event to another one of the terminal devices or a server shared by a plurality of the terminal devices; and reproduction means for receiving the voice information from the other terminal device or the server and reproducing the voice represented by the voice information.
[0046] According to this configuration, a user of a certain terminal device A can listen to the voice corresponding to a specific event related to the vehicle transmitted from another terminal device B (another vehicle). Therefore, the user of the terminal device can grasp the state of another vehicle while driving.
[0047] Also, in the terminal device described above, the detection means may detect that a specific driving operation has been performed by the user of the vehicle as the specific event, and the transmission means may transmit the voice information representing the voice for notifying that the specific driving operation has been performed when the detection means detects the specific driving operation. According to this configuration, the user of the terminal device can grasp while driving that a specific driving operation has been performed in another vehicle.
[0048] Also, in the terminal device described above, the detection means detects a sudden braking operation, and the transmission means may transmit the voice information representing the voice content notifying that a sudden braking operation has been performed when the detection means detects a sudden braking operation. According to this configuration, the user of the terminal device can grasp while driving that a sudden braking operation has been performed in another vehicle. Therefore, safe driving can be realized as compared with the case of confirming the state of another vehicle only by visual observation.
Brief Description of Drawings
[0049]
Figure 1
Figure 2
Figures 3A - 3C
Figure 4
Figure 5A
Figure 5B
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10A
Figure 10B
Figure 10C
Figure 11A
Figure 11B
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17A
Figure 17B
Figure 18
Figure 19
Figure 20
Figure 21A
Figure 21B
Figure 22A
Figure 22B
Figure 22C
Figure 22D
Figure 23
Figure 24
Embodiments for Carrying Out the Invention
[0050] Embodiments of the present invention will be described below with reference to the drawings. [Configuration of this Embodiment] FIG. 1 is a block diagram showing a schematic configuration of a communication system 1 to which the present invention is applied, and FIG. 2 is a block diagram showing a schematic configuration of an apparatus mounted on a vehicle. FIGS. 3A-3C are explanatory diagrams showing the configuration of the input unit 45. FIG. 4 is a block diagram showing the schematic configuration of the server 10.
[0051] The communication system 1 is a system having a function that allows users of the terminal devices 31 to 33 to communicate with each other by voice with a simple operation. The user here is an occupant (for example, a driver) of a vehicle. As shown in FIG. 1, the communication system 1 includes a plurality of terminal devices 31 to 33 that can communicate with each other, and a server 10 that can communicate with each of these terminal devices 31 to 33 via base stations 20 and 21. That is, the server 10 is shared by the plurality of terminal devices 31 to 33. In this embodiment, for convenience of explanation, three terminal devices 31 to 33 are illustrated, but more terminal devices may be used.
[0052] The terminal devices 31 to 33 are configured as in-vehicle devices mounted on vehicles such as passenger cars and trucks, and as shown in FIG. 2, are configured as well-known microcomputers including a CPU 41, a ROM 42, a RAM 43, and the like. The CPU 41 performs various processes such as voice transmission processing and voice reproduction processing described later based on programs stored in the ROM 42 and the like.
[0053] In addition, the terminal devices 31 to 33 further include a communication unit 44 and an input unit 45. The communication unit 44 has a function of communicating with the base stations 20 and 21 configured as wireless base stations of mobile phones (a communication unit for communicating with the server 10 and other terminal devices 31 to 33 via a mobile phone network), and other terminal devices 31 to 33 located within its line of sight (terminal devices 31 to 33 mounted on other vehicles such as a preceding vehicle, a following vehicle, and an oncoming vehicle) and a function of directly communicating with them (a communication unit for directly communicating with other terminal devices 31 to 33).
[0054] The input unit 45 is configured as buttons, switches, etc. for each user of the terminal devices 31 to 33 to input commands. Further, the input unit 45 also has a function for reading the biological characteristics of the user.
[0055] Furthermore, as shown in FIG. 3A, the input unit 45 is also configured as a plurality of touch panels 61, 62 on the user side surface of the vehicle steering wheel 60. And these touch panels 61, 62 are arranged around the horn switch 63 and at positions where they are not touched during the steering operation when the vehicle turns. Such touch panels 61, 62 detect the movement of the user's hand (finger) on the vehicle steering wheel 60.
[0056] Also, as shown in FIG. 2, the terminal devices 31 to 33 are connected via a communication line (for example, in-vehicle LAN) to a navigation device 48, an audio device 50, an ABS ECU 51, a microphone 52, a speaker 53, a brake sensor 54, a passing sensor 55, a turn signal sensor 56, a ground clearance sensor 57, a crosswind sensor 58, a collision / rollover sensor 59, and a plurality of other well-known sensors provided in the vehicle.
[0057] The navigation device 48 includes a current location detection unit for detecting the current location of the vehicle and a display for displaying an image, similar to a well-known navigation device. The navigation device 48 executes well-known navigation processing based on information such as the position coordinates (current location) of the vehicle detected by the GPS receiver 49. Also, when the navigation device 48 is requested for the current location from the terminal devices 31 to 33, it returns the latest current location information to the terminal devices 31 to 33. Further, when the navigation device 48 receives a command to display character information from the terminal devices 31 to 33, it displays characters corresponding to the character information on the display.
[0058] The audio device 50 is a well-known audio playback device for playing music and the like. The music and the like played by the audio device 50 are output from the speaker 53. The ABSECU51 is an electronic control unit (ECU) that controls an anti-lock braking system (ABS). The ABSECU51 suppresses brake lock by controlling the braking force (brake hydraulic pressure) so that the slip ratio of the wheel is within a preset range (a slip ratio at which the vehicle can be braked safely and promptly). That is, the ABSECU51 functions as a brake lock suppression device. Also, when the ABS is activated (when the control of the braking force is started), the ABSECU51 outputs a notification signal to the terminal devices 31 to 33.
[0059] The microphone 52 inputs the voice uttered by the user. The speaker 53 is configured as a 5.1ch surround sound system including, for example, five speakers and a subwoofer. These speakers are arranged so as to surround the user.
[0060] The brake sensor 54 detects a hard braking operation when the depression speed of the brake pedal by the user is equal to or higher than a determination reference value. Note that the determination of the hard braking operation is not limited to this, and for example, it may be determined based on the depression amount of the brake pedal, or it may be determined based on the acceleration (deceleration) of the vehicle.
[0061] The passing sensor 55 detects a passing operation (an operation of temporarily changing the headlight from low beam to high beam) by the user. The turn signal sensor 56 detects a turn signal operation (an operation of flashing either the left or right turn signal) by the user.
[0062] The following sensor 57 detects a state where the step of the road surface or the like is in contact with the lower surface of the vehicle or a state where there is a high risk of contact. In the present embodiment, a detection member that is easily deformed or displaced by contact with an external object (such as a step on the road surface) and returns to its original (before deformation or displacement) state when released from the contact state is provided on the lower surfaces of the front and rear bumpers (the lowermost part of the lower surface of the vehicle excluding the wheel parts). Then, the lower surface sensor 57 detects a state where the step of the road surface or the like is in contact with the lower surface of the vehicle or a state where there is a high risk of contact by detecting the deformation or displacement of this detection member. Note that a sensor that detects a state where the step of the road surface or the like is in contact with the lower surface of the vehicle or a state where there is a high risk of contact without contact with an external object may be used.
[0063] The crosswind sensor 58 detects a strong crosswind (for example, a crosswind with a wind pressure equal to or higher than a determination reference value) acting on the vehicle. Note that the crosswind sensor 58 may be a sensor that detects wind pressure, or may be a sensor that detects wind volume.
[0064] The collision / fall sensor 59 detects a state where the vehicle has collided with an external object (including other vehicles), a state where the vehicle has fallen, or a state where there is a high possibility that the vehicle has been in an accident. The collision / fall sensor 59 can be configured as a sensor that detects that a strong impact (acceleration) has occurred on the vehicle. Note that the collided state and the fallen state may be detected independently. For example, it may be determined whether or not the vehicle is in a fallen state according to the degree of inclination of the vehicle (vehicle body). In addition, other states where there is a high possibility that the vehicle has been in an accident (for example, a state where the airbag has been deployed) may be included.
[0065] As shown in FIG. 4, the server 10 has a hardware configuration as a well-known server including a CPU 11, a ROM 12, a RAM 13, a database (storage device) 14, a communication unit 15, and the like. This server 10 performs various processes such as voice distribution processing described later.
[0066] In addition, as the database 14 of the server 10, there are provided a name / nickname DB that associates the names of the users of the terminal devices 31 to 33 with their nicknames, a communication partner DB that describes the IDs of the terminal devices 31 to 33 and their communication partners, a prohibited word DB in which prohibited words are registered, an advertisement DB that describes the IDs of the terminal devices of the advertisers and the range for advertising, and the like. By registering their names, nicknames, communication partners, etc. with the server 10, the users can use the services of this system.
[0067] [Processing of this Embodiment] Regarding the processing implemented in such a communication system 1, it will be described with reference to the drawings such as FIGS. 5A and 5B below. FIGS. 5A and 5B are flowcharts showing the voice transmission processing executed by the CPUs 41 of the terminal devices 31 to 33.
[0068] The voice transmission processing starts when the power of the terminal devices 31 to 33 is turned on and is then repeatedly executed. Also, the voice transmission processing is a process of transmitting voice information representing the voice emitted by the user to other devices (such as the server 10 and other terminal devices 31 to 33). In the following description, the voice information for transmission to other devices is also referred to as "transmission voice information".
[0069] Specifically, as shown in FIGS. 5A and 5B, first, it is determined whether there has been an input operation via the input unit 45 (S101). Here, as the input operation, there are a normal input operation and an emergency input operation, which are registered as different input operations in a memory such as the own ROM.
[0070] As the normal input operation, for example, as shown in FIG. 3B, a hand movement of sliding in the radial direction of the handle 60 on the touch panels 61 and 62 is registered, and as the emergency input operation, for example, as shown in FIG. 3C, a hand movement of sliding in the circumferential direction of the handle 60 on the touch panels 61 and 62 is registered. In this process, it is determined whether any of these operations has been input via the touch panels 61 and 62.
[0071] It is preferable that the input operation is set to operate in a direction different from the operating direction (pushing direction) of the horn switch 63 in this way, because malfunction of the horn can be suppressed.
[0072] If there is no input operation (S101: NO), the process proceeds to the process of S110 described later. Also, if there is an input operation (S101: YES), recording of the voice (voice for transmission voice information) uttered by the user is started (S102).
[0073] Then, it is determined whether or not there has been an end operation via the input unit 45 (S103). The end operation is registered in advance in a memory such as a ROM in the same way as the input operation. This end operation may be the same operation as the input operation or a different operation.
[0074] If there is no end operation (S103: NO), after the input operation, it is determined whether or not a preset reference time (for example, 20 seconds) has elapsed (S104). Here, the reference time is provided to prohibit long-time voice input. When communicating with a specific communication partner, a reference time that is the same as or longer than that when communicating with an unspecified communication partner is set. Also, the reference time may be set to an infinite time.
[0075] If the reference time has not elapsed (S104: NO), the process returns to the process of S102. Also, when the reference time has elapsed (S104: YES) and when an end operation has been performed in the process of S103 (S103: YES), recording is ended, and it is determined whether or not the input operation is an operation input for emergency use (S105). If it is an operation input for emergency use (S105: YES), an emergency information flag indicating that it is voice information for emergency use is added (S106), and the process proceeds to the process of S107.
[0076] If it is not an emergency operation input (S105: NO), the process immediately proceeds to S107. Subsequently, the recorded voice (human voice) is converted into synthesized voice (mechanical voice) (S107). Specifically, first, the recorded voice is converted into text (character data) using well-known speech recognition technology. Next, the text obtained by the conversion is converted into synthesized voice using well-known speech synthesis technology. In this way, the recorded voice is converted into synthesized voice.
[0077] Subsequently, it is set to add operation content information to the transmission voice information (S108). Here, the operation content information refers to information indicating the state of the vehicle and the user's driving operations (operation states such as turn signal operation, passing operation, emergency brake operation, operating states of lights, wipers, TV, radio, etc., driving states such as driving speed, acceleration, traveling direction, etc., control states such as detection values by various sensors and the presence or absence of failures, etc.). Such information is acquired via a communication line in the vehicle (for example, in-vehicle LAN).
[0078] Subsequently, a classification process is executed in which the transmission voice information is classified into any one of a plurality of predetermined types (genres), and classification information indicating the classified type is set to be added to the transmission voice information (S109). Such classification is used, for example, when extracting (searching) necessary voice information from a large number of voice information in the server 10 or the terminal devices 31 to 33. The classification of the transmission voice information may be performed manually by the user or automatically. For example, selection items (a plurality of predetermined types) may be displayed on the display and the user may be allowed to make a selection. Also, for example, the content of the speech may be analyzed and classified into a type corresponding to the content. Also, for example, it may be classified into a type corresponding to the situation at the time of speech (such as the state of the vehicle and driving operations).
[0079] FIG. 6 is a flowchart showing an example of classification processing when automatically classifying transmitted voice information. First, it is determined whether or not a specific driving operation was being performed during the speech (S201). In this example, the passing operation detected by the passing sensor 55 and the sudden braking operation detected by the brake sensor 54 correspond to the specific driving operations referred to here. Also, in this example, the voice information is classified into three types: "normal", "attention", and "warning".
[0080] If it is determined that a specific driving operation was not being performed during the speech (S201: NO), the transmitted voice information is classified as "normal", and it is set to add classification information indicating "normal" to the transmitted voice information (S202).
[0081] On the other hand, if it is determined that a passing operation was being performed during the speech (S201: YES, S203: YES), the transmitted voice information is classified as "attention", and it is set to add classification information indicating "attention" to the transmitted voice information (S204). Also, if it is determined that a sudden braking operation was being performed during the speech (S203: NO, S205: YES), the transmitted voice information is classified as "warning", and it is set to add classification information indicating "warning" to the transmitted voice information (S206).
[0082] After such classification processing, the process proceeds to S114 (FIG. 5B) described later. Note that if it is determined in S205 that a sudden braking operation was not being performed during the speech (S205: NO), the process returns to S101 (FIG. 5A).
[0083] Incidentally, in this example, the transmitted voice information is classified into one type, but it is not limited to this. For example, like tag information, two or more types may be assigned to one piece of transmitted voice information.
[0084] Returning to Fig. 5A, if it is determined in the aforementioned S101 that there is no input operation (S101: NO), it is determined whether a specific driving operation has been performed (S110). The specific driving operation referred to here is an operation as a trigger for transmitting non-uttered voice of the user as transmission voice information, and may be different from the specific driving operation determined in the aforementioned S201. If it is determined that a specific driving operation has been performed (S110: YES), an operation voice determination process for determining transmission voice information corresponding to the driving operation is executed (S111).
[0085] Fig. 7 is a flowchart showing an example of the operation voice determination process. In this example, the turn signal operation detected by the turn signal sensor 56, the passing operation detected by the passing sensor 55, and the sudden braking operation detected by the brake sensor 54 correspond to the specific driving operation referred to in S110.
[0086] First, it is determined whether a turn signal operation has been detected as the specific driving operation referred to in S110 (S301). If it is determined that a turn signal operation has been detected (S301: YES), a voice (for example, a voice such as "Performed a turn signal operation", "Moving to the left (right)", "Please be careful", etc.) notifying that the turn signal operation has been performed is set as the transmission voice information (S302).
[0087] On the other hand, if it is determined that no turn signal operation has been detected (S301: NO), it is determined whether a passing operation has been detected as the specific driving operation referred to in S110 (S303). If it is determined that a passing operation has been detected (S303: YES), a voice (for example, a voice such as "Performed a passing operation", "Please be careful", etc.) notifying that the passing operation has been performed is set as the transmission voice information (S304).
[0088] When a turn signal operation is detected (S301: YES) or a passing operation is detected (S303: YES), the transmitted voice information is classified as "Attention", and classification information indicating "Attention" is set to be added to the transmitted voice information (S305). Then, the operation voice determination process is terminated, and the process proceeds to S114 (Fig. 5B) described later.
[0089] On the other hand, when it is determined that no passing operation is detected (S303: NO), it is determined whether or not an emergency braking operation is detected as the specific driving operation referred to in S110 (S306). When it is determined that an emergency braking operation is detected (S306: YES), voice information for notifying that an emergency braking operation has been performed (for example, voice such as "Emergency braking operation has been performed") is set as the transmitted voice information (S307). Then, the transmitted voice information is classified as "Warning" and classification information indicating "Warning" is set to be added to the transmitted voice information (S308). Then, the operation voice determination process is terminated, and the process proceeds to S114 (Fig. 5B) described later. When it is determined that no emergency braking operation is detected as the specific driving operation referred to in S110 (S306: NO), the process returns to S101 (Fig. 5A).
[0090] Returning to Fig. 5A, when it is determined that no specific driving operation has been performed in S110 described above (S110: NO), it is determined whether or not a specific state has been detected (S112). The specific state here is a state as a trigger for transmitting voice that the user has not uttered as transmitted voice information. When it is determined that no specific state has been detected (S112: NO), the voice transmission process is terminated. On the other hand, when it is determined that a specific state has been detected (S112: YES), a state voice determination process for determining transmitted voice information according to the state is executed (S113).
[0091] FIG. 8 is a flowchart showing an example of state voice determination processing. In this example, the state in which the ABS is activated, the state in which the lower surface of the vehicle contacts or is highly likely to contact a step on the road surface, the state in which a strong crosswind is detected for the vehicle, and the state in which the vehicle has collided or overturned correspond to the specific state referred to in S112.
[0092] First, based on the notification signal output from the ABS ECU 51, it is determined whether the ABS has been activated (S401). If it is determined that the ABS has been activated (S401: YES), a voice indicating that the ABS has been activated (for example, voices such as "The ABS has been activated", "Be careful of slipping", etc.) is set as the transmission voice information (S402).
[0093] On the other hand, if it is determined that the ABS is not activated (S401: NO), it is determined by the lower surface sensor 57 whether a state in which the lower surface of the vehicle contacts or is highly likely to contact a step on the road surface has been detected (S403). And if it is determined that a contacting state or a state with a high risk of contact has been detected (S403: YES), a voice indicating that there is a risk of rubbing the lower surface of the vehicle (for example, voices such as "There is a risk of rubbing the lower surface of the vehicle", "There is a step", etc.) is set as the transmission voice information (S404).
[0094] On the other hand, if it is determined that a contacting state or a state with a high risk of contact has not been detected (S403: NO), it is determined by the crosswind sensor 58 whether a strong crosswind has been detected for the vehicle (S405). And if it is determined that a strong crosswind has been detected (S405: YES), a voice indicating that a strong crosswind has been detected (for example, a voice such as "Be careful of the crosswind") is set as the transmission voice information (S406).
[0095] When the ABS is activated (S401: YES), when it is detected that the vehicle bottom surface is in contact with a road step or the like or is in a state where contact is highly likely (S403: YES), and when a strong crosswind is detected (S405: YES), in any of these cases, the transmitted voice information is classified as "Attention", and classification information indicating "Attention" is set to be added to the transmitted voice information (S407). After that, the state voice determination process is terminated, and the process proceeds to S114 (Fig. 5B) described later.
[0096] On the other hand, when it is determined that no strong crosswind is detected (S405: NO), it is determined by the collision / fall sensor 59 whether the vehicle is in a state of collision or fall (a state where the vehicle is highly likely to have an accident) (S408). When it is determined that the vehicle is in a state of collision or fall (S408: YES), voice information for notifying that the vehicle has collided or fallen (for example, voices such as "The vehicle has collided (or fallen)", "An accident has occurred", etc.) is set as the transmitted voice information (S409). Then, the transmitted voice information is classified as "Warning", and classification information indicating "Warning" is set to be added to the transmitted voice information (S410). Then After that, the state voice determination process is terminated, and the process proceeds to S114 (Fig. 5B) described later. If it is determined that the vehicle is not in a state of collision or fall (S408: NO), the process returns to S101 (Fig. 5A).
[0097] Returning to Fig. 5B, in S114, position information is acquired by requesting the navigation device 48 for position information, and the position information is set to be added to the transmitted voice information. The position information is information representing the current location of the navigation device 48 at the time of acquisition (in other words, the current location of the vehicle, the current locations of the terminal devices 31 to 33, and the current location of the user) in absolute positions (latitude, longitude, etc.). In particular, the position information acquired in S114 represents the position where the user uttered the voice (utterance position). The utterance position only needs to be able to identify the position where the user uttered the voice, and may have some variations due to factors such as the time difference between the timing of voice input and the timing of position information acquisition (change in position due to vehicle travel).
[0098] Subsequently, a voice packet is generated as transmission data, which includes transmission voice information, emergency information, location information, operation content information, classification information, and information including the user's nickname set in the authentication process described later (S115). Then, the generated voice packet is wirelessly transmitted to the communication partner's device (server 10 or other terminal devices 31 to 33) (S116), and the voice transmission process ends. That is, the transmission voice information representing the voice uttered by the user is transmitted in a form (associated state) in which the emergency information, location information (voice position), operation content information, classification information, and nickname can be specified. Note that all voice packets may be transmitted to the server 10, and specific voice packets (for example, those including transmission voice information with "Attention" or "Warning" added as classification information) may also be transmitted to other directly communicable terminal devices 31 to 33.
[0099] Also, in the terminal devices 31 to 33, the authentication process shown in FIG. 9 is performed when starting up or when a specific operation by the user is input. The authentication process is a process for identifying the user who performs voice input.
[0100] In the authentication process, as shown in FIG. 9, first, the user's biometric characteristics (for example, any one of fingerprint, retina, palm, voiceprint, etc.) are acquired via the input unit 45 of the terminal devices 31 to 33 (S501). Then, the user is identified by referring to a database in which the user name (ID) registered in advance in a memory such as RAM is associated with the biometric characteristics, and it is set to insert the nickname set by the user into the voice (S502). When such a process ends, the authentication process ends.
[0101] Next, the voice distribution process will be described using the flowcharts shown in FIGS. 10A - 10C. The voice distribution process is a process in which the server 10 performs a predetermined operation (process) on the voice packet transmitted from the terminal devices 31 to 33, and distributes the voice packet after the operation to the other terminal devices 31 to 33.
[0102] Specifically, as shown in FIGS. 10A - 10C, first, it is determined whether voice packets transmitted from the terminal devices 31 to 33 have been received (S701). If no voice packets have been received (S701: NO), the process proceeds to the process of S725 described later.
[0103] Also, if voice packets have been received (S701: YES), these voice packets are acquired (S702), and a nickname is added to the voice packets by referring to the name / nickname DB (S703). In this process, a new voice packet is generated by adding, in voice, the nickname, which is identification information for identifying the user, to the voice packet input by the user. At this time, the nickname is added before or after the voice input by the user so that the voice input by the user and the nickname do not overlap.
[0104] Subsequently, the voice represented by the voice information included in the new voice packet is converted into characters (S704). In this process, a well-known process for converting voice into characters that can be read by other users is used. At this time, the position information, operation content information, etc. are also converted into character data in the same manner. Note that when the conversion process from voice to characters is performed by the terminal devices 31 to 33 that are the transmission sources of the voice packets, the conversion process in the server 10 may be omitted (or simplified) by acquiring the character data obtained by the conversion process from the terminal devices 31 to 33.
[0105] Next, the name (personal name) in the voice (or the converted character data) is extracted, and it is determined whether there is a name (S705). Note that if the voice packet contains emergency information, the processes from S705 to S720 described later are omitted, and the process of S721 described later is proceeded to.
[0106] If there is no name in the voice (S705: NO), the process proceeds to the process of S709 described later. Also, if there is a name in the voice (S705: YES), it is determined whether this name exists in the name / nickname DB (S706). If the name in the voice exists in the name / nickname DB (S706: YES), this name is replaced with the corresponding nickname in voice (in character data, the name is replaced with the nickname in characters) (S707), and the process proceeds to the process of S709.
[0107] Also, if the name in the voice does not exist in the name / nickname DB (S706: NO), the name is replaced with the sound "pee" (in character data, the name is replaced with another character such as "***") (S708), and the process proceeds to the process of S709.
[0108] Subsequently, the preset prohibited words in the voice (or the converted character data) are extracted (S709). Prohibited words are generally words such as vulgar and lewd words that may make the listener uncomfortable, and these words are registered in the prohibited word DB. In this process, the prohibited words are extracted by comparing the words in the voice (or the converted character data) with the prohibited words in the prohibited word DB.
[0109] Regarding the prohibited words, they are replaced with the sound "boo" (in character data, the name is replaced with another character such as "xxx") (S710). Subsequently, intonation check and polite expression check are performed (S711, S712).
[0110] In these processes, whether the user's emotion when inputting the voice, whether it is a polite way of speaking, etc. are determined by detecting the amplitude (absolute value or change rate) of the input level to the microphone 52 and the polite expressions included in the voice from the voice information, and it is detected whether it is a situation where the emotion is heightened or the way of speaking is not polite.
[0111] If, as a result of the check, the input voice is in a normal emotion and is in a polite way of speaking (S713: YES), the process immediately proceeds to S715 described later. Also, if, as a result of the check, for the input voice, the emotion is heightened or the way of speaking is not polite (S713: NO), a voice distribution prohibition flag indicating that the distribution of this voice packet is prohibited is set (S714), and the process proceeds to the process of S715.
[0112] Subsequently, the behavior of the terminal devices 31 to 33 that transmitted the corresponding voice packet is identified by tracking the location information, and it is determined whether the behavior of the terminal devices 31 to 33 deviates from the main road to a side road or stops for a certain period of time (for example, about 5 minutes) or more (S715).
[0113] If the behavior of the terminal devices 31 to 33 deviates from the main road to a side road or stops (S715: YES), a rest flag is set assuming that they are resting (S716), and the process proceeds to the process of S717. Also, if the behavior of the terminal devices 31 to 33 does not deviate from the main road to a side road or does not correspond to stopping (S715: NO), the process immediately proceeds to the process of S717.
[0114] Then, a preset terminal device (for example, all terminal devices with power on, a part of terminal devices extracted based on the setting, etc.) is set as the communication partner (S717), and the result of a gaffe (the result of the presence or absence of a name or prohibited word, the result of intonation check, polite expression check, etc.) is notified to the terminal devices 31 to 33 that are the source of the voice input (S718).
[0115] Subsequently, it is determined whether the positions of the terminal devices 31 to 33 (destination vehicles) that perform distribution have passed a specific position (S719). Here, in the server 10, a plurality of advertisements are registered in the advertisement DB, and in this advertisement DB, an area for distributing each advertisement is set. In this process, it is detected that the terminal devices 31 to 33 have entered the area for distributing the advertisement (CM).
[0116] If the destination vehicle has passed a specific location (S719: YES), an advertisement corresponding to the specific location is added to the voice packet to be transmitted to the terminal devices 31 to 33 (S720). If the destination vehicle has not passed the specific location (S719: NO), the process proceeds to S721.
[0117] Subsequently, it is determined whether there are a plurality of voice packets to be distributed (S721). If there are a plurality of voice packets (S721: YES), they are arranged so as not to overlap (S722), and the process proceeds to S723. In this process, they are preferentially arranged so that the playback order of the emergency information comes first, and the voice packets with the rest flag set are arranged so that their order comes later.
[0118] If there is one voice packet (S721: NO), the process immediately proceeds to S723. Subsequently, it is determined whether a voice packet is already being transmitted (S723). If a voice packet is being transmitted (S723: YES), it is set to transmit the voice packet to be transmitted after the packet being transmitted (S724), and the process proceeds to S726. If a voice packet is not being transmitted (S723: NO), the process immediately proceeds to S726.
[0119] By the way, in the process of S725, it is determined whether there is a voice packet to be transmitted (S725). If there is a voice packet (S725: YES), the voice packet to be transmitted is transmitted (S726), and the voice distribution process ends. If there is no voice packet (S726: NO), the voice distribution process ends.
[0120] Next, the voice playback process in which the CPUs 41 of the terminal devices 31 to 33 play back voice will be described using the flowcharts shown in FIGS. 11A and 11B. In the voice playback process, as shown in FIGS. 11A and 11B, first, it is determined whether a packet has been received from another device such as other terminal devices 31 to 33 or the server 10 (S801). If no packet has been received (S801: NO), the voice playback process ends.
[0121] Also, if a packet is being received (S801: YES), it is determined whether this packet is an audio packet (S802). If it is not an audio packet (S802: NO), the process proceeds to the process of S812 described later.
[0122] Also, if it is an audio packet (S802: YES), the reproduction of the currently playing audio information is stopped, and the audio information that was being played is erased (S803). Then, it is determined whether a distribution prohibition flag is set (S804). If the distribution prohibition flag is set (S804: YES), an audio message (pre-recorded in a memory such as a ROM) indicating that audio reproduction is prohibited is generated, and a character message corresponding to this audio message is generated (S805). Subsequently, in order to make the audio easier to hear, the reproduction volume of music or the like being reproduced by the audio device 50 is turned off (S806). Note that the reproduction volume may be decreased, or the reproduction itself may be stopped.
[0123] Subsequently, a speaker setting process is performed to set the volume output when the audio is reproduced from the speaker 53 according to the positional relationship between the positions (voice generation positions) of the terminal devices 31 to 33 corresponding to the distribution source of the reproduced audio and the positions (current locations) of its own terminal devices 31 to 33 (S807), and the audio message generated in the process of S805 is reproduced (reproduction is started) at the set volume (S808). At this time, character data including the character message is output and displayed on the display of the navigation device 48 (S812).
[0124] Note that the character data displayed in this process includes character data indicating the content of the audio represented by the audio information included in the audio packet, character data indicating the operation content by the user, and the audio message.
[0125]
[0126] Also, in the process of S804, if the distribution prohibition flag is not set (S804: NO), similar to S806 described above, the playback volume of music or the like being played by the audio device 50 is turned off (S809). Further, speaker setting processing is performed (S810), the voice represented by the voice information included in the voice packet is played (S811), and the process proceeds to the process of S812. Since the voice represented by the voice information is synthesized voice, the playback mode can be easily changed. For example, it can be changed based on user operations or automatically, such as adjusting the playback speed, suppressing the emotional components included in the words, adjusting the pitch, and adjusting the voice timbre.
[0127] Note that in the process of S811, since the voices represented by the voice information included in the sorted voice packets are played in order, it is assumed that the voice packets do not overlap. However, in the case where there is a possibility of voice packet overlap, such as when receiving voice packets from a plurality of servers 10, the above-described process of S722 (the process of sorting voice packets) may be performed.
[0128] Subsequently, it is determined whether a voice representing a command has been input by the user (S813). In the present embodiment, during the playback of the voice, when a voice representing a command to skip the currently playing voice (for example, the voice "cut") is input (S813: YES), the currently playing voice information is skipped in a predetermined unit, and the next voice information is played (S814). That is, voices that the user does not want to hear can be skipped by voice operations and played. The predetermined unit referred to here may be, for example, a speech (comment) unit or a speaker unit. When the playback of the voice has ended for all received voice information (S815: YES), the playback volume of music or the like being played by the audio device 50 is turned back on (S816), and the voice playback process ends.
[0129] Next, the details of the speaker setting process executed in S807 and S810 will be described with reference to FIG. 12. FIG. 12 is a flowchart showing the speaker setting process in the voice playback process. In the speaker setting process, as shown in FIG. 12, first, position information indicating the position of the host vehicle (including past position information) is acquired from the navigation device 48 (S901).
[0130] Subsequently, the position information (voice generation position) included in the voice packet to be played back is extracted (S902), and the positions (distances) of the other terminal devices 31 to 33 with respect to the movement vector of the host vehicle are detected (S903). In this process, the azimuth in which the terminal device that is the transmission source of the voice packet to be played back is located and the distance to this terminal device are detected when the traveling direction of the own terminal devices 31 to 33 (vehicles) is used as a reference. Based on this, the azimuth in which the terminal device that is the transmission source of the voice packet to be played back is located and the distance to this terminal device are detected.
[0131] Subsequently, the volume ratio of the speakers is adjusted according to the azimuth in which the terminal device that is the transmission source of the voice packet is located (S904). In this process, the volume ratio is set so that the volume of the speaker located in the azimuth in which this terminal device exists becomes the largest so that the voice can be heard from the azimuth in which the transmission source terminal device is located.
[0132] Then, the maximum value of the volume for the speaker with the largest volume ratio is set according to the distance to the terminal device that is the transmission source of the voice packet (S905). By this process, the volume for the other speakers can also be set based on the volume ratio. When such a process is completed, the speaker setting process is terminated.
[0133] Next, the mode transition process will be described. FIG. 13 is a flowchart showing the mode transition process executed by the CPUs 41 of the terminal devices 31 to 33. The mode transition process is a process for switching between a distribution mode in which an unspecified number of communication partners are used and an interactive mode in which a specific partner is used as a communication partner, and is started when the power of the terminal devices 31 to 33 is turned on and is then repeatedly executed.
[0134] Specifically, as shown in FIG. 13, first, it is determined whether the current setting mode is the interactive mode or the distribution mode (S1001, S1004). If the setting mode is neither the interactive mode nor the distribution mode (S1001: NO, S1004: NO), the process returns to the process of S1001.
[0135] Also, if the setting mode is the interactive mode (S1001: YES), it is determined whether there is a mode transition operation via the input unit 45 (S1002). If there is no mode transition operation (S1002: NO), the mode transition process ends.
[0136] If there is a mode transition operation (S1002: YES), it transitions to the distribution mode (S1003). At this time, it notifies the server 10 to set an unspecified number of communication partners, and the server 10 receives this notification and changes the setting of the communication partners for these terminal devices 31 to 33 (description in the communication partner DB).
[0137] Also, if the setting mode is the distribution mode (S1001: NO, S1004: YES), it is determined whether there is a mode transition operation via the input unit 45 (S1005). If there is no mode transition operation (S1005: NO), the mode transition process ends.
[0138] If there is a mode transition operation (S1005: YES), a dialogue request is sent to the partner during which the voice packet is being played (S1006). Specifically, when a dialogue request is sent to the server 10 by specifying the voice packet (the terminal device of the transmission source), the server 10 sends a dialogue request to the corresponding terminal device. Then, the server 10 receives an answer to the dialogue request from this terminal device and sends the answer to the terminal devices 31 to 33 that are the dialogue request sources.
[0139] In the terminal devices 31 to 33, if there is an answer to accept the dialogue request (S1007: YES), it transitions to the dialogue mode with the terminal device that accepted the dialogue request as the communication partner (S1008). At this time, the terminal device designates and notifies the server 10 to set the communication partner to a specific terminal device, and the server 10 receives this notification and changes the setting of the communication partner.
[0140] Also, if there is no response indicating acceptance of the dialogue request (S1007: NO), the mode transition process is terminated. Next, FIG. 14 is a flowchart showing the replay process executed by the CPUs 41 of the terminal devices 31 to 33. The replay process is an interrupt process that starts when a replay operation indicating repeated replay is input during the replay of the voice represented by the voice information included in the voice packet. This process is executed by interrupting the voice playback process.
[0141] In the replay process, as shown in FIG. 14, first, a specified voice packet is selected (S1011). In this process, the currently playing voice packet or the voice packet immediately after playback (within the reference time (for example, 3 seconds) after playback ends) is selected according to the time after the start of playback of the currently playing voice packet.
[0142] Subsequently, the selected voice packet is played back from the beginning (S1012), and the replay process is terminated. When this process ends, the voice playback process is resumed. By the way, the server 10 may receive the voice packets transmitted from the terminal devices 31 to 33, accumulate the voice information, extract the voice information corresponding to the terminal devices 31 to 33 of the request source from the accumulated voice information when receiving a request from the terminal devices 31 to 33, and transmit the extracted voice information.
[0143] FIG. 15 is a flowchart showing the accumulation process executed by the CPU 11 of the server 10. The accumulation process is a process for accumulating (storing) the voice packets transmitted from the terminal devices 31 to 33 in the above-described voice transmission process (FIGS. 5A and 5B). Note that the accumulation process may be performed instead of the above-described voice distribution process (FIGS. 10A - 10C), or may be performed together with (for example, in parallel with) the voice distribution process.
[0144] In the accumulation process, first, it is determined whether or not a voice packet has been received from the terminal devices 31 to 33 (S1101). If no voice packet has been received (S1101: NO), the accumulation process is terminated.
[0145] When a voice packet is received (S1101: YES), the received voice packet is stored in the database 14 (S1102). Specifically, it is stored in a form that allows the voice information to be retrieved based on various information (voice information, emergency information, location information, operation content information, classification information, user's nickname, etc.) included in the voice packet. Then, the storage process ends. Note that voice information that satisfies a predetermined deletion condition (for example, old voice information, voice information of a canceled user, etc.) may be deleted from the database 14.
[0146] FIG. 16 is a flowchart showing the request process executed by the CPUs 41 of the terminal devices 31 to 33. The request process is executed at predetermined intervals (for example, several tens of seconds or several minutes). In the request process, first, it is determined whether or not route guidance processing is being executed by the navigation device 48 (in other words, whether a guidance route to the destination is set) (S1201). If it is determined that route guidance processing is being executed (S1201: YES), position information indicating the position of the own vehicle (current position information) and information on the guidance route (information on the destination and the route to the destination) are acquired from the navigation device 48 as vehicle position information (S1202). Then, the process proceeds to S1206. Note that the vehicle position information is information for causing the server 10 to set a peripheral area, which will be described later.
[0147] On the other hand, if it is determined that route guidance processing is not being executed (S1201: NO), it is determined whether or not the vehicle is in motion (S1203). Specifically, when the traveling speed of the vehicle exceeds a determination reference value, it is determined that the vehicle is in motion. This determination reference value is set to be higher than the speed during a general right or left turn (a speed lower than that during straight-ahead driving). Therefore, if it is determined that the vehicle is in motion, it is highly likely that the traveling direction of the vehicle will not change significantly. In other words, in S1203, it is determined whether or not the traveling state is such that the traveling direction can be easily predicted based on the change in the vehicle position information. However, the present invention is not limited to this, and for example, a value of 0 or a value close to 0 may be set as the determination reference value.
[0148] When it is determined that the vehicle is in motion (S1203: YES), the navigation device 48 acquires, as vehicle position information, position information indicating the position of the own vehicle (current position information) and one or more pieces of past position information (S1204). Thereafter, the process proceeds to S1206. When acquiring a plurality of pieces of past position information, for example, position information with equal time intervals such as the position information 5 seconds ago and the position information 10 seconds ago may be acquired.
[0149] On the other hand, when it is determined that the vehicle is not in motion (S1203: NO), the navigation device 48 acquires, as vehicle position information, position information indicating the position of the own vehicle (current position information) (S1205). Thereafter, the process proceeds to S1206.
[0150] Subsequently, favorite information is acquired (S1206). The favorite information is information that can be used to extract (search for) voice information required by the user from a large number of voice information, and is set manually or automatically by the user (for example, by analyzing the preferences from the user's speech, etc.). For example, when a specific user (for example, a pre-specified user) sets it as favorite information, the voice information transmitted by that specific user is preferentially extracted. Similarly, for example, when a specific keyword is set as favorite information, the voice information including that specific keyword is preferentially extracted. In the process of S1206, the already set favorite information may be acquired, or the user may be allowed to set it at this stage.
[0151] Subsequently, the type (genre) of the voice information to be requested is acquired (S1207). The type here refers to the types classified in the processes such as S109 described above (in this embodiment, "normal", "attention", and "warning").
[0152] Subsequently, an information request including vehicle position information, favorite information, the type of voice information, and the ID of itself (terminal devices 31 to 33) (information for the server 10 to identify the reply destination) is transmitted to the server 10 (S1208). Then, the request process ends.
[0153] FIGS. 17A and 17B are flowcharts showing response processing executed by the CPU 11 of the server 10. The response processing is processing for responding to the information request transmitted from the terminal devices 31 to 33 in the above-described request processing (FIG. 16). Note that the response processing may be performed instead of the above-described voice distribution processing (FIGS. 10A-10C), or may be performed together with (for example, in parallel with) the voice distribution processing.
[0154] In the response processing, first, it is determined whether or not an information request transmitted from the terminal devices 31 to 33 has been received (S1301). If the information request has not been received (S1301: NO), the response processing ends.
[0155] When the information request has been received (S1301: YES), a peripheral area corresponding to the vehicle position information included in the received information request is set (S1302 to S1306). In the present embodiment, as described above (S1201 to S1205, S1208), any one of (P1) position information indicating the position of the host vehicle, (P2) position information indicating the position of the host vehicle and one or more past position information, and (P3) position information indicating the position of the host vehicle and route guidance information is provided as the vehicle position information.
[0156] When it is determined that the vehicle position information included in the received information request is (P1) position information indicating the position of the host vehicle (S1302: YES), as illustrated in FIG. 18, the host vehicle position (The current location of the vehicle)C, an area A1 within a predetermined distance L1 is set as the peripheral area (S1303). Then, the process proceeds to S1307. In FIG. 18, the vertical and horizontal lines indicate roads on which the vehicle can travel, and "×" indicates the voice information pronunciation position stored in the server 10.
[0157] On the other hand, when it is determined that the vehicle position information included in the received information request is (P2) the position information indicating the position of the own vehicle and one or more pieces of past position information (S1302: NO, S1304: YES), as illustrated in FIG. 19, an area A2 within a predetermined distance L2 from a reference position Cb shifted toward the traveling direction side (the right side in FIG. 19) from the own vehicle position (the current location) C of the vehicle is set as the surrounding area (S1305). Then, the process proceeds to S1307.
[0158] Specifically, the traveling direction of the vehicle is estimated based on the positional relationship between the current position information C and the past position information Cp. Note that the distance from the current location C to the reference position Cb may be made larger than the radius L2 of the area. That is, an area A2 that does not include the current location C may be set as the surrounding area.
[0159] On the other hand, when it is determined that the vehicle position information included in the received information request is (P3) the position information indicating the position of the own vehicle and the information of the guidance route (S1304: NO), an area along the planned travel route is set as the surrounding area (S1306). Then, the process proceeds to S1307.
[0160] Specifically, for example, as shown in FIG. 20, among the planned travel route (thick line portion) from the departure point S to the destination G, an area A3 along the route from the current location C to the destination G (the area determined to be on the route) is set as the surrounding area. Here, the area along the route may be set, for example, as a range within a predetermined distance L3 from the route. In this case, the value of L3 may be changed according to the road attributes (for example, the width and type of the road). Also, instead of the range up to the destination G, it may be set as a range within a predetermined distance L4 from the current location C. Also, although the traveled route (the route from the departure point S to the current location C) is excluded here, it is not limited thereto, and the traveled route may also be included.
[0161] Subsequently, voice information whose voice position is within the peripheral area is extracted from the voice information stored in the database 14 (S1307). In the examples of FIGS. 18 to 20, voice information whose voice position is within the peripheral areas (regions surrounded by broken lines) A1, A2, and A3 is extracted.
[0162] Subsequently, voice information corresponding to the type of voice information included in the received information request is extracted from the voice information extracted in S1307 (S1308). Further, voice information corresponding to the favorite information included in the received information request is extracted from the voice information extracted in S1308 (S1309).
[0163] Through the above processing (S1302 to S1309), voice information corresponding to the information request received from the terminal devices 31 to 33 is extracted as voice information of transmission candidates to be transmitted to the terminal devices 31 to 33 which are the transmission sources of the information request.
[0164] Subsequently, for each of the voice information of transmission candidates, the content of the voice is analyzed (S1310). Specifically, the voice represented by the voice information is converted into text (character data), and keywords included in the converted text are extracted. Note that the text (character data) representing the content of the voice information may also be stored in the database 14 together with the voice information, and the stored information may be used.
[0165] Subsequently, for each of the voice information of transmission candidates, it is determined whether the content overlaps with other voice information (S1311). For example, when keywords included in a certain voice information A and keywords included in another voice information B match by a predetermined ratio or more (for example, more than half), it is determined that the contents of the voice information A and B overlap. Note that the determination of content overlap may be made based on context or the like, rather than simply based on the ratio of keywords.
[0166] Then, for all the voice information of the transmission candidates, if it is determined that the content does not overlap with other voice information (S1311: NO), the process proceeds to S1313, and the voice information of the transmission candidates is transmitted to the terminal devices 31 to 33 of the transmission source of the information request (S1313), and the response process ends. Note that for each of the destination terminal devices 31 to 33, the transmitted voice information may be stored, and the voice information that has already been transmitted may not be transmitted.
[0167] On the other hand, if it is determined that there is voice information in the transmission candidates whose content overlaps with other voice information (S1311: YES), among the multiple voice information with overlapping content, except for one voice information to be left as the transmission candidate voice information, the overlapping content is eliminated by excluding it from the transmission candidate voice information (S1312). For example, among the multiple voice information with overlapping content, the newest voice information may be left, or conversely, the oldest voice information may be left. Also, for example, the one with the most keywords included in the voice information may be left. After that, the voice information of the transmission candidates is transmitted to the terminal devices 31 to 33 of the transmission source of the information request (S1313), and the response process ends.
[0168] Note that the voice information transmitted from the server 10 to the terminal devices 31 to 33 is reproduced by the above-described voice reproduction process (Figs. 11A and 11B). When the terminal devices 31 to 33 receive a new voice packet during the reproduction of the voice represented by the already received voice packet, the voice represented by the newly received voice packet is preferentially reproduced. Also, the voice represented by the old voice packet is erased without being reproduced. As a result, the voice represented by the voice information corresponding to the position of the vehicle is preferentially reproduced.
[0169] [Effects of the Present Embodiment] As described above, according to the present embodiment, the following effects can be obtained. (1) In the vehicle, the terminal devices 31 to 33 receive the voice uttered by the user (S102) and acquire the position information (current location) (S114). Then, the terminal devices 31 to 33 transmit the voice information representing the input voice to the server 10 shared by the plurality of terminal devices 31 to 33 in a form in which the position information (utterance position) acquired at the time of input of the voice can be specified (S116). Further, the terminal devices 31 to 33 receive the voice information transmitted from other terminal devices 31 to 33 to the server 10 from the server 10 and reproduce the voice represented by the received voice information (S811). Specifically, the terminal devices 31 to 33 reproduce the voice represented by the voice information transmitted from other terminal devices 31 to 33 to the server 10, among which the utterance position is located within the peripheral area set based on the current location (S1303, S1305, S1306).
[0170] The voice uttered by the user at a certain location is often useful for other users present at that location (or heading towards that location). In particular, the position of a user riding in a vehicle can change significantly with the passage of time (travel of the vehicle) compared to a user not riding in the vehicle. Therefore, for a user riding in a vehicle, the voice of other users uttered within the peripheral area based on the current location can be useful information.
[0171] In this regard, in the present embodiment, the user of a certain terminal device 31 can listen to the voice uttered within the peripheral area set based on the current location of the terminal device 31 among the voices uttered by the users of other terminal devices 32, 33 (other vehicles). Therefore, according to the present embodiment, useful information can be provided to the user riding in the vehicle.
[0172] The sound generated within the peripheral area can include not only real-time sound but also past sound (e.g., sound emitted several hours ago). Such sound can be useful information compared to real-time sound emitted outside the peripheral area (e.g., a distant location). For example, when lane control or the like is being carried out at a specific point on a road, if a user mutters that information, another user heading towards that point can choose to avoid that point. Also, for example, when the scenery seen from a specific point on a road is wonderful, if a user mutters that information, another user heading towards that point can avoid missing that scenery.
[0173] (2) When vehicle position information including past position information or information on the planned driving route (planned progress route) in addition to the current position information of the vehicle is transmitted from the terminal devices 31 to 33 to the server 10, the peripheral area is set as an area biased towards the advancing direction with respect to the current location of the vehicle (FIGS. 19 and 20). Therefore, the sound emitted at the place already passed can be reduced from the reproduction target, and the sound emitted at the place to go can be increased. As a result, the usefulness of the information can be enhanced.
[0174] (3) When the peripheral area is set as an area along the planned driving route, the sound emitted outside the area along the planned driving route can be excluded, so that the usefulness of the information can be further improved.
[0175] (4) The terminal devices 31 to 33 automatically acquire vehicle position information at predetermined time intervals (S1202, S1204, S1205), and by transmitting it to the server 10 (S1208), acquire voice information corresponding to the current location of the vehicle. Therefore, even if the position changes as the vehicle travels, voice information corresponding to that position can be acquired. In particular, when the terminal devices 31 to 33 receive new voice information from the server 10, they delete the received voice information (S803) and play the new voice information (S811). That is, at predetermined time intervals, the voice information to be played is updated to the voice information corresponding to the latest vehicle position information. Therefore, the played voice can be made to follow the change in position as the vehicle travels.
[0176] (5) The terminal devices 31 to 33 transmit vehicle position information including the current location of the vehicle to the server 10 (S1208), causing the server 10 to extract voice information corresponding to the current location etc. (S1302 to S1312), and play the voice represented by the voice information received from the server 10 (S811). Therefore, compared with a configuration in which the terminal devices 31 to 33 extract voice information, the amount of voice information received from the server 10 can be reduced.
[0177] (6) The terminal devices 31 to 33 input the voice uttered by the user (S102) and convert the input voice into a synthesized voice (S107). By converting it into a synthesized voice and then playing it, it is easier to change the playback mode compared to the case of playing the voice as it is. For example, changes such as adjusting the playback speed, suppressing the emotional components contained in the words, adjusting the pitch, and adjusting the voice quality can be easily performed.
[0178] (7) When a voice representing a command is input during the playback of the voice, the terminal devices 31 to 33 change the playback mode of the voice according to the command. Specifically, when a voice representing a command to skip the currently playing voice is input (S813: YES), the currently playing voice is skipped and played in predetermined units (S814). Therefore, voices that are not wanted to be heard can be skipped and played in predetermined units such as, for example, in units of utterances (comments) or in units of speakers.
[0179] (8) The server 10 determines whether the content of each piece of voice information to be transmitted overlaps with that of other voice information (S1311). If it is determined that there is an overlap (S1311: YES), the voice information with overlapping content is deleted (S1312). For example, when multiple users make similar statements about an accident that occurred at a certain location, etc., the phenomenon of similar statements being repeatedly played is likely to occur. In this regard, according to the present embodiment, since the playback of overlapping voices is omitted, the phenomenon of similar statements being repeatedly played can be made less likely to occur. For example, when multiple users make similar statements about an accident that occurred at a certain location, etc., the phenomenon of similar statements being repeatedly played is likely to occur. In this regard, according to the present embodiment, since the playback of overlapping voices is omitted, the phenomenon of similar statements being repeatedly played can be made less likely to occur.
[0180] (9) The server 10 extracts the voice information corresponding to the favorite information from the voice information to be transmitted (S1309) and transmits it to the terminal devices 31 to 33 (S1313). Therefore, among a large amount of voice information transmitted by many other users, the voice information required by the user (meeting the priority conditions according to the user) can be efficiently extracted and preferentially played.
[0181] (10) The terminal devices 31 to 33 are configured to be communicable with the audio device mounted on the vehicle, and compare the volume of the audio device during voice playback with the volume before voice playback and lower it (S806). Therefore, the voice can be made easier to hear.
[0182] (11) The terminal devices 31 to 33 detect that a specific event related to the vehicle has occurred in the vehicle (S110, S112). Then, when the terminal devices 31 to 33 detect that a specific event has occurred (S110: YES or S112: YES), they transmit the voice information representing the voice corresponding to the specific event to other terminal devices 31 to 33 or the server 10 (S116). Also, the terminal devices 31 to 33 receive voice information from other terminal devices 31 to 33 or the server 10 and play the voice represented by the received voice information (S811).
[0183] According to this embodiment, a user of a certain terminal device 31 can listen to voices corresponding to specific events related to vehicles transmitted from other terminal devices 32 and 33 (other vehicles). Therefore, the users of terminal devices 31 to 33 can grasp the states of other vehicles while driving.
[0184] (12) The terminal devices 31 to 33 detect that a specific driving operation has been performed by a user of a vehicle as a specific event (S301, S303, S306). Then, when the terminal devices 31 to 33 detect a specific driving operation (S301: YES, S303: YES, or S306: YES), they transmit voice information representing a voice that notifies that the specific driving operation has been performed (S302, S304, S307, S116). For this reason, the users of terminal devices 31 to 33 can grasp while driving that a specific driving operation has been performed in other vehicles.
[0185] (13) When the terminal devices 31 to 33 detect an emergency braking operation (S306: YES), they transmit voice information representing a voice that notifies that the emergency braking operation has been performed (S307). For this reason, the users of terminal devices 31 to 33 can grasp while driving that an emergency braking operation has been performed in other vehicles. Therefore, compared with the case of confirming the states of other vehicles only by visual observation, safe driving can be realized.
[0186] (14) When the terminal devices 31 to 33 detect a passing operation (S303: YES), they transmit voice information representing a voice that notifies that the passing operation has been performed (S304). For this reason, the users of terminal devices 31 to 33 can grasp while driving that a passing operation has been performed in other vehicles. Therefore, compared with the case of confirming the states of other vehicles only by visual observation, it is easier to grasp a signal from other vehicles.
[0187] (15) When the terminal devices 31 to 33 detect a turn signal operation (S301: YES), they transmit voice information representing a voice that notifies that the turn signal operation has been performed (S30 2). Therefore, the users of the terminal devices 31 to 33 can grasp while driving that a turn signal operation has been performed in another vehicle. Therefore, compared with the case of checking the state of another vehicle only by visual inspection, it is easier to predict the movement of another vehicle.
[0188] (16) The terminal devices 31 to 33 detect a specific state related to the vehicle as a specific event (S401, S403, S405, S408). Then, when the terminal devices 31 to 33 detect a specific state (S401: YES, S403: YES, S405: YES or S408: YES), they transmit voice information representing a voice that notifies that the specific state has been detected (S402, S404, S406, S409, S116). Therefore, the users of the terminal devices 31 to 33 can grasp while driving that a specific state has been detected in another vehicle.
[0189] (17) When the terminal devices 31 to 33 detect the activation of the ABS (S401: YES), they transmit voice information representing a voice that notifies that the ABS has been activated (S402). Therefore, the users of the terminal devices 31 to 33 can grasp while driving that the ABS has been activated in another vehicle. Therefore, compared with the case of checking the road surface condition only by visual inspection, it is easier to predict a situation where slipping is likely.
[0190] (18) When the terminal devices 31 to 33 detect a state where the step of the road surface or the like is in contact with the lower surface of the vehicle or a state with a high risk of contact (S403: YES), they transmit voice information representing a voice that notifies that there is a risk of rubbing the lower surface of the vehicle (S404). Therefore, compared with the case of checking the road surface condition only by visual inspection, it is easier to predict a situation where there is a high risk of rubbing the lower surface of the vehicle.
[0191] (19) When the terminal devices 31 to 33 detect that the vehicle is under the influence of a crosswind (S405: YES), they transmit voice information representing a voice for notifying that the vehicle is under the influence of a crosswind (S406). Therefore, compared with the case of visually checking only the external situation, it is easier to predict a situation where the vehicle is likely to be affected by a crosswind.
[0192] (20) When the terminal devices 31 to 33 detect that the vehicle has fallen or collided (S408: YES), they transmit voice information representing a voice for notifying that the vehicle has fallen or collided (S409). Therefore, the users of the terminal devices 31 to 33 can grasp while driving that another vehicle has fallen or collided. Therefore, compared with the case of visually checking the state of another vehicle only, safe driving can be realized.
[0193] (21) In addition, the following effects can also be obtained. As described in detail above, the communication system 1 includes a plurality of terminal devices 31 to 33 capable of communicating with each other, and also includes a server 10 capable of communicating with each of the plurality of terminal devices 31 to 33. The CPU 41 of each terminal device 31 to 33 converts the voice input by the user into voice information when the user inputs a voice, and transmits the voice information to a device including other terminal devices 31 to 33.
[0194] Then, the CPU 41 receives the voice information transmitted by the other terminal devices 31 to 33 and reproduces the received voice information. Further, the server 10 receives the voice information transmitted from the plurality of terminal devices 31 to 33, arranges the voices corresponding to the respective voice information so that they do not overlap, and then distributes the voice information to the plurality of terminal devices 31 to 33.
[0195] According to such a communication system 1, since the content to be transmitted can be input by voice and the content transmitted from other terminal devices 31 to 33 can be reproduced by voice, even a user during driving can safely transmit and receive the content to be whispered etc. Further, in such a communication system 1, when reproducing the voice, the voices can be reproduced so that they do not overlap, so the The sound can be made easier to hear.
[0196] Also, in the communication system 1, when the CPU 41 receives a specific command from the user, it reproduces the voice information that has already been reproduced. In particular, the voice information reproduced immediately before may be reproduced. According to such a communication system 1, the voice information that could not be heard due to noise or the like can be reproduced again by the user's command.
[0197] Furthermore, in the communication system 1, the terminal devices 31 to 33 are mounted on the vehicle, and the CPUs 41 of the terminal devices 31 to 33 detect the movement of the user's hand on the steering wheel by the touch panels 61, 62, and when detecting a start specific operation that starts the input of voice by the touch panels 61, 62, start the process of generating a voice packet, and when detecting an end specific operation that ends the input of voice by the touch panels 61, 62, end the process of generating a voice packet. According to such a communication system 1, an operation for inputting voice can be performed on the steering wheel. Therefore, the user during vehicle driving can input voice more safely.
[0198] Also, in the communication system 1, the CPU 41 detects the start specific operation and the end specific operation input via the touch panels 61, 62 as the movement of the hand in a direction different from the operation direction when operating the horn. According to such a communication system 1, it is possible to suppress the horn from sounding during the start specific operation or the end specific operation.
[0199] Also, in the communication system 1, the server 10 generates new voice information in which identification information for identifying the user is added to the voice information input by the user, and distributes the generated voice information. In such a communication system 1, since the identification information for identifying the user can be added by voice, the user can notify other users of the user who is the voice transmission source without speaking his or her own identification information.
[0200] Furthermore, in the communication system 1, before playing the voice information, the server 10 extracts a preset keyword included in the voice information from the voice information, and replaces the keyword included in the voice information with the voice of another word or a preset sound. According to such a communication system 1, it is possible to register words that are not preferable for distribution (such as obscene words, words indicating personal names, and barbaric words) as keywords, and replace them with words or sounds that can be distributed and then perform the distribution.
[0201] Also, in the communication system 1, the CPUs 41 of the terminal devices 31 to 33 acquire their own location information, associate the location information with the voice information transmitted from their own terminal devices 31 to 33, and transmit it to other terminal devices 31 to 33, and acquire the location information transmitted by the other terminal devices 31 to 33. Then, according to the positional relationship between the position of the terminal devices 31 to 33 corresponding to the voice to be played and the position of their own terminal devices 31 to 33, the volume output when playing the voice is controlled.
[0202] In particular, the CPU 41 controls the volume to decrease as the distance to other terminal devices 31 to 33 increases, and when outputting voice from a plurality of speakers, according to the azimuth in which the other terminal devices 31 to 33 are located with respect to the position of its own terminal devices 31 to 33, controls the volume so that the volume from the direction in which the other terminal devices 31 to 33 are located is larger as seen from the user.
[0203] According to such a communication system 1, since the volume is controlled according to the positional relationship between the position of the terminal devices 31 to 33 corresponding to the voice to be played and the position of its own terminal devices 31 to 33, the user listening to the voice can consciously grasp the positional relationship between the other terminal devices 31 to 33 and its own terminal devices 31 to 33.
[0204] Furthermore, in the communication system 1, the server 10 converts the voice information into characters and outputs the converted characters to a display device. According to such a communication system 1, it is possible to confirm the information missed by voice in characters.
[0205] Furthermore, in the communication system 1, the CPUs 41 of the terminal devices 31 to 33 transmit information indicating the state of the vehicle and the driving operations of the user (such as the operating states of lights, wipers, TVs, radios, etc., the driving states such as traveling speed and traveling direction, and the control states such as detection values by vehicle sensors and the presence or absence of failures) together with voice. In such a communication system 1, information indicating the driving operations of the user can also be sent to other terminal devices 31 to 33.
[0206] Also, in the communication system 1, advertisements are played by voice according to the positions of the terminal devices 31 to 33. The content of the advertisements is inserted by the server 10, and the server 10 transmits the advertisements to the terminal devices 31 to 33 that play the voice information. According to such a communication system 1, a business model capable of obtaining advertising revenue can be provided.
[0207] Furthermore, in the communication system 1, a distribution mode for an unspecified large number of communication partners and an interactive mode for making a call with a specific communication partner are set to be changeable. According to such a communication system 1, it is also possible to enjoy a call with a specific partner.
[0208] Furthermore, in the communication system 1, the server 10 determines whether the user is emotional or has a polite way of speaking when inputting voice, etc., by detecting from the voice information the amplitude of the input level and the polite expressions included in the voice, and does not distribute the voice by a user who is emotional or does not have a polite way of speaking. According to such a communication system 1, it is possible to exclude the speech from the user that causes arguments.
[0209] Also, in the communication system 1, the server 10 notifies a user who speaks ill of others, or a user who is emotional or does not have a polite way of speaking, of a slip of the tongue or the like. In this case, the determination is made based on whether or not a word registered as a keyword as a curse word or the like is included in the voice information. According to such a communication system 1, it is possible to prompt the user who has made a slip of the tongue not to make a slip of the tongue.
[0210] Note that S102 corresponds to an example of the processing as the input means, S114, S1202, S1204, and S1205 correspond to examples of the processing as the acquisition means, S116 corresponds to an example of the processing as the transmission means, and S811 corresponds to an example of the processing as the reproduction means.
[0211] [Other Embodiments] The embodiments of the present invention are not limited to the above-described embodiments at all, and various forms can be adopted as long as they belong to the technical scope of the present invention.
[0212] (1) In the above embodiment, the extraction process of extracting the voice information corresponding to the terminal devices 31 to 33 that reproduce the voice from among the plurality of accumulated voice information is executed by the server 10 (S1302 to S1312), but it is not limited thereto. For example, part or all of such extraction processing may be executed by the terminal devices 31 to 33 that reproduce the voice. For example, the server 10 may transmit the voice information to the terminal devices 31 to 33 in a form that may also include extra voice information that may not be reproduced by the terminal devices 31 to 33, and the terminal devices 31 to 33 may extract the voice information to be reproduced (discard unnecessary voice information) from the received voice information. Note that the transmission of the voice information from the server 10 to the terminal devices 31 to 33 may be performed in response to a request from the terminal devices 31 to 33, or may be performed regardless of the request. It may be done in this way.
[0213] (2) In the above embodiment, the server 10 determines whether the content of each piece of voice information to be transmitted overlaps with that of other voice information (S1311), and if it is determined that there is an overlap (S1311: YES), the content overlap is resolved (S1312). By doing so, it is possible to make it difficult for the phenomenon of the same speech being repeatedly reproduced to occur. However, from a different perspective, the more the number of pieces of voice information with overlapping content, the higher the credibility of the content is considered to be. Therefore, it may be a condition for reproducing the voice of the content that the number of pieces of voice information with overlapping content is equal to or more than a certain number.
[0214] For example, the response processing shown in FIGS. 21A and 21B replaces the processing of S1311 to S1312 in the response processing of FIGS. 17A and 17B described above with the processing of S1401 to S1403. That is, in the response processing shown in FIGS. 21A and 21B, for each piece of voice information of the transmission candidates, it is determined whether or not it contains a specific keyword set in advance, such as "accident" or "traffic control" (S1401).
[0215] And, when it is determined that there is voice information containing a specific keyword among the voice information of the transmission candidates (S1401: YES), it is determined whether or not there is voice information with the same content (duplicate content) less than a certain number (for example, 3) among the voice information containing the specific keyword (S1402). Here, when it is determined that there is voice information with the same content less than a certain number (S1402: YES), the voice information of that content is excluded from the voice information of the transmission candidates (S1403). Thereafter, the voice information of the transmission candidates is transmitted to the terminal devices 31 to 33 of the transmission source of the information request (S1313), and the response processing is terminated.
[0216] On the other hand, when it is determined that there is no voice information containing a specific keyword among the voice information of the transmission candidates (S1401: NO), or when it is determined that there is no voice information with the same content less than a certain number (S1402: NO), the process directly proceeds to S1313, and the voice information of the transmission candidates is transmitted to the terminal devices 31 to 33 of the transmission source of the information request (S1313), and the response processing is terminated.
[0217] That is, according to the response processing illustrated in FIGS. 21A and 21B, for voice information containing a specific keyword, it is not transmitted to the terminal devices 31 to 33 until a certain number of voice information with the same content is accumulated, and it is transmitted to the terminal devices 31 to 33 on the condition that the certain number has been reached. By doing so, it is possible to suppress the reproduction of voices with low credibility of content, such as false information, by the terminal devices 31 to 33.
[0218] Note that, in the examples of FIGS. 21A and 21B, the condition for transmission is that there are a certain number of voice information items with the same content, targeting voice information including specific keywords, but it is not limited to this. For example, regardless of whether specific keywords are included or not, it may be set as a condition for transmission that there are a certain number of voice information items with the same content for all voice information items that are candidates for transmission. Also, in determining whether there are a certain number of voices with the same content, even if there are multiple pieces of voice information with the same source, they may be counted as one (not counted as multiple). In this way, it is possible to suppress the reproduction of voices with low credibility of content on the terminal devices 31 to 33 of other users due to the same user speaking the same voice multiple times.
[0219] (3) In the above embodiment, the terminal devices 31 to 33 that transmit voice information perform the process of converting the recorded voice into synthesized voice (S107), but it is not limited to this. For example, part or all of such conversion processing may be executed by the server 10, or may be executed by the terminal devices 31 to 33 that play the voice.
[0220] (4) In the above embodiment, the terminal devices 31 to 33 execute the request process (FIG. 16) at every predetermined time, but it is not limited to this, and it may be executed at other predetermined timings. For example, it may be executed every time a predetermined distance is traveled, at the time of a request operation by the user, at the arrival of a specified time, at the passing of a specified location, at the exit from the surrounding area, etc.
[0221] (5) In the above embodiment, voice information within the peripheral area is extracted and played back, but it is not limited to this. For example, an enlarged area (wider than the peripheral area) that includes the peripheral area may be set, voice information within the enlarged area may be extracted, and voice information within the peripheral area may be further extracted from among them and played back. Specifically, in S1307 of the above embodiment, voice information whose voice generation position is within the enlarged area is extracted. Then, in the terminal devices 31 to 33, the voice information within the enlarged area received from the server 10 is stored (cached), and at the stage of playing back the voice represented by the voice information, voice information whose voice generation position is within the peripheral area is extracted and played back. In this way, even if the peripheral area based on the current location of the vehicle changes as the vehicle travels, as long as the changed peripheral area is within the enlarged area at the time of receiving the voice information, the voice represented by the voice information corresponding to the peripheral area can be played back.
[0222] (6) In the above embodiment, the peripheral area is set based on the current location (actual position) of the vehicle, but it is not limited to this. Location information other than the current location may be acquired, and the peripheral area may be set based on the acquired location information. Specifically, location information set (input) by the user may be acquired. In this way, for example, by setting the peripheral area based on the destination, voice information transmitted near the destination can be confirmed in advance at a location far from the destination.
[0223] (7) In the above embodiment, as the terminal device, the terminal devices 31 to 33 configured as in-vehicle devices mounted on the vehicle are exemplified, but it is not limited to this. For example, a device (such as a mobile phone) that can be carried by the user and used in the vehicle may also be used. Also, the terminal device may or may not share at least a part of its configuration (for example, a configuration for detecting the position, a configuration for inputting voice, a configuration for playing back voice, etc.) with a device mounted on the vehicle.
[0224] (8) In the above-described embodiment, among the voice information transmitted from other terminal devices 31 to 33 to the server 10, an example of the configuration is illustrated in which the voice represented by the voice information whose voice generation position is located within the peripheral area set based on the current location is reproduced. However, the present invention is not limited to this, and the voice information may be extracted and reproduced regardless of the location information.
[0225] (9) The information input by the terminal device from the user is not limited to the voice uttered by the user. For example, it may be character information input by the user. That is, the terminal device may input character information from the user instead of (or together with) the voice uttered by the user. In this case, the terminal device may display the character information so that the user can visually recognize it, or may convert the character information into synthesized voice and reproduce it.
[0226] (10) The above-described embodiment is merely an example of an embodiment to which the present invention is applied. The present invention can be realized in various forms such as a system, a device, a method, a program, a recorded recording medium, and the like.
[0227] (11) In addition, the following may be done. For example, in the above communication system 1, the content to be added as identification information according to the communication mode (for example, real name, first nickname, second nickname, etc.) may be changed. For example, information for identifying each other may be exchanged in advance, and the identification information may be changed depending on whether the communication partner is a pre-registered partner or not. That is, if it is a pre-registered partner, the identification information may be changed.
[0228] Also, in the above-described embodiment, the keyword is configured to be replaced with another word or the like in any case. However, for example, depending on conditions such as the communication partner and the communication mode, the process of replacing the keyword with another word or the like may not be performed. This is because when communicating with a specific partner, such a process is not necessary unlike when distributing to the public.
[0229] Also, for words such as "um" and "uh" that are specific to spoken language and have no meaning, it may be possible to generate voice information with these words omitted. In this case, the playback time of the voice information can be shortened by the amount omitted.
[0230] Also, in the above embodiment, the CPU 41 of each terminal device 31 to 33 outputs information indicating the state of the vehicle and the driving operation of the user in characters, but it may also be output in voice. Further, for example, various information such as traffic jam information may be generated based on information obtained from other terminal devices 31 to 33, and the generated information may be output.
[0231] Furthermore, in the above communication system 1, a configuration in which each of the terminal devices 31 to 33 directly exchanges voice information and a configuration in which each of the terminal devices 31 to 33 exchanges voice information via the server 10 may be switched according to preset conditions (for example, the distance to the base station connected to the server 10, the communication state between the terminal devices 31 to 33, or settings by the user, etc.).
[0232] Also, in the above embodiment, advertisements are configured to be played in voice according to the positions of the terminal devices 31 to 33, but advertisements may also be played in voice at regular intervals or for every certain number of voice information reproductions. Also, the content of the advertisement may be transmitted from the terminal devices 31 to 33 of the advertiser to other terminal devices 31 to 33. At this time, it may be possible to enable a call (communication) with the advertiser by the user inputting a predetermined operation to the terminal devices 31 to 33, or to guide the user of the terminal devices 31 to 33 to the advertiser's store.
[0233] Furthermore, in the above embodiment, a mode for communicating with a specific plurality of communication partners may be set. For example, in the case of a configuration for exchanging position information, terminal devices 31 to 33 within a certain distance may be used as communication partners, or specific users (such as a preceding vehicle or an oncoming vehicle (identified by license plate imaging using a camera, identified by GPS), users within a preset group) may be used as communication partners. Also, it may be possible to change the communication partner according to the setting (to change the communication mode).
[0234] Also, when the terminal devices 31 to 33 are mounted on a vehicle, the terminal devices 31 to 33 moving in the same direction on the same road may be used as communication partners, or the terminal devices 31 to 33 with consistent vehicle behavior (such as removing them from the communication partners when deviating to a side road) may be used as communication partners.
[0235] Also, it may be configured to be able to specify other users in advance and set favorite registration, exclusion registration, etc., so that a user can select a preferred partner and communicate. In this case, when transmitting voice, communication partner information for specifying the communication partner may be transmitted, or when there is a server, the communication partner information may be registered with the server, so that voice information can be transmitted only to a predetermined partner.
[0236] Also, directions may be associated with a plurality of switches such as buttons respectively. When the user designates the direction in which the communication partner is located with this switch, only the user in that direction may be set as the communication partner.
[0237] Furthermore, in the communication system 1 described above, although it is configured to exchange voice information, it may be configured to acquire information consisting of characters, convert this information into voice, and play it back. Also, when transmitting voice, it may be configured to convert the voice into characters, then transmit it, and restore it to voice before playing it back. With this configuration, the amount of data to be transmitted by communication can be reduced.
[0238] Furthermore, in the configuration where data is transmitted by characters in this way, the data may be transmitted after being translated into the language with the least amount of data among a plurality of preset languages.
[0239] Further, some of the functions of the server 10 in the above embodiment may be implemented by the terminal devices 31 to 33. For example, in the CPUs 41 of the terminal devices 31 to 33, when there is a plurality of voice information for which playback has not been completed, they may be sorted so that the voices corresponding to each voice information do not overlap and then played back. In the CPU 41, new voice information obtained by adding voice identification information for identifying the user to the voice information input by the user may be generated, and the generated voice information may be transmitted.
[0240] Specifically, in the voice transmission process and the voice playback process, a part of the voice distribution process may be implemented. For example, when implementing a part of the voice distribution process in the voice playback process, as shown in FIGS. 22A to 22D, between S802 and S804 of the voice playback process, the processes of S703 to S722 (excluding S717) of the voice distribution process may be implemented.
[0241] Even in such a communication system, when playing back voices, the voices can be played back so as not to overlap, so that the voices can be made easier to hear. Also, even in such a communication system 1, since identification information for identifying the user can be added by voice, the user can notify other users of the user who is the voice transmission source without speaking their own identification information.
[0242] Further, in the above embodiment, it is prohibited for voices to be played back overlappingly, but a part of the voices may be deliberately played back overlappingly. In this case, only the last few seconds of the previous voice and the first few seconds of the subsequent voice may be overlapped. In this case, a fade-out effect of gradually decreasing the volume at the end of the previous voice and a fade-in effect of gradually increasing the voice at the beginning of the subsequent voice may be combined.
[0243] Furthermore, it may be determined whether the driver is drowsy based on the presence or absence of the voice uttered by the user. To implement this configuration, for example, the drowsiness determination process shown in FIG. 23 may be implemented. The drowsiness determination process is a process that is executed in parallel with the voice call process in the terminal devices 31 to 33. When the user does not speak for a certain period of time or even if the user speaks, it is estimated that the level of consciousness is low, an alarm is issued or vehicle control such as stopping is performed. Specifically, as shown in FIG. 23, first, the time from the final utterance by the user is counted (S1401).
[0244] Then, it is determined whether or not a preset first reference time (for example, about 10 minutes) has elapsed since the user's last utterance (S1402). If the first reference time has not elapsed (S1402: NO), the drowsiness determination process ends.
[0245] If the first reference time has elapsed (S1402: YES), an inquiry to the user is performed (S1403). In this process, a question is asked to prompt the user to speak in order for the user to confirm the level of consciousness. For this purpose, a direct question such as "Are you awake?" or an indirect question such as "Please tell me the current time" may be used. It may also be an indirect question such as "Please tell me the current time".
[0246] Subsequently, it is determined whether or not a response to such a question has been obtained (S1404). If a response to the question has been obtained (S1404: YES), the level of the response is determined (S1405).
[0247] Here, the level of the response indicates items such as the volume of the voice and the smoothness of the speech, and the level (the volume of the sound, the accuracy by voice recognition) decreases when getting sleepy. In this process, any of these items is compared with a preset threshold value for judgment.
[0248] If the level of the response is sufficiently high (S1405: YES), the drowsiness determination process ends. Also, when the level of the response is insufficient (S1405: NO), and when no response to the question is obtained in the process of S1404 (S1404: NO), it is determined whether or not a second reference time (for example, about 10 minutes and 10 seconds) set to a value equal to or greater than the above-mentioned first reference time has elapsed since the user's last utterance (S1406).
[0249] If the second reference time has not elapsed (S1406: NO), the process returns to S1403. If the second reference time has elapsed (S1406: YES), an alarm is issued via the speaker 53 or the like (S1407).
[0250] Next, it is determined whether or not a response to the alarm (for example, a voice such as "Waah" or an operation via input unit 45 to cancel the alarm) has been obtained (S1408). If a response to the alarm has been obtained (S1408: YES), the level of the response is determined (S1409) in the same manner as in the process of S1405. Note that if an operation has been input via input unit 45, it is determined that the level of the response is sufficient.
[0251] If the response level is sufficiently high (S1409: YES), the dozing judgment process is terminated. If the response level is insufficient (S1409: NO), or if no response to the question is obtained in the process of S1408 (S1408: NO), it is judged whether a third reference time (for example, about 10 minutes and 30 seconds) set to a value equal to or greater than the above-mentioned second reference time has elapsed since the last utterance by the user (S1410).
[0252] If the third reference time has not elapsed (S1410: NO), the process returns to S1407. If the third reference time has elapsed (S1410: YES), vehicle control such as stopping the vehicle is performed (S1411), and the drowsiness determination process ends.
[0253] By carrying out such a process, it is possible to prevent the user (driver) from falling asleep at the wheel. Furthermore, when an advertisement is inserted by voice (S720) or when voice is also displayed as text (S812), additional text information may be displayed on the display of the navigation device 48 or the like.
[0254] In order to implement such a configuration, for example, the additional information display process shown in FIG. 24 may be performed. Note that the additional information display process is a process that is performed in parallel with the voice distribution process on the server 10.
[0255] In the additional information display process, as shown in FIG. 24, first, it is determined whether an advertisement (CM) is inserted in the voice distribution process (S1501), and whether characters are displayed (S1502). If an advertisement is inserted (S1501: YES), or if characters are displayed (S1502: YES), information related to the characters to be displayed is acquired (S1503).
[0256] Here, the information related to the characters corresponds to information associated with this character string, such as an advertisement or regional information corresponding to the character string (keyword) to be displayed, or information for explaining information related to the character string. This information may be acquired from within the server 10, or may be acquired from outside the server 10 (such as on the Internet).
[0257] When the information related to the characters is acquired, this information is displayed on the display as characters or images (S1504), and the additional information display process ends. Note that if no advertisement is inserted and no characters are displayed in the processes of S1501 and S1502 (S1501: NO and S1502: NO), the additional information display process ends.
[0258] According to such a configuration, not only can the voice be displayed as characters, but also related information can be displayed. Note that the display for displaying character information is not limited to the display of the navigation device 48, and may be any display such as a head-up display in the vehicle.
Explanation of Reference Numerals
[0259] 1... Communication system, 10... Server, 20... Base station, 31 - 33... Terminal devices, 41... CPU, 42... ROM, 43... RAM, 44... Communication unit, 45... Input unit, 48... Navigation device, 49... GPS receiver, 52... Microphone, 53... Speaker
Claims
【Claim 1】 A communication system comprising a plurality of terminal devices capable of communicating with each other, wherein each terminal device has voice input conversion means for converting voice into a voice signal when a user inputs voice, voice transmission means for transmitting the voice signal to a device including other terminal devices, voice reception means for receiving a voice signal transmitted by another terminal device, voice reproduction means for reproducing the received voice signal, and when there are a plurality of voice signals for which reproduction is not complete, the voice reproduction means reproduces the signals after arranging them so that the voices corresponding to the respective voice signals do not overlap. A communication system.
Citation Information
Patent Citations
Information terminal device and information server
JP2003106850A
Communication method and system, and communication device constituting the communication system
JP2004343784A
Radio terminal, management server and PTT system
JP2006246201A
Radio communication system, and information deployment method
JP2011114738A
Terminal, system and method for displaying information, and program
JP2011215861A
Cited By
Game machine
JP2025129264A
Game machine
JP2025134968A