Systems, methods, and programs for transmitting sound
The voice transmission system adjusts voice parameters with vehicle motion or driver input to distinguish artificial voice from human voice, addressing misinterpretation and unnaturalness, ensuring transparency and user satisfaction.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- TOYOTA JIDOSHA KK
- Filing Date
- 2023-07-21
- Publication Date
- 2026-05-19
AI Technical Summary
Existing systems that use machine learning models to generate artificial voice for vehicle users struggle to ensure ease of distinction between artificial and human voice, leading to potential misinterpretation and unnaturalness.
A voice transmission system that adjusts parameters like pitch and volume of artificial voice in conjunction with vehicle motion or driver input to differentiate it from human voice, using a machine learning model to generate speech that resembles human speech.
This approach enhances the distinguishability of artificial voice from human voice, reducing unnaturalness and preventing misinterpretation, while maintaining a human-like quality.
Smart Images

Figure 0007861716000001 
Figure 0007861716000002 
Figure 0007861716000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a technology for transmitting voice to a user of a vehicle, specifically, a technology for transmitting artificial voice generated by a machine learning model.
Background Art
[0002] Patent Document 1 discloses an in-vehicle device that reproduces voice for a user riding in a vehicle based on the voice information data for speech received from a roadside unit. The in-vehicle device includes a speech speed determination means and a voice signal generation means. The speech speed determination means determines a speech speed according to the situation when reproducing the voice from the received voice information data for speech. Also, the voice signal generation means generates a voice signal based on the voice information data for speech so as to be the determined speech speed. Then, the voice based on the generated voice signal is output to the user from a speaker or the like. According to the speech speed determination means, when the voice information data for speech is highly urgent, the speech speed is set high in order to more surely arouse the attention of the user.
[0003] Thus, a device for transmitting voice to a user of a vehicle is known. In addition to the above Patent Document 1, the following Patent Document 2 and Patent Document 3 can be exemplified as documents showing the technical level at the time of filing in the technical field of the present disclosure.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Patent Document 3
Summary of the Invention
Problems to be Solved by the Invention
[0005] Consider applying a machine learning model to a device that transmits voice to vehicle users, and outputting the artificial voice generated by the machine learning model to the user. While applying a machine learning model makes it possible to make the artificial voice closer to a real human voice, it may become difficult for the user listening to the artificial voice to distinguish it from a real human voice. However, from the perspective of transparency regarding the system to which the machine learning model is applied, it is desirable that it be easy for the user to distinguish between the artificial voice generated by the machine learning model and a human voice.
[0006] One objective of this disclosure is to provide a technology for ensuring ease of distinction between artificial voice and human voice in a system that transmits artificial voice generated by a machine learning model to a vehicle user. [Means for solving the problem]
[0007] This disclosure provides a voice transmission system. The voice transmission system of this disclosure comprises one or more processors. One or more processors generate artificial speech that resembles human speech using a machine learning model. The generated artificial speech is then transmitted to the user of a vehicle. At this time, the processors change a second parameter relating to the properties of the artificial speech in conjunction with a first parameter relating to at least one of the vehicle's motion and the vehicle driver's input.
[0008] In the voice transmission system of this disclosure, the first parameter may be the acceleration of the vehicle, and the second parameter may be the pitch of the artificial voice. The processor may then link the first and second parameters such that the pitch of the artificial voice increases as the acceleration increases.
[0009] This disclosure provides a method for voice transmission. The voice transmission method of this disclosure includes the following steps: The first step is to generate an artificial voice that resembles a human voice using a machine learning model. The second step is to transmit the artificial voice to a vehicle user. The third step is to change a second parameter relating to the properties of the artificial voice in conjunction with a first parameter relating to at least one of the vehicle's motion and the vehicle driver's input.
[0010] This disclosure provides a voice communication program that is executed by a computer. The voice communication program causes the computer to execute the voice communication method described above. [Effects of the Invention]
[0011] According to the technology disclosed herein, parameters relating to the properties of the artificial voice can be changed in conjunction with parameters related to the vehicle's motion or the driver's input. Since this change in voice in conjunction with vehicle-related parameters does not occur with the voice of a real human, it is possible to prevent the vehicle user from misinterpreting the artificial voice emitted by the voice transmission system as a real human voice. Furthermore, since it is no longer necessary to make the artificial voice sound mechanical to prevent user misinterpretation, the sense of unnaturalness that users may feel towards the artificial voice can be reduced.
[0012] Thus, the technology disclosed herein makes it possible to make artificial speech generated by a machine learning model more easily distinguishable from a real human voice while minimizing the sense of unnaturalness experienced by the user. [Brief explanation of the drawing]
[0013] [Figure 1] This is a block diagram showing an example of the vehicle configuration according to this embodiment. [Figure 2] This is a conceptual diagram illustrating an example of a scenario in which voice transmission occurs using the voice transmission system according to this embodiment. [Figure 3] This block diagram shows an example of processing by a voice transmission system. [Figure 4] This table shows examples of combinations of the first and second parameters. [Figure 5] This is a block diagram showing another example of processing by a voice transmission system. [Figure 6] This block diagram shows yet another example of processing by a voice transmission system. [Modes for carrying out the invention]
[0014] Embodiments of this disclosure will be described with reference to the attached drawings.
[0015] 1. Overview of the voice transmission system The voice transmission system according to this embodiment is a system that generates artificial voice using a machine learning model and transmits the generated artificial voice to the vehicle user. The user to whom the artificial voice is transmitted may be the vehicle driver or a passenger other than the driver. Figure 1 is a diagram showing an example configuration of a vehicle 1 to which the voice transmission system 100 according to this embodiment is applied. Vehicle 1 may be a vehicle owned by an individual or a ride-sharing vehicle used for various mobility services. Vehicle 1 may also be a vehicle that is automatically driven by an autonomous driving system.
[0016] Vehicle 1 includes a computer 10 that generates artificial voice, a group of sensors 20, and a speaker 30. The computer 10 is connected to the group of sensors 20 and the speaker 30 by an in-vehicle network. The voice transmission system 100 includes at least the computer 10. However, the voice transmission system 100 may also include the speaker 30 in addition to the computer 10. Furthermore, the voice transmission system 100 may also include the group of sensors 20.
[0017] The computer 10 is, for example, an ECU (Electronic Control Unit) mounted on the vehicle 1, or an aggregate of a plurality of ECUs. Alternatively, some or all of the functions of the computer 10 may be arranged in an external server. In this case, the vehicle 1 and the external server are connected by a wireless communication network. In any case, the computer 10 includes one or more processors 11 (hereinafter simply referred to as the processor 11) and one or more storage devices 12 (hereinafter simply referred to as the storage device 12).
[0018] The processor 11 executes various processes. Examples of the processor 11 include a CPU (Central Processing Unit), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), etc. The processor 11 may be one type of processor, or may include a plurality of types of processors.
[0019] Examples of the storage device 12 include an HDD (Hard Disk Drive), an SSD (Solid State Drive), a volatile memory, a non-volatile memory, etc. The storage device 12 may be one type of storage device, or may include a plurality of types of storage devices. The storage device 12 includes at least a program storage area 13 and a model data storage area 14. Each storage area may be realized by a single storage device, or may be realized by separate storage devices.
[0020] One or more programs are stored in the program storage area 13. Each program is composed of a plurality of instructions. Through the cooperation of the processor 11 that executes the program and the storage device 12, various processes by the voice transmission system 100 are realized. The program stored in the program storage area 13 includes at least a voice transmission program for causing the processor 11 to execute processes related to the transmission of artificial voices. The voice transmission program may be recorded on a computer-readable recording medium.
[0021] The model data storage area 14 stores model data. The model data stored in the model data storage area 14 is data for a machine learning model used to generate artificial speech.
[0022] The storage device 12 may include storage areas other than the program storage area 13 and the model data storage area 14. For example, map information may be stored in the storage areas other than the program storage area 13 and the model data storage area 14.
[0023] The sensor group 20 includes a vehicle state sensor 21. The vehicle state sensor 21 detects the state of vehicle 1. Specifically, it detects parameters such as the vehicle's speed, acceleration, steering angle, roll angular velocity, pitch angular velocity, yaw angular velocity, and tilt angle. Examples of vehicle state sensors 21 include a vehicle speed sensor, acceleration sensor, yaw rate sensor, steering angle sensor, roll angular velocity sensor, pitch angular velocity sensor, yaw angular velocity sensor, and tilt angle sensor. The sensor group may also include a recognition sensor 22 and a position sensor 23. The recognition sensor 22 recognizes the surrounding conditions of vehicle 1. Examples of recognition sensors 22 include a camera, LiDAR (Laser Imaging Detection and Ranging), radar, etc. Examples of cameras include a front camera that images the area in front of vehicle 1, a rear camera that images the area behind vehicle 1, and a side camera that images the area to the side of vehicle 1. The position sensor 23 detects the position and orientation of vehicle 1. An example of a position sensor 23 is a GNSS (Global Navigation Satellite System) sensor.
[0024] Computer 10 acquires information detected by these sensors via the in-vehicle network. For example, computer 10 acquires the status of vehicle 1 from vehicle status sensor 21. Computer 10 may also acquire information detected by recognition sensor 22 and position sensor 23. Furthermore, based on the information obtained from recognition sensor 22, computer 10 may perform processes such as detecting objects present around vehicle 1, measuring the relative position and relative speed of the detected objects with respect to vehicle 1, etc.
[0025] Speaker 30 is a speaker that transmits artificial voice generated by computer 10 to the user of vehicle 1. Speaker 30 may include multiple speakers. The artificial voice data generated by computer 10 is sent to speaker 30 via the in-vehicle network. Then, speaker 30 plays the voice based on the received data, thereby transmitting the artificial voice to the user.
[0026] The above describes an example of the configuration of the voice transmission system 100. Next, we will explain the transmission of artificial voice by the voice transmission system 100 using a specific scenario as an example. Figure 2 is a conceptual diagram showing an example of a scenario in which artificial voice is transmitted.
[0027] In the example shown in Figure 2, speaker 30 includes three speakers: speaker 30a, speaker 30b, and speaker 30c. Of these three speakers, speaker 30a is used to transmit artificial voice to the user. Thus, when speaker 30 includes multiple speakers, only one speaker may emit sound at a time. Alternatively, sound may be emitted from multiple speakers simultaneously.
[0028] Furthermore, in the example shown in Figure 2, Vehicle 1 is an autonomous vehicle, and a voice message saying "Starting autonomous driving" is transmitted to inform the user that autonomous driving of Vehicle 1 is about to begin. In this way, information can be transmitted to the user of Vehicle 1 through the transmission of artificial voice by the voice transmission system 100. Other examples of the content of the transmitted artificial voice include messages such as "Starting to move," "Changing lanes," "Arriving at your destination shortly," "Approaching a sharp curve," "There is an intersection ahead. Please watch out for pedestrians," "Approaching the vehicle in front. Please maintain a safe distance," and "Entering a road with a different speed limit. Changing speed." However, the transmitted artificial voice is not limited to providing information; it may also include messages that provide some kind of warning, such as a voice message that alerts the user when drowsiness is detected.
[0029] Computer 10 uses a machine learning model to generate artificial speech that resembles human speech in order to convey these messages. Thanks to recent technological advancements, it is now possible to make the artificial speech generated using machine learning models sound very similar to a real human voice. By making the artificial speech sound more like a real human voice, the sense of unnaturalness experienced by the user receiving the speech can be reduced compared to when a mechanical voice is played. On the other hand, it may become more difficult for the user to distinguish whether the speech they hear is artificial or a real human voice. For example, in the scenario shown in Figure 2, the driver of vehicle 1 who hears the speech coming from speaker 30a may mistakenly perceive it as being spoken by another person in vehicle 1, such as a passenger sitting in the front passenger seat. Alternatively, they may mistakenly perceive it as speech transmitted wirelessly from a person outside vehicle 1.
[0030] From the perspective of transparency regarding the system to which the machine learning model is applied, it is desirable that the user of Vehicle 1 can easily distinguish the artificial voice generated by the machine learning model from a real human voice. However, in order to reduce the sense of unnaturalness that users may feel when hearing the artificial voice, the artificial voice itself should be close to a human voice rather than a mechanical voice.
[0031] People determine whether the sound they hear is a real human voice or an artificial voice based on the properties of the sound. of It can be identified by the naturalness of the changes. The properties of a voice, such as volume, pitch, and tone quality, do change in normal human conversation, but these changes do not follow any specific rules. For example, factors that change the properties of a human voice in conversation include the topic of conversation, the attitude of the person being spoken to, and changes in the surrounding environment, but these do not give regularity to the changes in voice properties. Therefore, if there is some regularity in the changes in a heard voice, people instinctively perceive these changes as unnatural and unconsciously determine that it is not a real human voice, but rather an artificial voice. but can.
[0032] Therefore, the voice transmission system 100 according to this embodiment acquires parameters related to the vehicle's motion as first parameters. Then, it changes second parameters related to the properties of the artificial voice in conjunction with the first parameters. By regularly changing the second parameters in conjunction with the first parameters, it becomes easier to distinguish the artificial voice transmitted by the voice transmission system 100 from an actual human voice.
[0033] 2. Processing Example Figure 3 is a block diagram showing an example of the processes executed by the processor 11 when the voice transmission program is executed. Upon execution of the voice transmission program, the processor 11 executes processes 111, 112, 113, 114, and 115.
[0034] First, processor 11 executes process 111. Process 111 detects a scene in which artificial voice transmission takes place. A scene in which artificial voice transmission takes place is, for example, the scene in which vehicle 1 starts autonomous driving. In this case, processor 11 determines the timing of the start of autonomous driving based on information received from the autonomous driving device 200 that performs autonomous driving of vehicle 1. If it is determined that it is the timing for autonomous driving to start, it detects the scene in question as a scene in which artificial voice transmission takes place.
[0035] The processor 11 may also detect scenes in which artificial voice transmission is performed based on information received from devices other than the autonomous driving system 200, such as the recognition sensor 22. For example, the processor 11 calculates the distance between vehicle 1 and the preceding vehicle based on information obtained when the recognition sensor 22 recognizes the preceding vehicle. Then, if the distance between vehicle 1 and the preceding vehicle is smaller than a predetermined distance, it is detected as a scene in which artificial voice transmission is performed.
[0036] Alternatively, the processor 11 may detect scenes for transmitting artificial voice based on map information stored in the storage device 12 and location information detected by the location sensor 23. For example, the processor 11 obtains the current location of vehicle 1 from the location sensor 23. The processor 11 also obtains the destination of vehicle 1 registered in the map information. Based on this information, the processor 11 detects a scene for transmitting artificial voice when it is determined that vehicle 1 is approaching its destination. As another example, the processor 11 can similarly detect scenes for transmitting artificial voice based on the map information and the location information of vehicle 1 when vehicle 1 is approaching an intersection or a sharp curve.
[0037] The processor 11 may also perform processing 111 using a machine learning model. In other words, a machine learning model may be used to detect scenes in which artificial speech is transmitted.
[0038] Next, processor 11 executes process 112. Process 112 determines the content of the speech produced by the artificial voice. The content of the artificial voice is determined according to the scene detected in process 111. The content of the artificial voice corresponding to each scene is pre-stored, for example, in the memory device 12. Alternatively, processor 11 may perform process 112 using a machine learning model. In other words, the content of the speech produced by the artificial voice may be determined by a machine learning model.
[0039] Processor 11 executes processes 113 and 114 in parallel with processes 111 and 112. Process 113 acquires parameters related to the motion of vehicle 1 as first parameters. Examples of parameters related to the motion of vehicle 1 include the vehicle's speed, acceleration, steering angle, yaw rate, roll angular velocity, pitch angular velocity, yaw angular velocity, etc. Processor 11 can acquire these parameters from the vehicle state sensor 21.
[0040] Next, processor 11 executes process 114. Process 114 determines the second parameter. The second parameter is a parameter relating to the properties of the artificial voice generated by the voice transmission system 100. Examples of the second parameter include parameters representing the loudness (volume, sound pressure, amplitude), pitch (frequency), sound quality (waveform), etc. of the artificial voice. The second parameter is determined in conjunction with the first parameter obtained in process 113.
[0041] Examples of combinations of the first and second parameters are shown in the table in Figure 4. For example, the processor 11 links the acceleration of vehicle 1 with the pitch of the artificial voice, determining the pitch of the artificial voice so that the higher the acceleration of vehicle 1, the higher the pitch of the voice. Alternatively, for example, it links the speed of vehicle 1 with the volume of the artificial voice, determining the volume of the artificial voice so that the higher the speed of vehicle 1, the louder the voice. Alternatively, for example, it links the yaw rate of vehicle 1 with the sound quality of the artificial voice, changing the sound quality of the artificial voice as the yaw rate of vehicle 1 increases or decreases.
[0042] Note that the combinations shown in the table in Figure 4 are just examples, and it is possible to select and combine any parameters as the first and second parameters. Furthermore, the parameters that are linked do not have to be in a one-to-one relationship; multiple second parameters may be linked to a single first parameter. For example, both the pitch and volume of the artificial voice may be changed in conjunction with the acceleration of vehicle 1.
[0043] The combination of the first and second parameters, and the correspondence indicating how the two parameters should work together, are pre-stored in, for example, the memory device 12. Alternatively, the processor 11 may perform processing 114 using a machine learning model. In other words, the second parameter may be determined by the machine learning model in conjunction with the first parameter.
[0044] Refer to Figure 3 again. After the execution of processes 112 and 114, the processor 11 executes process 115. Process 115 generates artificial speech based on the content determined in process 112 and the parameters determined in process 114. In process 115, the processor 11 uses a machine learning model to generate artificial speech that resembles human speech. Note that the method of synthesizing artificial speech that reads fixed lines using a machine learning model is well known and will be omitted here. The generated artificial speech data is sent to the speaker 30, and the artificial speech is output from the speaker 30.
[0045] Through the above processing, the voice transmission system 100 transmits sound. As shown in the above processing, according to the voice transmission system 100 of this embodiment, when transmitting sound, a second parameter relating to the properties of the artificial voice is changed in conjunction with a first parameter relating to the movement of the vehicle 1. This makes it easier for the user to determine that the sound transmitted from the voice transmission system 100 is an artificial voice. In other words, by causing a change in the properties of the sound that does not occur in actual human speech, it becomes easier for the user to notice that the sound emitted by the voice transmission system 100 is an artificial voice, and it is possible to prevent the user from mistakenly perceiving it as human speech.
[0046] Furthermore, according to the voice transmission system 100 of this embodiment, the second parameter is not simply changed, but is changed in conjunction with the first parameter. This makes it easier for the user to recognize that the voice transmitted by the voice transmission system 100 is an artificial voice.
[0047] The movement of vehicle 1, which determines the first parameter, is something that users riding in vehicle 1 can feel. Therefore, users can easily notice that the change in the second parameter is linked to the first parameter. Furthermore, the fact that the properties of the voice change in conjunction with the parameters related to the behavior of vehicle 1 is something that would not normally occur with human voice. Therefore, it becomes easier for users to understand that the change in the second parameter is not merely a matter of chance but is caused by the voice transmission system 100.
[0048] One way to make it easier to distinguish the artificial voice generated by the system from a real human voice is to change the artificial voice into a mechanical voice, making it far removed from a human voice. However, for users receiving the voice, a voice that is closer to a human voice is more likely to be pleasant to the ear than a mechanical voice. In the voice transmission system 100 according to this embodiment, it is possible to improve the ease of distinguishing the artificial voice from a human voice while keeping the artificial voice close to a human voice, thereby ensuring the transparency of the system and increasing user satisfaction.
[0049] 3. Variant 3-1. First variation (Variation with respect to the first parameter) The first parameter may be a parameter other than one related to the motion of vehicle 1. For example, a parameter related to the driver's input of vehicle 1 can be used as the first parameter. Figure 5 is a block diagram showing an example of processing by the voice transmission program when a parameter related to the driver's input of vehicle 1 is used as the first parameter. Here, the computer 10 is connected to the input device 40 and the in-vehicle network in addition to the sensor group 20 and speaker 30. The input device 40 is a device in which the driver of vehicle 1 inputs an input amount. Examples of input amounts that the driver inputs include accelerator input, brake input, steering input, etc. Examples of input devices 40 include the accelerator pedal, brake pedal, steering wheel, etc.
[0050] Processes 221 and 222 are the same as processes 111 and 112 in Figure 3. Also, as in Figure 3, processes 223 and 224 are executed in parallel with processes 121 and 122, and the first parameter is obtained by process 223. However, the first parameter obtained by process 223 is a parameter related to the driver's input of vehicle 1. The processor 11 obtains the driver's input from the input device 40 and uses it as the first parameter, or calculates the first parameter based on the obtained input.
[0051] The first parameter may also represent the driver's input amount itself. For example, the first parameter may be a parameter indicating the magnitude of the accelerator input that the driver inputs to the input device 40.
[0052] Alternatively, the difference between the ideal input and the input actually entered by the driver may be set as the first parameter. The processor 11 can calculate the ideal input based on the information obtained from the sensor group 20. For example, the ideal input may be an input that keeps the speed of vehicle 1 constant. In this case, the processor 11 can calculate the ideal input based on the acceleration of vehicle 1 obtained from the vehicle state sensor 21. Alternatively, for example, the ideal input may be an input that keeps the distance to the preceding vehicle constant. The processor 11 calculates the relative distance between vehicle 1 and the preceding vehicle from the information obtained when the recognition sensor 22 recognizes the preceding vehicle. Then, the accelerator input that keeps the relative distance to the preceding vehicle constant can be set as the ideal input. Alternatively, the processor 11 may calculate an input that takes into account the change in gradient, based on the inclination angle of vehicle 1 obtained from the vehicle state sensor 21, such that the accelerator input is larger when vehicle 1 is traveling on an uphill road surface and the brake input is larger when vehicle 1 is traveling on a downhill road surface.
[0053] The fact that the first parameter is related to the amount of driver input means that the first parameter is related to the operation that the driver, who is the user of vehicle 1, inputs himself. In other words, the first parameter is a parameter related to the amount that the driver can perceive himself. Therefore, it becomes easier for the driver to notice that the second parameter, which is related to the artificial voice, is being changed in conjunction with the first parameter. In this way, it becomes easier for the user to determine that the voice transmitted by the voice transmission system 100 is an artificial voice using a machine learning model, and the transparency of the system can be increased.
[0054] Furthermore, if the difference between the ideal amount of input and the amount of input actually entered by the driver is considered the first parameter, the following effects can also be expected. For the purpose of explanation, consider a specific scenario in which the voice transmission system 100 transmits an artificial voice to alert the driver that their accelerator input is too little or too much. For example, when vehicle 1 approaches an uphill slope, the driver may not realize that their accelerator input is insufficient, causing vehicle 1 to decelerate. In such a scenario, it is conceivable that the system would use an artificial voice to alert the driver to press the accelerator.
[0055] Here, the first parameter is defined as the difference between the ideal accelerator input and the actual accelerator input, and the second parameter is defined as the pitch of the artificial voice. The pitch of the artificial voice will change according to the difference between the driver's accelerator input and the ideal accelerator input. The voice transmission system 100 not only communicates that the accelerator input is insufficient through the content of the artificial voice, but also communicates to the driver the amount of accelerator input the driver should press through the change in the pitch of the artificial voice. In other words, the driver gains the benefit of being able to intuitively grasp the amount of input that is insufficient.
[0056] Another example of the first parameter may be the distance between vehicle 1 and a target located in front of vehicle 1. An example of a target located in front of vehicle 1 is a preceding vehicle.
[0057] However, it is desirable that the first parameter be a parameter that the user can perceive. This is because if the first parameter is one that the user can easily perceive as changing through their senses, the user will be more likely to notice that the first and second parameters are linked, and this will be more effective in preventing the user from mistakenly perceiving the artificial voice as a human voice. For this reason, the first parameter should ideally be a parameter related to the movement of vehicle 1 or the amount of input from the driver of vehicle 1.
[0058] 3-2. Second variation (Variation with respect to the second parameter) The second parameter may be a parameter other than one relating to the properties of the artificial voice. For example, the speech rate of the artificial voice may be used as the second parameter. In this case, for example, the first parameter may be the speed of vehicle 1, and the first and second parameters may be linked so that the speech rate increases as the vehicle speed increases.
[0059] Alternatively, if speaker 30 includes multiple speakers, the second parameter may be a parameter representing the position of the speaker that plays the artificial voice. For example, consider the case where multiple speakers 30 are installed on the left, right, and center of the vehicle interior, as shown in Figure 2. In this case, the first parameter is the steering angle of vehicle 1, and the second parameter is the position of the speaker that plays the artificial voice. The steering angle and the speaker position can be linked so that the sound source moves left or right as the steering angle changes left or right. Specifically, it is assumed that when the steering angle increases to the right, the artificial voice is emitted from the right speaker 30a; when it increases to the left, it is emitted from the left speaker 30c; and when the steering angle is close to 0, it is emitted from the center speaker 30b. This makes it easier for the user to understand the changes in the second parameter, which is linked to the first parameter, and improves the ease with which the user can identify that the voice transmitted from the voice transmission system 100 is an artificial voice.
[0060] 3-3. Third variation (Voice communication as dialogue) The above embodiment describes a case in which the artificial voice transmitted to the user by the voice transmission system 100 is a one-way transmission of voice for purposes such as notifying information. However, the artificial voice transmitted by the voice transmission system 100 may also be intended for dialogue with the user.
[0061] Figure 6 shows an example of processing performed by a voice transmission program when voice intended for interaction with a user is transmitted. In this modified example, the voice transmission system 100 includes a microphone 50 for collecting the voice emitted by the user. When the processor 11 executes processing 311, the user's voice is collected by the microphone 50 and the collected voice is analyzed. The processor 11 may perform voice analysis using a machine learning model.
[0062] Next, processor 11 executes process 312. In process 312, a machine learning model determines the content of the dialogue to respond to the user's voice analyzed in process 311.
[0063] Processes 313, 314, and 315 are the same as processes 113, 114, and 115 in Figure 3. Process 315 generates artificial speech based on the content of the dialogue determined in process 312, and transmits it to the user through speaker 30.
[0064] If the artificial voice is intended for dialogue, the transmission of voice by the voice transmission system 100 is likely to continue for a certain period of time. Therefore, users will be more likely to notice changes in the second parameter of the artificial voice, and the ease with which they can identify that the voice transmitted from the voice transmission system 100 is an artificial voice will be increased.
[0065] 3-4. A fourth variation (Voice communication system as an autonomous driving system) The voice transmission system 100 may be part of the vehicle 1's autonomous driving system. In this case, the autonomous driving device 200 of the vehicle 1 is composed of part or all of the processor 11 and the storage device 12. The autonomous driving device 200 may be a device that performs autonomous driving of the vehicle 1 using a machine learning model.
[0066] 3-5. Fifth variation (upper and lower limits of the second parameter) Upper and lower limits may be set for the second parameter. By setting upper and lower limits, it is possible to prevent the second parameter from changing excessively even if the first parameter changes significantly. In particular, setting upper and lower limits is effective when the second parameter is the volume or pitch of the artificial voice. That is, if the second parameter is the volume of the artificial voice, setting upper and lower limits for the second parameter prevents the volume of the artificial voice from becoming excessively loud or quiet in conjunction with the first parameter, making it difficult for the user to hear. Similarly, if the second parameter is the pitch of the artificial voice, it is possible to prevent the pitch of the artificial voice from becoming excessively high or low, making it difficult for the user to hear.
[0067] The upper and lower limits of the second parameter are set in advance, taking into consideration factors such as ease of listening for the user, and are stored in the storage device 12. If the second parameter exceeds the upper limit or falls below the lower limit, the second parameter will not change even if the first parameter changes significantly further. In this case, the nature of the artificial voice has changed to a level that is unnatural for a human voice, so it is considered unlikely that the user will mistakenly perceive the voice from the voice transmission system 100 as a human voice.
[0068] 4. Examples of application 4-1. First example Finally, a specific example of a scenario in which the voice transmission system 100 according to this embodiment is applied will be described. In the first example, consider a scenario in which vehicle 1 is traveling on a road that is tilted sideways. The first parameter is the roll angle of vehicle 1, and the second parameter is the sound quality of the artificial voice. If the vehicle speed of vehicle 1 is too high, the roll angle becomes large and there is a risk of vehicle 1 overturning, it is assumed that the voice transmission system 100 will transmit a warning voice to the driver. At this time, the driver will be able to more easily recognize that the voice transmitted from the voice transmission system 100 is an artificial voice because the sound quality of the artificial voice changes in conjunction with the roll angle of vehicle 1. At the same time, because the sound quality of the artificial voice changes in conjunction with the size of the roll angle of vehicle 1, the driver will also be able to perceive the size of the roll angle through the change in sound quality, and will be able to more easily recognize how much maneuver is required to return the vehicle to the correct posture.
[0069] 4-2. Second example As a second example, consider a case where vehicle 1 is a vehicle towing a trailer, and the recognition sensor 22 includes a rearview camera that captures images of the area behind vehicle 1. In this case, the first parameter may be a parameter related to the swing of the trailer towed by vehicle 1, and the second parameter may be the volume of the artificial voice. As parameters related to the trailer's swing, for example, the speed, magnitude, and period of the swing may be set as the first parameter. The processor 11 can obtain these parameters related to the trailer's swing by analyzing the images captured by the rearview camera.
[0070] When the trailer towed by vehicle 1 begins to swing from side to side, the voice transmission system 100 is expected to transmit an artificial voice to the user informing them that the trailer is swinging. At this time, for example, the volume of the artificial voice can be linked to the magnitude of the swing, such that it increases as the swing of the trailer increases. For the user, it becomes easy to determine that the voice transmitted from the voice transmission system 100 is an artificial voice by the change in the second parameter, which is linked to the first parameter. At the same time, the magnitude of the swing becomes easier to recognize by the volume of the artificial voice. In other words, as the swing increases and the importance of taking action increases, the volume of the artificial voice increases, so the user can more easily grasp the importance of the situation notified by the artificial voice. [Explanation of symbols]
[0071] 1 Vehicle, 10 Computer, 11 Processor, 12 Storage device, 13 Program storage area, 14 Model data storage area, 20 Sensor group, 21 Vehicle status sensor, 22 Recognition sensor, 23 Position sensor, 30 Speaker, 40 Input device, 50 Microphone, 100 Voice transmission system, 200 Autonomous driving system
Claims
1. A voice transmission system, Equipped with one or more processors, The one or more processors described above are: By generating artificial voices that resemble human voices using machine learning models, The aforementioned artificial voice is transmitted to the vehicle user. A second parameter relating to the properties of the artificial voice is changed in conjunction with a first parameter relating to at least one of the motion of the vehicle and the amount of input from the driver of the vehicle. The aforementioned artificial voice is a voice intended for dialogue with the user, The one or more processors described above are: The voice spoken by the aforementioned user is collected, Determine the content of the dialogue in response to the voice uttered by the user that was collected. The artificial voice is generated based on the content of the determined dialogue. A voice transmission system characterized by the following:
2. A voice transmission system according to claim 1, The first parameter is one of the vehicle's speed, acceleration, or steering angle. A voice transmission system characterized by the following:
3. A voice transmission system according to claim 1 or 2, The second parameter includes at least one of the volume, pitch, and sound quality of the artificial voice. A voice transmission system characterized by the following:
4. A voice transmission system according to claim 3, The first parameter is the acceleration of the vehicle, The second parameter includes the pitch of the artificial voice, The machine learning model increases the pitch of the artificial voice as the acceleration increases. A voice transmission system characterized by the following:
5. A voice transmission system according to claim 3, The second parameter is given an upper and lower limit. A voice transmission system characterized by the following:
6. A voice transmission system according to claim 1, The first parameter is, This is the difference between the ideal amount of input and the amount of input that the driver actually receives. A voice transmission system characterized by the following:
7. A method of voice transmission, The process involves generating artificial speech that resembles human speech using a machine learning model, and The steps include: transmitting the artificial voice to the vehicle user; The steps include changing parameters relating to the properties of the artificial voice in conjunction with parameters relating to at least one of the motion of the vehicle and the amount of input from the driver of the vehicle, and Includes, The aforementioned artificial voice is a voice intended for dialogue with the user, The aforementioned voice transmission method further, The steps include: collecting the voice spoken by the user; The steps include determining the content of the dialogue in response to the collected voice uttered by the user, The steps include generating the artificial voice based on the content of the determined dialogue and including A method of voice transmission characterized by the following features.
8. A voice communication program executed by a computer, The process involves generating artificial speech that resembles human speech using a machine learning model, and The steps include: transmitting the artificial voice to the vehicle user; The steps include changing parameters relating to the properties of the artificial voice in conjunction with parameters relating to at least one of the motion of the vehicle and the amount of input from the driver of the vehicle, and Have the computer run it, The aforementioned artificial voice is a voice intended for dialogue with the user, The aforementioned voice transmission program further, The steps include: collecting the voice spoken by the user; The steps include determining the content of the dialogue in response to the collected voice uttered by the user, The steps include generating the artificial voice based on the content of the determined dialogue and Make the computer execute it. A voice communication program characterized by the following features.