Sound output device and sound output method
By determining whether the recipient can hear the message normally in the sound output device, and only outputting sound when the conditions are met, the problem of missed messages in the prior art is solved, and message notification is realized at the appropriate time.
Patent Information
- Application Number
- CN202010689914.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-17
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2040-07-17
AI Technical Summary
Existing sound output devices continue to output sound even when the recipient cannot hear or concentrate on the message, leading to missed messages.
When a message is received, the sound output device determines whether the recipient is capable of hearing the message normally. It only outputs sound if the condition is met; otherwise, it retains the message sound output.
It effectively prevents messages from being missed, ensuring that the recipient can hear the message normally before outputting their voice, thus avoiding message omission in other situations.
Smart Images

Figure CN114089943B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a sound output device and a sound output method, and particularly to a sound output device and a sound output method having the function of outputting received messages as sound. Background Technology
[0002] Previously, there existed sound output devices that were installed in the interior space of a vehicle or a room in a house, etc., and had the function of receiving messages from chat applications or email systems. When a message was received, it would output the content as sound. Such sound output devices typically output the sound of a message immediately upon receipt, so that the recipient could instantly recognize the content of the message in response to its reception. Furthermore, Patent Document 1 describes an in-vehicle information device that recognizes the content of a vehicle passenger's speech and outputs the speech content in accordance with a method corresponding to the vehicle's surrounding environment or the vehicle's operating status.
[0003] Existing technical documents
[0004] Patent documents
[0005] Patent Document 1: International Publication No. 2014 / 002128 Summary of the Invention
[0006] The problem that the invention aims to solve
[0007] However, in the aforementioned conventional sound output devices, since the sound of the message is output immediately upon receipt, even when the recipient is unable to listen to the message or cannot concentrate on listening to the message, the sound is output in accordance with the receipt of the message, resulting in the recipient missing the message.
[0008] This invention was made to solve such a problem, with the aim of preventing the recipient from missing messages.
[0009] Methods used to solve problems
[0010] In order to solve the above problems, in the present invention, regarding a sound output device provided with a sound output section for sound output and set in a specified space, when a message is received, it determines whether a start condition is met in which an object in the specified space can normally hear the sound. If the start condition is met, the sound output section starts to output the sound of the message. If the start condition is not met, the sound output of the message to be processed by the sound output section is retained.
[0011] Invention Effects
[0012] According to the present invention configured as described above, sound output is not performed immediately in response to the receipt of a message, but only when the recipient is in a state where they can normally hear the message. Otherwise, the sound output of the message is retained, thus preventing the recipient from missing the message. Attached Figure Description
[0013] Figure 1 This is a block diagram illustrating a functional structure example of a sound output device and a portable terminal according to the first embodiment of the present invention.
[0014] Figure 2 This is a flowchart illustrating an example of the operation of the sound output device according to the first embodiment of the present invention.
[0015] Figure 3 This is a block diagram illustrating a functional structure example of a sound output device and portable terminal according to the second embodiment of the present invention.
[0016] Figure 4 This is a flowchart illustrating an example of the operation of the sound output device according to the second embodiment of the present invention.
[0017] Figure 5 This is a block diagram illustrating a functional structure example of a sound output device and portable terminal according to a third embodiment of the present invention.
[0018] Figure 6 This is a flowchart illustrating an example of the operation of the sound output device according to the third embodiment of the present invention.
[0019] Figure 7 This is a block diagram illustrating a functional structure example of a sound output device and portable terminal according to the fourth embodiment of the present invention.
[0020] Figure 8 This is a flowchart illustrating an example of the operation of the sound output device according to the fourth embodiment of the present invention.
[0021] Figure 9 This is a block diagram illustrating a functional structure example of a sound output device and portable terminal according to the fifth embodiment of the present invention.
[0022] Figure 10 This is a flowchart illustrating an example of the operation of the sound output device according to the fifth embodiment of the present invention.
[0023] Explanation of reference numerals in the attached figures
[0024] 1. 1A, 1B, 1C, 1D sound output devices
[0025] 11 Message Receiving Department
[0026] 12, 12A, 12B, 12C, 12D Sound Output Control Unit
[0027] 13 Message Receiving Department Detailed Implementation
[0028] <First Implementation>
[0029] Hereinafter, the first embodiment of the present invention will be described based on the accompanying drawings. Figure 1 This is a block diagram illustrating a functional structure example of the sound output device 1 and the portable terminal 2 connected to the sound output device 1 according to this embodiment. The sound output device 1 is a device installed in the interior space of a vehicle (equivalent to the "specified space" in the claims), and as an example, it enables a so-called in-vehicle navigation system to function as the sound output device 1. In this embodiment, it is assumed that the vehicle equipped with the sound output device 1 is not an autonomous vehicle, but a vehicle that is driven essentially by a driver (however, this includes vehicles that perform a certain level of autonomous driving in a limited environment). Hereinafter, the vehicle equipped with the sound output device 1 will be referred to as "this vehicle".
[0030] As in Figure 1 As shown in the diagram, the sound output device 1 is connected to the sound processing device 3 installed in the vehicle interior. The sound processing device 3 includes a D / A converter and an amplifier, as well as speakers installed in the vehicle. It performs D / A conversion on the input sound signal, amplifies it, and outputs it as sound from the speakers. In addition, the sound output device 1 is connected to the in-vehicle camera 4. The in-vehicle camera 4 is a photographic device installed in the vehicle interior that takes pictures of the area including the driver's seat area at predetermined intervals and outputs the photographic image data based on the photographic results to the sound output control unit 12 (described later) of the sound output device 1.
[0031] Portable terminal 2 is a portable terminal brought into the vehicle by a passenger. For example, it can be a smartphone, a portable phone other than a smartphone, or a tablet computer without telephone functionality. In this embodiment, for simplicity, it is assumed that portable terminal 2 is owned by the driver. Portable terminal 2 is equipped with a text chat application (hereinafter referred to as the "chat application") that allows for the exchange of messages with others in a chat-like manner. In addition to providing a user interface related to text chat on portable terminal 2, the chat application also has the function of receiving message data and sending voice output as text data to voice output device 1 (this function will be described in detail later).
[0032] like Figure 1As shown, the sound output device 1, as a functional structure, includes a sound output device-side communication unit 10, a message receiving unit 11, a sound output control unit 12, and a sound output unit 13. Furthermore, the portable terminal 2, as a functional structure, includes a portable terminal-side communication unit 20 and a chat application execution unit 21. The aforementioned functional blocks 10-13, 20, and 21 can be constructed using hardware, a DSP (Digital Signal Processor), or software. For example, in the case of software construction, the aforementioned functional blocks 10-13, 20, and 21 are actually constructed using a computer's CPU, RAM, ROM, etc., and are implemented through program operations stored in a recording medium such as RAM or ROM, a hard disk, or a semiconductor memory. The same applies to other embodiments described later.
[0033] The audio output device-side communication unit 10 of the audio output device 1 and the portable terminal-side communication unit 20 of the portable terminal 2 perform wireless communication according to the communication specifications stipulated in the relevant wireless communication regulations. The relevant wireless communication specifications may be, for example, Bluetooth (registered trademark) or wireless LAN communication specifications. Alternatively, the audio output device 1 and the portable terminal 2 may be wired together, with the audio output device-side communication unit 10 and the portable terminal-side communication unit 20 performing wired communication according to the communication specifications stipulated in the relevant wired communication regulations. The chat application execution unit 21 of the portable terminal 2 reads and executes the chat application and its associated programs (including programs and APIs part of the OS) through hardware such as a CPU, performing various processing tasks.
[0034] The voice output device 1 has the function of outputting voice information (hereinafter referred to as "message voice") related to the received text chat message data from the portable terminal 2 when the operation mode is message notification mode. The driver (the recipient), as the owner of the portable terminal 2, can confirm the content of the message by listening to the message voice output by the voice output device 1. Furthermore, when both the voice output device 1 and the portable terminal 2 are powered on, the transition to message notification mode is performed by explicit instruction from the driver. The following describes in detail the processing of the portable terminal 2 and the voice output device 1 after the portable terminal 2 receives message data when the operation mode is message notification mode.
[0035] Portable terminal 2 can access network N, including the Internet. Access to network N can be made either directly through the mobile communication network or indirectly through the tethering function of a mobile router. If message data related to text chat is sent to portable terminal 2 from another terminal with a chat application installed, the chat application execution unit 21 of portable terminal 2 receives it. In response to the receipt of message data, the chat application execution unit 21 may sound a tone to notify the user that a message has been received, or display necessary information on the display unit of portable terminal 2, performing pre-set processing, etc.
[0036] In addition to this processing, the chat application execution unit 21 of this embodiment also performs the process of generating text data for voice output and sending it to the portable terminal 2. The text data for voice output is data in which the message text contained in supplementary information and message data is written in text. As will become clear later, the information written as text in the text data for voice output is ultimately output as message voice through the voice output device 1. Furthermore, the supplementary information is information transmitted to supplement matters related to the message before the message text is output as voice. In this embodiment, the supplementary information includes the sender's name and the date and time the message data was received.
[0037] For example, suppose the chat application execution unit 21 receives a "Hello" message from a sender named "AA" at "April 1st, 1:30 PM". In this case, the information recorded as text in the voice output text data would be "October 1st, 1:30 PM. From AA. Hello." Furthermore, if the message text contains markings, stamps, or pictorial characters, the chat application execution unit 21 will perform omissions or transformations into textual information according to predetermined rules.
[0038] The message receiving unit 11 of the voice output device 1 receives the voice output text data sent by the chat application execution unit 21 and saves it to the receive buffer 22. The receive buffer 22 is a storage area formed in the working area of RAM or the like. The voice output text data is the object of message reception by the message receiving unit 11, equivalent to the "message" in the claims. As a result of the above processing between the chat application execution unit 21 and the message receiving unit 11, the voice output text data is saved to the receive buffer 22 in real time corresponding to the message data received by the chat application execution unit 21. Hereinafter, there will be cases where the message receiving unit 11 receives the voice output text data and saves the voice output text data to the receive buffer 22, simply as "the message receiving unit 11 receives a message".
[0039] If the message receiving unit 11 saves the audio output text data into the receiving buffer 22 (i.e., if the message receiving unit 11 receives a message), the audio output control unit 12 performs the following processing. That is, the audio output control unit 12 analyzes the photographic image data input from the in-vehicle camera 4 to determine whether the first starting condition that the driver is present in the vehicle interior space is met.
[0040] Here, while the message notification mode is active, the driver is not always seated in the driver's seat inside the vehicle; there are instances where the driver leaves the vehicle. For example, the driver may leave the vehicle to load (or unload) goods from the trunk, or to make a short shopping trip. Furthermore, if the driver leaves the vehicle and is no longer inside (i.e., the first initial condition is not met), the driver cannot properly hear the sound output from the sound output device 1. On the other hand, if the driver is inside the vehicle (i.e., the first initial condition is met), from the perspective that the driver is within the range capable of hearing the message sound output from the sound output device 1, the driver can properly hear the message sound. Therefore, the first initial condition can be said to be a condition that is met when the driver (the recipient) can properly hear the sound while inside the vehicle.
[0041] The voice output control unit 12 determines whether the first start condition is met using the following method: When a person is seated in the driver's seat, the driver is considered to be present in the vehicle interior; when a person is not seated in the driver's seat, the driver is considered not to be present in the vehicle interior. Accordingly, the voice output control unit 12 determines the driver's seat area from the input photographic image data. The driver's seat area in the photographic image data is preset. Next, the voice output control unit 12 uses existing facial recognition technology to determine whether a face image is contained in the driver's seat area. If a face image is contained, the first start condition is determined to be met; otherwise, the first start condition is determined to be unmet.
[0042] When the first start condition is met, the sound output control unit 12 controls the sound output unit 13 to begin sound output based on the message sound of the text data for sound output stored in the receive buffer 22. Specifically, the sound output control unit 12 outputs a start notification signal to the sound output unit 13. If a start notification signal is input, the sound output unit 13 generates sound data for outputting information that will be written as text in the text data for sound output. For example, the sound data is data of a sound waveform sampled at a predetermined sampling period. The generation of the sound data is performed appropriately using sound synthesis technology or other existing technologies. Next, the sound output unit 13 outputs a sound signal based on the sound data to the sound processing device 3, causing the sound processing device 3 to output the sound (= message sound) based on the sound data. As a result, when the first start condition is met, the sound output of the message sound begins immediately in response to the message received by the message receiving unit 11. Therefore, the driver can immediately recognize the content of the message in response to the message received.
[0043] On the other hand, if the first start condition is not met, the voice output control unit 12 retains the voice output of the message voice from the voice output unit 13. Specifically, the voice output control unit 12 does not output a start notification signal to the voice output unit 13 at that point. Furthermore, the voice output control unit 12 continuously inputs photographic image data from the in-vehicle camera 4 and monitors whether the first start condition is met by continuously analyzing the photographic image data. That is, the voice output control unit 12 monitors whether the driver is seated in the driver's seat. For example, the voice output control unit 12 continuously analyzes the intermittently input photographic image data and monitors whether a new image containing a face is found in the driver's seat area. When such a state is reached, the first start condition is determined to be met.
[0044] When the first starting condition is met, the sound output control unit 12 controls the sound output unit 13 to begin sound output of the sound output text data stored in the receive buffer 22. Specifically, the sound output control unit 12 outputs a start notification signal to the sound output unit 13.
[0045] As described above, if the first start condition is not met when the message receiving unit 11 receives the text data for sound output, the sound output control unit 12 retains the sound output of the message sound performed by the sound output unit 13. Then, when the first start condition is met, the sound output of the message sound begins. Therefore, it is possible to prevent the message sound from being output when the driver is not in the vehicle and thus prevent the driver from missing the message sound. In other words, according to this embodiment, sound output is not performed immediately in response to message reception, but only when the recipient can normally hear the message. In other cases, the sound output of the message is retained, thus preventing the recipient from missing the message.
[0046] Next, a flowchart will be used to explain an example of the operation of the sound output device 1 of this embodiment. Figure 2 This is a flowchart illustrating an example of the operation of the sound output device 1 after receiving text data for sound output from the portable terminal 2. For example... Figure 2 As shown, the message receiving unit 11 of the sound output device 1 receives the text data for sound output and saves it to the receiving buffer 22 (step SA1). If the text data for sound output is saved in the receiving buffer 22, the sound output control unit 12 determines whether the first start condition is met based on the photographic image data of the in-vehicle camera 4 (step SA2). As described above, the first start condition is that a person is sitting in the driver's seat.
[0047] If the first start condition is met (step SA2: YES), the sound output control unit 12 controls the sound output unit 13 to start the sound output of the message sound (step SA3). If the first start condition is not met (step SA2: NO), the sound output control unit 12 retains the sound output of the message sound from the sound output unit 13 (step SA4). Next, the sound output control unit 12 monitors whether the first start condition is met (step SA5). If the first start condition is met (step SA5: YES), the sound output control unit 12 controls the sound output unit 13 to start the sound output of the message sound (step SA6).
[0048] <First Variation of the First Embodiment>
[0049] Next, a first variation of the first embodiment described above will be explained. Furthermore, in the following description, the sound output control unit 12 (or the component corresponding to the sound output control unit 12 in embodiments other than the first embodiment) will cause the sound output unit 13 to start outputting message audio, which will only appear as "the sound output control unit 12 starts outputting audio." In addition, there are cases where the sound output control unit 12 retains the output of message audio from the sound output unit 13, which will only appear as "the sound output control unit 12 retains the audio output."
[0050] In the first embodiment described above, the recipient of the message audio is the driver. In this modified example, the recipient is someone other than the driver. Hereinafter, a brief description of an example of the structure and processing of the audio output device 1 in this modified example will be provided. First, the imaging area of the in-vehicle camera 4 is set to cover the entire interior space of the vehicle. In particular, the in-vehicle camera 4 is positioned so that it can capture a person's face from the front, regardless of which seat the person is seated in the vehicle. Multiple in-vehicle cameras 4 may also be installed. Furthermore, image data (hereinafter referred to as "face image data") storing an image of the face of the recipient is pre-registered in the audio output device 1. Additionally, the recipient is essentially the owner of the portable terminal 2.
[0051] Furthermore, when the message receiving unit 11 receives a message, the voice output control unit 12 determines whether the first starting condition—that the target person exists in the vehicle interior space—is met. Specifically, the voice output control unit 12 determines whether the photographic image data input from the in-vehicle camera 4 contains an image of the face of the same person as the face represented by the registered face image data. This determination is performed based on existing face recognition technology. As a simple example, the voice output control unit 12 performs existing multi-dimensional vector comparison processing on the feature vectors of the face image represented by the registered face image data and the feature vectors of the face images contained in the photographic image data to calculate the approximation. If the photographic image data contains a face image with an approximation of a threshold or higher, it determines that the photographic image data contains an image of the face of the same person as the face represented by the registered face image data. If the photographic image data contains an image of the face of the same person as the face represented by the registered face image data, the voice output control unit 12 determines that the first starting condition is met; otherwise, it determines that the first starting condition is not met.
[0052] The operation of the sound output device 1 after determining whether the first starting condition is met is the same as in the first embodiment described above. According to the structure of this modified example, similar to the first embodiment described above, the sound output is not performed immediately in response to the reception of the message, but only when the recipient can normally hear the message. Otherwise, the sound output of the message is retained, thus preventing the recipient from missing the message.
[0053] <Other variations of the first embodiment>
[0054] Next, other variations of the first embodiment will be described. In the first embodiment described above, it is assumed that the vehicle is not an autonomous vehicle, but it may be. Furthermore, in the first embodiment described above, the voice output control unit 12 determines whether the first start condition is met by analyzing the photographic image data from the in-vehicle camera 4. In this regard, it is also possible to set up a structure that determines whether the first start condition is met using other methods. As an example, it is also possible to set up a sensor (weight sensor, switch sensor, optical sensor, etc.) on the driver's seat to detect whether the driver is seated, and the voice output control unit 12 determines whether the first start condition is met or not based on the sensor's detection value.
[0055] <Second Implementation Method>
[0056] Next, the second embodiment will be described. Figure 3 This is a block diagram illustrating a functional structure example of the sound output device 1A according to this embodiment. Figure 1 and Figure 3 A comparison reveals that the sound output device 1A of this embodiment replaces the sound output control unit 12A of the first embodiment. Furthermore, in this embodiment, similar to the first embodiment, it is assumed that the vehicle is not an autonomous vehicle, and the recipient of the message sound is the driver.
[0057] When the voice output control unit 12A receives a message from the message receiving unit 11, it performs the following processing: The voice output control unit 12A determines whether the vehicle is parked. This determination is made, for example, by determining whether the handbrake is engaged. If the vehicle is not parked, the voice output control unit 12A does not perform the determination of whether the second start condition described below is met and begins voice output.
[0058] On the other hand, when the vehicle is parked, the sound output control unit 12A analyzes the photographic image data input from the in-vehicle camera 4 to determine whether the second start condition—that the driver sitting in the driver's seat is not asleep—is met. At the time of determining whether the second start condition is met, the vehicle is parked, and the driver may be sleeping to rest. Furthermore, if the driver is asleep (i.e., the first start condition is not met), the driver cannot normally hear the sound output by the sound output device 1A. On the other hand, if the driver is not asleep (i.e., the first start condition is met), the driver is awake and can normally hear the message sound output by the sound output device 1A. Therefore, the second start condition can be said to be a condition that is met when the driver (the subject) can normally hear the sound in the vehicle interior.
[0059] The voice output control unit 12A determines whether the second start condition is met based on existing technology. In a simple example, the voice output control unit 12A, based on existing facial recognition technology, identifies an image of a person's face in the driver's seat area from the input photographic image data and treats this identified face image as the driver's face. Furthermore, the voice output control unit 12A analyzes the identified face image to determine whether the driver is asleep. For example, the voice output control unit 12A continuously analyzes photographic image data for a certain period (e.g., 5 seconds) to determine whether the closed eyes state has persisted for a certain period of time. If the closed eyes state has persisted for a certain period of time, the voice output control unit 12A determines that the driver is asleep and the second start condition is not met; otherwise, the second start condition is met.
[0060] If the second start condition is met, the sound output control unit 12A begins sound output. Conversely, if the second start condition is not met, the sound output control unit 12A retains sound output. Furthermore, the sound output control unit 12A continuously analyzes the photographic image data input from the in-vehicle camera 4, monitoring whether the second start condition is met. That is, the sound output control unit 12A monitors whether the driver is not asleep (i.e., awake). If the second start condition is met, the sound output control unit 12A begins sound output.
[0061] As described above, if the second start condition is not met when the message is received by the message receiving unit 11, the sound output control unit 12A retains the sound output, and then starts the sound output when the second start condition is met. Therefore, it is possible to prevent the message sound from being output when the driver is asleep, thus preventing the driver from missing the message sound. That is, according to this embodiment, the sound output is not performed immediately in response to the message reception, but only when the recipient can normally hear the message, and the sound output of the message is retained in other situations, so it is possible to prevent the recipient from missing the message.
[0062] Next, a flowchart will be used to explain an example of the operation of the sound output device 1A of this embodiment. Figure 4 This is a flowchart illustrating an example of the operation of the sound output device 1A after receiving text data for sound output from the portable terminal 2. For example... Figure 4 As shown, the message receiving unit 11 of the sound output device 1A receives the text data for sound output and saves it to the receiving buffer 22 (step SB1). If the text data for sound output is saved in the receiving buffer 22, the sound output control unit 12A determines whether the vehicle is parked (step SB2). If the vehicle is not parked (step SB2: NO), the sound output control unit 12A controls the sound output unit 13 to start the sound output of the message (step SB3). After the processing in step SB3, the flowchart ends. On the other hand, if the vehicle is parked (step SB2: YES), the sound output control unit 12A determines whether the second start condition is met based on the photographic image data from the in-vehicle camera 4 (step SB4). As described above, the second start condition is that the driver is not asleep.
[0063] If the second start condition is met (step SB4: YES), the sound output control unit 12A controls the sound output unit 13 to start outputting a message sound (step SB5). If the second start condition is not met (step SB4: NO), the sound output control unit 12A retains the message sound output by the sound output unit 13 (step SB6). Next, the sound output control unit 12A monitors whether the second start condition is met (step SB7). If the second start condition is met (step SB7: YES), the sound output control unit 12A controls the sound output unit 13 to start outputting a message sound (step SB8).
[0064] <First variation of the second embodiment>
[0065] Next, a first variation of the second embodiment will be described. In this variation, the vehicle is an autonomous vehicle with fully automated driving capabilities. Furthermore, it is assumed that the physical, technical, or legal conditions related to fully automated driving are in place, and the vehicle can operate in fully automated driving mode on public roads. The driver is also permitted to sleep during fully automated driving.
[0066] Incidentally, regarding this modified example, if the message receiving unit 11 receives a message when the vehicle is driving in fully automated mode, the voice output control unit 12A does not determine whether the vehicle is parked, but instead determines whether the second starting condition is met. This is because, in this modified example, not only when the vehicle is parked, but also when the vehicle is driving in fully automated mode, the driver may be asleep. The operation of the voice output device 1A after the determination is the same as in the second embodiment. According to this modified example, when the vehicle is driving in fully automated mode, it is possible to prevent the driver from missing message sounds because the driver is asleep.
[0067] <Other variations of the second embodiment>
[0068] Next, other variations of the second embodiment will be described. Regarding the second embodiment, it is also possible to configure the voice output control unit 12A to start voice output only when both the first and second start conditions are met when the message receiving unit 11 receives a message, and to retain voice output otherwise. Furthermore, in the second embodiment described above, the driver is the recipient of the message voice, but the recipient is not limited to the driver. When someone other than the driver is the recipient, the voice output control unit 12A can use the technology of the first variation of the first embodiment to determine whether the recipient, who is not the driver, is asleep based on the photographic image data from the in-vehicle camera 4. Furthermore, in the second embodiment described above, the voice output control unit 12A determines whether the second start condition is met by analyzing the photographic image data from the in-vehicle camera 4. Regarding this, it is also possible to configure it to determine whether the second start condition is met using other methods. For example, the voice output control unit 12A could obtain the driver's biometric information (pulse waves or brain waves, etc.), determine whether the driver is asleep based on the biometric information, and determine the fulfillment / non-fulfillment of the second start condition based on this.
[0069] <Third Implementation Method>
[0070] Next, the third embodiment will be described. Figure 5This is a block diagram illustrating a functional structure example of the sound output device 1B according to this embodiment. In this embodiment, it is assumed that the vehicle is not an autonomous vehicle. Furthermore, in this embodiment, the person listening to the message sound is the driver, and the subject conducting the hands-free call described later is also the driver. Figure 1 and Figure 5 A comparison reveals that the sound output device 1B of this embodiment replaces the sound output control unit 12B of the first embodiment. Furthermore, the sound output device 1B of this embodiment includes a hands-free calling execution unit 30. The hands-free calling execution unit 30 is a functional block that works in conjunction with the portable terminal 2 to implement hands-free calling. Appropriately provided equipment necessary for implementing hands-free calling (e.g., a microphone for inputting voice). During hands-free calling, the hands-free calling execution unit 30 outputs a signal indicating that a hands-free call is in progress to the sound output control unit 12B.
[0071] When the voice output control unit 12B receives a message from the message receiving unit 11, it performs the following processing: The voice output control unit 12B determines whether the third starting condition—that the driver is not currently on a phone call—is met. Furthermore, in this embodiment, the driver may make hands-free calls while driving the vehicle (and may also make hands-free calls when not driving), or may make calls using a device with telephone functionality other than their own portable terminal 2 when the vehicle is parked. Additionally, in this embodiment, the portable terminal 2 can be used as a telephone while maintaining its original operating mode as message notification mode.
[0072] Here, when the driver is on a phone call (i.e., when condition 3 is not met), the driver cannot concentrate on listening to the sound output from the sound output device 1B. On the other hand, when the driver is not on a phone call (i.e., when condition 3 is met), from the perspective of being able to concentrate on listening to the message sound without being hindered by responding to the phone call, the driver can listen to the message sound normally. Therefore, condition 3 can be said to be a condition that is met when the driver (the subject) can listen to the sound normally in the vehicle space.
[0073] The voice output control unit 12B determines whether the third start condition is met using the following method: If a signal indicating a hands-free call is input from the hands-free call execution unit 30 at the time the message receiving unit 11 receives a message, the voice output control unit 12B determines that the driver is on a hands-free call (i.e., in a phone call), and the third start condition is not met. Furthermore, the voice output control unit 12B analyzes the photographic image data input from the in-vehicle camera 4 to determine whether the driver is on a phone call. Here, when the driver is on a phone call, actions typically involve the driver holding the phone to their ear or moving their mouth to speak while wearing headphones—actions unique to being on a phone call. Accordingly, the voice output control unit 12B analyzes the photographic image data using image analysis technology that corresponds to the style images of actions unique to being on a phone call to determine whether the driver is on a phone call. If the voice output control unit 12B determines that the driver is on a phone call based on the analysis of the photographic image data, the third start condition is not met. Based on the two points above, the sound output control unit 12B determines that the third start condition is met if it is not determined that the third start condition is not met.
[0074] If the third start condition is met, the voice output control unit 12B begins voice output. Conversely, if the third start condition is not met, the voice output control unit 12B maintains voice output while monitoring whether the third start condition is met. That is, the voice output control unit 12B monitors whether the driver is not on a phone call. If the third start condition is detected to be met, the voice output control unit 12B begins voice output.
[0075] As described above, if the third start condition is not met when the message is received by the message receiving unit 11, the sound output control unit 12B retains the sound output. Then, when the third start condition is met, the sound output begins. This structure prevents situations where the message sound is output while the driver is responding to a phone call, and the driver cannot concentrate on listening to the message sound and thus misses it. In other words, according to this embodiment, sound output is not performed immediately in response to message reception, but only when the recipient can normally listen to the message; otherwise, the message sound output is retained, thus preventing the recipient from missing the message.
[0076] Next, a flowchart will be used to explain an example of the operation of the sound output device 1B of this embodiment. Figure 6 This is a flowchart illustrating an example of the operation of the sound output device 1B after receiving text data for sound output from the portable terminal 2. For example... Figure 6As shown, the message receiving unit 11 of the voice output device 1B receives the voice output text data and saves it to the receiving buffer 22 (step SC1). If the voice output text data is saved in the receiving buffer 22, the voice output control unit 12B determines whether the third start condition is met based on the input from the hands-free calling execution unit 30 and the photographic image data from the in-vehicle camera 4 (step SC2). As described above, the third start condition is that the driver is not in a telephone conversation.
[0077] If the third start condition is met (step SC2: YES), the sound output control unit 12B controls the sound output unit 13 to start the sound output of the message sound (step SC3). If the third start condition is not met (step SC2: NO), the sound output control unit 12B retains the sound output of the message sound from the sound output unit 13 (step SC4). Next, the sound output control unit 12B monitors whether the third start condition is met (step SC5). If the third start condition is met (step SC5: YES), the sound output control unit 12B controls the sound output unit 13 to start the sound output of the message sound (step SC6).
[0078] <Modifications of the Third Embodiment>
[0079] Next, a variation of the third embodiment will be described. In the third embodiment described above, the vehicle is not an autonomous vehicle, but it could of course be an autonomous vehicle. Furthermore, regarding the third embodiment described above, it can also be structured as follows: when the message receiving unit 11 receives a message, the voice output control unit 12B immediately starts voice output only when any combination of the third start condition and the first and second start conditions (the combination includes a combination consisting of one condition and a combination consisting of two conditions) is met, and retains voice output otherwise. Furthermore, in the third embodiment described above, the driver is the recipient of the message voice, but the recipient is not limited to the driver. When a person other than the driver is the recipient, the voice output control unit 12B can use the technology of the first variation of the first embodiment to determine whether a person other than the driver is on a phone call based on the photographic image data of the in-vehicle camera 4. Alternatively, it is also possible to determine whether the driver is on a phone call using a method other than the method described in the third embodiment described above.
[0080] <Fourth Implementation Method>
[0081] Next, the fourth embodiment will be described. Figure 7 This is a block diagram illustrating a functional structure example of the sound output device 1C according to this embodiment. In this embodiment, it is assumed that the person listening to the message sound is a driver, and the vehicle is not an autonomous vehicle. Figure 1 and Figure 7 As can be seen from the comparison, the sound output device 1C of this embodiment replaces the sound output control unit 12 of the first embodiment and has a sound output control unit 12C.
[0082] When a message is received by the message receiving unit 11, the sound output control unit 12C performs the following processing: The sound output control unit 12C determines whether the fourth start condition—that the driver is not engaged in conversation—is met. "The driver is engaged in conversation" means that there is a passenger other than the driver in the vehicle, and the driver is talking to that passenger. Here, when the driver is engaged in conversation (i.e., the fourth start condition is not met), the driver cannot concentrate on listening to the sound output by the sound output device 1C. On the other hand, when the driver is not engaged in conversation (i.e., the fourth start condition is met), the driver can normally listen to the message sound from the perspective of being able to concentrate on listening to the message sound without being hindered by conversation. Therefore, the fourth start condition can be said to be a condition that is met when the driver (the recipient) can normally listen to the sound in the vehicle interior space.
[0083] The voice output control unit 12C determines whether the fourth start condition is met by the following method: The voice output control unit 12C determines an image of the driver's face from the photographic image data input from the in-vehicle camera 4 using the method described in the first embodiment. Next, the voice output control unit 12C determines the mouth region in the face image and tracks and analyzes the mouth region for a certain period (e.g., 5 seconds). Based on the analysis of the mouth region, if the mouth does not move continuously for a certain period (either in its most closed state or its open state), the voice output control unit 12C determines that the driver is not talking, and the fourth start condition is met; otherwise, if the driver talks, the fourth start condition is not met.
[0084] If the fourth start condition is met, the sound output control unit 12C begins sound output. Conversely, if the fourth start condition is not met, the sound output control unit 12C maintains sound output while monitoring whether the fourth start condition is met. That is, the sound output control unit 12C monitors whether the driver is not engaged in conversation. For example, the sound output control unit 12C continuously analyzes photographic image data and continuously monitors the driver's mouth movements. If the mouth does not move continuously for a certain period of time, it is determined that the fourth start condition is met. Upon detecting that the fourth start condition is met, the sound output control unit 12C begins sound output.
[0085] As described above, if the fourth start condition is not met when the message is received by the message receiving unit 11, the sound output control unit 12C retains the sound output. Then, when the fourth start condition is met, the sound output begins. Because of this structure, it is possible to prevent the message sound from being output during driver conversation, thus preventing the driver from missing the message sound due to a lack of focus. In other words, according to this embodiment, sound output is not performed immediately in response to message reception, but only when the recipient can normally hear the message. Otherwise, the message sound output is retained, thus preventing the recipient from missing the message.
[0086] Next, a flowchart will be used to explain an example of the operation of the sound output device 1C of this embodiment. Figure 8 This is a flowchart illustrating an example of the operation of the sound output device 1C after receiving text data for sound output from the portable terminal 2. For example, by... Figure 8 As indicated, the message receiving unit 11 of the sound output device 1C receives the text data for sound output and saves it to the receiving buffer 22 (step SD1). If the text data for sound output is saved to the receiving buffer 22, the sound output control unit 12C determines whether the fourth start condition is met based on the image data captured by the in-vehicle camera 4 (step SD2). As described above, the fourth start condition is that the driver is not engaged in conversation.
[0087] If the fourth start condition is met (step SD2: YES), the sound output control unit 12C controls the sound output unit 13 to start outputting the message sound (step SD3). If the fourth start condition is not met (step SD2: NO), the sound output control unit 12C retains the message sound output by the sound output unit 13 (step SD4). Next, the sound output control unit 12C monitors whether the fourth start condition is met (step SD5). If the fourth start condition is met (step SD5: YES), the sound output control unit 12C controls the sound output unit 13 to start outputting the message sound (step SD6).
[0088] <First variation of the fourth embodiment>
[0089] Next, a first variation of the fourth embodiment will be described. In the fourth embodiment described above, the sound output control unit 12C determines whether the fourth start condition—that the driver is not talking—is met. Regarding this, the sound output control unit 12C in this variation determines whether the fourth start condition—that the driver is not talking loudly—is met. Specifically, the sound output device 1C is connected to a microphone that receives the driver's voice. For the sound output control unit 12C, a sound pressure level is input from the sound processing circuit that processes the input from the microphone; this sound pressure level is the sound pressure level of the sound input to the microphone.
[0090] Furthermore, when the message receiving unit 11 receives a message, the sound output control unit 12C determines whether the fourth start condition—that the driver is not talking loudly—is met. Specifically, if the driver is talking (whether the driver is talking is determined by the method described in the fourth embodiment above), and if the input sound pressure level is above a threshold, the sound output control unit 12C considers the fourth start condition not met; otherwise, it determines that the fourth start condition is met. Additionally, if the driver is talking and the sound pressure level of the sound input to the microphone is above a threshold, it can be considered that the driver is talking at a certain volume or higher.
[0091] Here, when the driver speaks softly, compared to speaking loudly, the level of concentration on the conversation is lower. It can be said that the driver is more likely to hear the audio message without missing it. Therefore, according to this variation, when the driver speaks softly, since the audio message is output immediately in response to message reception, it prevents the driver from experiencing the unpleasant feeling that might occur if the audio message is retained even though it is audible.
[0092] <Other variations of the fourth embodiment>
[0093] Next, a variation of the fourth embodiment will be described. Regarding the fourth embodiment, it can also be structured as follows: when the message receiving unit 11 receives a message, the voice output control unit 12C immediately starts voice output only when any combination of the fourth start condition and any combination of the first to third start conditions (including combinations consisting of one condition, two conditions, and three conditions) is met; otherwise, the voice output is retained. Alternatively, the content of the fourth start condition can be set to the content of the fourth start condition of the first variation of the fourth embodiment. Furthermore, in the above-described fourth embodiment, the vehicle is not an autonomous vehicle, but it could certainly be one. Furthermore, in the above-described fourth embodiment, the driver is the recipient of the message voice, but the recipient is not limited to the driver. In this case, the voice output control unit 12C can use the technology of the first variation of the first embodiment to determine whether a recipient other than the driver is engaged in conversation based on the photographic image data from the in-vehicle camera 4. Furthermore, the method for determining whether the driver is engaged in conversation is not limited to the method exemplified in the above-described fourth embodiment; any method employing existing technology can be used.
[0094] <Fifth Implementation>
[0095] Next, the fifth embodiment will be described. Figure 9 This is a block diagram illustrating a functional structure example of the sound output device 1D according to this embodiment. In this embodiment, it is assumed that the person listening to the message sound is a driver, and that the vehicle is not an autonomous vehicle. Figure 1 and Figure 9 A comparison reveals that the sound output device 1D of this embodiment replaces the sound output control unit 12D of the first embodiment. When the message receiving unit 11 receives a message, the sound output control unit 12D performs the following processing: It determines whether the fifth start condition—that the driver is in a relaxed state—is met. Here, if the driver is not in a relaxed state (i.e., the fifth start condition is not met), the driver may not be able to concentrate on listening to the sound output by the sound output device 1D. On the other hand, if the driver is in a relaxed state (i.e., the fifth start condition is met), from the viewpoint that the driver can concentrate on listening to the message sound output by the sound output device 1D in a calm state, the driver can listen to the message normally. Therefore, the fifth start condition can be said to be a condition that is met when the driver (the recipient) can listen to the sound normally in the vehicle interior space.
[0096] The voice output control unit 12D determines whether the fifth start condition is met using the following method: The voice output control unit 12D, based on the photographic image data input from the in-vehicle camera 4, determines an image of the driver's face using the method described in the first embodiment, and determines whether the driver is in a relaxed state based on existing facial expression recognition technology. Furthermore, the voice output control unit 12D determines that the fifth start condition is met if the driver is in a relaxed state, and determines that the fifth start condition is not met if the driver is not in a relaxed state. Additionally, in this embodiment, given the characteristic that the driver is driving the vehicle, there are times when the driver is focused on driving; in such cases, it is expected that the fifth start condition will be determined to be not met. Accordingly, various parameters of the facial expression recognition module based on facial expression recognition technology are appropriately adjusted so that when the driver is focused on driving, the driver is determined to be in a non-relaxed state.
[0097] If the fifth start condition is met, the sound output control unit 12D begins sound output. Conversely, if the fifth start condition is not met, the sound output control unit 12D maintains sound output while monitoring whether the fifth start condition is met. That is, the sound output control unit 12D monitors whether the driver is in a relaxed state. If the fifth start condition is detected to be met, the sound output control unit 12D begins sound output.
[0098] As described above, if the fifth start condition is not met when the message is received by the message receiving unit 11, the sound output control unit 12D retains the sound output. Then, when the fifth start condition is met, the sound output begins. This structure prevents the message sound from being output when the driver is not relaxed, thus preventing the driver from missing the message sound due to inattention. In other words, according to this embodiment, since the sound output is not performed immediately in response to message reception, but only when the recipient is able to hear the message normally, and the sound output is retained in other situations, it is possible to prevent the recipient from missing the message.
[0099] Next, a flowchart will be used to explain an example of the operation of the sound output device 1D of this embodiment. Figure 10 This is a flowchart illustrating an example of the operation of the sound output device 1D after receiving text data for sound output from the portable terminal 2. For example, in... Figure 10As indicated in the diagram, the message receiving unit 11 of the sound output device 1D receives the text data for sound output and saves it to the receiving buffer 22 (step SE1). If the text data for sound output is saved to the receiving buffer 22, the sound output control unit 12D determines whether the fifth start condition is met based on the photographic image data from the in-vehicle camera 4 (step SE2). As described above, the fifth start condition is that the driver is in a relaxed state.
[0100] If the fifth start condition is met (step SE2: YES), the sound output control unit 12D controls the sound output unit 13 to start outputting the message sound (step SE3). If the fifth start condition is not met (step SE2: NO), the sound output control unit 12D retains the message sound output by the sound output unit 13 (step SE4). Next, the sound output control unit 12D monitors whether the fifth start condition is met (step SE5). If the fifth start condition is met (step SE5: YES), the sound output control unit 12D controls the sound output unit 13 to start outputting the message sound (step SE6).
[0101] <Modifications of the 5th Embodiment>
[0102] Next, a variation of the fifth embodiment will be described. Regarding the fifth embodiment, it can also be structured as follows: when the message receiving unit 11 receives a message, the sound output control unit 12D immediately starts sound output only when any combination of the fifth start condition and any combination of the first to fourth start conditions (including combinations consisting of one condition, two conditions, three conditions, and four conditions) is met; otherwise, sound output is retained. In this case, the content of the fourth start condition can also be set to the content of the fourth start condition of the first variation of the fourth embodiment.
[0103] Furthermore, in the fifth embodiment described above, the vehicle is not an autonomous vehicle, but it could certainly be one. In the case of an autonomous vehicle, for example, when the driver is focused on reading, the message audio can be prevented from being output. Also, in the fifth embodiment described above, the driver is the recipient of the message audio, but the recipient is not limited to the driver. In this case, the audio output control unit 12D can utilize the technology of the first variation of the first embodiment to determine whether the recipient, who is not the driver, is in a relaxed state based on the photographic image data from the in-vehicle camera 4.
[0104] Furthermore, the method for determining whether the driver is in a relaxed state is not limited to the method exemplified in the fifth embodiment described above, and any method employing existing technology can be used. For example, it could be a structure that obtains the driver's biometric information (pulse waves or brain waves, etc.) and determines whether the driver is in a relaxed state based on the biometric information. Alternatively, it could be a structure where the voice output control unit 12D identifies the vehicle's condition, and based on the vehicle's condition, if the driver is in a focused driving state (or is required to be focused on driving), it considers the driver not to be in a relaxed state and determines that the fifth starting condition is not met. The vehicle's condition could be a traffic jam, or the vehicle is about to enter an intersection, is entering an intersection, is in motion for parking, or is frequently accelerating and decelerating, etc.
[0105] The above description illustrates embodiments (including variations) of the present invention. However, each of the above embodiments is merely a specific example of implementing the present invention and is not intended to limit the scope of the invention. That is, the present invention can be implemented in various forms without departing from its spirit or main features.
[0106] For example, in the first embodiment described above, the sound output control unit 12 starts sound output when the first starting condition "the driver is present in the vehicle interior" is met, and retains sound output when the first starting condition is not met. This processing order is referred to as the "first processing order." In this regard, the processing order in which the condition "not present in the vehicle interior" is defined as the negation of the first starting condition, and the sound output control unit 12 retains sound output when this condition is met, and starts sound output when this condition is not met, is synonymous with the first processing order. That is, such a processing order is naturally included within the scope of the present invention. The same applies to other embodiments (including variations).
[0107] Furthermore, in the first embodiment, a chat application is installed in the portable terminal 2, but it could also be structured as follows: the chat application is installed in the voice output device 1, and the voice output device 1 has a function to access the network N, and the message receiving unit 11 of the voice output device 1 directly receives message data related to a message sent by a specified terminal. Furthermore, in the first embodiment, the chat application execution unit 21 of the portable terminal 2 generates text data for voice output, but it could also be structured as follows: when the chat application execution unit 21 receives message data, it does not generate text data for voice output, but instead sends the message data to the voice output device 1. In these structures, the message data received by the message receiving unit 11 is equivalent to the "message" in the claims. In these cases, for example, it could be structured such that when the message receiving unit 11 receives message data, it generates text data for voice output and saves it to the receiving buffer 22, or it could be structured such that when the message receiving unit 11 receives message data, it saves the message data to the receiving buffer 22, and the voice output unit 13 reads the message data when a start notification signal is input from the voice output control unit 12, and generates text data for voice output based on the message data. The same applies to other implementation methods (including variations).
[0108] Furthermore, in the first embodiment, if the first start condition is not met when the message receiving unit 11 receives the message, the sound output control unit 12 temporarily holds the sound output, and then starts the sound output when the first start condition is met. Regarding this, the following structure is also possible: after holding the sound output, the sound output control unit 12 does not automatically start the sound output, but instead outputs a warning that a message has been received, and starts the sound output only when there is an explicit instruction from the driver (however, it could also be someone other than the driver). The warning is, for example, by displaying predetermined information on the display unit when the sound output device 1 has a display unit, or by lighting / flashing the LED in a predetermined pattern when the housing of the sound output device 1 has an LED. The same applies to other embodiments (including variations).
[0109] Furthermore, in the above embodiment, the space where the sound output device 1 is installed is a vehicle interior space, which corresponds to the "specified space" in the claims. However, the space where the sound output device 1 is installed and where the message sound delivered by the recipient is heard is not limited to a vehicle interior space. As an example, it could also be a room in a residence or a room in an office, etc. The same applies to other embodiments (including variations).
[0110] Furthermore, regarding the first embodiment described above, it is also possible to configure a structure in which all or part of the processing performed by each functional block described as the sound output device 1 is performed by an external device connected to the sound output device 1. The external device could be, for example, a portable terminal 2, or, for example, a cloud server connected to a network N. In this case, the sound output device 1 and the external device cooperate to function as a "sound output device." The same applies to other embodiments (including variations).
[0111] Furthermore, in the above embodiments, messages are exchanged in text chat, but the messages are not limited to this; for example, they could also be emails.
Claims
1. A sound output device, characterized in that, It is placed in a designated space and equipped with a sound output unit for sound output. The aforementioned sound output device also includes: The message receiving unit receives messages; and When the message receiving unit receives a message, the sound output control unit determines an image of the face of the subject in the specified space using face recognition technology, analyzes the determined face image, identifies actions specific to the subject when they are unable to hear normally, and determines whether a start condition is met for the subject to be able to hear normally in the specified space. If the start condition is met, the sound output unit starts outputting the message; otherwise, the sound output of the message to be processed by the sound output unit is retained.
2. The sound output device as described in claim 1, characterized in that, After the sound output control unit retains the sound output of the message to be processed by the sound output unit, it monitors whether the start condition is met. If the condition is met, it causes the sound output unit to start the sound output of the message.
3. The sound output device as described in claim 1, characterized in that, As a criterion for determining whether the aforementioned starting condition is met, the aforementioned sound output control unit determines whether the condition that the aforementioned object exists in the aforementioned defined space is met.
4. The sound output device as described in claim 3, characterized in that, After the sound output control unit retains the sound output of the message to be processed by the sound output unit, it monitors whether the start condition is met. If the condition is met, it causes the sound output unit to start the sound output of the message.
5. The sound output device as described in claim 1, characterized in that, As a criterion for determining whether the above-mentioned starting condition is met, the above-mentioned sound output control unit determines whether the condition that the subject is not in a sleeping state is met.
6. The sound output device as described in claim 5, characterized in that, After the sound output control unit retains the sound output of the message to be processed by the sound output unit, it monitors whether the start condition is met. If the condition is met, it causes the sound output unit to start the sound output of the message.
7. The sound output device as described in claim 1, characterized in that, As a criterion for determining whether the aforementioned starting condition is met, the aforementioned voice output control unit determines whether the condition that the subject is not making a phone call is met.
8. The sound output device as described in claim 7, characterized in that, After the sound output control unit retains the sound output of the message to be processed by the sound output unit, it monitors whether the start condition is met. If the condition is met, it causes the sound output unit to start the sound output of the message.
9. The sound output device as described in claim 1, characterized in that, As a criterion for determining whether the aforementioned starting condition is met, the aforementioned voice output control unit determines whether the condition that the subject is not in conversation is met.
10. The sound output device as described in claim 9, characterized in that, As a criterion for determining whether the aforementioned starting condition is met, the aforementioned sound output control unit determines whether the condition that the subject is not speaking loudly is met.
11. The sound output device as described in claim 9, characterized in that, After reserving the audio output of the message to be processed by the audio output unit, the audio output control unit monitors whether the start condition is met. If the condition is met, the audio output unit starts the audio output of the message.
12. The sound output device as claimed in claim 1, characterized in that, As a criterion for determining whether the aforementioned starting condition is met, the aforementioned sound output control unit determines whether the condition that the subject is in a relaxed state is met.
13. The sound output device as described in claim 12, characterized in that, After reserving the audio output of the message to be processed by the audio output unit, the audio output control unit monitors whether the start condition is met. If the condition is met, the audio output unit starts the audio output of the message.
14. The sound output device as claimed in claim 1, characterized in that, The space specified above refers to the interior space within the vehicle.
15. A sound output method, characterized in that, It is a sound output method based on a sound output device that is set in a specified space and has a sound output section for sound output. include: The steps of receiving messages by the message receiving unit of the aforementioned sound output device; When the message receiving unit receives a message, the sound output control unit of the aforementioned sound output device determines an image of the face of the subject in the specified space based on face recognition technology, analyzes the determined face image, identifies actions specific to the subject when the subject is unable to hear normally, and determines whether the starting condition for the subject to be able to hear normally in the specified space is met. If the starting condition is met, the sound output unit starts outputting the message; otherwise, the sound output of the message to be performed by the sound output unit is reserved.
Citation Information
Patent Citations
On-board information device
WO2014002128A1
Device message playing control method and apparatus, message playing device and storage medium
CN107465595A
Onboard voice outputting device, voice outputting device, voice outputting method, and medium
CN109951764A