Sound data processing device and sound data processing method

By processing sound data of sound images positioning sounds inside the vehicle, emphasizing the sounds related to the occupants' attention to the object, the problem of difficulty for occupants to hear internal sounds is solved, and the clear transmission of sound is achieved.

CN115315374BActive Publication Date: 2025-08-26NISSAN MOTOR CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080098932.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-03-25
Publication Date
2025-08-26
Estimated Expiration
2040-03-25

AI Technical Summary

Technical Problem

Inside the vehicle, specific sounds are difficult to clearly hear by the occupants, and the prior art cannot effectively emphasize internal sounds for occupants to pay attention.

Method used

By acquiring and processing the sound data of the audio-visual positioning sound data inside the vehicle, the occupant is determined to pay attention to the object, and generate sound data that emphasizes the attention to the object, and outputs it to the output device for occupants to hear it.

Benefits of technology

The occupants are able to hear specific sounds inside the vehicle more easily, improving the clarity and audibility of the sound.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115315374B_ABST
    Figure CN115315374B_ABST
Patent Text Reader

Abstract

The sound data processing device (5) comprises: a sound data acquisition unit (150) for acquiring first sound data, which is data of sound image-localized in the interior of a vehicle; an object determination unit (160) for determining an object to which a vehicle occupant directs his attention, i.e., an attention object; a sound data processing unit (170) for generating second sound data, which is sound data in which the sound related to the attention object is emphasized compared to the first sound data; and a sound data output unit (180) for outputting the second sound data to an output device (4) for outputting sound to the occupant.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a sound data processing device and a sound data processing method. Background Art

[0002] A known ambient condition notification device collects ambient sounds outside a vehicle and reproduces the resulting sound information as localized sounds within the vehicle (Patent Document 1). This ambient condition notification device identifies a direction of attention, a direction in the vehicle's surroundings that the driver is particularly attentive to. Furthermore, the ambient condition notification device reproduces sounds localized in the localized direction in a manner that emphasizes them over sounds in directions around the vehicle other than the localized direction.

[0003] Prior art literature

[0004] Patent Literature

[0005] Patent Document 1: Japanese Patent Application Laid-Open No. 2005-316704 Summary of the Invention

[0006] Problems to be solved by the invention

[0007] In conventional technology, specific sounds outside the vehicle are reproduced in an emphasized manner compared to other sounds outside the vehicle, while sounds inside the vehicle are transmitted to the vehicle occupants as they are. As a result, for example, even if the occupants want to pay attention to specific sounds inside the vehicle, they may find it difficult to hear them.

[0008] The problem to be solved by the present invention is to provide a sound data processing device and a sound data processing method that make it easy for a vehicle occupant to hear a specific sound inside the vehicle.

[0009] Solutions for solving problems

[0010] The present invention solves the above-mentioned problem by the following scheme, which is: obtaining first sound data, which is data of sound that has been sound-image localized in the interior of a vehicle; determining an object to which an occupant's attention is directed, namely, an attention object; generating second sound data, which is sound data that emphasizes sounds related to the attention object compared to the first sound data and is sound-image localized; and outputting the second sound data to an output device for outputting sound to the occupant.

[0011] Effects of the Invention

[0012] According to the present invention, the occupant of the vehicle can easily hear a specific sound inside the vehicle. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 This is a block diagram showing an example of a sound output system including the sound data processing device according to the first embodiment.

[0014] Figure 2 It shows Figure 1 A block diagram of the functions of the control device shown.

[0015] Figure 3 This is an example of position information of an audio source located inside a vehicle.

[0016] Figure 4 This is a diagram for explaining a method of specifying an attention target and an audio source corresponding to the attention target using position information of an audio source.

[0017] Figure 5 This is an example of in-car spatial information.

[0018] Figure 6 This is a diagram for explaining a method of specifying a cautionary object and an audio source corresponding to the cautionary object using vehicle interior space information.

[0019] Figure 7 This is a flowchart showing the processing performed by the sound data processing device.

[0020] Figure 8 yes Figure 7 The subroutine of step S5 is shown.

[0021] Figure 9 yes Figure 7 The subroutine of step S6 is shown.

[0022] Figure 10 This is an example of a scene where a passenger wearing a head-mounted display is communicating with an icon.

[0023] Figure 11 is Figure 10 An example of candidate attention objects presented to the occupant in the scene shown.

[0024] Figure 12 yes Figure 7 The subroutine of step S5 shown is a subroutine according to the second embodiment. DETAILED DESCRIPTION

[0025] Hereinafter, embodiments of a sound data processing device and a sound data processing method according to the present invention will be described with reference to the accompanying drawings.

[0026] First Implementation Method

[0027] In this embodiment, a sound output system mounted on a vehicle is described as an example. Figure 1 This is a block diagram showing an example of a sound output system 100 including the sound data processing device 5 according to the first embodiment.

[0028] like Figure 1 As shown, sound output system 100 includes a sound collection device 1, a camera device 2, a database 3, an output device 4, and a sound data processing device 5. These devices are connected via a CAN (Controller Area Network) or other in-vehicle LAN to exchange information. Furthermore, connections between the devices are not limited to in-vehicle LANs such as CAN; they can also be connected via other wired or wireless LANs.

[0029] The sound output system 100 outputs sound to occupants of the vehicle. The sound output by the sound output system 100 will be described later. Although not shown, the vehicle is also equipped with a voice dialogue system, a notification system, a warning system, a vehicle audio system, and other systems. For ease of explanation, the vehicle's occupants will be referred to as simply "occupants" in the following description.

[0030] The voice dialogue system uses speech recognition and speech synthesis technologies to communicate with occupants. The notification system uses notification sounds to inform occupants of information related to onboard equipment. The warning system uses warning sounds to warn occupants of predicted dangers to the vehicle. The vehicle audio system connects to a recording medium, for example, to play back music recorded on the recording medium. The sound data processing device 5, described later, is connected to these onboard systems via a predetermined network.

[0031] In the present embodiment, the position of the seat where the passenger sits is not particularly limited. In addition, the number of passengers is also not particularly limited, and the sound output system 100 outputs sound to one or more passengers.

[0032] Each component included in the sound output system 100 will be described.

[0033] The sound collecting device 1 is set in the interior of the vehicle and collects the sounds heard by the passengers in the interior of the vehicle. The sounds collected by the sound collecting device 1 are mainly sounds whose audio sources are located in the interior of the vehicle. As the sounds collected by the sound collecting device 1, for example, conversations between passengers, conversations between the voice dialogue system and passengers, voice guidance by the voice dialogue system, notification sounds issued by the notification system, warning sounds issued by the warning system, audio sounds issued by the audio system, etc. can be listed. In addition, the sounds collected by the sound collecting device 1 may also include sounds whose audio sources are located outside the vehicle (such as the engine sounds of other vehicles). In addition, in the following description, "the interior of the vehicle" can also be replaced by "inside the vehicle" in words. In addition, "the outside of the vehicle" can also be replaced by "outside the vehicle" in words.

[0034] The sound collection device 1 collects sound that has been localized within the vehicle's interior. Localized sound refers to sound that, when heard by a person, allows the direction and distance of the sound source to be determined. In other words, when a sound image is localized at a predetermined position relative to a person, the person hears the sound as if the sound source is at the predetermined position and is outputting the sound from that position. Binaural recording is an example of a technique for collecting such localized sound. In binaural recording, sound is recorded as it reaches the person's eardrum.

[0035] An example of the sound collection device 1 is a binaural microphone, but the form of the sound collection device 1 is not particularly limited. For example, if the sound collection device 1 is an earphone type, the sound collection device 1 is worn on the left and right ears of the occupant. In the case of an earphone type, the earphones are equipped with microphones to collect the sound captured by the occupant's left and right ears. Alternatively, the sound collection device 1 may be a headphone type that can be worn on the occupant's head.

[0036] Furthermore, for example, if the sound collection device 1 is a dummy head type, it is installed in a location that corresponds to the seated occupant's head. An example of a location that corresponds to the occupant's head is near the headrest. A dummy head is a sound recorder shaped like a human head. In the case of a dummy head type, microphones are installed in the ear areas of the dummy head, enabling the collection of sounds as if they were being captured at the occupant's left and right ears.

[0037] As described above, localized sounds are sounds that allow humans to discern the direction and distance from the audio source. Therefore, even when the same sound is output from an audio source, the perceived direction and distance from the audio source vary depending on the positional relationship between the person and the audio source. Therefore, in this embodiment, the vehicle is equipped with a number of sound collection devices 1 equal to the number of seats in the vehicle. Furthermore, in this embodiment, the sound collection devices 1 are positioned at the same locations as the vehicle seats. This allows for the acquisition of sound data containing information about the direction and distance from the audio source perceived by each occupant, regardless of the placement or number of audio sources.

[0038] For example, a case where there are two seats in the front of the vehicle (a driver's seat and a passenger seat) and two seats in the rear of the vehicle (rear seats) is used as an example for explanation. A sound collecting device 1 is provided at each seat. In addition, speakers are provided in the front and on the left and right sides of the vehicle interior, for example, music is played indoors. In this example, when the front speaker is closer to the passenger sitting in the driver's seat (the front right seat) than the speakers on the left and right sides, the passenger feels that the audio source of the sound coming from the front is closer than the audio source of the sound coming from the left and right. In addition, the passenger feels that the audio source of the sound coming from the right side is closer than the audio source of the sound coming from the left side. The sound collecting device 1 provided at the driver's seat can collect the sound reaching the eardrum of the passenger sitting in the driver's seat.

[0039] The sound collection device 1 converts the collected sound into a predetermined sound signal and outputs the converted sound signal as sound data to the sound data processing device 5. The sound data processing device 5 then processes the collected sound. The sound data output from the sound collection device 1 to the sound data processing device 5 includes information that allows passengers to determine the direction and distance from the sound source. Furthermore, if a sound collection device 1 is installed for each seat, sound data is output from each sound collection device 1 to the sound data processing device 5. The sound data processing device 5 is configured to be able to identify the seat in which the sound data originates.

[0040] The imaging device 2 captures the interior of the vehicle. The captured images by the imaging device 2 are output to the audio data processing device 5. An example of the imaging device 2 is a camera equipped with a CCD element. The type of image captured by the imaging device 2 is not limited; the imaging device 2 only needs to be able to capture at least one of still images and moving images.

[0041] For example, the camera device 2 is installed at a position within the vehicle interior where it can capture images of the occupants, capturing the occupants' appearances. Furthermore, the location where the camera device 2 is installed and the number of camera devices 2 are not particularly limited. For example, a camera device 2 may be installed for each seat, or at a position where the entire interior can be viewed.

[0042] Database 3 stores positional information of audio sources within the vehicle's interior and in-vehicle spatial information related to the audio sources within the vehicle's interior. In the following description, for ease of explanation, the positional information of audio sources within the vehicle's interior will be referred to simply as audio source positional information. Furthermore, for ease of explanation, in-vehicle spatial information related to audio sources within the vehicle's interior will also be referred to simply as in-vehicle spatial information.

[0043] Audio sources within the vehicle's interior include speakers and people (passengers). Audio source location information indicates the speaker's installation location or the position of a passenger's head while seated. Specific examples of audio source location information are described below. In-vehicle spatial information is information used to associate objects within the vehicle's interior that passengers direct their attention with audio sources within the vehicle's interior. Specific examples of in-vehicle spatial information are described below. Database 3 outputs audio source location information and in-vehicle spatial information to sound data processing device 5 in response to access from the sound data processing device 5.

[0044] The output device 4 receives audio data from the audio data processing device 5. The output device 4 generates reproduced sound based on the audio data and outputs the reproduced sound in stereo.

[0045] For example, when the sound data output from the sound data processing device 5 to the output device 4 includes a signal of stereo recording, the output device 4 outputs the reproduced sound in stereo. In this case, a speaker can be cited as the output device 4. The location and number of the output device 4 are not particularly limited. The number of output devices 4 that can output the reproduced sound in stereo is set in the interior of the vehicle. In addition, the output device 4 is set at a specified position in the interior of the vehicle so that the reproduced sound can be output in stereo. For example, in order to give different stereo sounds to each passenger, an output device 4 is set for each seat. In this way, it is possible to reproduce the sound as if it were captured by the left ear and the right ear of each passenger respectively.

[0046] Furthermore, output device 4 may be a device other than a speaker. For example, if the sound data output from sound data processing device 5 to output device 4 includes a binaural recording signal, output device 4 outputs the reproduced sound using a binaural format. In this case, examples of output device 4 include earphones that can be worn on both ears or headphones that can be worn on the head. For example, to provide each passenger with a different stereo sound experience, each passenger may wear or wear output device 4. This allows for the reproduction of sound as if it were captured in each passenger's left and right ears.

[0047] The sound data processing device 5 is comprised of a computer equipped with hardware and software. The computer comprises a ROM (Read Only Memory) storing programs, a CPU (Central Processing Unit) executing the programs stored in the ROM, and RAM (Random Access Memory) functioning as an accessible storage device. Furthermore, an MPU (Micro Processing Unit), a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), or an FPGA (Field Programmable Gate Array) can be used as an operating circuit in place of or in addition to the CPU. Figure 1 The control device 50 shown corresponds to a CPU. Figure 1 The storage device 51 shown corresponds to the ROM and RAM. In this embodiment, the sound data processing device 5 is provided in the vehicle as a module.

[0048] Figure 2 It shows Figure 1 The block diagram of each function of the control device 50 shown in FIG. Figure 2 The functions of the control device 50 will be described below. Figure 2 As shown, the control device 50 includes a sound data acquisition unit 150 , a cautionary object identification unit 160 , a sound data processing unit 170 , and a sound data output unit 180 . These blocks implement various functions described below by software built in a ROM.

[0049] The sound data acquisition unit 150 acquires sound data from the sound collection device 1. In addition, when the sound data can be acquired from a system other than the sound output system 100, the sound data acquisition unit 150 acquires the sound data from the system. Examples of systems other than the sound output system 100 include a voice dialogue system, a notification system, a warning system, and a vehicle audio system. In the following description, for the sake of convenience, the sound data acquired by the sound data acquisition unit 150 will also be referred to as first sound data. In addition, in the following description, the processing for one passenger is cited as an example for explanation, but in the case where there are multiple passengers, that is, when there are multiple first sound data, it is assumed that the processing described below is performed for each first sound data.

[0050] As described above, the sound data output from the sound collection device 1 contains information that allows the occupant to determine the direction and distance from the sound source. The sound data acquisition unit 150 uses the audio source location information stored in the database 3 to determine the location of the audio source for one or more sounds heard by the occupant. For example, if the first sound data contains the voice of another occupant, the sound data acquisition unit 150 determines that the audio source is the occupant and determines the occupant's location by referring to the audio source location information.

[0051] For example, when acquiring sound data from a voice dialogue system, etc., sound data acquisition unit 150 determines that the audio source is a speaker and identifies the speaker's location by referring to the audio source's location information. In this case, sound data acquisition unit 150 analyzes the first sound data and identifies the speaker closest to the occupant's perception of all speakers installed in the vehicle's interior as the audio source.

[0052] The attention object identification unit 160 identifies an object to which the vehicle occupant is paying attention, that is, an attention object. Furthermore, the attention object identification unit 160 identifies an audio source corresponding to the attention object. The attention object is a device or a person inside the vehicle.

[0053] like Figure 2 As shown, the attention object identification unit 160 includes a motion recognition unit 161, a gaze recognition unit 162, and a speech content recognition unit 163 as functional blocks for determining whether the occupant is directing their attention toward the object. These blocks are used to identify the occupant's behavior. Furthermore, the attention object identification unit 160 includes an audio source identification unit 164 as a functional block for identifying the attention object and the audio source of the sound related to the attention object.

[0054] First, the functional blocks for determining whether the passenger is directing his attention toward an object are described. The motion recognition unit 161 recognizes the passenger's motion based on the camera image captured by the camera device 2. For example, the motion recognition unit 161 recognizes the passenger's gesture by performing image processing on the camera image to analyze the appearance of the passenger's hand. Furthermore, in the case where the passenger's gesture is to point with a finger, the motion recognition unit 161 recognizes the position indicated by the finger or the direction indicated by the finger. In addition, in the following description, for the sake of convenience, the position indicated by the finger is also referred to as the indicated position, and the direction indicated by the finger is also referred to as the indicated direction.

[0055] For example, assume that the characteristic points of a person's finger pointing gesture (e.g., the positional relationship of each finger) are pre-set and stored in a storage medium such as a hard disk drive (HDD) or ROM. In this case, the motion recognition unit 161 extracts the characteristic points of the hand from the portion of the camera image captured by the camera device 2 that reflects the passenger's hand. The motion recognition unit 161 then compares the extracted characteristic points with the characteristic points stored in the storage medium to determine whether the passenger's gesture matches a finger pointing gesture. For example, if a predetermined number or more of the extracted characteristic points match the characteristic points stored in the storage medium, the motion recognition unit 161 determines that the passenger's gesture is a finger pointing gesture. On the other hand, if a predetermined number of the extracted characteristic points match the characteristic points stored in the storage medium, the motion recognition unit 161 determines that the passenger's gesture is a gesture other than a finger pointing gesture. The predetermined number is a threshold used to determine whether the passenger's gesture is a finger pointing gesture and is a predetermined threshold. The above-described determination method is an example, and the motion recognition unit 161 can use a technique known at the time of filing this application to determine whether the passenger's gesture is an indication with a finger.

[0056] The line of sight recognition unit 162 recognizes the line of sight of the occupant based on the camera image captured by the camera device 2. For example, the line of sight recognition unit 162 recognizes the line of sight direction of the occupant by performing image processing on the camera image for analyzing the appearance of the occupant's eyes. Furthermore, when the occupant is looking, the line of sight recognition unit 162 recognizes the position or direction of the occupant's gaze. The gaze position is a prescribed position in the interior of the vehicle, and the gaze direction is a prescribed direction in the interior of the vehicle. In addition, in the following description, for the sake of convenience, the position where the occupant is looking is also referred to as the gaze position, and the direction where the occupant is looking is also referred to as the gaze direction.

[0057] For example, the sight line recognition unit 162 continuously monitors the portion of the camera image captured by the camera device 2 that reflects the eyes of the occupant. For example, when the occupant's sight line does not move but shows the same direction and remains in the same direction for more than a certain period of time, the sight line recognition unit 162 determines that the occupant is looking. On the other hand, when the occupant's sight line moves within a certain period of time, the sight line recognition unit 162 determines that the occupant is not looking. The certain period of time is a threshold for determining whether the occupant is looking, and is a predetermined threshold. In addition, the above determination method is an example, and the sight line recognition unit 162 can use technology known at the time of filing this application to determine whether the occupant is looking.

[0058] Speech content recognition unit 163 obtains the occupant's speech from a device for collecting speech within the vehicle's interior and recognizes the content of the occupant's speech based on the occupant's speech. The device for collecting the occupant's speech may be the sound collection device 1 or a different sound collection device. For example, speech content recognition unit 163 identifies the occupant's speech content by performing speech recognition processing on sound data corresponding to the occupant's speech to identify the occupant's speech. Furthermore, speech content recognition unit 163 can use speech recognition technology known at the time of filing this application to recognize the occupant's speech content.

[0059] The attention object identification unit 160 determines whether the occupant is directing their attention toward the object using at least one of the results obtained by the action recognition unit 161, the gaze recognition unit 162, and the speech content recognition unit 163. Furthermore, when making a determination using multiple results, the attention object identification unit 160 may also perform the determination by assigning priorities or weights to the results of each block.

[0060] For example, when the action recognition unit 161 determines that the passenger's gesture is pointing with a finger, the attention object determination unit 160 determines that the passenger is directing his attention to the object. In addition, for example, when the line of sight recognition unit 162 determines that the passenger is looking, the attention object determination unit 160 determines that the passenger is directing his attention to the object. In addition, for example, when the speech content recognition unit 163 determines that the passenger's voice contains specific keywords or specific key phrases, the attention object determination unit 160 determines that the passenger is directing his attention to the object. The specific keywords or specific key phrases are keywords or key phrases used to determine whether the passenger is directing his attention to the object, and are pre-set. As specific keywords, for example, words related to equipment installed in the vehicle can be listed. In addition, as specific key phrases, for example, phrases expressing wishes such as "Let me listen to X" and "I want to see Y" can be listed.

[0061] Next, the functional blocks used to identify a caution object and the audio source corresponding to the caution object will be described. The audio source identification unit 164 identifies a caution object and the audio source corresponding to the caution object based on at least one of the results obtained by the action recognition unit 161, the gaze recognition unit 162, and the speech content recognition unit 163, as well as the audio source's location information or vehicle interior space information stored in the database 3. Furthermore, when using multiple results to identify a caution object and the audio source corresponding to the caution object, the caution object identification unit 160 may also use a process that prioritizes or weights the results of each module to determine the caution object and the audio source corresponding to the caution object.

[0062] The audio source identification unit 164 identifies a cautionary object and an audio source corresponding to the cautionary object based on the occupant's pointing position or pointing direction, as well as the audio source's positional information or vehicle interior space information. Furthermore, the audio source identification unit 164 identifies a cautionary object and an audio source corresponding to the cautionary object based on the occupant's gaze position or gaze direction, as well as the audio source's positional information or vehicle interior space information. Furthermore, the audio source identification unit 164 identifies a cautionary object and an audio source corresponding to the cautionary object based on the occupant's speech content, as well as the audio source's positional information or vehicle interior space information.

[0063] use Figures 3 to 6 A method of identifying an attention target and an audio source corresponding to the attention target will be described. Figure 3 This is an example of the position information of the audio source in the interior of the vehicle stored in the database 3 . Figure 3 FIG. 1 shows a top view of the interior of a vehicle V. The vehicle V has two seats in the front and two seats in the rear. Figure 3 In FIG, the traveling direction of the vehicle V is set to the upper side of the drawing. Figure 3 In, P 11 ~P 15 Indicates the location where the speaker is located, P 22 ~P 25 Indicates the position of the occupant's head when sitting in the seat. 22 ~P 25 are shown overlapping the seats. Figure 3 In FIG, D represents a display embedded in the instrument panel. The display D displays a menu screen of the navigation system, guidance information to the destination, and the like. The navigation system is a system included in the voice dialogue system.

[0064] Figure 4 This is a diagram for explaining a method of specifying an attention target and an audio source corresponding to the attention target using position information of an audio source. Figure 4 The P shown 11 ~P 15、P 22 、P 23 and P 25 and Figure 3 The P shown 11 ~P 15 、P 22 、P 23 and P 25 In addition, Figure 4 In , U1 represents the passenger sitting in the driver's seat. Passenger U1 is facing left relative to the direction of travel of the vehicle V. Figure 4 In FIG, the line of sight of the passenger U1 is indicated by a dotted arrow. In addition, the passenger U1 is pointing to the left with his finger relative to the traveling direction of the vehicle V. Figure 4 In FIG, the solid arrow indicates the direction indicated by the passenger U1. Figure 4 In the example, it is assumed that the vehicle V is stopped or parked at a predetermined location, or the vehicle V is automatically or autonomously driven by a so-called automatic driving function, and even if the passenger U1 is facing left relative to the direction of travel, it does not affect the driving of the vehicle V. In addition, for the sake of convenience, although not shown in the figure, Figure 4 In FIG, passenger U2 is sitting in the passenger seat of vehicle V, and passenger U1 and passenger U2 are having a conversation.

[0065] exist Figure 4 In the example of FIG. 1 , the audio source identification unit 164 compares the indicated position of the occupant U1 with the position of each audio source (position P 11 ~ Position P 15 , Position P 22 and position P 25 ) is compared. If the audio source identification unit 164 determines that the indicated position of the passenger U1 is the position P 22 , the audio source identification unit 164 identifies occupant U2 as the attention target. In this case, the audio source identification unit 164 determines that the sound that occupant U1 wants to pay attention to is the sound of occupant U2. Furthermore, because the sound that occupant U1 wants to pay attention to is emitted by occupant U2, the audio source identification unit 164 identifies occupant U2 as the audio source corresponding to the attention target. Furthermore, in the above-described identification method, the audio source identification unit 164 can identify the attention target and the audio source corresponding to the attention target using the same method as the identification method using the indication position, by replacing the indication position with an indication direction, gaze position, or gaze direction.

[0066] Figure 5 This is an example of the vehicle interior space information stored in the database 3 . Figure 5 and Figure 3 and Figure 4 Likewise, a top view showing the interior of the vehicle V is shown.

[0067] exist Figure 5 In the figure, R1 represents the area associated with the notification sound. The area associated with the notification sound includes, for example, the speedometer, fuel gauge, water temperature gauge, odometer, etc. located in front of the driver's seat. In addition, the area associated with the notification sound may also include the center console, gear lever, air conditioning operating unit located between the driver's seat and the passenger seat, and may also include the storage space called the front panel located in front of the passenger seat. Area R1 is the same as that in Figure 3 Configured at position P 11 ~P 15 corresponding to the speakers.

[0068] In addition, Figure 5 In FIG, R2 represents an area related to the voice dialogue between the navigation system and the passenger. The area related to the voice dialogue includes a display that displays the menu screen of the navigation system, etc. Figure 5 In the case of R2 and Figure 3 The display D shown in FIG. 2 corresponds to the display D shown in FIG. Figure 3 Configured at position P 11 ~P 15 corresponding to the speakers.

[0069] In addition, Figure 5 In the figure, R3 represents the area associated with the passenger's speech. The area associated with the passenger's speech includes the seat where the passenger sits. Figure 3 Sitting in P 22 ~P 25 corresponding to the crew members.

[0070] Figure 6 This is a diagram for explaining a method of specifying a cautionary object and an audio source corresponding to the cautionary object using vehicle interior space information. Figure 6 R1~R3 shown in the figure are Figure 5 The R1 to R3 shown in the figure correspond to each other. Figure 6 In the figure, U1 represents the passenger sitting in the driver's seat. Passenger U1 is watching the display D (refer to Figure 3 ).exist Figure 6 In FIG, the dashed arrow indicates the line of sight of the passenger U1. Figure 6 The scene shown in the example is Figure 4 Similarly to the scene shown in the example of , it is assumed that even if the passenger U1 watches the display D, it does not affect the driving of the vehicle V.

[0071] exist Figure 6In the example, the audio source determination unit 164 compares the gaze position of the occupant U1 with each area (area R1 to area R3). If the audio source determination unit 164 determines that the gaze position of the occupant U1 is near the area R2, the display D is determined as the object of attention. At this time, the audio source determination unit 164 determines that the sound that the occupant U1 wants to pay attention to is the output sound from the speaker based on the correspondence between the area R2 and the speaker. And, since the sound that the occupant U2 wants to pay attention to is output from these speakers, the audio source determination unit 164 determines these speakers as the audio sources corresponding to the object of attention. In addition, the determined speakers are Figure 3 Configured at position P 11 ~P 15 Multiple speakers.

[0072] In addition, the audio source determination unit 164 determines the speaker closest to the occupant among the multiple speakers as the audio source. The audio source determination unit 164 determines the speaker closest to the occupant U1 and its position among the multiple speakers by analyzing the first sound data. Figure 6 In the example, for example, the audio source determination unit 164 analyzes the first sound data and determines that the occupant U1 feels that the nearest speaker is located at Figure 3 The position P shown 14 As described above, in this embodiment, when there are multiple audio sources corresponding to the attention object, the audio source that the occupant perceives as closest is identified. Since the sound output from this audio source is emphasized by the sound data processing unit 170, described later, the emphasized sound can be effectively transmitted to the occupant.

[0073] Back again Figure 2 ,right Figure 1 The functions of the control device 50 shown in FIG. 1 are described below. The sound data processing unit 170 processes the sound data collected by the sound collection device 1 to emphasize specific sounds over other sounds, thereby generating sound data that localizes the sound image within the vehicle interior. For ease of explanation, the sound data generated by the sound data processing unit 170 will also be referred to as second sound data.

[0074] In this embodiment, the sound data processing unit 170 generates second sound data that emphasizes sounds related to the attention object, compared to the first sound data acquired by the sound data acquisition unit 150. The attention object is the object to which the occupant directs his or her attention, as determined by the attention object identification unit 160. In other words, when comparing the first and second sound data, while the number of audio sources of the sounds heard by the occupant and the position of the sound image localized relative to the occupant are the same, the volume or intensity of the sounds related to the attention object in the second sound data is relatively higher than that of other sounds in the first sound data.

[0075] The sound associated with the attention object is any of the following: a sound output from the attention object, a sound output from another related object that is related to the attention object but different from the attention object, or both. In other words, the sound associated with the attention object includes at least one of the sound output from the attention object and the sound output from the related object.

[0076] For example, if two passengers are conversing and the object of attention for one passenger is the other passenger, as described above, the audio source corresponding to the object of attention is also the passenger. In this case, the audio data processing unit 170 will emphasize the sound uttered by the other passenger.

[0077] For example, if a passenger's attention object is a navigation system's route guidance screen, as described above, the attention object is a display, and the audio source associated with the attention object is a speaker. In this embodiment, if no sound is being output from the attention object, an object associated with the attention object and outputting sound is identified as a related object. In this example, sound data processing unit 170 identifies the speaker as the related object and, further, emphasizes the sound output from the speaker.

[0078] Furthermore, for example, if a conversation is taking place between three or more passengers, and the attention target for one passenger is a specific passenger among the multiple passengers, the audio source corresponding to the attention target is the specific passenger. In this embodiment, when an object of the same category as the attention target is identified, the identified object is identified as a related object. In this example, the audio data processing unit 170 identifies passengers other than the specific passenger among the multiple passengers, i.e., other passengers, as related objects. Furthermore, the audio data processing unit 170 emphasizes not only the speech of the specific passenger but also the speech of other passengers.

[0079] like Figure 2 As shown, the audio data processing unit 170 includes a type determination unit 171 and an audio signal processing unit 172 .

[0080] Type determination unit 171 determines whether the type of sound related to the attention object, which is the target of emphasis processing, is controllable by a system different from sound output system 100. Examples of systems different from sound output system 100 include voice dialogue systems, notification systems, warning systems, and vehicle audio systems. Examples of the objects controlled by these systems include volume and sound intensity.

[0081] For example, if the sound associated with the attention object is a sound programmed into a voice dialogue system, the type determination unit 171 determines that the type of sound associated with the attention object is controllable by the system. In other words, the type determination unit 171 determines that the sound data associated with the attention object can be acquired from a system different from the sound output system 100. The sound signal processing unit 172 acquires the sound data associated with the attention object from the corresponding system and performs processing to superimpose the acquired data on the first sound data, thereby generating second sound data. For convenience, the sound data associated with the attention object will be referred to as third sound data in the following description. Furthermore, when generating the second sound data, the sound signal processing unit 172 may also perform processing to increase the volume or enhance the intensity of the acquired third sound data and then superimpose the processed third sound data on the first sound data.

[0082] The aforementioned method of emphasizing the sound associated with the attention object relative to other sounds is merely an example. The sound signal processing unit 172 can utilize known sound emphasis processing at the time of filing this application to emphasize the sound associated with the attention object relative to other sounds. For example, when a device worn by an occupant, such as headphones, is used as the output device 4, the sound signal processing unit 172 can also perform processing to increase the volume of the sound associated with the attention object relative to other sounds. In this case, the sound signal processing unit 172 uses the sound data after the volume adjustment as the second sound data.

[0083] Furthermore, for example, if the sound related to the caution object is the voice of an occupant, type determination unit 171 determines that the sound related to the caution object is a type that cannot be controlled by the system. In other words, type determination unit 171 determines that the sound data related to the caution object cannot be acquired from the specified system. Sound signal processing unit 172 generates second sound data by extracting the sound data related to the caution object from the first sound data and performing emphasis processing on the extracted sound data.

[0084] The audio data output unit 180 outputs the second audio data generated by the audio data processing unit 170 to the output device 4 .

[0085] Next, use Figure 7 and Figure 8 The processing executed by the sound data processing device 5 in the sound output system 100 will be described. Figure 7 This is a flowchart showing the processing executed by the audio data processing device 5 according to this embodiment.

[0086] In step S1, the sound data processing device 5 acquires first sound data from the sound collection device 1. This first sound data contains information that allows the occupant to determine the direction and distance from the sound source. In step S2, the sound data processing device 5 acquires a camera image from the camera device 2 that depicts the interior of the vehicle.

[0087] In step S3, the sound data processing device 5 identifies the behavior of the passenger based on the first sound data obtained in step S1 or the camera image obtained in step S2. For example, the sound data processing device 5 determines whether the passenger is pointing with a finger based on the camera image. If the sound data processing device 5 determines that the passenger is pointing with a finger, it determines the indicated position or indicated direction indicated by the passenger's finger based on the camera image. In addition, the sound data processing device 5 may determine whether the passenger is looking based on the camera image, and if it determines that the passenger is looking, it determines the gaze position or gaze direction of the passenger. In addition, the sound data processing device 5 may also identify the content of the passenger's speech based on the first sound data. The sound data processing device 5 identifies the behavior of the passenger by performing one or more of these processes. The processes in steps S1 to S3 described above are also continuously performed after step S5 described later.

[0088] In step S4, the audio data processing device 5 determines whether the passenger is directing their attention to the object based on the passenger's behavior identified in step S3. For example, if the passenger is pointing with a finger in step S4, the audio data processing device 5 determines that the passenger is directing their attention to the object. In this case, the process proceeds to step S5.

[0089] On the other hand, if the audio data processing device 5 determines in step S4 that the passenger is not pointing with a finger, it determines that the passenger is not directing their attention to the object. In this case, the process returns to step S1. The above determination method is an example, and the audio data processing device 5 can determine whether the passenger is directing their attention to the object based on other determination results obtained in step S3 or a combination of these determination results.

[0090] If it is determined in step S4 that the occupant is directing his attention toward the object, the process proceeds to step S5. Figure 8 In the subroutine shown, the sound data processing device 5 performs processing such as specifying an attention object. Figure 8 yes Figure 7 The subroutine of step S5 is shown.

[0091] In step S51, the sound data processing device 5 obtains the position information of the audio source in the vehicle interior from the database 3. As the position information of the audio source, for example, Figure 3 In step S52, the sound data processing device 5 obtains the vehicle interior space information from the database 3. As the vehicle interior space information, for example, Figure 5 The position information of the audio source and the in-vehicle space information may be information showing the interior of the vehicle, and the form thereof is not limited to a plan view.

[0092] In step S53 , the sound data processing device 5 specifies an attention object, which is an object to which the occupant directs his or her attention, based on the position information of the audio source acquired in step S51 or the vehicle interior space information acquired in step S52 .

[0093] For example, in Figure 4 As shown, when the passenger in the driver's seat points a finger toward the passenger seat, the sound data processing device 5 identifies the passenger in the passenger seat as a cautionary object based on the passenger's pointing position or pointing direction and the positional information of the audio source. Furthermore, the sound data processing device 5 identifies the passenger as the audio source corresponding to the cautionary object.

[0094] In addition, for example, in Figure 6 When the passenger sitting in the driver's seat is looking at the display as shown, the sound data processing device 5 determines the display as the attention object based on the passenger's gaze position or gaze direction and the vehicle interior space information. Figure 6 When the audio source corresponding to the illustrated region R2 is a speaker, the audio data processing device 5 specifies the corresponding speaker as the audio source corresponding to the attention object.

[0095] When the processing in step S53 is completed, the Figure 7 In step S6, the second audio data generation process is performed. Figure 9 yes Figure 7 The subroutine of step S6 is shown.

[0096] In step S61, the sound data processing device 5 determines whether the Figure 8 The audio source corresponding to the caution object, as determined in step S53, is of a type controllable by a system separate from sound output system 100. For example, if third sound data, representing sound data related to the caution object, can be obtained from a system separate from sound output system 100, sound data processing device 5 determines that the audio source corresponding to the caution object is of a type controllable by the other system. Examples of sounds that fall into this category include voices programmed into a voice dialogue system, notification sounds set by a notification system, warning sounds set by a warning system, and audio sounds set by a vehicle audio system.

[0097] On the other hand, if the third sound data can only be obtained from the sound collection device 1 included in the sound output system 100, the sound data processing device 5 determines that the audio source corresponding to the attention object is of a type that cannot be controlled by other systems. An example of a sound that meets this type of classification is the voice of an occupant.

[0098] In step S62, the sound data processing device 5 performs emphasis processing based on the determination result in step S61, emphasizing the sound related to the attention object over other sounds. For example, if the speech programmed by the voice dialogue system is a sound related to the attention object, the sound data processing device 5 obtains third sound data from the voice dialogue system and superimposes the third sound data on the first sound data obtained in step S1. Alternatively, if the occupant's speech is a sound related to the attention object, the sound data processing device 5 extracts third sound data from the first sound data obtained in step S1 and performs emphasis processing on the extracted third sound data.

[0099] In step S63, the sound data processing device 5 generates second sound data in which the sound related to the attention object is emphasized based on the execution result of step S62. When the processing in step S63 is completed, the process proceeds to step S63. Figure 7 Step S7 is shown.

[0100] In step S7, the audio data processing device 5 outputs the second audio data generated in step S6 to the output device 4. This step indicates that the output of the second audio data from the audio data processing device 5 to the output device 4 starts.

[0101] In step S8, the sound data processing device 5 determines whether the occupant's attention has left the attention object. If the sound data processing device 5 determines based on the occupant's behavior recognition result in step S3 that the occupant's attention is not directed toward the attention object identified in step S5, it determines that the occupant's attention has left the attention object. In this case, the process proceeds to step S9. In step S9, the sound data processing device 5 stops the generation process of the second sound data and ends. Figure 7 The processing is shown in the flowchart.

[0102] For example, based on the occupant's indicated position and the positional information of the audio source, the sound data processing device 5 determines that the occupant's attention has shifted away from the attention object if no attention object exists at or near the occupant's indicated position. Furthermore, the above determination method is merely an example. For example, the sound data processing device 5 may pre-set a gesture for determining that attention has shifted away from the attention object, and upon recognizing that the occupant has performed this gesture, determine that the occupant's attention has shifted away from the attention object.

[0103] On the other hand, if the sound data processing device 5 determines that the occupant's attention is directed toward the attention object identified in step S5, it determines that the occupant's attention has not shifted from the attention object. In this case, the process proceeds to step S10. For example, based on the occupant's indicated position and the positional information of the audio source, the sound data processing device 5 determines that the occupant's attention has not shifted from the attention object if the attention object is located at or near the occupant's indicated position. Furthermore, the above determination method is merely an example. For example, the sound data processing device 5 may pre-set a gesture for determining that attention has shifted from the attention object, and if the occupant does not perform this gesture, it may determine that the occupant's attention has not shifted from the attention object.

[0104] If it is determined in step S8 that the occupant's attention has not left the attention object, the process proceeds to step S10. In step S10, the sound data processing device 5 determines whether a sound related to the attention object is being output. For example, if the sound data processing device 5 fails to confirm the output from the audio source corresponding to the attention object within a predetermined time period, it determines that no sound related to the attention object is being output. In this case, the process proceeds to step S9. In step S9, the sound data processing device 5 stops the generation process of the second sound data and ends. Figure 7 The predetermined time is a time for determining whether or not a sound related to the attention object is being output, and is a time set in advance.

[0105] On the other hand, for example, if the sound data processing device 5 can confirm the output from the audio source corresponding to the attention object during the predetermined time, it determines that the sound related to the attention object is being output. In this case, the process returns to step S8.

[0106] As described above, in this embodiment, the sound data processing device 5 includes: a sound data acquisition unit 150 that acquires first sound data, which is sound image localized within the vehicle cabin; an attention object identification unit 160 that identifies an attention object, which is an object to which the occupant directs their attention; a sound data processing unit 170 that generates second sound data, which is sound data in which the sound related to the attention object is emphasized compared to the first sound data; and a sound data output unit 180 that outputs the second sound data to the output device 4. Thus, since the sound that the vehicle occupant wants to pay attention to is reproduced in an emphasized state, the occupant can easily hear the sound that they want to pay attention to.

[0107] Furthermore, in this embodiment, the caution object identification unit 160 acquires an image of the occupant from the imaging device 2, identifies the occupant's indicated position or indicated direction based on the acquired image, acquires audio source location information or vehicle interior space information from the database 3, and identifies a caution object based on the identified indicated position or indicated direction and the audio source location information or vehicle interior space information. The occupant can communicate the caution object to the audio data processing device 5 using an intuitive and efficient method such as gestures. The audio data processing device 5 can identify the caution object with high accuracy.

[0108] Furthermore, in this embodiment, the caution object identification unit 160 identifies the occupant's gaze position or gaze direction based on the captured image from the imaging device 2, obtains audio source location information or vehicle interior space information from the database 3, and identifies the caution object based on the identified gaze position or gaze direction and the audio source location information or vehicle interior space information. This allows the occupant to communicate the caution object to the audio data processing device 5 intuitively and efficiently through their line of sight. The audio data processing device 5 can then identify the caution object with high accuracy.

[0109] In addition, in this embodiment, the caution object identification unit 160 acquires the occupant's speech from the sound collection device 1 or another sound collection device, recognizes the occupant's speech content based on the acquired occupant's speech, and identifies caution objects based on the recognized speech content. This allows the occupant to communicate caution objects to the sound data processing device 5 intuitively and efficiently using the speech content. The sound data processing device 5 can then identify caution objects with high accuracy.

[0110] In this embodiment, the sound related to the attention object is the sound output from the attention object. For example, if the attention object is another passenger, the sound is emphasized from the direction of the passenger's attention, making it easier for the passenger to hear the sound they want to pay attention to.

[0111] Furthermore, in this embodiment, the sound associated with the attention object is a sound output from another related object that is related to the attention object but different from the attention object. For example, if the attention object is the voice guidance of a navigation system, even if the occupant is directing their attention to the display, the voice guidance corresponding to the information displayed on the display is emphasized. Therefore, even if the occupant is directing their attention to an object that does not output sound, the sound associated with that object is easily heard.

[0112] Furthermore, in this embodiment, the sounds associated with the attention object include both the sounds output from the attention object and the sounds output from related objects. For example, in a scene where three or more passengers are having a conversation, if the attention object is one of the passengers and the related objects are other passengers, not only is the object to which the passenger is directing their attention emphasized, but the sounds associated with that object are also emphasized. Even if the passenger is not directing their attention to an object related to the object to which the passenger is directing their attention, the emphasized sounds from that object can still be heard. This provides a sound output system 100 that is highly convenient for the passengers.

[0113] Furthermore, in this embodiment, when third sound data, representing a sound related to an attention object, can be acquired from a system separate from the sound output system 100, the sound data processing unit 170 generates second sound data by superimposing the third sound data on the first sound data. For example, when an object, such as a navigation system's voice guidance whose volume and intensity can be directly controlled, is the target of emphasis processing, a simple method can be used to emphasize the sound that the occupant wants to pay attention to.

[0114] Furthermore, in this embodiment, if the third sound data cannot be obtained from a system separate from the sound output system 100, the sound data processing unit 170 generates the second sound data by performing sound emphasis processing on the third sound data included in the first sound data. For example, even for objects whose volume or intensity cannot be directly controlled, such as the passenger's voice, it is possible to emphasize only such objects. Regardless of the object being emphasized, the sound that the passenger wants to pay attention to can be emphasized.

[0115] In addition, in this embodiment, the sound data acquisition unit 150 acquires first sound data from the sound collection device 1 that performs binaural recording of sounds generated in the interior of the vehicle. As a result, the first sound data contains information that allows the occupant to determine the direction of the audio source and the distance from the audio source. After performing processing to emphasize the sound related to the object of attention, the sound that has been localized can be transmitted to the occupant without performing sound image localization processing on the sound. Complex processing such as sound image localization processing can be omitted, and the computational load of the sound data acquisition unit 150 can be reduced. In addition, based on the information that allows the occupant to determine the position of the audio source and the distance from the audio source, the audio source and its position in the interior of the vehicle can be easily determined. In addition, the sound can be reproduced as if it were captured by the left and right ears of the occupant respectively.

[0116] Furthermore, in this embodiment, after identifying the attention object, the attention object identification unit 160 determines whether the occupant is directing their attention toward the attention object. If the attention object identification unit 160 determines that the occupant's attention is not directed toward the attention object, the sound data processing unit 170 stops generating the second sound data. This prevents the occupant from emphasizing sounds related to the attention object even when their attention is not directed toward the object, thus preventing the occupant from feeling uncomfortable. In other words, sounds that the occupant wants to pay attention to can be emphasized in appropriate situations when the occupant's attention is directed toward the attention object.

[0117] Second Implementation Method

[0118] Next, the sound data processing device 5 according to the second embodiment will be described. In this embodiment, the structure is the same as that of the first embodiment except for the following aspects: Figure 1 The illustrated sound collection device 1 and output device 4 are provided in a head-mounted display-type device. In this embodiment, icons, known as avatars, are included as attention objects and corresponding audio sources. Furthermore, in this embodiment, some functions of the sound data processing device 5 are different. Therefore, regarding the same configuration as in the first embodiment, the description of the first embodiment is incorporated herein by reference.

[0119] In the sound output system 100 involved in this embodiment, a head-mounted display type device is used. The head-mounted display type device is equipped with AR (Augmented Reality) technology. Icons (also called virtual images) are displayed on the display of the head-mounted display type device. The passengers wearing the device can visually confirm the icons (also called virtual images) through the display and can communicate with the icons. In the following description, such a device is also referred to as a head-mounted display. In this embodiment, the object includes not only the equipment or people in the vehicle interior, but also the icons presented to the passengers through the head-mounted display.

[0120] When the sound collection device 1 and the output device 4 are integrally provided as a head mounted display as in this embodiment, for example, the sound collection device 1 and the output device 4 may be headphones capable of binaural recording.

[0121] Another example of a scenario where a passenger can communicate with an icon inside a vehicle via the head-mounted display is a conversation with a person outside the vehicle. The remote location is not specifically limited to any location outside the vehicle, as long as it is outside the vehicle. In this scenario, when the passenger in the driver's seat looks at the passenger seat via the display, an icon corresponding to the remote person appears at a position corresponding to the passenger seat. The voice of the remote person is also output through headphones.

[0122] use Figures 10 to 12 The functions of the audio data processing device 5 according to this embodiment will be described. Figure 10 This is an example of a scene where a passenger wearing a head-mounted display is communicating with an icon. Figure 10 The same as that used in the description of the first embodiment Figure 4 The scene corresponds to. Figure 10 In the figure, the passenger U1 wears a head mounted display (HD). Figure 10 In FIG, the dotted arrow indicates the gaze direction of the passenger U1.

[0123] The audio data processing device 5 according to the present embodiment has a function of presenting candidates of the attention object to the occupant via the head mounted display during the process of specifying the attention object, and allowing the occupant to select the candidate of the attention object. Figure 11 is Figure 10 An example of a candidate object of attention presented to the occupant in the scene shown. Figure 11 , I represents an icon corresponding to a person who is far away.

[0124] For example, the sound data processing device 5 obtains a camera image corresponding to the passenger's field of view from a camera device mounted on a head-mounted display. Figure 11 As shown in the example of FIG. 5 , the sound data processing device 5 makes P representing the position of the audio source 12 and P 22 It is displayed in a superimposed manner on the camera image showing the passenger seat. Figure 11 As shown, the sound data processing device 5 presents the position of the audio source to the occupant in such a manner that the occupant can be identified as the audio source.

[0125] The sound data processing device 5 determines whether there are multiple candidate caution objects in the image visually recognized by the occupant. Furthermore, if the sound data processing device 5 determines that there are multiple candidate caution objects, it determines whether there are multiple categories. Examples of categories include occupants, icons, and speakers. Alternatively, categories may be categorized based on whether they can be controlled by a system other than the sound output system 100.

[0126] When the voice data processing device 5 determines that there are multiple categories of candidates for the caution object, it requests the occupant to select the caution object. The voice data processing device 5 determines one candidate for the caution object selected by the occupant as the caution object. Figure 11 The P shown 12 and P 22 and Figure 3 The P shown 12 and P 22 correspond.

[0127] Figure 12 yes Figure 7 The subroutine of step S5 shown is the subroutine involved in this embodiment. Figure 12 1 is a diagram for explaining a method of determining a cautionary object executed by the audio data processing device 5 according to this embodiment. Figure 12 In the first embodiment, Figure 7 The same processing as the subroutine of step S5 is marked as Figure 7 The same reference numerals are used and reference is made to their description.

[0128] In step S151, the sound data processing device 5 presents candidates for attention objects. For example, the sound data processing device 5 presents candidates for attention objects by Figure 11 A plurality of attention object candidates are presented in the manner shown in the example.

[0129] In step S152, the audio data processing device 5 determines whether there are multiple candidates for the attention object presented in step S151. If it is determined that there are multiple candidates for the attention object, the process proceeds to step S153. If it is determined that there are not multiple candidates for the attention object, the process proceeds to step S54.

[0130] If it is determined in step S152 that there are multiple candidates for the attention object, the process proceeds to step S153. In step S153, the sound data processing device 5 determines whether there are multiple candidate categories for the attention object. If it is determined that there are multiple candidate categories for the attention object, the process proceeds to step S154. If it is determined that there are not multiple candidate categories for the attention object, the process proceeds to step S54.

[0131] If it is determined in step S153 that there are multiple candidate categories for the caution object, the process proceeds to step S154. In step S154, the audio data processing device 5 receives a selection signal from the passenger. For example, the passenger selects a candidate for the caution object from the multiple candidate categories by making a gesture such as pointing with a finger. When the process in step S154 is completed, the process proceeds to step S54, and the caution object is determined.

[0132] Thus, in this embodiment, the sound data processing device 5 is applied to a head-mounted display device equipped with AR technology. This allows the occupant to select the desired attention object, thereby accurately emphasizing and outputting the sound that the occupant wants to pay attention to. Furthermore, objects that do not actually exist but output sound, such as icons, can be included in the attention objects. As a result, the sound objects that the occupant wants to pay attention to can be emphasized.

[0133] In addition, the embodiments described above are described to facilitate understanding of the present invention and are not described to limit the present invention. Therefore, the gist of the various elements disclosed in the above embodiments is intended to include all design changes and equivalents that fall within the technical scope of the present invention.

[0134] For example, in the first embodiment described above, the method of identifying a cautionary object was described using either audio source position information or vehicle interior space information. However, it is sufficient to use at least one of the audio source position information and vehicle interior space information to identify a cautionary object. For example, a cautionary object may be identified using only audio source position information or only vehicle interior space information. Furthermore, for example, a method may be employed in which, when the audio source position information cannot be used to identify a cautionary object, vehicle interior space information is used to identify the cautionary object.

[0135] For example, in the second embodiment described above, some of the functions of the audio data processing device 5 can also utilize functions of a head-mounted display device. For example, if the head-mounted display is equipped with a camera to capture the surroundings and a microphone to capture the occupant's voice, the audio data processing device 5 can also obtain information related to the occupant's movements, line of sight, and voice from these devices or equipment. Furthermore, the audio data processing device 5 can use this information to perform processing such as occupant movement recognition, occupant line of sight recognition, or occupant speech recognition.

[0136] Description of Reference Numerals

[0137] 1: Sound collecting device; 2: Camera device; 3: Database; 4: Output device; 5: Sound data processing device; 50: Control device; 150: Sound data acquisition unit; 160: Attention object determination unit; 161: Action recognition unit; 162: Eyesight recognition unit; 163: Speech content recognition unit; 164: Audio source determination unit; 170: Sound data processing unit; 171: Type determination unit; 172: Sound signal processing unit; 180: Sound data output unit; 51: Storage device; 100: Sound output system.

Claims

1. A sound data processing device comprising: a sound data acquisition unit that acquires first sound data, the first sound data being data on sound inside a vehicle; an object identifying unit that identifies an attention object, which is an object to which the occupant of the vehicle directs his attention; a sound data processing unit configured to generate second sound data, the second sound data being data in which the sound related to the attention object is emphasized compared to the first sound data; as well as a sound data output unit configured to output the second sound data to an output device for outputting sound to the occupant; wherein the object identification unit identifies the occupant as the audio source corresponding to the attention object, and identifies a speaker as the audio source corresponding to the attention object, When a conversation is conducted between three or more passengers and the attention object for one passenger is a specific passenger among the plurality of passengers, and another passenger other than the specific passenger is identified, the voice data processing unit identifies the other passenger as a related object. The sound data processing unit generates the second sound data in which the sound related to the attention object and the sound related to the related object are emphasized. When the first sound data and the second sound data are compared, the number of audio sources of the sound heard by the occupant and the position of the sound image localized with respect to the occupant are the same, and in the second sound data, the volume or intensity of the sound is relatively greater than the volume or intensity of other sounds in comparison with the first sound data.

2. The sound data processing device according to claim 1, wherein: The object identification unit acquires a captured image of the occupant from an imaging device capturing the interior of the room. The object identification unit recognizes the indicated position or indicated direction indicated by the occupant based on the captured image, The object identifying unit acquires position information of the object in the room from a storage device, The object identification unit identifies the caution object based on the indicated position or indicated direction and the position information.

3. The sound data processing device according to claim 1 or 2, wherein: The object identification unit acquires a captured image of the occupant from an imaging device capturing the interior of the room. The object identification unit identifies a gaze position or a gaze direction of the occupant based on the captured image. The object identifying unit acquires position information of the object in the room from a storage device, The object identification unit identifies the attention object based on the gaze position or the gaze direction and the position information.

4. The sound data processing device according to claim 1 or 2, wherein: The object identification unit acquires the voice of the passenger from a device that collects sounds in the room. The object identification unit recognizes the speech content of the passenger based on the voice of the passenger, The object identification unit identifies the attention object based on the speech content.

5. The sound data processing device according to claim 1 or 2, wherein: The sound related to the attention object is the sound output from the attention object.

6. The sound data processing device according to claim 1 or 2, wherein: The sound related to the attention object is the sound output from the related object.

7. The sound data processing device according to claim 1 or 2, wherein: The sound related to the attention object is the sound output from the attention object and the sound output from the related object.

8. The sound data processing device according to claim 1 or 2, wherein: The sound data processing unit generates the second sound data by performing a process of superimposing the third sound data on the first sound data when third sound data, which is the sound data related to the attention object, can be acquired from a predetermined system.

9. The sound data processing device according to claim 1 or 2, wherein: The sound data processing unit generates the second sound data by performing sound emphasis processing on the third sound data included in the first sound data when third sound data, which is the sound data related to the attention object, cannot be acquired from a predetermined system.

10. The sound data processing device according to claim 1 or 2, wherein: The sound data acquisition unit acquires the first sound data from a device that performs binaural recording of the sound generated in the room.

11. The sound data processing device according to claim 1 or 2, wherein: After identifying the attention object, the object identification unit determines whether the occupant is directing attention toward the attention object. The sound data processing unit stops generating the second sound data when the object identification unit determines that the occupant's attention is not directed toward the attention object.

12. A method for processing sound data, the method being executed by a processor, comprising the following steps: acquiring first sound data, the first sound data being data of sound inside a vehicle interior; determining an object to which the occupant of the vehicle is directing attention, ie, an attention object; When an object of the same category as the attention object is identified, recognizing the identified object as a related object; generating second sound data in which the sound related to the attention object and the sound related to the related object are emphasized compared to the first sound data; determining the occupant as an audio source corresponding to the attention object, and determining a speaker as an audio source corresponding to the attention object; as well as When a conversation is conducted between three or more passengers, the attention object for one passenger is a specific passenger among the plurality of passengers, and another passenger other than the specific passenger is identified, the other passenger is identified as a related object, and the second sound data is output to an output device for outputting sound to the passenger. In the sound data processing method, when the first sound data and the second sound data are compared, the number of audio sources of the sound heard by the occupant and the position of the sound image localized with respect to the occupant are the same, and in the second sound data, the volume or intensity of the sound is relatively greater than the volume or intensity of other sounds in comparison with the first sound data.

Citation Information

Patent Citations

  • Surrounding state notification device and method for notifying surrounding state

    JP2005316704A

  • Conversation support device, conversation support method, and conversation support program

    JP2015071320A

  • Conversation support device, conversation support system, and conversation support method

    JP2019068237A