Information processing system, information processing device, control method for information processing system, and program
The information processing system enhances voice recognition accuracy in vehicles by using real-time noise data and pre-stored user voice information to separate passenger voices from vehicle noise, addressing the low signal-to-noise ratio challenge.
Patent Information
- Application Number
- JP2025506457
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-03-15
- Filing Date
- 2023-06-20
- Publication Date
- 2026-01-14
- Estimated Expiration
- 2043-06-20
AI Technical Summary
The low signal-to-noise ratio inside vehicles due to engine noise and other disturbances hampers effective voice recognition of passengers, as existing technologies rely on pre-stored noise data that does not account for real-time noise variations.
An information processing system that includes a user voice acquisition unit, noise acquisition unit, superimposing unit, filter setting unit, and sound source separation unit to separate and process passenger voices from noise in real-time, using superimposed sound information to set a sound source separation filter.
Improves speech recognition accuracy by effectively separating passenger voices from vehicle noise using real-time noise data and pre-stored user voice information, enhancing the accuracy of voice recognition systems.
Smart Images

Figure 0007799139000001 
Figure 0007799139000002 
Figure 0007799139000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing system, an information processing device, a control method for an information processing system, and a program. [Background technology]
[0002] Patent Document 1 discloses a voice recognition device for a vehicle that can dramatically improve the recognition rate of input voice while the vehicle is running.
[0003] Patent Document 2 discloses a characteristic setting method for a noise reduction device, in which coefficients corresponding to a large number of noise reduction characteristics are stored in advance, a predetermined noise reduction characteristic is set in a noise reduction unit, and a driving noise signal is removed from a voice signal detected by a voice detection means and input to a voice recognition device. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 7-146698 [Patent Document 2] Japanese Patent Application Laid-Open No. 2002-221986 Summary of the Invention [Problem to be solved by the invention]
[0005] Since sounds inside a vehicle contain noise such as engine noise, the signal-to-noise ratio of the sounds inside the vehicle is low (noise accounts for a large proportion). Therefore, in order to properly recognize the voice of a passenger, it is necessary to separate and process the passenger's voice, which is the target of voice recognition, from the noise. The technologies described in Patent Documents 1 and 2 above perform voice recognition by superimposing pre-stored noise data on the user's voice data and performing predetermined processing.
[0006] However, because noise data (noise) changes in type and volume in real time, real-time noise data must be used for more effective speech recognition.
[0007] One example of a problem to be solved by the present invention is to effectively improve speech recognition accuracy. [Means for solving the problem]
[0008] The invention described in claim 1 is An information processing system that uses at least a communication terminal in a mobile body, a user voice acquisition unit that acquires user voice information related to the voices of passengers from a storage unit in which the user voice information is stored in advance; a noise acquisition unit that acquires noise information related to noise inside the moving object; a superimposing unit that generates superimposed sound information by superimposing the user voice information and the noise information acquired by the user voice acquisition unit from the storage unit; a filter setting unit that sets a sound source separation filter using the superimposed sound information; a moving body sound acquisition unit that acquires moving body sound information related to sounds inside the moving body; a sound source separation unit that separates the internal moving body sound information using the sound source separation filter.
[0009] The invention described in claim 11 is a user voice acquisition unit that acquires user voice information related to the voices of passengers from a storage unit in which the user voice information is stored in advance; a noise acquisition unit that acquires noise information related to noise inside the moving body; a superimposing unit that generates superimposed sound information by superimposing the user voice information and the noise information acquired by the user voice acquisition unit from the storage unit; a filter setting unit that sets a sound source separation filter using the superimposed sound information; a moving body sound acquisition unit that acquires moving body sound information related to sounds inside the moving body; a sound source separation unit that separates the moving body sound information using the sound source separation filter.
[0010] The invention described in claim 12 is One or more computers that realize an information processing system that uses at least a communication terminal in a mobile body, acquiring user voice information relating to the voice of a passenger from a storage unit in which the user voice information is stored in advance; Acquire noise information regarding noise inside the moving body; generating superimposed sound information by superimposing the user voice information and the noise information acquired from the storage unit; Using the superimposed sound information, a sound source separation filter is set; acquiring internal sound information relating to sounds inside the moving body; The method for controlling an information processing system includes separating the in-moving body sound information using the sound source separation filter.
[0011] The invention described in claim 13 is One or more computers that realize an information processing system that uses at least a communication terminal in a mobile body, a step of acquiring user voice information from a storage unit in which user voice information relating to the voice of a passenger is stored in advance; a step of acquiring noise information regarding noise inside the moving body; a step of generating superimposed sound information by superimposing the user voice information and the noise information acquired from the storage unit; a step of setting a sound source separation filter using the superimposed sound information; a step of acquiring internal sound information relating to sounds inside the moving body; The program is for executing a procedure for separating the in-moving body sound information using the sound source separation filter. [Brief explanation of the drawings]
[0012] [Figure 1] 1 is a block diagram showing functions of an information processing system according to a first embodiment. [Figure 2] FIG. 1 is a diagram illustrating an example of a hardware configuration of an information processing device. [Figure 3] FIG. 10 is a flow diagram showing how a filter setting unit sets a sound source separation filter. [Figure 4] FIG. 4 is a flowchart showing details of step S100 according to the first embodiment. [Figure 5] FIG. 4 is a flowchart showing details of step S200 according to the first embodiment. [Figure 6] FIG. 10 is a block diagram showing functions of an information processing device according to a second embodiment. [Figure 7] FIG. 10 is a flowchart showing details of step S100 according to the second embodiment. [Figure 8] FIG. 11 is a flowchart showing details of step S200 according to the third embodiment. [Figure 9] FIG. 4 is a schematic diagram for explaining an example of a method for a filter setting unit to determine whether a first criterion is satisfied. [Figure 10] FIG. 10 is a block diagram showing functions of an information processing device according to a fourth embodiment. [Figure 11] FIG. 11 is a block diagram showing functions of an information processing device according to a fifth embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0013] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings, like components are designated by like reference numerals, and the description thereof will be omitted as appropriate.
[0014] In the following description, each component of each device represents a functional block, not a hardware configuration. Each component of each device is realized by any combination of hardware and software, centered around the CPU of any computer, memory, a program loaded into the memory, a storage medium such as a hard disk that stores the program, and a network connection interface. There are many variations in the realization method and device.
[0015] First Embodiment Fig. 1 is a block diagram showing the functions of an information processing system 100 according to the first embodiment. The information processing system 100 according to the first embodiment will be described with reference to Fig. 1. The information processing system 100 uses at least a communication terminal 2 located inside a mobile body. In the first embodiment, the information processing system 100 uses an information processing device 1 and the communication terminal 2.
[0016] The information processing system 100 is configured to function as a voice recognition system when a predetermined keyword is detected. The predetermined keyword may be set arbitrarily and is stored in advance in the storage unit 4. After detecting the predetermined keyword, the information processing system 100 receives input of a predetermined command and then executes processing corresponding to the command.
[0017] For example, a case will be described in which the predetermined keyword is "yyy" and the predetermined command "search for the nearest convenience store" is input. In this case, when a passenger utters "yyy, find the nearest convenience store," the information processing system 100 detects that the predetermined keyword is "yyy," functions as a voice recognition system, searches for the nearest convenience store, and provides guidance to the passenger, such as "There is a convenience store 500 meters ahead."
[0018] (Information processing device 1) In the first embodiment, the information processing device 1 is provided outside the mobile body. The information processing device 1 may be a so-called cloud server. In the first embodiment, the mobile body is a vehicle 3. The information processing device 1 is configured to function as at least a part of a voice recognition device when a predetermined keyword is detected.
[0019] The information processing device 1 is configured to be able to communicate with a communication terminal 2 (described later) via a communication network 101. The communication network 101 includes, for example, communication over a 4G or 5G line. The information processing device 1 is configured to include an external vehicle communication unit (not shown) for configuring the communication network 101.
[0020] (Communication terminal 2) The communication terminal 2 is located inside the vehicle 3. In the first embodiment, the communication terminal 2 is mounted on the vehicle 3. The communication terminal 2 may be brought into the vehicle 3 from outside the vehicle 3. The communication terminal 2 is configured to be able to communicate with the information processing device 1 via the communication network 101. The communication terminal 2 is configured to include an external vehicle communication unit (not shown) for configuring the communication network 101.
[0021] In the first embodiment, the communication terminal 2 includes an audio output unit 2a, a position sensor unit 2b, an imaging unit 2c, and an audio input unit 2d. Although not shown, the communication terminal 2 may also include a display.
[0022] In the first embodiment, the audio output unit 2a includes a speaker (not shown). In the first embodiment, the audio output unit 2a outputs at least one of a mechanical voice and a sound effect via the speaker.
[0023] The position sensor unit 2b acquires the position information of the mobile object, for example, by using the Global Navigation Satellite System (GNSS).
[0024] The imaging unit 2c has an in-camera and an out-camera (not shown). The in-camera faces the interior of the vehicle, and the driver's seat is included in the imaging range. The in-camera captures images of the interior of the vehicle 3 so that at least the driver is visible. The out-camera faces outside the vehicle, and captures images of the situation outside the vehicle.
[0025] In the first embodiment, the voice input unit 2d includes a microphone (not shown). Voices uttered by passengers (driver and fellow passengers) in the vehicle 3 are input to the voice input unit 2d via the microphone.
[0026] (User voice acquisition unit 10) The information processing device 1 according to the first embodiment includes a user voice acquisition unit 10, a noise acquisition unit 20, a superimposition unit 30, a filter setting unit 40, an in-vehicle sound acquisition unit 50, and a sound source separation unit 60. The user voice acquisition unit 10 acquires user voice information from the storage unit 4. The user voice information is information relating to the voices of passengers (driver and passengers). The user voice information is stored in advance in the storage unit 4.
[0027] The user voice information includes at least one of information regarding the sound pressure, volume, and frequency of the passenger's voice. The user voice information may include voice waveform data. The user voice information may include, for example, text information converted from the voice waveform data.
[0028] The user voice acquisition unit 10 identifies user voice information corresponding to a passenger in the vehicle 3 from among a plurality of pieces of user voice information stored in the storage unit 4, and acquires the identified user voice information from the storage unit 4. For example, if there is one passenger (driver), the user voice acquisition unit 10 may recognize the driver using image data of the driver captured and generated by the imaging unit 2c, identify user voice information corresponding to (linked to) the driver from the storage unit 4, and acquire the user voice information from the storage unit 4.
[0029] In the first embodiment, the user voice acquisition unit 10 uses keyword sound information related to the sound of a predetermined keyword to identify user voice information corresponding to a passenger who spoke inside the vehicle 3. The keyword sound information includes at least one of information related to the sound pressure, volume, and frequency of the sound of the predetermined keyword. The keyword sound information may include voice waveform data. The keyword sound information may include, for example, text information converted from the voice waveform data.
[0030] A method in which the user voice acquisition unit 10 uses keyword sound information to identify user voice information corresponding to a passenger who spoke inside the vehicle 3 will be specifically described. For example, when user voice information of a certain driver U is stored in advance in the storage unit 4, it is assumed that the driver U utters a predetermined keyword inside the vehicle 3. At that time, the user voice acquisition unit 10 analyzes keyword sound information (such as a voice waveform, sound pressure, volume, frequency, and text information of the keyword) related to the sound of the predetermined keyword uttered by the driver U. Then, the user voice acquisition unit 10 uses the analysis result to identify the user voice information of the driver U stored in advance in the storage unit 4.
[0031] In the first embodiment, the user voice acquiring unit 10 may identify the user voice information corresponding to the passenger who first uttered a predetermined keyword after the information processing system 100 was started up.
[0032] (Noise acquisition unit 20) The noise acquisition unit 20 acquires noise information via the voice input unit 2d (microphone). The noise acquisition unit 20 acquires noise information in real time while the vehicle 3 is traveling. The noise information is information related to noise (acoustic noise, disturbances) of the vehicle 3. The noise information includes, for example, information related to engine operation sounds, wind from the air conditioner, opening and closing sounds of windows, running sounds of the vehicle 3, sounds caused by the environment outside the vehicle 3 (noise at a construction site, etc.), and other noises. In the first embodiment, the noise information includes voice information of passengers that is not subject to voice recognition (conversations between passengers, utterances that do not match specified keyword sounds, etc.).
[0033] (Superimposing unit 30) The superimposing unit 30 generates superimposed sound information by superimposing the user voice information acquired by the user voice acquiring unit 10 from the storage unit 4 on noise information acquired while the vehicle 3 is traveling. The superimposed sound information is information about a signal obtained by superimposing a signal related to the voice of a passenger in the user voice information on a signal related to a noise sound in the noise information.
[0034] (Filter setting unit 40) The filter setting unit 40 sets a sound source separation filter using the superimposed sound information. The sound source separation filter is used to remove noise components (such as noise inside the vehicle 3 and conversations between passengers) that are unnecessary for speech recognition from the sound source components. In the first embodiment, the filter setting unit 40 calculates predetermined parameters for generating a sound source separation filter using the superimposed sound information, and sets the sound source separation filter to which the parameters are applied.
[0035] (Moving internal sound acquisition unit 50) The moving body sound acquisition unit 50 acquires moving body sound information via the voice input unit 2d (microphone). The moving body sound information is information related to sounds inside the vehicle 3. The sounds inside the vehicle 3 include at least noise inside the vehicle 3 (noise and conversations between passengers that are not subject to voice recognition) and the spoken voices of passengers that are used for voice recognition.
[0036] (Sound source separation section 60) The sound source separation unit 60 separates the internal sound information of the vehicle using a sound source separation filter. Specifically, the sound source separation unit 60 separates the speech of the passenger to be used for voice recognition from the internal sound information of the vehicle including noise information inside the vehicle 3. In this way, the sound source separation unit 60 extracts the speech information of the passenger to be used for voice recognition from the internal sound information of the vehicle.
[0037] (Hardware configuration example) 2 is a diagram showing an example of the hardware configuration of the information processing device 1. The information processing device 1 includes a bus 1010, a processor 1020, a memory 1030, a storage device 1040, an input / output interface 1050, and a network interface 1060.
[0038] The bus 1010 is a data transmission path for transmitting and receiving data among the processor 1020, memory 1030, storage device 1040, input / output interface 1050, and network interface 1060. However, the method of connecting the processor 1020 and the like to each other is not limited to bus connection.
[0039] The processor 1020 is implemented by a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or the like.
[0040] The memory 1030 is a main storage device realized by a RAM (Random Access Memory) or the like.
[0041] The storage device 1040 is an auxiliary storage device realized by removable media such as a hard disk drive (HDD), a solid state drive (SSD), or a memory card, or a read only memory (ROM), and has a recording medium. The recording medium of the storage device 1040 stores program modules that realize each function of the information processing device 1 (for example, a user voice acquisition unit 10, a noise acquisition unit 20, a superimposition unit 30, a filter setting unit 40, an in-moving body sound acquisition unit 50, and a sound source separation unit 60). The processor 1020 loads each of these program modules into the memory 1030 and executes them, thereby realizing each function corresponding to the program module. The storage device 1040 may also function as a memory unit 4 included in the information processing device 1.
[0042] The input / output interface 1050 is an interface for connecting the information processing device 1 to various input / output devices.
[0043] The network interface 1060 is an interface for connecting the information processing device 1 to a network. This network is, for example, a LAN (Local Area Network) or a WAN (Wide Area Network). The network interface 1060 may be connected to the network wirelessly or by wire. The information processing device 1 may communicate with the communication terminal 2 via the network interface 1060.
[0044] The hardware configuration of the communication terminal 2 may also be the same as above.
[0045] (Operation example 1 of the first embodiment) 3 is a flow diagram showing the process up to when the filter setting unit 40 sets a sound source separation filter. In step S100, the user voice acquisition unit 10 acquires user voice information to be used for setting the sound source separation filter from the storage unit 4. In step S200, the filter setting unit 40 sets the sound source separation filter using the user voice information.
[0046] (Operation example 2 of the first embodiment) Fig. 4 is a flow diagram showing details of step S100 according to the first embodiment. Using Fig. 4, the flow up to when the user voice acquisition unit 10 acquires user voice information used to set the sound source separation filter from the storage unit 4 will be described.
[0047] In step S110, when a passenger in the vehicle 3 utters a predetermined keyword, the user voice acquisition unit 10 acquires keyword sound information relating to the sound of the predetermined keyword.
[0048] In step S120, the user voice acquisition unit 10 uses the keyword sound information to identify user voice information corresponding to a passenger who uttered the predetermined keyword inside the vehicle 3. Specifically, for example, it is assumed that user voice information of a driver U's voice "yyy" is pre-stored in the storage unit 4, and the driver U utters the predetermined keyword "yyy" inside the vehicle 3. At that time, the user voice acquisition unit 10 analyzes the voice information of the predetermined keyword ("yyy") uttered by the driver U. Then, the user voice acquisition unit 10 determines whether the similarity (how similar) between the voice information of "yyy" uttered by the driver U and the user voice information of the voice of "yyy" pre-stored in the storage unit 4 satisfies a predetermined criterion (for example, how similar the text information of "yyy" is), and if the predetermined criterion is satisfied, the user voice information of the driver U may be identified.
[0049] In step S130, the user voice acquisition unit 10 acquires, from the storage unit 4, the user voice information of the driver U identified in step S120.
[0050] (Operation example 3 of the first embodiment) 5 is a flow diagram showing details of step S200 according to the first embodiment. Using FIG. 5, the flow up to when the filter setting unit 40 sets the sound source separation filter will be described.
[0051] In step S210, the noise acquisition unit 20 acquires real-time noise information inside the vehicle 3, and stores the noise information in the storage unit 4 (or memory 1030).
[0052] In step S220, the superimposing unit 30 superimposes the user voice information of the driver U acquired in step S130 and the noise information stored in the storage unit 4 (or memory 1030) in step S210 to generate superimposed sound information.
[0053] In step S230, the filter setting unit 40 uses the superimposed sound information to calculate predetermined parameters for generating a sound source separation filter, and sets the sound source separation filter to which the parameters are applied.
[0054] In the first embodiment, the timing for setting the sound source separation filter is arbitrary in step S200 of Fig. 3. However, when setting the sound source separation filter in S200, it is necessary to specify in advance which user voice information to use.
[0055] It should be noted that which user voice information is used to set the sound source separation filter may be switched by a predetermined action (such as when a driver is changed and the new driver utters a predetermined keyword). For example, in the first embodiment, after the information processing system 100 is started (or after the engine of the vehicle 3 is started), if a passenger utters a predetermined keyword, the user voice information corresponding to the passenger is identified, and the user voice information used to set the sound source separation filter may be switched from the user voice information used last time to the newly identified user voice information.
[0056] As described above, according to the first embodiment, the information processing device 1 includes a user voice acquisition unit 10, a noise acquisition unit 20, a superimposition unit 30, a filter setting unit 40, an internal moving body sound acquisition unit 50, and a sound source separation unit 60.
[0057] By calculating and setting a sound source separation filter using user voice information stored in advance in the storage unit 4 and superimposed sound information obtained by superimposing noise information acquired in real time, sound sources can be separated with high accuracy, thereby improving the accuracy of speech recognition.
[0058] Furthermore, the user voice acquisition unit 10 identifies user voice information corresponding to passengers in the vehicle 3 from among the multiple pieces of user voice information, and acquires the identified user voice information from the storage unit 4. Then, the superimposing unit 30 generates superimposed sound information using the user voice information identified by the user voice acquisition unit 10. Then, the filter setting unit 40 sets a sound source separation filter using the superimposed sound information. By setting the sound source separation filter using the user voice information corresponding to the passengers in the vehicle 3, sound sources can be separated with high accuracy.
[0059] Furthermore, the user voice acquisition unit 10 uses keyword sound information related to the sound of a predetermined keyword to identify user voice information corresponding to a passenger who has spoken inside the vehicle 3. Then, the user voice acquisition unit 10 acquires the identified user voice information from the storage unit 4. Then, the filter setting unit 40 uses the user voice information to set a sound source separation filter.
[0060] Sound sources can be separated with high accuracy by setting a sound source separation filter using user voice information corresponding to a passenger who uttered a predetermined keyword inside the vehicle 3. This effectively improves the accuracy of voice recognition.
[0061] Furthermore, the user voice acquisition unit 10 identifies user voice information corresponding to the passenger who first uttered a predetermined keyword after the information processing system 100 was started up, thereby making it possible to effectively identify and acquire user voice information to be used for setting the sound source separation filter.
[0062] Second Embodiment 6 is a block diagram showing the functions of the information processing device 1 according to the second embodiment. Unlike the information processing device 1 according to the first embodiment, the information processing device 1 according to the second embodiment further includes a detection unit .
[0063] (Detection unit 70) The detection unit 70 detects that the occupant (driver, passenger) of the vehicle 3 has changed. In the second embodiment, the occupant is the driver. In the second embodiment, the detection unit 70 detects that the driver of the vehicle 3 has been replaced by another driver.
[0064] In the second embodiment, the detection unit 70 may detect a change in occupants using at least one of the captured image captured by the imaging unit 2c mounted in the vehicle 3, the opening and closing of the doors of the vehicle 3 (including the sound of the doors opening and closing), and the occupants fastening and unfastening their seat belts.
[0065] In the second embodiment, after the detection unit 70 detects that the passenger (driver) has changed, the user voice acquisition unit 10 identifies the user voice information corresponding to the passenger who uttered a specified keyword, and acquires the identified user voice information from the memory unit 4.
[0066] (Operation example of the second embodiment) Fig. 7 is a flow diagram showing details of step S100 according to the second embodiment. Using Fig. 7, a flow up to when the user voice acquisition unit 10 according to the second embodiment acquires user voice information used for setting a sound source separation filter from the storage unit 4 will be described.
[0067] In step S106, the detection unit 70 detects that there has been a change in the occupant of the vehicle 3. For example, when the user U1 who was the driver switches driving to a user U2 in the passenger seat, the detection unit 70 detects that the driver has changed from the user U1 to the user U2, using the captured image captured by the imaging unit 2c of the user U2 who is now sitting in the driver's seat.
[0068] In step S111, when the user U2 who has switched driving roles with the user U1 utters a predetermined keyword, the user voice acquiring unit 10 acquires keyword sound information relating to the sound of the predetermined keyword.
[0069] In step S121, the user voice acquisition unit 10 uses the keyword sound information to identify the user voice information corresponding to the user U2.
[0070] In step S131, the user voice acquisition unit 10 acquires, from the storage unit 4, the user voice information of the user U2 identified in step S121.
[0071] Then, in step S210 and subsequent steps described in the first embodiment, the user voice information acquired in step S131 is used to set a sound source separation filter.
[0072] As described above, according to the second embodiment, the information processing device 1 further includes a detection unit 70. Then, the user voice acquisition unit 10 identifies user voice information corresponding to the driver after the change who uttered a predetermined keyword after the passenger (driver) was changed. This makes it possible to effectively identify user voice information to be used for setting a sound source separation filter.
[0073] Furthermore, the detection unit 70 detects a change in occupants using at least one of an image captured by the imaging unit 2c mounted in the vehicle 3, the opening and closing of the doors of the vehicle 3, and the fastening and unfastening of the seat belts of the occupants. This allows for accurate detection of a change in occupants.
[0074] Third Embodiment The information processing device 1 according to the third embodiment differs from the first embodiment in that some of the functions of the filter setting unit 40 are different. The filter setting unit 40 according to the third embodiment updates the sound source separation filter by using information about the sound inside the moving body. Updating the sound source separation filter means switching the previously used sound source separation filter to another sound source separation filter (calculating predetermined parameters again).
[0075] The above-mentioned in-vehicle sound information includes at least one of information on the sound pressure, information on the volume, and information on the frequency of the sound inside the vehicle 3. The filter setting unit 40 according to the third embodiment updates the sound source separation filter when at least one of the sound pressure, volume, and frequency inside the vehicle 3 satisfies a predetermined criterion.
[0076] (Operation example of the third embodiment) Fig. 8 is a flow diagram showing details of step S200 according to the third embodiment. A flow up to when the filter setting unit 40 according to the third embodiment updates the sound source separation filter will be described using Fig. 8. It should be noted that in Fig. 8, it is assumed that the user voice information used to update the sound source separation filter has been acquired in advance by the user voice acquisition unit 10.
[0077] In step S202, the filter setting unit 40 (or the moving body sound acquisition unit 50) acquires moving body sound information from the voice input unit 2d (microphone) while the moving body is running.
[0078] In step S205, the filter setting unit 40 determines whether or not at least one of the sound pressure, volume, and frequency related to the moving body sound information acquired in step S202 satisfies a predetermined criterion (first criterion). If the predetermined criterion (first criterion) is satisfied (YES in step S205), the process proceeds to step S211. If the predetermined criterion (first criterion) is not satisfied (NO in step S205), the process ends. The processes in steps S211, S221, and S231 are the same as those in the first embodiment.
[0079] FIG. 9 is a schematic diagram illustrating an example of a method for determining whether the filter setting unit 40 satisfies the first criterion. The horizontal axis of FIG. 9 represents time t, and the vertical axis represents sound pressure level Lp. As shown on the horizontal axis of FIG. 9, the time axis is divided into predetermined time periods (time N-1, time N, time N+1, time N+2, ...). For example, time N=1 [s], time N+1=2 [s], time N+2=3 [s], .... Furthermore, in FIG. 9, the unit of sound pressure level is dB (decibels), and the average of the sound pressure levels over the predetermined time periods is expressed as the average sound pressure level. For example, the average sound pressure level between time N and time N+1 is 52 [dB].
[0080] A specific method by which the filter setting unit 40 updates the sound source separation filter using moving body sound information will be described below. In the third embodiment, the filter setting unit 40 detects the difference between the average sound pressure between a certain time period (for example, between time N+2 and time N+3 in FIG. 9) and the average sound pressure between the previous time periods (for example, between time N+1 and time N+2 in FIG. 9). Then, the filter setting unit 40 determines whether the difference between the average sound pressures satisfies a first criterion.
[0081] In the third embodiment, the sound pressure inside the vehicle 3 satisfies the first criterion when the difference in the average sound pressure exceeds 3 dB. In this case, the filter setting unit 40 determines that the first criterion is satisfied at time N+2, for example. This is because the difference between the average sound pressure (52 dB) between time N and time N+1 and the average sound pressure (70 dB) between time N+1 and time N+2 is 18 dB.
[0082] Then, when the filter setting unit 40 determines that the sound pressure (difference in sound pressure) inside the vehicle 3 satisfies the first criterion (YES in step S205 in FIG. 8), the noise acquiring unit 20 acquires noise information at that time (time N+2) in step S211. Then, in the flow from step S212 onwards, a sound source separation filter is calculated, and the sound source separation filter used previously is updated (changed) to the sound source separation filter after calculation.
[0083] Furthermore, the filter setting unit 40 also determines that the sound pressure inside the vehicle 3 satisfies the first criterion at time N+4 (for the same reason as described above), and therefore updates the sound source separation filter in the same manner as above.
[0084] If the sound pressure inside the vehicle 3 does not satisfy the first criterion, the filter setting unit 40 does not update the sound source separation filter. That is, the sound source separation filter is not updated at time N+1 and time N+3.
[0085] As described above, the filter setting unit 40 according to the third embodiment updates the sound source separation filter by using the in-vehicle sound information. When the sound inside the vehicle 3 changes by more than a predetermined standard, this means that the type of noise has changed, etc., so by updating the sound source separation filter when the sound inside the vehicle 3 has changed by more than the predetermined standard, the accuracy of the sound source separation filter can be improved.
[0086] More specifically, the timing when the sound inside the vehicle 3 changes by a predetermined standard or more means that the noise information (noise data) has changed significantly, so by using the real-time noise information (noise data) at that time and pre-stored user voice information to set the sound source separation filter, the accuracy of the sound source separation filter can be improved, thereby further improving the voice recognition accuracy.
[0087] Furthermore, the vehicle internal sound information includes at least one of information regarding the sound pressure, volume, and frequency of the sound inside the vehicle 3. This can further improve the accuracy of the sound source separation filter.
[0088] The filter setting unit 40 updates the sound source separation filter when at least one of the sound pressure, volume, and frequency satisfies a predetermined criterion (first criterion), thereby further improving the accuracy of the sound source separation filter.
[0089] <Fourth embodiment> 10 is a block diagram showing the functions of the information processing device 1 according to the fourth embodiment. The information processing device 1 according to the fourth embodiment further includes a device information acquisition unit 80, unlike the first embodiment.
[0090] (Device information acquisition unit 80) The device information acquisition unit 80 acquires driver identification information from a communication device of the driver in the vehicle 3. The communication device includes, for example, at least one of a smartphone, a tablet, and a PC. The driver identification information is information that can identify the driver.
[0091] The user voice acquisition unit 10 identifies user voice information corresponding to the driver identification information acquired by the device information acquisition unit 80. Then, the user voice acquisition unit 10 acquires the identified user voice information from the storage unit 4.
[0092] For example, when a driver gets into the vehicle 3, the driver's smartphone (communication device) communicates with the communication terminal 2 via a predetermined communication network (e.g., Bluetooth (registered trademark)). At this time, the device information acquisition unit 80 acquires an ID identifying the driver (e.g., a terminal ID of the driver's smartphone) from the driver's smartphone via the communication terminal 2. The device information acquisition unit 80 then identifies user voice information associated with the ID identifying the driver that is stored in advance in the storage unit 4. In this way, the user voice acquisition unit 10 may identify user voice information corresponding to the driver identification information (ID identifying the driver) acquired by the device information acquisition unit 80.
[0093] As described above, according to the fourth embodiment, the information processing device 1 further includes the device information acquisition unit 80. This makes it possible to easily identify the user voice information corresponding to the driver.
[0094] Fifth Embodiment 11 is a block diagram showing the functions of an information processing device 1 according to the fifth embodiment. Unlike the first embodiment, the information processing device 1 according to the fifth embodiment further includes a control unit 90 and a storage processing unit 95. Some of the functions of the user voice acquisition unit 10, filter setting unit 40, and sound source separation unit 60 according to the fifth embodiment are different from those of the first embodiment.
[0095] As in the first embodiment, the user voice acquisition unit 10 in the fifth embodiment uses keyword sound information related to the sound of a predetermined keyword to identify user voice information corresponding to a passenger who uttered a predetermined keyword inside the vehicle 3 from the user voice information previously stored in the memory unit 4, and acquires the identified user voice information from the memory unit 4.
[0096] The flow up to when the filter setting unit 40 in the fifth embodiment sets the sound source separation filter is the same as the flow described in Figures 3, 4, and 5, but the filter setting unit 40 in the fifth embodiment updates the sound source separation filter every time a passenger in the vehicle 3 speaks a predetermined keyword.
[0097] Specifically, each time a passenger in the vehicle 3 utters a predetermined keyword, the user voice acquisition unit 10 acquires user voice information corresponding to the passenger who made the utterance from the storage unit 4, from among multiple pieces of user voice information, using keyword sound information for the passenger. Then, the superimposition unit 30 superimposes the user voice information newly acquired by the user voice acquisition unit 10 on noise information to generate new superimposed sound information. Then, the filter setting unit 40 updates the sound source separation filter using the superimposed sound information newly generated by the superimposition unit 30. In this way, the filter setting unit 40 according to the fifth embodiment updates the sound source separation filter each time a passenger in the vehicle 3 utters a predetermined keyword.
[0098] As a result, even when multiple passengers each utter a predetermined keyword or when the passenger who utters the predetermined keyword changes, the sound source separation filter is updated each time (a filter is set for each passenger who speaks), improving the accuracy of sound source separation. This makes it possible to effectively improve the accuracy of speech recognition.
[0099] (Control unit 90, memory processing unit 95) After detecting the predetermined keyword, the control unit 90 receives a predetermined command and executes a process corresponding to the command. The predetermined command is a word for executing the process corresponding to the command. In the fifth embodiment, the predetermined command is spoken to the passenger following the predetermined keyword.
[0100] (Memory Processing Unit 95) The storage processing unit 95 stores command sound information relating to the sound of a predetermined command uttered by the passenger after a predetermined keyword in the storage unit 4 or the memory 1030. The command sound information includes voice information corresponding to the command uttered by the passenger and noise sound information.
[0101] The filter setting unit 40 sets a sound source separation filter using keyword sound information related to the sound of a predetermined keyword detected by the control unit 90. Then, the sound source separation unit 60 uses the sound source separation filter to separate command sound information (stored in the storage unit 4 or the memory 1030) related to the sound of a predetermined command uttered by the passenger after the predetermined keyword detected by the control unit 90.
[0102] More specifically, for example, a case will be described in which the predetermined keyword is "zzz" and a predetermined command such as "Get traffic congestion information to destination" is input. When the passenger utters "zzz, tell me traffic congestion information to destination G," the filter setting unit 40 first executes a process of setting a sound source separation filter using "zzz" as keyword sound information detected by the control unit 90 (the process of setting the sound source filter is the same as that described in the first embodiment). Simultaneously with this process, the storage processing unit 95 stores "Tell me traffic congestion information to destination G" as command sound information in the storage unit 4 or the memory 1030. After the filter setting unit 40 completes the process of setting the sound source separation filter, the sound source separation unit 60 applies the sound source separation filter to the command sound information stored in the storage unit 4 or the memory 1030 to separate a sound source related to the command sound information. The sound source separation unit 60 then extracts voice information to be used for voice recognition from the command sound information.
[0103] Normally, the process of setting a sound source separation filter takes a certain amount of time. Therefore, when a passenger utters a predetermined command after a predetermined keyword, the process of generating a sound source separation filter using the predetermined keyword and separating the sound source of the predetermined command using the just-generated sound source separation filter will not be fast enough.
[0104] However, as described above, the generation process of the sound source separation filter is executed when the passenger utters a predetermined keyword, and the just-generated sound source separation filter is applied to the command sound information stored in the storage unit 4 or the memory 1030. Therefore, even if a predetermined command is to be executed for the first time after a change of passenger (driver), the sound source separation process can be applied to the command. This improves the speech recognition accuracy.
[0105] Furthermore, in the fifth embodiment, after the control unit 90 detects a predetermined keyword, if it takes more than a predetermined time (for example, more than 10 seconds) for the filter setting unit 40 to update the sound source separation filter using keyword sound information related to the sound of the predetermined keyword, the sound source separation unit 60 uses the sound source separation filter used last time to separate command sound information (stored in the storage unit 4 or memory 1030) related to the sound of a predetermined command uttered by the passenger after the predetermined keyword detected by the control unit 90.
[0106] When a passenger utters a predetermined command following a predetermined keyword, there may not be enough time for the filter setting unit 40 to update the sound source separation filter using keyword sound information related to the sound of the predetermined keyword and then apply the newly updated sound source separation filter to the predetermined command.
[0107] However, if the time is insufficient (if it takes more than a predetermined time to update the sound source separation filter), the sound source separation process for the command sound information is performed using the sound source separation filter that was used last time. This makes it possible to prevent the sound source separation accuracy from deteriorating.
[0108] Although the embodiments have been described above with reference to the drawings, these are merely examples of the present invention, and various other configurations can also be adopted.
[0109] If user voice information corresponding to a passenger in the vehicle 3 is not pre-stored in the storage unit 4, the information processing system 100 may prompt the passenger to register the passenger's voice information. Specifically, for example, a mechanical voice such as "How are you feeling today?" is output from the voice output unit 2a, and by the passenger replying to the mechanical voice, voice data relating to the reply is stored in the storage unit 4.
[0110] Furthermore, even if user voice information corresponding to a passenger in the vehicle 3 is already stored in the storage unit 4, when the information processing system 100 is started up, the information processing system 100 may prompt the passenger to register the passenger's voice information. The reason for this is that, since the voice of the same passenger may change from day to day (for example, when the passenger has a cold), always storing the latest voice information improves the accuracy of sound source separation.
[0111] The information processing system 100 is realized using one or more computers such as the information processing device 1 and the communication terminal 2. The number of computers used to realize the information processing system 100 is arbitrary.
[0112] Furthermore, in the first embodiment, the user voice acquiring unit 10, the noise acquiring unit 20, the superimposing unit 30, the filter setting unit 40, the moving body sound acquiring unit 50, and the sound source separating unit 60 have been described as being provided in the information processing device 1, but at least some of the functions of the user voice acquiring unit 10, the noise acquiring unit 20, the superimposing unit 30, the filter setting unit 40, the moving body sound acquiring unit 50, the sound source separating unit 60, and the storage unit 4 may be provided in the communication terminal 2. Note that the communication terminal 2 including the user voice acquiring unit 10, the noise acquiring unit 20, the superimposing unit 30, the filter setting unit 40, the moving body sound acquiring unit 50, and the sound source separating unit 60 also achieves the same effects as those of the first embodiment.
[0113] In addition, the user voice acquisition unit 10, the noise acquisition unit 20, the superimposition unit 30, the filter setting unit 40, the internal moving body sound acquisition unit 50, the sound source separation unit 60, and the memory unit 4 may all be configured to be mounted on the communication terminal 2.
[0114] Furthermore, the filter setting unit 40 according to the third embodiment updates the sound source separation filter using the moving body sound information, but the filter setting unit 40 can remove keyword sound information relating to the sound of a predetermined keyword uttered by the passenger from the moving body sound information. Specifically, if keyword sound information relating to the sound of a predetermined keyword uttered by the passenger is mixed into the moving body sound information, the keyword sound information may be used to erroneously update the sound source separation filter, and to prevent this, the keyword sound information can be removed from the moving body sound information.
[0115] Below, examples of reference forms are added. 1. An information processing system that uses at least a communication terminal in a mobile body, a user voice acquisition unit that acquires user voice information related to the voices of passengers from a storage unit in which the user voice information is stored in advance; a noise acquisition unit that acquires noise information related to noise inside the moving object; a superimposing unit that generates superimposed sound information by superimposing the user voice information and the noise information acquired by the user voice acquisition unit from the storage unit; a filter setting unit that sets a sound source separation filter using the superimposed sound information; a moving body sound acquisition unit that acquires moving body sound information related to sounds inside the moving body; a sound source separation unit that separates the moving body's internal sound information using the sound source separation filter. 2. In the information processing system described in 1., The user voice acquisition unit identifies the user voice information corresponding to a passenger in the vehicle from among the multiple pieces of user voice information, and acquires the identified user voice information from the memory unit. 3. In the information processing system described in 2., the information processing system is configured to function as a voice recognition system when a predetermined keyword is detected; The user voice acquisition unit uses keyword sound information relating to the sound of the predetermined keyword to identify the user voice information corresponding to the passenger who spoke inside the vehicle. 4. In the information processing system described in 3., The user voice acquisition unit identifies the user voice information corresponding to the passenger who first uttered the predetermined keyword after the information processing system was started. 5. In the information processing system described in 3. or 4., A detection unit is further provided to detect a change in the occupant of the moving body, An information processing system in which, after the detection unit detects that the passenger has changed, the user voice acquisition unit identifies the user voice information corresponding to the passenger who uttered the specified keyword and acquires the identified user voice information from the memory unit. 6. In the information processing system described in 5., The detection unit detects a change in the occupant using at least one of an image captured by an imaging unit mounted inside the vehicle, the opening and closing of a door of the vehicle, and the fastening and unfastening of a seat belt by the occupant. 7. In the information processing system according to any one of 1. to 6., The filter setting unit updates the sound source separation filter using the moving body sound information. 8. In the information processing system described in 7., An information processing system, wherein the internal sound information of the moving body includes at least one of information regarding sound pressure, volume, and frequency of sound inside the moving body. 9. In the information processing system according to 8., The filter setting unit updates the sound source separation filter when at least one of the sound pressure, the volume, and the frequency satisfies a predetermined criterion. 10. In the information processing system according to any one of 1. to 9., a device information acquisition unit that acquires driver identification information capable of identifying the driver from a communication device of the driver in the vehicle; The user voice acquisition unit identifies the user voice information corresponding to the driver identification information acquired by the device information acquisition unit, and acquires the identified user voice information from the storage unit. 11. A user voice acquisition unit that acquires user voice information related to the voice of a passenger from a storage unit in which the user voice information is stored in advance; a noise acquisition unit that acquires noise information related to noise inside the moving body; a superimposing unit that generates superimposed sound information by superimposing the user voice information and the noise information acquired by the user voice acquisition unit from the storage unit; a filter setting unit that sets a sound source separation filter using the superimposed sound information; a moving body sound acquisition unit that acquires moving body sound information related to sounds inside the moving body; a sound source separation unit that separates the moving body sound information using the sound source separation filter. 12. One or more computers that implement an information processing system using at least a communication terminal in a mobile body, acquiring user voice information relating to the voice of a passenger from a storage unit in which the user voice information is stored in advance; Acquire noise information regarding noise inside the moving body; generating superimposed sound information by superimposing the user voice information and the noise information acquired from the storage unit; Using the superimposed sound information, a sound source separation filter is set; acquiring internal sound information relating to sounds inside the moving body; A control method for an information processing system, comprising: separating the in-moving body sound information using the sound source separation filter. 13. One or more computers that implement an information processing system using at least a communication terminal in a mobile body, a step of acquiring user voice information from a storage unit in which user voice information relating to the voice of a passenger is stored in advance; a step of acquiring noise information regarding noise inside the moving body; a step of generating superimposed sound information by superimposing the user voice information and the noise information acquired from the storage unit; a step of setting a sound source separation filter using the superimposed sound information; a step of acquiring internal sound information relating to sounds inside the moving body; A program for executing a procedure for separating the in-moving body sound information using the sound source separation filter.
[0116] This application claims priority based on Japanese Patent Application No. 2023-040235, filed March 15, 2023, the disclosure of which is incorporated herein by reference in its entirety. [Explanation of symbols]
[0117] 1. Information processing equipment 2. Communication terminals 3 vehicles 4 Storage section 10 User voice acquisition unit 20 Noise acquisition section 30 Superimposed section 40 Filter setting section 50 Mobile internal sound acquisition unit 60 Sound source separation section 70 Detection unit 80 Device information acquisition unit 90 Control Unit 95 Memory Processing Unit 100 Information Processing Systems
Claims
1. An information processing system that uses at least a communication terminal in a mobile body, a user voice acquisition unit that acquires user voice information related to the voices of passengers from a storage unit in which the user voice information is stored in advance; a noise acquisition unit that acquires noise information related to noise inside the moving object; a superimposing unit that generates superimposed sound information by superimposing the user voice information and the noise information acquired by the user voice acquisition unit from the storage unit; a filter setting unit that sets a sound source separation filter using the superimposed sound information; a moving body sound acquisition unit that acquires moving body sound information related to sounds inside the moving body; a sound source separation unit that separates the moving body's internal sound information using the sound source separation filter.
2. 2. The information processing system according to claim 1, The user voice acquisition unit identifies the user voice information corresponding to a passenger in the vehicle from among the multiple pieces of user voice information, and acquires the identified user voice information from the memory unit.
3. 3. The information processing system according to claim 2, the information processing system is configured to function as a voice recognition system when a predetermined keyword is detected; The user voice acquisition unit uses keyword sound information relating to the sound of the predetermined keyword to identify the user voice information corresponding to the passenger who spoke inside the vehicle.
4. 4. The information processing system according to claim 3, The user voice acquisition unit identifies the user voice information corresponding to the passenger who first uttered the predetermined keyword after the information processing system was started.
5. 5. The information processing system according to claim 3, A detection unit is further provided to detect a change in the occupant of the moving body, An information processing system in which, after the detection unit detects that the passenger has changed, the user voice acquisition unit identifies the user voice information corresponding to the passenger who uttered the specified keyword and acquires the identified user voice information from the memory unit.
6. 6. The information processing system according to claim 5, The detection unit detects a change in the occupant using at least one of an image captured by an imaging unit mounted inside the vehicle, the opening and closing of a door of the vehicle, and the fastening and unfastening of a seat belt by the occupant.
7. 5. The information processing system according to claim 1, The filter setting unit updates the sound source separation filter using the moving body sound information.
8. 8. The information processing system according to claim 7, An information processing system, wherein the internal sound information of the moving body includes at least one of information regarding sound pressure, volume, and frequency of sound inside the moving body.
9. 9. The information processing system according to claim 8, The filter setting unit updates the sound source separation filter when at least one of the sound pressure, the volume, and the frequency satisfies a predetermined criterion.
10. 5. The information processing system according to claim 1, a device information acquisition unit that acquires driver identification information capable of identifying the driver from a communication device of the driver in the vehicle; The user voice acquisition unit identifies the user voice information corresponding to the driver identification information acquired by the device information acquisition unit, and acquires the identified user voice information from the storage unit.
11. a user voice acquisition unit that acquires user voice information related to the voices of passengers from a storage unit in which the user voice information is stored in advance; a noise acquisition unit that acquires noise information related to noise inside the moving body; a superimposing unit that generates superimposed sound information by superimposing the user voice information and the noise information acquired by the user voice acquisition unit from the storage unit; a filter setting unit that sets a sound source separation filter using the superimposed sound information; a moving body sound acquisition unit that acquires moving body sound information related to sounds inside the moving body; a sound source separation unit that separates the moving body sound information using the sound source separation filter.
12. One or more computers that realize an information processing system using at least a communication terminal in a mobile body, acquiring user voice information relating to the voice of a passenger from a storage unit in which the user voice information is stored in advance; Acquire noise information regarding noise inside the moving body; generating superimposed sound information by superimposing the user voice information and the noise information acquired from the storage unit; Using the superimposed sound information, a sound source separation filter is set; acquiring internal sound information relating to sounds inside the moving body; A control method for an information processing system, comprising: separating the in-moving body sound information using the sound source separation filter.
13. One or more computers that realize an information processing system that uses at least a communication terminal in a mobile body, a step of acquiring user voice information from a storage unit in which user voice information relating to the voice of a passenger is stored in advance; a step of acquiring noise information regarding noise inside the moving body; a step of generating superimposed sound information by superimposing the user voice information and the noise information acquired from the storage unit; a step of setting a sound source separation filter using the superimposed sound information; a step of acquiring internal sound information relating to sounds inside the moving body; A program for executing a procedure for separating the in-moving body sound information using the sound source separation filter.
Citation Information
Patent Citations
Voice recognizing device for vehicle
JP1995146698A
Characteristic setting method of noise reduction device
JP2002221986A
Speech detection device
JP2008299221A
Information processor
JP2009282645A
Sound source extracting device
JP2010054728A