Determination device
The determination device predicts silence in a vehicle interior by analyzing sound frequencies for evoked otoacoustic emissions, addressing the lack of silence prediction in existing technologies and providing a method to determine quietness through onomatopoeia detection.
Patent Information
- Application Number
- JP2024059076
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-01
- Publication Date
- 2025-10-14
AI Technical Summary
Existing technologies lack the ability to predict whether a vehicle interior is silent based on human perception of quietness, as indicated by the presence or absence of onomatopoeia.
A determination device that utilizes a determination model to analyze sound data for specific frequencies indicative of evoked otoacoustic emissions between 1800 Hz and 2200 Hz, predicting silence by detecting or absence of onomatopoeia through a display device.
Enables prediction of silence in a vehicle interior based on human perception, utilizing a determination device that displays the presence or absence of onomatopoeia, thereby determining quietness.
Smart Images

Figure 2025155310000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a determination device. [Background technology]
[0002] Patent Document 1 discloses a vehicle interior sound field evaluation device that can accurately evaluate the quietness of a vehicle interior. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2019-045824 Summary of the Invention [Problem to be solved by the invention]
[0004] However, Patent Document 1 does not disclose any technology for predicting whether the interior of a vehicle is silent. Generally, if an onomatopoeia that indicates the silence perceived by humans occurs in a quiet environment, it can be predicted that the space in which the onomatopoeia occurs is silent.
[0005] Therefore, an object of the present disclosure is to provide a determination device that can predict whether a specified space is silent or not based on the presence or absence of onomatopoeia that indicate the quietness perceived by humans in a quiet environment. [Means for solving the problem]
[0006] The determination device according to claim 1 comprises: an input unit that inputs predetermined sound data obtained by frequency analysis of a predetermined sound and converting it into numerical data into a determination model that, when sound data is input, outputs whether a specific frequency of 1800 Hz or more and 2200 Hz or less, which indicates evoked otoacoustic emissions, is detected in addition to the frequency of the input sound data; and a control unit that, when the determination model to which the predetermined sound data has been input outputs that the specific frequency has been detected, displays on a display device that an onomatopoeia indicating quietness as perceived by humans in a quiet environment is generated; and, when the determination model outputs that the specific frequency has not been detected, displays on the display device that the onomatopoeia is not generated.
[0007] In the determination device according to claim 1, an input unit inputs predetermined sound data, which is obtained by frequency analysis of the predetermined sound and converting it into numerical data, into a determination model that, when inputting sound data, outputs whether a specific frequency between 1800 Hz and 2200 Hz inclusive, which indicates evoked otoacoustic emissions, is detected in addition to the frequency of the input sound data. If the determination model to which the predetermined sound data has been input outputs that a specific frequency has been detected, the control unit displays on the display device that the onomatopoeia has occurred, and if the control unit outputs that the specific frequency has not been detected, displays on the display device that the onomatopoeia has not occurred. This makes it possible for the determination device to predict whether a specific space is silent based on the presence or absence of an onomatopoeia that indicates the silence perceived by humans in a quiet environment. [Effects of the Invention]
[0008] The determination device according to the present disclosure can predict whether a predetermined space is silent or not based on the presence or absence of onomatopoeia that indicate the quietness perceived by humans in a quiet environment. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 2 is a block diagram showing a schematic configuration of a determination device. [Figure 2] 10 is a flowchart showing the flow of a specification process. [Figure 3] FIG. 2 is an explanatory diagram showing a display example of a display device. DETAILED DESCRIPTION OF THE INVENTION
[0010] The determination device 50 according to this embodiment will be described below. 1 is a block diagram showing the hardware configuration of the determination device 50. As an example, the determination device 50 may be a general-purpose computer device such as a server computer or a PC (Personal Computer), or a mobile terminal such as a smartphone or a tablet terminal. In this embodiment, the determination device 50 is a "PC."
[0011] 1, the determination device 50 includes a CPU (Central Processing Unit) 51, a ROM (Read Only Memory) 52, a RAM (Random Access Memory) 53, a storage 54, an input device 55, a display device 56, and a communication unit 57. Each component is connected to each other via a bus 58 so as to be able to communicate with each other.
[0012] The CPU 51 is a central processing unit that executes various programs and controls each part. That is, the CPU 51 reads programs from the ROM 52 or the storage 54 and executes the programs using the RAM 53 as a work area. The CPU 51 controls each of the above components and performs various arithmetic processing in accordance with the programs stored in the ROM 52 or the storage 54.
[0013] The ROM 52 stores various programs and various data. The RAM 53 serves as a working area for temporarily storing programs or data.
[0014] The storage 54 is configured by a storage device such as a hard disk drive (HDD), a solid state drive (SSD), or a flash memory, and stores various programs and various data.
[0015] A determination program 54A and a determination model 54B are stored in the storage 54. The determination program 54A is a program for causing the CPU 51 to execute specific processing, which will be described later.
[0016] The determination model 54B is a numerical model that, when input with sound data, outputs whether a specific frequency between 1800 Hz and 2200 Hz, inclusive, indicating evoked otoacoustic emissions, is detected in addition to the frequency of the input sound data. Specifically, the determination model 54B is a numerical model of the human cochlea, which is responsible for frequency analysis and evoked otoacoustic emissions. The determination model 54B reproduces the behavior of the basilar membrane, which causes evoked otoacoustic emissions, and after inputting sound data, calculates whether a specific frequency other than the frequency of the sound data is detected and outputs the result.
[0017] The input device 55 includes, for example, a pointing device such as a mouse, various buttons, a keyboard, a microphone, a camera, and the like, and is used to perform various inputs.
[0018] The display device 56 is, for example, a liquid crystal display, and displays various information. The display device 56 may function as the input device 55 by adopting a touch panel system.
[0019] The communication unit 57 is an interface for communicating with other devices. For this communication, for example, a wired communication standard such as Ethernet (registered trademark) or FDDI, or a wireless communication standard such as 4G, 5G, or WI / Fi (registered trademark) is used.
[0020] The CPU 51 of the determination device 50 has, as functional components, an acquisition unit 51A, an input unit 51B, and a control unit 51C. Each functional component is realized by the CPU 51 reading and executing a determination program 54A stored in the storage 54.
[0021] The acquisition unit 51A acquires various types of sound data stored in the storage 54. The storage 54 stores a plurality of pieces of sound data indicating various types of recorded sounds.
[0022] The input unit 51B performs frequency analysis on the sound data indicating the predetermined sound acquired by the acquisition unit 51A from among the plurality of sound data, converts it into numerical data, and inputs the predetermined sound data, which is the numerical data, to the determination model 54B.
[0023] When the determination model 54B, to which the predetermined sound data has been input, outputs that a specific frequency has been detected, the control unit 51C displays on the display device 56 that an onomatopoeia indicating the quietness perceived by humans in a quiet environment has been generated. The onomatopoeia is, for example, "silence" or "scene." Hereinafter, the occurrence of the onomatopoeia may be referred to as "a sense of scene is generated," and the absence of the onomatopoeia may be referred to as "no sense of scene is generated."
[0024] It is speculated that the reason why humans perceive a sense of scene is due to evoked otoacoustic emissions (OTOEs), which are sounds that are emitted from the cochlea of the inner ear due to changes in sound pressure or quiet environments. It is also speculated that the perception of a sense of scene is due to the nonlinear behavior of outer hair cells on the basilar membrane of the cochlea. The applicant then discovered that a numerical model of the cochlea can reproduce phenomena related to evoked OTOEs. Specifically, the applicant discovered that when sound data near 100 Hz is input into a numerical model (e.g., determination model 54B) that reproduces the behavior of the basilar membrane of the cochlea, a basilar membrane vibration response occurs at a frequency between 1800 Hz and 2200 Hz, indicative of evoked OTOEs, in addition to the 100 Hz frequency. Therefore, it can be predicted that after inputting sound data into determination model 54B, the occurrence of a basilar membrane vibration response between 1800 Hz and 2200 Hz will trigger a perception of a sense of scene.
[0025] Furthermore, when the determination model 54B to which the predetermined sound data has been input outputs that a specific frequency has not been detected, the control unit 51C displays a message indicating that a sense of scene is not generated on the display device 56. An example of the display on the display device 56 will be described later with reference to FIG.
[0026] 2 is a flowchart showing the flow of the identification process executed by the determination device 50. The CPU 51 reads the determination program 54A from the storage 54, expands it in the RAM 53, and executes it, thereby performing the identification process. As an example, the identification process is performed when a user operates the determination device 50 and a predetermined application installed in the determination device 50 is executed.
[0027] 2, the CPU 51 acquires sound data indicating a predetermined sound selected by the user from among a plurality of pieces of sound data stored in the storage 54. Then, the CPU 51 proceeds to step S11.
[0028] In step S11, the CPU 51 performs frequency analysis on the sound data indicating the predetermined sound acquired in step S10 and converts the converted sound data into numerical data, and inputs the numerical data to the determination model 54B. Then, the CPU 51 proceeds to step S12.
[0029] In step S12, the CPU 51 inputs the predetermined sound data into the determination model 54B and then determines whether a sense of scene is created. If the CPU 51 determines that a sense of scene is created (step S12: YES), the process proceeds to step S13. On the other hand, if the CPU 51 determines that a sense of scene is not created (step S12: NO), the process proceeds to step S14.
[0030] In step S13, the CPU 51 displays a message indicating that a scene feeling will occur on the display device 56. Then, the CPU 51 ends the specification process.
[0031] In step S14, the CPU 51 displays a message that no scene feeling is generated on the display device 56. Then, the CPU 51 ends the specification process.
[0032] Next, a display example of the display device 56 will be described with reference to Fig. 3. Fig. 3(A) shows a display example when the CPU 51 determines that a sense of scenery will be created, and Fig. 3(B) shows a display example when the CPU 51 determines that a sense of scenery will not be created.
[0033] As shown in FIG. 3(A), the display device 56 displays basilar membrane behavior information 60 and scene sensation generation information 90.
[0034] The basilar membrane behavior information 60 shows the behavior of the basilar membrane reproduced by the determination model 54B to which the predetermined sound data has been input. The basilar membrane behavior information 60 shown in FIG. 3A shows a case in which predetermined sound data obtained by frequency-analyzing the door closing sound generated when a vehicle door is closed and converting the sound data into numerical data is input to the determination model 54B as the predetermined sound. The basilar membrane behavior information 60 displays first data 70, which is waveform data of the frequency of the door closing sound, and second data 80, which is waveform data of the frequency of a sound different from the door closing sound that occurred a predetermined time (e.g., 0.5 to 0.7 seconds) after the first data 70 was generated. The frequency (basilar membrane position) indicated by the first data 70 is 100 Hz, and the frequency (basilar membrane position) indicated by the second data 80 is between 1800 Hz and 2200 Hz (e.g., 2000 Hz). In this case, the CPU 51 determines that a scene is generated. Therefore, the scene feeling generation information 90 shown in FIG. 3(A) is displayed with a circle 92, indicating that a scene feeling will be generated.
[0035] Next, the basilar membrane behavior information 60 shown in FIG. 3(B) shows a case where predetermined sound data obtained by frequency analysis of a predetermined sound other than the door closing sound and converting it into numerical data is input to the determination model 54B as the predetermined sound. The basilar membrane behavior information 60 displays third data 75, which is waveform data of the frequency of the other predetermined sound. The frequency (basilar membrane position) indicated by the third data 75 is 300 Hz. Unlike FIG. 3(A), the basilar membrane behavior information 60 does not display waveform data corresponding to the second data 80 after a predetermined time (e.g., 0.5 to 0.7 seconds) has elapsed since the occurrence of the third data 75. In this case, the CPU 51 determines that no sense of scene is generated. Therefore, the sense of scene generation information 90 shown in FIG. 3(B) displays a cross 94, indicating that no sense of scene is generated.
[0036] As described above, in the determination device 50, the CPU 51 inputs predetermined sound data, which is obtained by frequency analysis of the predetermined sound and converting it into numerical data, into the determination model 54B, which, upon inputting sound data, outputs whether a specific frequency between 1800 Hz and 2200 Hz, inclusive, indicating evoked otoacoustic emissions, is detected in addition to the frequency of the input sound data. If the determination model 54B, to which the predetermined sound data has been input, outputs that a specific frequency has been detected, the CPU 51 displays on the display device 56 that a sense of scene has occurred. If the determination model 54B outputs that a specific frequency has not been detected, the CPU 51 displays on the display device 56 that a sense of scene has not occurred. Thus, the determination device 50 can predict whether a specific space is silent based on whether a sense of scene has occurred. It has been confirmed that the sense of scene perceived by a person in a vehicle cabin after closing a vehicle door is correlated with evoked otoacoustic emissions. Therefore, the determination device 50 can be used to predict whether a vehicle cabin is silent.
[0037] (others) The determination model 54B may be a predictive numerical model based on AI (Artificial Intelligence). In this case, the determination model 54B predicts whether a specific frequency will be detected by learning the causal relationship between the input predetermined sound data and the predicted data based on an algorithm using deep learning or the like. In this case, the determination model 54B may be machine-learned using, for example, the envelope area 85 of the waveform data corresponding to the second data 80 in FIG. 3(A) as a feature. Alternatively, the determination model 54B may be machine-learned using feature amounts such as Mel-frequency cepstrum coefficients and chromagrams.
[0038] In the above embodiment, the specific process executed by the CPU 51 after reading the software (program) may be executed by various processors other than a CPU. Examples of such processors include programmable logic devices (PLDs) (such as field-programmable gate arrays (FPGAs)) whose circuit configuration can be changed after fabrication, and dedicated electrical circuits such as application-specific integrated circuits (ASICs) that are processors with circuit configurations specifically designed to execute specific processes. The specific process may be executed by one of these processors, or by a combination of two or more processors of the same or different types (e.g., multiple FPGAs, or a combination of a CPU and an FPGA). The hardware structure of these processors is, more specifically, an electrical circuit that combines circuit elements such as semiconductor devices.
[0039] In the above embodiment, the determination program 54A is pre-stored (installed) in the storage 54, but the present invention is not limited to this. The determination program 54A may be provided in a form recorded on a recording medium such as a CD-ROM (Compact Disk Read Only Memory), a DVD-ROM (Digital Versatile Disk Read Only Memory), or a USB (Universal Serial Bus) memory. The determination program 54A may also be downloaded from an external device via a network. [Explanation of symbols]
[0040] 50 Judgment device 51B Input section 51C Control Unit 54B Judgment Model 56 Display device
Claims
[Claim 1] an input unit that inputs predetermined sound data obtained by frequency analysis of the predetermined sound and converting it into numerical data into a determination model that outputs whether or not a specific frequency between 1800 Hz and 2200 Hz, which indicates evoked otoacoustic emissions, is detected in addition to the frequency of the input sound data; a control unit that displays on a display device, when the determination model to which the predetermined sound data has been input, outputs that the specific frequency has been detected, a message indicating that an onomatopoeia indicating the quietness perceived by humans in a quiet environment is generated, and, when the determination model outputs that the specific frequency has not been detected, a message indicating that the onomatopoeia is not generated; A determination device comprising:
Citation Information
Patent Citations
Vehicle interior sound field evaluation device, vehicle interior sound field evaluation method, vehicle interior sound field control device, and interior sound field evaluation device
JP2019045824A
Cited By
Device for measuring gear parameters with manual gear roll tester
US20250027841A1