Sound control method, sound control device, and program
The sound control device adjusts sound presentation based on cognitive load using biosignals, ensuring users remain attentive to both task and external sounds, thereby enhancing task performance.
Patent Information
- Application Number
- PCT/JP2024/036249
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-10-10
- Publication Date
- 2026-04-16
AI Technical Summary
Existing technologies fail to effectively adjust sound presentation based on cognitive load, affecting task performance and sound perception during tasks, as they do not consider how to maintain attention to sounds while performing tasks under varying cognitive loads.
A sound control device that utilizes biosignal acquisition to adjust sound volume and presentation based on cognitive load, using cyber-real fusion acoustics to overlay cyber sounds with external sounds through open-ear earphones, tailored to the listener's cognitive state.
Enables users to concentrate on tasks while being aware of external sounds, maintaining task performance by adjusting sound perception according to cognitive load.
Smart Images

Figure JP2024036249_16042026_PF_FP_ABST
Abstract
Description
Sound control method, sound control device, and program
[0001] The disclosed technology relates to a technique that adjusts the sound to be perceived based on the cognitive load calculated from biosignals.
[0002] Cognitive load refers to the amount of brain resources required to perform a task. The more difficult the task, the higher the cognitive load, which is reflected in biosignals such as brainwave amplitude and heart rate. It has been found that when a sound is presented during a high cognitive load, the perceived volume is lower than the actual sound. For this reason, previous research has investigated the characteristics of sounds that are easily noticed during a task. Non-patent document 1 describes a test in which participants were presented with sounds while performing a task and were given a listening test. The results showed that increasing the signal-to-noise ratio (SNR) improved the accuracy of listening to the sounds. However, it has not been investigated whether presenting sounds with a high SNR affects task performance. Non-patent document 2 reports that when sounds are presented during a task, cognitive load affects sound perception and thus affects task performance. When cognitive load is low, attention is more easily drawn to sounds, but task performance tends to decrease, while when cognitive load is high, attention is less easily drawn to sounds, but task performance does not decrease as much.
[0003] Yataka, et al., "Context-Dependent Speech Information Presentation Method for Wearable Computing," Transactions of the Information Processing Society of Japan, vol. 51, No. 12, pp. 2384-2395, 2010. Hughes et al., "Cognitive Control of Auditory Distraction: Impact of Task Difficulty, Foreknowledge, and Working Memory Capacity Supports Duplex-Mechanism Account," Journal of Experimental Psychology: Human Perception and Performance, vol. 39, No. 2, pp. 539-553, 2013.
[0004] Non-patent document 1 focuses on the perception of sound during task execution and does not consider task performance. Non-patent document 2 shows that sound perception affects task performance under different cognitive loads, but it does not clarify how to adjust and present sounds according to the cognitive load in order to maintain task performance while still being aware of the sound.
[0005] To solve the above problems, the sound control device relating to the disclosed technology comprises a biosignal acquisition unit, a sound processing unit, and a sound presentation unit. The biosignal acquisition unit acquires biosignal information of a listener performing a predetermined task. The sound processing unit sets the listening volume so that the change in the numerical value of the biosignal information falls within a predetermined range. The sound presentation unit emits a control sound from open-ear earphones to the listener at the set listening volume.
[0006] According to the disclosed technology, by adjusting external sounds based on cognitive load, users can concentrate on the sounds presented by the earphones while also being aware of external sounds.
[0007] A functional block diagram showing an example configuration of the ambient sound control device according to the first embodiment. A flowchart explaining an example of operation of the ambient sound control device according to the first embodiment. A functional block diagram showing an example configuration of the ambient sound control device according to the second embodiment. A flowchart explaining an example of operation of the ambient sound control device according to the second embodiment. A functional block diagram showing an example configuration of the ambient sound control device according to the third embodiment. A flowchart explaining an example of operation of the ambient sound control device according to the third embodiment. A diagram showing an example of the functional configuration of a computer.
[0008] The embodiments of the disclosed technology will be described in detail below. Components with the same function will be numbered identically, and redundant explanations will be omitted.
[0009] The disclosed technology uses a technique called cyber-real fusion acoustics, which alters the impression of external sounds (real sounds) by overlaying them with cyber sounds (artificially synthesized sounds) and presenting them to the listener. Cyber sounds are generated based on cognitive load calculated from biosignals. By playing cyber sounds through open-ear earphones and overlaying them with external sounds (real sounds), cyber-real fusion acoustics are provided that are tailored to the cognitive load. Because the sound is adjusted at the listener's ear using open-ear earphones, this technology can also be used for sounds that the listener cannot directly control, such as other people's voices or car horns.
[0010] [First Embodiment] Figure 1 is a functional block diagram showing an example configuration of an ambient sound control device according to the first embodiment, shown together with the Cyber-Real Fusion Sound System 1. The video conferencing device 101 communicates with a remote video conferencing device via, for example, a communication network 114. The video output of the video conferencing device 101 is input to the monitor 102. The audio output of the video conferencing device 101 is input to the ambient sound control device 106 and output from an open-ear type earphone 105 connected to the ambient sound control device 106. The listener 104 is assumed to be wearing the open-ear type earphone 105. The ringtone of the mobile phone 103 is assumed to be output into the air from the mobile phone without going through the earphone. The following explanation will take as an example a scenario in which the ringtone of the mobile phone 103 (ambient sound) is heard while the listener 104 is paying attention to the sound of the video conference (task sound).
[0011] The ambient sound control device 106 comprises a task sound acquisition unit 107, a biosignal acquisition unit 108, a cognitive load calculation unit 109, a sound processing unit 110, an ambient sound acquisition unit 111, a sound presentation unit 112, and a microphone 113. Figure 2 is a flowchart illustrating an example of the operation of the ambient sound control device 106. The first embodiment will be described below using Figures 1 and 2.
[0012] The task sound acquisition unit 107 of the ambient sound control device 106 acquires the task sound (sound of a video conference) that the listener is paying attention to from the video conferencing device 101 (step S201).
[0013] The biosignal acquisition unit 108 acquires brainwave and heart rate data from an electroencephalograph (not shown) attached to the listener (step S202). A commercially available sensor can be used for the electroencephalograph.
[0014] The cognitive load calculation unit 109 calculates cognitive load from electroencephalogram (EEG) and heart rate data (step S204). Conventional machine learning models or software provided by the EEG manufacturer can be used for the calculation. In the first embodiment, a machine learning model is used. When a machine learning model is used, the input to the model is the waveform data of the EEG and heart rate. The output of the model is the level of cognitive load, and it outputs a binary classification result (high / low). The machine learning model can utilize structures such as VGG or ResNet, as used in Reference 1 below. When actually training this model, the dataset corresponding to EEG / heart rate data and cognitive load can be prepared independently or publicly available datasets can be used.
[0015] Reference 1: Angkan et al., "Multimodal Brain-Computer Interface for In-Vehicle Driver Cognitive Load Measurement: Dataset and Baselines", IEEE Transactions on Intelligent Transportation Systems, 2024.
[0016] The ambient sound acquisition unit 111 acquires the ringtone of the mobile phone 103 using the microphone 113 (step S203). The microphone is positioned so that the volume of the ringtone that the listener hears directly and the volume of the ringtone acquired by the microphone are as close as possible.
[0017] The sound processing unit 110 obtains cognitive load (high / low) from the cognitive load calculation unit 109 and ambient sound from the ambient sound acquisition unit 111, and determines the volume of the ambient sound to be perceived by the listener 104 (in this case, the volume of the ringtone) based on the cognitive load and ambient sound (step S205). The method for calculating the volume is shown below. The ringtone volume acquired by the microphone 113 is V callTherefore, the ringtone volume V' that listener 104 should perceive. call Calculate it as follows: The coefficient α is predetermined, with 0 < α ≤ 1 when the cognitive load level is low, and α > 1 when it is high. According to equation (1), the volume of the mobile phone's ringtone is lowered when the cognitive load is low, and raised when it is high. Finally, the sound processing unit 110 plays V' near the listener's ear. call As perceived, the ringtone output of the earphones is superimposed on the external ringtone of open-ear earphones. call_cyber (A control sound) is generated (step S206).
[0018] The sound presentation unit 112 superimposes the task sound acquired from the task sound acquisition unit 107 with the control sound generated by the sound processing unit 110 (step S207), and outputs it from the open-ear earphone 105 (step S208).
[0019] As described above, by overlaying cyber sounds (control sounds) generated according to the listener's cognitive load state onto external sounds (ambient sounds), listeners can more easily notice their mobile phone's ringtone while concentrating on a video conference. Note that task sounds, biosignals, and ambient sounds are acquired sequentially, and the level of cognitive load is monitored in real time. Also note that when no ambient sounds are acquired, the earphone output is nothing more than the task sound. This concludes the description of the first embodiment.
[0020] [Second Embodiment] In the first embodiment, the cognitive load of the listener was determined, the desired ringtone volume was determined using the cognitive load, and a control tone was generated based on the ringtone volume. In addition to cognitive load, the ease with which sound is perceived may change depending on mental state and other factors. Therefore, in the second embodiment, while considering cognitive load, sound processing is performed from biosignals, taking into account multiple factors, without directly calculating the level of cognitive load.
[0021] Figure 3 is a functional block diagram showing an example configuration of the ambient sound control device according to the second embodiment, along with the cyber-real fusion sound system 3. It differs from Figure 1 in that the cognitive load calculation unit 109 is omitted in the ambient sound control device 106. Hereinafter, as with the first embodiment, we will explain using as an example a scenario in which the listener 104 hears the ringtone of a mobile phone 103 (ambient sound) while paying attention to the sound of a video conference (task sound).
[0022] Figure 4 is a flowchart illustrating an example of the operation of the ambient sound control device 106. The second embodiment will be described below using Figures 3 and 4.
[0023] Steps S201, S202, and S203 in Figure 4 are the same as in the first embodiment. The sound processing unit 301 uses electroencephalogram and heart rate data acquired by the biosignal acquisition unit 108 and ambient sound (V) acquired by the ambient sound acquisition unit. call The input is the mobile phone ringtone (V) output from open-ear earphones. call_cyber (Step S401) generates a control sound.
[0024] The sound processing unit 301 of the second embodiment is configured as an end-to-end model. The model learns to optimize two performances: for the listener to notice the ringtone of their mobile phone while participating in a video conference, referencing the architecture of conventional machine learning models. Specifically, brainwave and heart rate data are used as training data, and data corresponding to the volume of sounds that are easily perceived by the listener, which represent these biosignals, are used as ground truth data. This end-to-end model makes it possible to adjust the volume of ambient sounds without depending solely on the level of cognitive load.
[0025] Steps S207 and S208 are the same as in the first embodiment. This concludes the description of the second embodiment.
[0026] [Third Embodiment] In the third embodiment, the ringtone acquired in step S203 is saved, and if the listener does not notice the ringtone on the mobile phone, it is snoozed and played again through the earphones to draw their attention.
[0027] FIG. 5 is a functional block diagram showing a configuration example of an environmental sound control device according to the third embodiment, shown together with the cyber-real fusion sound system 5. It is different from FIG. 1 in that a response determination unit 501 and a recording unit 502 are added to the environmental sound control device 106. FIG. 6 is a flowchart for explaining an example of the operation of the environmental sound control device 106. Hereinafter, the third embodiment will be described using FIGS. 5 and 6.
[0028] Steps S201 to S208 are the same as those in the first embodiment. However, the environmental sound acquisition unit 111 records the acquired incoming call sound (environmental sound) in the recording unit 502. The response determination unit 501 determines whether the listener 104 is aware of the incoming call sound of the mobile phone. For example, it is determined by checking whether the listener shows a reaction such as picking up the mobile phone or whether the listener's cognitive load changes from high to low. If it is determined that the listener is not aware of the incoming call sound of the mobile phone (No in step S601), the sound processing unit 110 updates V' call (the volume to be perceived) (step S602). Two examples of update timing are shown.
[0029] <Update after a predetermined time has elapsed> When it is determined that the listener has not noticed the incoming call sound and after t seconds, V' is updated by the following formula. call The value of t is determined in advance. The coefficient α has the same value as in step S205. The coefficient β is a value of 0 or more determined in advance. <Update at the timing when the cognitive load becomes low> After it is determined that the listener has not noticed the incoming call sound, at the timing when the listener's cognitive load becomes low (that is, when the load of the video conference has decreased), V' is updated by formula (2). The coefficient α has the same value (the value corresponding to low) as in step S205. The coefficient β is a value of 0 or more determined in advance.
[0030] The sound processing unit 110 generates an incoming call sound V call (control sound) of the earphone output to be superimposed on the incoming call sound outside the open-ear type earphone so that V' is perceived near the listener's ear (step S206).
[0031]
[0032] call call_cyber
[0001]
[0032] The sound presentation unit 112 superimposes the task sound acquired from the task sound acquisition unit 107 with the control sound generated by the sound processing unit 110 (step S207), and outputs it from the open-ear earphone 105 (step S208).
[0033] The above is a description of the third embodiment.
[0034] [Supplement] In the first and third embodiments, examples were shown of calculating cognitive load from electroencephalogram (EEG) and heart rate data. Cognitive load can also be calculated from other biological signals such as eye movements (gaze, blinking, etc.), electromyography (EMG), sweating, cerebral blood flow, and skin electrical activity, in addition to EEG and heart rate data. For this reason, the biological signal acquisition unit 108 may acquire biological signals other than EEG and heart rate data, such as eye movements, EMG, sweating, cerebral blood flow, and skin electrical activity, and use them to calculate cognitive load. Cognitive load can also be estimated from the user's behavior. Cognitive load may be determined using the user's behavior instead of, or in conjunction with, biological signals.
[0035] It should be noted that the cases described in each embodiment are merely examples. In the first embodiment, we described controlling the ringtone of a mobile phone perceived by the user in an environment where the user is participating in a video conference, according to the cognitive load. However, this can be applied to any case in which the volume is controlled to attract the user's attention in accordance with the user's cognitive load, which changes as the user is performing some action. For example, the cognitive load of a user who is watching television or operating a handset may be obtained, and the volume of a doorbell (a sound that notifies of a visitor, etc.) may be controlled according to the cognitive load.
[0036] Furthermore, in the above embodiment, the volume of the ringtone was controlled according to the calculated cognitive load, but it is also possible to control smartphone screen notifications or vibration notifications from wearable devices. For example, in the case of smartphone screen notifications, the area and text of the notification displayed on the smartphone are enlarged when the user's cognitive load is high, and the area and text are reduced when the user's cognitive load is low. In the case of vibration notifications from wearable devices, the vibration of the wearable device is controlled to be stronger the greater the cognitive load.
[0037] [Program, Recording Medium] The functions realized by the components described in this specification may be implemented in circuitry or processing circuitry including a general-purpose processor, a specific-purpose processor, an integrated circuit, ASICs (Application Specific Integrated Circuits), a CPU (a Central Processing Unit), a conventional circuit, and / or a combination thereof, programmed to realize the described functions. The processor includes transistors and other circuits and is regarded as circuitry or processing circuitry. The processor may be a programmed processor that executes a program stored in a memory.
[0038] In this specification, circuitry, unit, and means are hardware programmed to realize the described functions or hardware that executes them. The hardware may be any hardware disclosed in this specification or any hardware known to be programmed or execute to realize the described functions.
[0039] When the hardware is a processor regarded as a type of circuitry, the circuitry, means, or unit is a combination of hardware and software used to configure the hardware and / or the processor.
[0040] The above various processes can be implemented by causing the recording unit 2020 of the computer 2000 shown in FIG. 7 to read a program that causes each step of the above method to be executed and operating it on the control unit 2010, the input unit 2030, the output unit 2040, the display unit 2050, etc.
[0041] The program describing this processing content can be recorded on a computer-readable recording medium. As the computer-readable recording medium, for example, any of a magnetic recording device, an optical disk, a magneto-optical recording medium, a semiconductor memory, etc. may be used.
[0042] Also, the distribution of this program can be carried out, for example, by selling, transferring, lending, etc. a portable recording medium such as a DVD or CD-ROM on which the program is recorded. Further, it is also possible to configure the distribution of this program by storing the program in a storage device of a server computer and transferring the program from the server computer to other computers via a network.
[0043] A computer that executes such a program first stores, for example, the program recorded on a portable recording medium or the program transferred from a server computer once in its own storage device. Then, at the time of executing the process, this computer reads the program stored in its own recording medium and executes the process according to the read program. Also, as another execution form of this program, it is also possible that the computer directly reads the program from the portable recording medium and executes the process according to the program. Further, every time a program is transferred from the server computer to this computer, it is also possible to sequentially execute the process according to the received program. Also, it is also possible to configure the process to be executed by a so-called ASP (Application Service Provider) type service that realizes the processing function only by the execution instruction and result acquisition without transferring the program from the server computer to this computer. Furthermore, it is also possible to configure the execution of the process of the terminal by using a so-called SaaS (Software as a Service) type service in which a part of the server computer is used by the user together with the program. Note that the program in this embodiment includes information used for processing by an electronic computer and similar to the program (data, etc. that are not direct instructions for the computer but have the property of defining the processing of the computer).
[0044] Also, in this embodiment, the present device is configured by causing a predetermined program to be executed on a computer, but at least a part of these processing contents may be realized hardware-wise.
Claims
1. A sound control method comprising: a biosignal acquisition unit acquiring biosignal information of a listener performing a predetermined task; a sound processing unit setting the listening volume so that the change in the numerical value of the biosignal information falls within a predetermined range; and a sound presentation unit emitting a control sound to the listener at the specified listening volume.
2. A sound control method according to claim 1, wherein the listener listens to sounds associated with the task (task sounds) using open-ear acoustic equipment, and is also able to listen to both the task sounds and sounds occurring around the listener (ambient sounds), the listening volume is the volume of the ambient sounds that the listener is to perceive, and the control sound is generated by the sound processing unit and emitted by the sound presentation unit from the open-ear acoustic equipment.
3. A sound control method according to claim 2, wherein a cognitive load calculation unit estimates the cognitive load of the listener from the biological information, an ambient sound acquisition unit acquires the ambient sound, and the sound processing unit sets the listening volume to a low setting volume that is less than the ambient sound volume when the cognitive load is low, and sets the listening volume to a high setting volume that is greater than the ambient sound volume when the cognitive load is high.
4. A sound control method according to claim 3, wherein, after a predetermined time has elapsed since the emission of the control sound, the sound processing unit increases the low set volume or the high set volume by a predetermined value and updates it, and the sound presentation unit emits the control sound to the listener so that the updated listening volume is achieved.
5. A sound control method according to claim 3, wherein, after emitting the control sound, when the cognitive load changes from a high state to a low state, the sound processing unit increases and updates the low set volume by a predetermined value, and the sound presentation unit emits the control sound to the listener so that the updated listening volume becomes the listening volume.
6. A sound control method according to claim 2, wherein an ambient sound acquisition unit acquires ambient sound, and a sound processing unit generates a control sound based on the biological information and ambient sound so as to optimize the efficiency of the task and the efficiency of recognizing the ambient sound.
7. A sound control device comprising: a biosignal acquisition unit that acquires biosignal information of a listener performing a predetermined task; a sound processing unit that sets the listening volume so that the change in the numerical value of the biosignal information falls within a predetermined range; and a sound presentation unit that emits a control sound to the listener so that the listening volume is set to the predetermined level.
8. A program for causing a computer to perform the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Sound reproducing device
JP2005037649A
Information processing system and program
JP2021090136A
Information processing apparatus and program
JP2022042227A
Volume controlled headphone using brain wave detection
KR102370081B1