Determination device, determination method, and program
The determination device uses synthesized sounds and pulse rate analysis to assess auditory discrimination ability, addressing the need for therapist-assisted evaluation by providing a self-evaluation method for patients with attention disorders.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- CITIZEN WATCH CO LTD
- Filing Date
- 2022-12-09
- Publication Date
- 2026-05-26
AI Technical Summary
Existing rehabilitation apparatuses for patients with attention disorders, such as those described in Patent Document 1, require the assistance of a speech therapist to determine the cocktail party effect, making it difficult to assess patients with poor communication skills.
A determination device that outputs synthesized sounds combining background and predetermined sounds at random timings, detects pulse rate, and determines auditory discrimination ability based on pulse rate and emotional responses, without requiring subject interaction.
Enables easy and accurate determination of a subject's auditory discrimination ability by analyzing pulse rate and emotional responses, facilitating assessment without therapist intervention.
Smart Images

Figure 0007865869000001 
Figure 0007865869000002 
Figure 0007865869000003
Abstract
Description
Technical Field
[0001] The present invention relates to a determination device, a determination method, and a program.
Background Art
[0002] There is known a brain function called the cocktail party effect in which important information for oneself is unconsciously selected from various everyday sounds such as human voices and music. Since the cocktail party effect is a function of the brain rather than the ear, it is known that its function deteriorates due to mental stress and improves due to training.
[0003] Patent Document 1 describes an apparatus for rehabilitating a patient's attention disorder by having the patient listen to degraded voice data created by superimposing noise data on voice data and having the patient answer the content of the voice data.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] The apparatus described in Patent Document 1 is used when a speech therapist performs rehabilitation on a patient. In this apparatus, since the patient must answer on their own that they were able to hear the test voice data with the assistance of a speech therapist, there is a problem that patients with poor communication skills cannot be properly examined. Therefore, there is a need to enable easier determination of the cocktail party effect without the assistance of a speech therapist.
[0006] The present invention was made to solve the above-mentioned problems, and aims to provide a determination device, determination method, and program that enable easy determination of the auditory discrimination ability of a subject's brain. [Means for solving the problem]
[0007] The determination device according to an embodiment of the present invention is characterized by comprising: an output unit that outputs a synthesized sound obtained by combining a background sound and a predetermined sound output at random timings such that the difference between the frequency of the main sound of the background sound and the frequency of the main sound of the predetermined sound satisfies a predetermined condition; a pulse rate detection unit that detects the pulse rate of the subject while the synthesized sound is being output; and a determination unit that determines the acoustic discrimination ability of the subject's brain based on the pulse rate.
[0008] Furthermore, the determination device preferably includes an emotion detection unit that detects the subject's emotions based on their pulse rate, and the determination unit determines the subject's brain's ability to distinguish acoustics based on the subject's emotions.
[0009] Furthermore, it is preferable that the designated sound is the voice of the subject's name, and the background sound is music selected by the subject.
[0010] Furthermore, the determination device preferably includes a selection unit that accepts the selection of one of a plurality of pre-set conditions regarding the difference in the frequencies of the main sounds in response to the subject's operation, and the output unit outputs a synthesized sound that satisfies the selected condition in terms of the difference between the frequency of the main sound of the background sound and the frequency of the main sound of the predetermined sound.
[0011] A determination method according to an embodiment of the present invention is a determination method performed by a determination device, characterized in that it outputs a synthesized sound which is produced by combining a background sound and a predetermined sound which is output at random timings such that the difference between the frequency of the main sound of the background sound and the frequency of the main sound of the predetermined sound satisfies a predetermined condition, the pulse rate of the subject while the synthesized sound is being output is detected, and the acoustic discrimination ability of the subject's brain is determined based on the pulse rate.
[0012] A program according to an embodiment of the present invention is characterized by causing a computer to perform the following actions: output a synthesized sound, which is created by combining a background sound and a predetermined sound output at random timings such that the difference between the frequency of the main sound of the background sound and the frequency of the main sound of the predetermined sound satisfies a predetermined condition; detect the pulse rate of the subject while the synthesized sound is being output; and determine the subject's brain's ability to distinguish acoustics based on the pulse rate. [Effects of the Invention]
[0013] The determination device, determination method, and program according to the present invention enable the easy determination of a subject's auditory discrimination ability by the brain. [Brief explanation of the drawing]
[0014] [Figure 1] This diagram shows the schematic configuration of the determination device 1. [Figure 2] This diagram shows an example of the settings screen G1. [Figure 3] This figure shows an example of the results screen G2. [Figure 4] This is a flowchart showing the flow of the decision-making process. [Figure 5] This is a schematic diagram illustrating the method for detecting pulse rate. [Figure 6] (A) and (B) are schematic diagrams illustrating a method for determining the auditory discrimination ability of a subject's brain. [Modes for carrying out the invention]
[0015] Various embodiments of the present invention will be described below with reference to the drawings. Please note that the technical scope of the present invention is not limited to these embodiments, but extends to the invention described in the claims and its equivalents.
[0016] FIG. 1 is a functional block diagram of a determination device 1 according to an embodiment of the present invention. The determination device 1 is an information processing terminal such as a smartphone, a mobile phone, a tablet terminal, a PC (Personal Computer), or a portable game machine. The determination device 1 outputs a synthesized sound in which a name voice that reads out the name of a subject and background sound are synthesized. Further, the determination device 1 determines the acoustic discrimination ability of the subject by the brain by detecting the pulse rate of the subject when the name voice is output. The determination device 1 includes a storage unit 11, an imaging unit 12, an audio output unit 13, a display unit 14, an operation unit 15, and a processing unit 16.
[0017] The storage unit 11 is a configuration for storing programs and data, and includes, for example, a semiconductor memory. The storage unit 11 stores, as programs, an operating system program, a driver program, an application program, etc. used for processing by the processing unit 16. The program is installed in the storage unit 21 from a computer-readable and non-temporary portable storage medium such as a CD-ROM (Compact Disc Read Only Memory) or a DVD-ROM (Digital Versatile Disc Read Only Memory).
[0018] The imaging unit 12 is a configuration for imaging a learner who views emotion induction content, and includes a camera. The camera includes an imaging optical system for forming an image on a light receiving surface, a photoelectric conversion element such as a CCD (Charge Coupled Device) sensor that is two-dimensionally arranged on the light receiving surface and outputs an electric signal according to the amount of incident light, and an image generation circuit that generates an image based on the output of the photoelectric conversion element. The imaging unit 12 generates an image based on a control signal supplied from the processing unit 16, and supplies the generated image to the processing unit 16.
[0019] The audio output unit 13 is a configuration for outputting audio, and includes a speaker. The audio output unit 13 outputs audio by converting an electric signal supplied from the processing unit 16 into mechanical vibration.
[0020] The display unit 14 is a component for displaying images, and includes, for example, a liquid crystal display or an organic EL (Electro Luminescence) display. The display unit 14 displays an image based on the display data supplied from the processing unit 16.
[0021] The operation unit 15 is a component for receiving user operations, and includes, for example, a keyboard, a keypad, and a mouse. The operation unit 15 may include a touch panel and may be integrated with the display unit 14. The operation unit 15 generates an operation signal corresponding to the user operation and supplies it to the processing unit 16.
[0022] The processing unit 16 is a device that comprehensively controls the operation of the determination device 1, and includes one or more processors and their peripheral circuits. The processing unit 16 includes, for example, a CPU (Central Processing Unit). The processing unit 16 may include a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), an LSI (Large Scale Integration), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), etc. The processing unit 16 controls the operation of each component and executes various processes so that various processes of the determination device 1 are executed in an appropriate procedure based on the program stored in the storage unit 11 and the operation signal from the operation unit 15.
[0023] The processing unit 16 has reception unit 161, generation unit 162, determination unit 163, image acquisition unit 164, output control unit 165, pulse rate detection unit 166, and determination unit 167 as functional blocks. Each of these units is a functional module realized by a program executed by the processing unit 16. Each of these units may be implemented in the determination device 1 as firmware. Note that the determination unit 167 is an example of an emotion detection unit and a determination unit.
[0024] Figure 2 shows an example of the settings screen G1 displayed on the display unit 14. The settings screen G1 is a screen for the subject to input settings related to synthesized sound. The settings screen G1 is displayed when the processing unit 16 executes a judgment application program for determining the brain's ability to distinguish sounds. The settings screen G1 includes a name setting object G11, a background sound selection object G12, a difficulty level selection object G13, a volume setting object G14, and a start object G15.
[0025] The name setting object G11 is an object for setting the subject's name, and includes, for example, a text box. The set subject's name is used to generate the name sound. To ensure that the name sound is generated appropriately, it is preferable that the subject's name be entered using characters whose pronunciation can be uniquely determined, such as kana or Roman letters.
[0026] The background sound selection object G12 is an object used to select a background sound to be synthesized with a name voice from among several pre-stored background sounds. In the example shown in Figure 2, the multiple background sounds are music, but it is not limited to this example and can also be human speech, everyday sounds, white noise, colored noise, etc.
[0027] The difficulty selection object G13 is used to select the difference between the dominant frequency of the background sound and the dominant frequency of the name voice. The smaller the difference between the dominant frequency of the background sound and the dominant frequency of the name voice, the more difficult it becomes for the subject listening to the synthesized sound (a combination of the background sound and the name voice) to identify the name voice, thus increasing the difficulty level. The dominant frequency is the average frequency of the voice. The dominant frequency may also be the peak frequency of the voice spectrum.
[0028] The volume setting object G14 is an object used to set the volume of synthesized sound.
[0029] The start object G15 is the object used to initiate the output of synthesized sound. When the start object G15 is selected, the output of synthesized sound for a predetermined period of time begins.
[0030] Figure 3 shows an example of the results screen G2 displayed on the display unit 15. The results screen G2 is displayed when the output of the synthesized sound is finished. The results screen G2 includes a message indicating that the output of the synthesized sound has finished, and the result of the assessment of the subject's brain's ability to distinguish acoustics. In the example shown in Figure 3, the assessment result indicates that the subject's brain's ability to distinguish acoustics was judged to be satisfactory.
[0031] Figure 4 is a flowchart showing the flow of the judgment process performed by the judgment device 1. The judgment process is performed in response to the execution of a judgment application program for determining the brain's ability to distinguish acoustics. The judgment process is realized by the processing unit 16 working in cooperation with other components of the judgment device 1, based on the program stored in the memory unit 11.
[0032] First, the reception unit 161 accepts the subject's name, background sound selection, difficulty level selection, and volume setting in response to the subject's operation (step S101). The reception unit 161 displays the setting screen G1 by supplying the display data of the setting screen G1 to the display unit 14. The setting acquisition unit 161 acquires the set subject's name, selected background sound, selected difficulty level, and set volume in response to the selection of the start object G15 on the setting screen G1.
[0033] Next, the generation unit 162 generates a voice recording of the subject's name (step S102). The generation unit 162 applies text-to-speech synthesis technology based on a Hidden Markov Model, a deep neural network, etc., to the acquired subject's name to generate a voice recording of the subject's name being read aloud.
[0034] Furthermore, the generation unit 162 modifies the main frequency of the name voice so that the difference between the main frequency of the background sound and the main frequency of the name voice satisfies predetermined conditions. For example, the generation unit 162 modifies the main frequency of the name voice by applying resampling and time stretching to the name voice. The predetermined conditions are based on the selected difficulty level. For example, the predetermined conditions are that the absolute value of the difference is less than 101 Hz when the difficulty level is high, the absolute value of the difference is 101 Hz or more and less than 201 Hz when the difficulty level is medium, and the absolute value of the difference is 201 Hz or more when the difficulty level is low.
[0035] Next, the determination unit 163 determines the timing of the name voice output during the period from when the background sound output starts until it ends (step S103). The name voice is output a predetermined number of times (for example, 10 times) during the period from when the background sound output starts until it ends. The determination unit 163 randomly determines the timing of each name voice output.
[0036] Next, the image acquisition unit 164 starts acquiring images of the subject (step S104). The image acquisition unit 164 controls the imaging unit 12 to image the subject. The image acquisition unit 164 sequentially stores the motion image data of the subject generated by the imaging unit 12 in the storage unit 11.
[0037] Next, the output control unit 165 starts outputting background sound (step S105). The output control unit 165 acquires the audio data of the selected background sound, which has been previously stored in the storage unit 11. The output control unit 165 outputs the background sound by supplying the audio data to the audio output unit 13.
[0038] Next, the output control unit 165 determines whether or not the timing for outputting the name voice has arrived (step S106). The output control unit 165 determines whether or not the timing for outputting the name voice has arrived based on the elapsed time since the background sound output started.
[0039] When the output timing arrives (step S106-Yes), the output control unit 165 outputs the name voice superimposed on the background sound (step S107). The output control unit 165 outputs the name voice by supplying the generated name voice voice data to the voice output unit 13. In this way, the output control unit 165 outputs a synthesized sound which is a combination of the background sound and the name voice output at random timings.
[0040] If the output timing has not yet arrived (step S106-No), or immediately after step S107, the output control unit 165 determines whether the background sound output has ended (step S108). The output control unit 165 determines whether the background sound output has ended based on the elapsed time since the background sound output started.
[0041] If the background sound output has not finished (step S108-No), the determination process returns to step S105, and the output control unit 165 continues to output the background sound.
[0042] When the output of background sound ends (step S108-Yes), the image acquisition unit 164 controls the imaging unit 12 to end the acquisition of the subject's image (step S109).
[0043] Next, the pulse rate detection unit 166 detects the subject's pulse rate from the start to the end of background sound output based on the subject's image (step S110).
[0044] Figure 5 is a schematic diagram illustrating the method for obtaining pulse rate. In the graph in Figure 5, the vertical axis represents the amplitude A of the pulse wave, and the horizontal axis represents time t. The pulse rate detection unit 166 extracts multiple frame images from the user's video data. The pulse rate detection unit 166 applies a contour detection algorithm or a feature point extraction algorithm to the multiple frame images to identify the region corresponding to the user's forehead in each frame image. The pulse rate detection unit 166 extracts pixels included in the region corresponding to the forehead in each frame image and calculates the time change of the G (green) pixel value by obtaining the G pixel value from the RGB pixel values of the extracted pixels. Since capillaries are concentrated in the forehead, the user's blood flow is reflected in the G pixel value of the region corresponding to the forehead, and therefore the pulse wave is detected based on the time change of the G pixel value. The pulse rate detection unit 166 extracts the pulse wave signal PW shown in Figure 5 by applying a bandpass filter with a transmission bandwidth of 0.5Hz to 3Hz corresponding to a human pulse wave to the signal showing the time change of the G pixel value. The pulse rate detection unit 166 may also extract regions in the frame image that correspond to exposed skin areas of other body parts other than the forehead.
[0045] The pulse rate detection unit 166 calculates the user's pulse wave interval and pulse rate based on the extracted pulse wave signal PW. For example, the pulse rate detection unit 166 identifies the peak point P(n) of the pulse wave signal PW. The pulse rate detection unit 166 calculates and obtains the interval between two adjacent peak points (P(n), P(n+1)) as the pulse wave interval d(n). The pulse rate detection unit 166 calculates and obtains the pulse rate per minute as 60 times the reciprocal of the pulse wave interval.
[0046] Returning to Figure 4, the determination unit 167 then determines the subject's brain's ability to distinguish acoustics based on the subject's pulse rate (step S111). For example, the emotion detection unit 167 detects the subject's emotions based on the subject's pulse rate and determines the subject's brain's ability to distinguish acoustics based on those emotions.
[0047] Figure 6 is a schematic diagram illustrating a method for determining acoustic discrimination ability. In the graphs of Figures 6(A) and (B), the vertical axis represents pulse rate B, and the horizontal axis represents time t. The determination unit 167 detects the subject's emotions based on the time change of the subject's pulse rate. For example, the determination unit 167 extracts the pulse rate b2 at the point where the time change PR of the subject's pulse rate is at its maximum, and the pulse rate b1 a predetermined time before the point of maximum value. The determination unit 167 detects the subject's emotions at time t1 of pulse rate b1 if pulse rate b2 is greater than the average pulse rate bav of a typical person, and the difference Δb between pulse rate b2 and pulse rate b1 is greater than or equal to a predetermined value.
[0048] The determination unit 167 determines whether the subject recognized the name voice based on the relationship between the period during which the name voice was output and the time when the subject's emotion was detected. For example, the determination unit 167 determines that the subject recognized the name voice if the subject's emotion was detected during the period P1 during which the name voice was output, or during the period P2 from the end of period P1 until a predetermined time has elapsed. In the example shown in Figure 6(A), the time t1 of the pulse rate b1 is included in period P1, so the determination unit 167 determines that the subject recognized the name voice.
[0049] Generally, when a person is called by their name, they react consciously or unconsciously, and emotions are detected at that time. Because it takes time for a person's brain to recognize their name being called and for a physical reaction to occur, emotions are detected during the period when the name is being called or afterward. The determination unit 167 determines whether the subject recognized the name sound if emotions are detected during the period P1 in which the name sound is output or during the period P2 afterward, thereby enabling an appropriate determination of whether the subject recognized the name sound. Note that period P2 is generally set based on the time it takes for a person to react after their name is called, and is set to, for example, 1 second.
[0050] Furthermore, multiple emotional responses may be detected within a predetermined time frame when a person's name is called. In the example shown in Figure 6(B), emotional responses are detected at times t1, t2, and t3, respectively. When multiple emotional responses are detected within a predetermined time frame in this way, the determination unit 167 determines whether the subject recognized the name voice based on the relationship between the period during which the name voice was output and the time when the first of the multiple emotional responses was detected. For example, the determination unit 167 determines whether the subject recognized the name voice based on whether the time t1 when the first emotional response was detected falls within period P1 or P2. In the example shown in Figure 6(B), since the time t1 when the first emotional response was detected is before period P1, the determination unit 167 determines that the subject did not recognize the name voice.
[0051] Returning to Figure 4, the determination unit 167 determines the subject's auditory discrimination ability by comparing the number of times the subject recognizes the name sound between the start and end of background sound output with a threshold. For example, if the name sound is output 10 times, the threshold is set to 8 times. If the number of times the subject recognizes the name sound is equal to or greater than the threshold, the determination unit 167 determines that the subject's auditory discrimination ability is satisfactory. If the number of times the subject recognizes the name sound is less than the threshold, the determination unit 167 determines that the subject's auditory discrimination ability is unsatisfactory.
[0052] The determination unit 167 generates display data for the result screen G2, including the determination result, and supplies it to the display unit 15, thereby outputting the determination result. This completes the determination process.
[0053] As explained above, the determination device 1 outputs a synthesized sound that combines background sound and a name voice, and determines the subject's auditory discrimination ability based on the subject's pulse rate while the synthesized sound is being output. Because the determination device 1 determines auditory discrimination ability based on the subject's pulse rate, it makes it possible to easily determine the subject's auditory discrimination ability based on the subject's brain without requiring a response from the subject.
[0054] Furthermore, the determination device 1 detects the subject's emotions based on their pulse rate and determines the subject's auditory discrimination ability based on their emotions. The determination device 1 enables the appropriate determination of the subject's auditory discrimination ability by detecting the emotions that naturally arise when a person responds to a name sound.
[0055] Furthermore, the judgment device 1 outputs a synthesized sound that combines background noise with a voice reading the subject's name. Since the voice reading one's own name is one of the sounds most susceptible to the cocktail party effect, the judgment device 1 enables a more accurate determination of the subject's brain's ability to recognize speech.
[0056] Furthermore, the judgment device 1 accepts the selection of one of several difficulty levels in response to the subject's operation, and outputs a synthesized sound that satisfies the conditions based on the selected difficulty level, such that the difference between the main frequency of the background sound and the main frequency of the predetermined sound satisfies the conditions. In this way, the judgment device 1 outputs a synthesized sound suitable for determining the subject's acoustic discrimination ability, enabling a more accurate determination of the subject's speech discrimination ability.
[0057] In the above explanation, step S102 of the judgment process is assumed to be an example in which the generation unit 162 generates a voice reading out the subject's name. However, the example is not limited to this. The generation unit 162 may generate any voice that detects the subject's emotions. For example, the generation unit 162 may generate a voice reading out the names of the subject's close relatives or any string of characters entered by the subject.
[0058] In the above explanation, it was assumed that in steps S106-S107 of the determination process, the output control unit 165 outputs the name voice superimposed on the background sound when the timing for outputting the name voice arrives, but the example is not limited to this. The output control unit 165 may also generate synthesized voice data in advance, which is a combination of the background sound and the name voice output at a predetermined timing, before the output of the background sound begins.
[0059] In the above description, the pulse rate detection unit 166 was assumed to detect the subject's pulse rate based on the subject's image, but this is not the only example. For example, the pulse rate detection unit 166 may detect the subject's pulse rate by acquiring pulse rate data from a sensor such as a pulse meter worn by the subject via an interface.
[0060] Those skilled in the art will understand that various changes, substitutions, and modifications can be made without departing the scope of the present invention. For example, the embodiments and modifications described above may be combined as appropriate within the scope of the present invention. [Explanation of symbols]
[0061] 1 Judgment device 161 Setting Acquisition Section 162 Generation part 163 Decision Section 164 Image acquisition unit 165 Output Control Unit 166 Pulse rate detection unit 167 Judgment section
Claims
1. An output unit outputs a synthesized sound in which a background sound and a predetermined sound output at random timings are combined such that the difference between the frequency of the main sound of the background sound and the frequency of the main sound of the predetermined sound satisfies a predetermined condition. A pulse rate detection unit detects the pulse rate of the subject while the synthesized sound is being output, A determination unit that determines the subject's brain's ability to distinguish sounds based on the subject's pulse rate while the synthesized sound is being output, It has, The system further includes an emotion detection unit that detects the subject's emotions based on the time change in the subject's pulse rate while the synthesized sound is being output. The determination unit determines that the subject has recognized the predetermined sound when the emotion detection unit detects the subject's emotion. A determination device characterized by the following features.
2. The emotion detection unit further detects the subject's emotions based on the time change in the subject's pulse rate within a predetermined period after the synthesized sound is output. The determination unit determines the subject's auditory discrimination ability by comparing the number of times the subject is determined to have recognized the predetermined sound with a threshold value. The determination device according to claim 1.
3. The emotion detection unit detects the emotion of the subject when the second pulse rate at the point where the temporal change in the subject's pulse rate reaches its maximum is greater than a predetermined pulse rate, and the difference between the second pulse rate and the first pulse rate at a predetermined time prior to the point where the maximum value is reached is greater than or equal to a predetermined value. The determination device according to claim 1 or 2.
4. The aforementioned predetermined sound is the voice of the subject's name. The aforementioned background sound is music selected by the subject. The determination device according to claim 1.
5. The system further includes a selection unit that accepts the selection of one of a set of conditions regarding the difference in the frequencies of the main sound in response to the subject's operation, The output unit outputs a synthesized sound such that the difference between the frequency of the main sound of the background sound and the frequency of the main sound of the predetermined sound satisfies the selected condition. The determination device according to claim 1.
6. A determination method performed by a determination device, A step of outputting a synthesized sound which is produced by combining a background sound and a predetermined sound output at random timings such that the difference between the frequency of the main sound of the background sound and the frequency of the main sound of the predetermined sound satisfies a predetermined condition, The steps include detecting the subject's pulse rate while the synthesized sound is being output, The steps include determining the subject's brain's ability to distinguish sounds based on the subject's pulse rate while the synthesized sound is being output, The step includes detecting the subject's emotions based on the time change in the subject's pulse rate while the synthesized sound is being output, The determination step involves determining that the subject has recognized the predetermined sound if the subject's emotion is detected in the detection step. A determination method characterized by the above.
7. A step of outputting a synthesized sound which is produced by combining a background sound and a predetermined sound output at random timings such that the difference between the frequency of the main sound of the background sound and the frequency of the main sound of the predetermined sound satisfies a predetermined condition, The steps include detecting the subject's pulse rate while the synthesized sound is being output, The steps include determining the subject's brain's ability to distinguish sounds based on the subject's pulse rate while the synthesized sound is being output, The computer is instructed to perform the following steps: detect the subject's emotions based on the time change in the subject's pulse rate while the synthesized sound is being output; The determination step involves determining that the subject has recognized the predetermined sound if the subject's emotion is detected in the detection step. A program characterized by the following features.