Portable device and method for controlling portable device
The portable device addresses the challenge of selectively transmitting a specific person's voice in noisy environments by registering a wake word and vibrating upon recognition, effectively alerting the wearer.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SOLIDSONIC CO LTD
- Filing Date
- 2025-10-09
- Publication Date
- 2026-05-21
AI Technical Summary
Existing sound amplifiers and hearing aids fail to selectively transmit the voice of a specific person in environments with multiple people and ambient sounds, posing challenges for individuals with hearing impairments and elderly people, as well as pets, who may not hear important voices or ambient sounds.
A portable device with a mounting unit, sound collection unit, vibration unit, and control unit that registers a predetermined wake word, collects and matches voice patterns, and vibrates upon recognition to alert the wearer of a specific person's voice.
Enables the wearer to be aware of a specific person's voice even in noisy environments by directly transmitting the voice through bone conduction vibrations.
Smart Images

Figure JP2025035901_21052026_PF_FP_ABST
Abstract
Description
Mobile device and method for controlling a mobile device
[0001] The present invention relates to a mobile device and a method for controlling a mobile device.
[0002] Conventionally, there are technologies related to mobile devices having a voice recognition function. For example, Japanese Patent Application Laid-Open No. 2017-009956 (Patent Document 1) discloses a data recording device including a storage unit, a microphone unit, an identification unit, and a text conversion unit. Here, the storage unit stores reference voiceprint data of a plurality of persons, and the microphone unit detects voice. The identification unit collates the voiceprint of the voice detected by the microphone unit with the reference voiceprint data of a plurality of persons to identify the matching speaker, and the text conversion unit converts the voice detected by the microphone unit into text. The data recording device stores the text-converted text data in the storage unit in association with the speaker identified by the identification unit. Thereby, it is said that the work of converting voice into text and the management of data are facilitated.
[0003] Also, Japanese Patent Application Laid-Open No. 2017-151735 (Patent Document 2) discloses a portable device including a vibration detection unit, a vibration authentication means, and a specific processing means. The vibration detection unit is attached to the human body to detect vibration related to the sound emitted by the person, and the vibration authentication means refers to a vibration pattern storage unit storing data related to vibrations that permit authentication based on the pattern of the vibration detected by the vibration detection unit to perform authentication. The specific processing means performs a specific process when authenticated by the vibration authentication means. Thereby, it is said that the security is further improved.
[0004] Also, Japanese Patent Application Laid-Open No. 2020-205485 (Patent Document 3) discloses a telephone system including a voiceprint database, a voiceprint authentication unit, and a notification unit. Here, the voiceprint database associates a pre-registered speaker with their voiceprint, and the voiceprint authentication unit accesses the voiceprint database based on the voiceprint information obtained by analyzing the voice of the calling party of an incoming call to authenticate the calling party. Also, the notification unit notifies the responder of the incoming call with a message when the calling party is not registered. Thereby, it is said that special fraud can be suppressed.
[0005] Furthermore, Japanese Patent Publication No. 2018-087838 (Patent Document 4) discloses a speech recognition device comprising a microphone, a speech recognition filter, a voiceprint filter, a priority determination unit, and an instruction content recognition unit. The microphone records ambient sounds, and the speech recognition filter selects sounds with frequencies corresponding to speech from the ambient sounds recorded by the microphone. The voiceprint filter selects the voice of a registered user from the sounds with frequencies corresponding to speech selected by the speech recognition filter, and the priority determination unit determines the priority of the registered user's voice selected by the voiceprint filter. The instruction content recognition unit analyzes the voice of the registered user determined to have the highest priority by the priority determination unit and recognizes the instruction content of this registered user's voice. This allows for accurate recognition of voice instructions even in environments where multiple users are using the device simultaneously.
[0006] Furthermore, U.S. Patent Application Publication No. 2013 / 0030242 (Patent Document 5) discloses a waterproof hearing device and system for transmitting sound via transcutaneous bone conduction to provide therapeutic auditory stimulation music to dogs. The waterproof hearing device comprises a transducer and a support. The transducer generates music from a musical signal, and the support holds the transducer in vibratory contact with the dog's head. Each transducer can be positioned at multiple locations on the support, which includes an adjustable dog collar component and a removable and adjustable muzzle component. This allows for the delivery of high-fidelity therapeutic auditory stimulation music to dogs via transcutaneous bone conduction.
[0007] Japanese Patent Publication No. 2017-009956, Japanese Patent Publication No. 2017-151735, Japanese Patent Publication No. 2020-205485, Japanese Patent Publication No. 2018-087838, U.S. Patent Application Publication No. 2013 / 0030242
[0008] Currently, sound amplifiers and hearing aids that transmit sound to people such as able-bodied individuals, people with hearing impairments, people with hearing loss, elderly people with reduced hearing ability, and animals such as pets are widely available on the market. These sound amplifiers and hearing aids typically collect ambient sounds and transmit the collected sounds to the wearer, but generally do not identify or select the sounds they collect. Therefore, in environments with multiple people and ambient sounds, all people's voices and ambient sounds are collected, which presents the challenge of not being able to directly transmit only the necessary sounds to the wearer.
[0009] Furthermore, there is a challenge in that hearing-impaired individuals, those with hearing loss, and the elderly may experience difficulties in their daily lives if they are unable to hear the voices of those around them or ambient sounds. For example, if a person with hearing loss is walking around and an acquaintance who knows them calls out to them, the voice may not reach the person with hearing loss, making it difficult for the acquaintance to communicate easily with them.
[0010] Furthermore, even pet dogs and cats, as they age, may lose the ability to hear the voices of people around them or ambient sounds. For example, even if an owner calls their elderly dog by name, the dog may not be able to hear the owner's voice and therefore will not notice the owner.
[0011] Here, there is a need for technology that can directly transmit the voice of a specific person to animals, or to make them aware of a specific person's voice when that person speaks to them, not just humans.
[0012] Here, the technology described in Patent Document 1 transcribes the voice of a speaker that matches a reference voiceprint data from multiple speakers. The technology described in Patent Document 2 performs user authentication in a specific system using the user's own voice. The technology described in Patent Document 3 prevents special fraud by performing voiceprint authentication on the voice of the person on a phone call and identifying only the person who has been pre-registered. The technology described in Patent Document 4 extracts only the voice of a specific person in an environment with multiple people and ambient sounds and transmits it to the user. The technology described in Patent Document 5 transmits the audio signal of pre-registered music to a dog via bone conduction. However, the technologies in Patent Documents 1-5 cannot make people or animals aware of a specific person's voice or directly transmit a specific person's voice to them.
[0013] Therefore, the present invention has been made to solve the aforementioned problems, and aims to provide a portable device and a method for controlling the portable device that can make the wearer aware of a specific person's voice even in an environment where multiple people and ambient noises are present.
[0014] The portable device according to the present invention comprises a mounting unit, a sound collection unit, a vibration unit, a registration determination control unit, a voice registration control unit, a voice sound collection control unit, a voice determination control unit, a sound collection continuation control unit, and a vibration control unit. The mounting unit can be attached to a part of the body of a person or animal. The sound collection unit is attached to the mounting unit and collects sounds from people in the surrounding area. The vibration unit is attached to the mounting unit. When the portable device is activated, the registration determination control unit determines whether a predetermined wake word is registered in a predetermined storage unit. If the wake word is not registered in the storage unit, the voice registration control unit uses the sound collection unit to store a sound from a person as the wake word in the storage unit. If the wake word is registered in the storage unit, the voice sound collection control unit uses the sound collection unit to collect a sound from a person. The voice determination control unit determines whether the collected sound matches the wake word. If the sound does not match the wake word, the sound collection continuation control unit uses the sound collection unit to continue collecting the sound. The vibration control unit vibrates the vibrating part when the sound matches the wake word.
[0015] The method for controlling a portable device according to the present invention comprises a mounting step, a registration determination control step, a voice registration control step, a voice sound collection control step, a voice determination control step, a sound collection continuation control step, and a vibration control step. The mounting step involves mounting a portable device, which includes a mounting unit, a sound collection unit, and a vibration unit, onto a part of a person's or animal's body. The registration determination control step determines whether a predetermined wake word is registered in a predetermined storage unit when the portable device is activated. If the wake word is not registered in the storage unit, the voice registration control step uses the sound collection unit to store a voice from a person as the wake word in the storage unit. If the wake word is registered, the voice sound collection control step uses the sound collection unit to collect a voice from a person. The voice determination control step determines whether the collected voice matches the wake word. If the voice does not match the wake word, the sound collection continuation control step uses the sound collection unit to continue collecting the voice. The vibration control process involves vibrating the vibrating part when the sound matches the wake word.
[0016] According to the present invention, even in environments where multiple people and ambient noises are present, it is possible to make the wearer aware of a specific person's voice.
[0017] Figure 4A shows an example of a portable device according to an embodiment of the present invention being worn around a dog's neck, and Figure 5B shows an example of the device being worn around a dog's neck.Figure 6A shows an example of a portable device according to an embodiment of the present invention being worn around a person's wrist, and Figure 6B shows an example of the device being worn around a person's wrist, and Figure 7A shows an example of a case in which a portable device according to an embodiment of the present invention extracts feature quantities from the voice of a specific person and registers the voice and feature quantities as a wake word, and Figure 7B shows an example of a case in which a collected voice and the feature quantities extracted from the voice are compared with the registered wake word. This figure shows an example of a case in which a portable device according to an embodiment of the present invention is used in an environment where ambient sounds and other people are present in addition to a specific person.
[0018] The following describes embodiments of the present invention with reference to the attached drawings to facilitate understanding of the invention. Note that the following embodiments are merely examples of the present invention and do not limit the technical scope of the invention.
[0019] As shown in Figure 1, the portable device 1 according to an embodiment of the present invention comprises a mounting section 10, a sound collection section 20, a vibration section 30, a control section 40, and a storage section 50.
[0020] Here, the attachment part 10 can be attached to a part of a person's or animal's body. There are no particular limitations on the configuration of the attachment part 10, but for example, as shown in Figure 1, the attachment part 10 can be a collar that can be wrapped around an animal's neck or a belt that can be wrapped around a person's arm, but it is not limited to these, and is not particularly limited as long as it can be attached to a part of a person's or animal's body, such as clothing, a hat, or gloves.
[0021] Furthermore, the sound-collecting unit 20 is attached to the mounting unit 10 and collects ambient sounds. Here, the sound-collecting unit 20 receives the vibrations of the air, which are ambient sounds, as mechanical vibrations, converts the mechanical vibrations into electrical signals, and transmits them to the control unit 40. For example, as shown in Figure 1, the sound-collecting unit 20 is attached to the front mounting unit 10a of the collar of the mounting unit 10, which is located in front of the pet dog to be worn. This makes it easy to collect voices from people in front of the pet dog. Furthermore, there are no particular limitations on the configuration of the sound-collecting unit 20, but examples include a dynamic microphone and a condenser microphone. Furthermore, there are no particular limitations on the sound collection range of the sound-collecting unit 20, but for example, it can be within the range of 10m to 50m. The sound collection range may also be changed as appropriate by user operation.
[0022] Furthermore, the vibrating unit 30 is attached to the attachment unit 10. Here, the vibrating unit 30 vibrates in response to an electrical signal from the control unit 40. For example, as shown in Figure 1, the vibrating unit 30 is attached to the rear attachment unit 10b of the collar of the attachment unit 10, which is located at the rear of the pet dog to be worn, and is positioned to contact a part of the back of the dog's neck. Here, since the dog's ears are located near the back of its neck, the vibration of the vibrating unit 30 makes it easy for the dog to notice the vibration. Also, if the vibrating unit 30 vibrates in response to sound, the vibration is transmitted to the dog's ears as bone conduction, so the dog can recognize the sound by vibration. Furthermore, there are no particular limitations on the configuration of the vibrating unit 30, but examples include a diaphragm, a vibrator, and a vibrating element. Specifically, examples include a piezoelectric ceramic vibrator, an electromagnetic vibrator, and a supermagnetostrictive vibrator. Furthermore, the vibrating unit 30 may switch between multiple vibration patterns or change the vibration time.
[0023] Furthermore, the control unit 40 is attached to the mounting section 10 and controls various parts of the portable device 1 and controls the operation of the portable device 1. The control unit 40 is, for example, built into the front mounting section 10a. There are no particular limitations on the configuration of the control unit 40, but examples include a microcontroller such as Arduino®, which is composed of a CPU and dedicated circuits, or a single-board computer such as Rhaspberry Pi®.
[0024] Furthermore, the storage unit 50 is attached to the mounting unit 10 and stores data required for control of the portable device 1. The storage unit 50 is, for example, built into the front mounting unit 10a. There are no particular limitations on the configuration of the storage unit 50, but examples include recording media such as HDDs, SSDs, and flash memory.
[0025] Furthermore, the portable device 1 may also include a power supply unit 60, a power switch unit 70, a display unit 80, and a communication unit 90.
[0026] Here, the power supply unit 60 is attached to the mounting unit 10 and supplies power to the sound collection unit 20, vibration unit 30, control unit 40, memory unit 50, display unit 80, and communication unit 90. The power supply unit 60 is, for example, built into the front mounting unit 10a. There are no particular limitations on the configuration of the power supply unit 60, but examples include primary batteries and secondary batteries. Here, the power supply unit 60 may, for example, have a rechargeable secondary battery such as an alkaline storage battery or a lithium-ion battery built into the mounting unit 10, or it may have a non-rechargeable primary battery such as an alkaline dry cell battery built into the mounting unit 10. Also, when using a rechargeable secondary battery, a terminal for connecting the secondary battery and an external power supply by an electrical cable may be provided on the mounting unit 10. When using a primary battery, a battery case that can be opened and closed for battery replacement may be provided on the mounting unit 10.
[0027] Furthermore, the power switch unit 70 is attached to the mounting unit 10 and electrically connected to the power supply unit 60, and the power supply unit 60 supplies or stops power to each unit through user operation such as pressing. The power switch unit 70 is installed, for example, on the surface of the front mounting unit 10a. There are no particular limitations on the configuration of the power switch unit 70, but examples include a push switch, a tact switch, a slide switch, etc.
[0028] Furthermore, the display unit 80 is attached to the mounting unit 10 and, in response to the user's operation of the mobile device 1, displays various information and conveys the status of the mobile device 1 to the user. The display unit 80 is installed, for example, on the surface of the front mounting unit 10a. There are no particular limitations on the configuration of the display unit 80, but examples include LEDs and small liquid crystal displays.
[0029] Furthermore, the communication unit 90 is attached to the mounting unit 10, and the portable device 1 communicates with external devices such as terminal devices and mobile terminal devices via wireless communication such as a network or wired communication. The communication unit 90 is installed, for example, on the surface of the front mounting unit 10a. There are no particular limitations on the configuration of the communication unit 90, but examples include a wired connection unit such as USB (Universal Serial Bus) and a wireless communication unit via Wi-Fi (trademark registered) or Bluetooth (registered trademark). Here, the communication unit 90 may, for example, have a USB terminal provided on the mounting unit 10. This allows the portable device 1 to connect to an external device via USB, and the registered wake word 50a to be displayed or edited on the external device.
[0030] Furthermore, the control unit 40 incorporates a CPU, ROM, RAM, etc. (not shown), and the CPU, for example, uses the RAM as a working area to execute programs stored in the ROM, etc. Similarly, each control unit, described later, is realized by the CPU executing programs.
[0031] Next, the configuration and execution procedure of an embodiment of the present invention will be described with reference to Figures 2-6. First, the user attaches the attachment part 10 of the portable device 1 to a part of the body of a person or animal. Here, the attachment part 10 of the portable device 1 is made up of a collar, and as shown in Figure 4A, the user, who is the owner, attaches the collar of the attachment part 10 to the neck of their pet dog.
[0032] Next, when the user presses the power switch 70 of the mobile device 1 (Figure 3: S101YES), the power switch 70 turns ON, which allows the power supply unit 60 to supply power, and the power supply unit 60 starts supplying power to each part. As a result, the mobile device 1 starts up.
[0033] Here, for example, if the display unit 80 is an LED, the display unit 80, having received power from the power supply unit 60, may light up in a color (for example, blue) to indicate that it is starting up. Alternatively, if the display unit 80 is a liquid crystal display, it may display a message (for example, "Starting Up") to indicate that it is starting up. This allows the user to be notified that the portable device 1 is starting up.
[0034] Furthermore, if the user does not press the power switch unit 70 of the mobile device 1 (Figure 3: S101NO), the power switch unit 70 is OFF, the power switch unit 70 stops the power supply to the power supply unit 60, and the mobile device 1 does not start up.
[0035] When the mobile device 1 is started up, the registration determination control unit 101 of the control unit 40 determines whether or not the wake word 50a is registered in the storage unit 50 (Figure 3: S102).
[0036] Here, there are no particular limitations on the determination method of the registration determination control unit 101. For example, the registration determination control unit 101 refers to the storage unit 50 and determines whether or not text (character information) indicating speech is registered as a wake word 50a in the referenced storage unit 50. Here, we assume that no text is registered in the storage unit 50.
[0037] If the determination shows that the wake word 50a is not registered (Figure 3: S102NO), the registration determination control unit 101 determines that the wake word 50a is not registered and notifies the voice registration control unit 102 of the control unit 40 of this. Upon receiving this notification, the voice registration control unit 102 uses the sound collection unit 20 to store a voice from a person as the wake word 50a in the storage unit 50 (Figure 3: S103).
[0038] There are no particular limitations on the storage method of the voice registration control unit 102. For example, the voice registration control unit 102 activates the sound collection unit 20 for a predetermined time (for example, 5 to 30 seconds) to collect voice from a person. Here, for example, as shown in Figure 4A, when the owner calls the name of their dog (for example, "Shiro"), the voice registration control unit 102 uses the sound collection unit 20 to collect the name of the dog ("Shiro"), extracts the text from the collected name, and stores it in the storage unit 50 as a wake word 50a. In this way, the owner can store the name of their dog ("Shiro") as a wake word 50a.
[0039] Furthermore, for example, when the voice registration control unit 102 is activating the sound collection unit 20, the LED display unit 80 may light up in a color (for example, green) to indicate that sound collection is in progress. Alternatively, the liquid crystal display unit 80 may display a message (for example, "Sound Collection") indicating that sound collection is in progress. This allows the user to be notified that the wake word 50a has been stored.
[0040] Furthermore, when the voice registration control unit 102 collects voice using the sound collection unit 20, the display unit 80 of the liquid crystal display may display the collected voice (for example, "Shiro"). This allows the user to be informed of the specific wake word 50a.
[0041] Once the voice registration control unit 102 has finished storing the information, the voice sound collection control unit 103 of the control unit 40 uses the sound collection unit 20 to collect voice from a person (Figure 3: S104).
[0042] On the other hand, in S102, if as a result of the determination, the wake word 50a is registered, the registration determination control unit 101 determines that the wake word 50a is registered and notifies the voice collection control unit 103 to that effect. Upon receiving the notification, the voice collection control unit 103 collects the voice from a person using the collection unit 20 as described above (Fig. 3: S104).
[0043] Here, there is no particular limitation on the voice collection method of the voice collection control unit 103. For example, the voice collection control unit 103 activates the collection unit 20 to collect the voice from a person. Here, as described above, when the voice collection control unit 103 activates the collection unit 20, the LED display unit 80 may light up with a color (green) indicating that collection is in progress. Alternatively, the display unit 80 of the liquid crystal display may display a message ("Collection") indicating that collection is in progress. Thereby, it is possible to notify the user of the collection.
[0044] Here, for example, as shown in Fig. 4B, assume that the owner uttered a name ("Kuro") different from the name of the pet dog ("Shiro"). Then, the voice collection control unit 103 collects the name ("Kuro") from the owner and notifies the voice determination control unit 104 of the control unit 40 to that effect. Upon receiving the notification, the voice determination control unit 104 determines whether the collected voice matches the wake word 50a (Fig. 3: S105).
[0045] Here, there is no particular limitation on the determination method of the voice determination control unit 104. For example, when the name ("Kuro") from the owner is collected, the voice determination control unit 104 extracts text from the collected name ("Kuro"). Next, the voice determination control unit 104 refers to the storage unit 50 to obtain the wake word 50a ("Shiro") and determines whether the extracted text ("Kuro") matches the wake word 50a ("Shiro").
[0046] Here, for example, the voice determination control unit 104 may compare each character of the text ("Kuro") with each character of the wake word 50a ("Shiro") to determine whether the text ("Kuro") completely matches the wake word 50a ("Shiro"). Also, for example, when the voice from the owner is a voice containing the name ("Kuro") (e.g., "Ah, Kuro"), the voice determination control unit 104 compares each character of the text ("Ah, Kuro") with each character of the wake word 50a ("Shiro") to determine whether the wake word 50a ("Shiro") is included in the text ("Ah, Kuro") or whether there is a partial match. Also, when extracting text from the collected voice, the voice determination control unit 104 may use only the voice with high intensity as the text and determine whether the text matches the wake word 50a ("Shiro"). Thereby, for example, when the owner intentionally calls the name "Shiro" of the pet dog, voice determination becomes possible.
[0047] Now, if as a result of the determination, the text ("Kuro") does not match the wake word 50a ("Shiro") (Fig. 3: S105 NO), the voice determination control unit 104 determines that the collected voice does not match the wake word 50a and notifies the voice collection continuation control unit 105 of the control unit 40 to that effect. Upon receiving the notification, the voice collection continuation control unit 105 returns to S104 and continues to collect the voice from the person using the voice collection unit 20 (Fig. 3: S104).
[0048] On the other hand, in S104, as shown in Fig. 5A, when the owner utters the name ("Shiro") of the pet dog, the voice collection control unit 103 collects the name ("Shiro") from the owner (Fig. 3: S104), and the voice determination control unit 104 determines whether the text ("Shiro") of the collected voice matches the wake word 50a ("Shiro") (Fig. 3: S105).
[0049] If the judgment determines that the text ("Shiro") matches the wake word 50a ("Shiro") (Figure 3: S105 YES), the voice judgment control unit 104 determines that the collected voice matches the wake word 50a and notifies the vibration control unit 106 of the control unit 40 of this. Upon receiving this notification, the vibration control unit 106 vibrates the vibration unit 30 (Figure 3: S106).
[0050] There are no particular limitations on the vibration method of the vibration control unit 106. For example, as shown in Figure 5A, the vibration control unit 106 activates the vibration unit 30 and vibrates the vibration unit 30 according to a preset vibration pattern. This allows the owner to alert the dog by correctly pronouncing the dog's name. The dog then recognizes that it has been called and can pay attention to the owner.
[0051] Furthermore, the vibration control unit 106 can, for example, vibrate the vibration unit 30 with a vibration pattern corresponding to a matching voice (for example, the owner's name "Shiro"), thereby directly transmitting the name spoken by the owner to the dog through bone conduction.
[0052] Here, the vibration control unit 106 may, for example, activate the sound collection unit 20 for a predetermined time (for example, 5 to 30 seconds) to collect sounds that the person has further uttered, and then vibrate the vibration unit 30 with a vibration pattern corresponding to the collected sounds. Specifically, if the owner says the dog's name ("Shiro") and then says another word (for example, "sit"), as shown in Figure 5B, the vibration control unit 106 will vibrate the vibration unit 30 with a vibration pattern corresponding to the other word ("sit"). This makes it possible for the owner to directly communicate various words to their dog, starting from the wake word 50a.
[0053] Now, once the vibration control unit 106 has completed the vibration, it determines whether the portable device 1 has finished (Figure 3: S107). Here, the portable device 1 has finished, for example, when the power switch unit 70 is pressed by the user or when the elapsed time since the voice sound collection control unit 103 started sound collection exceeds a predetermined threshold time. For example, the vibration control unit 106 determines whether the power switch unit 70 has been pressed by the user (Figure 3: S107). Here, if the user does not press the power switch unit 70 again (Figure 3: S107NO), the process returns to S104, and the voice sound collection control unit 103 uses the sound collection unit 20 to collect voice from the person (Figure 3: S104).
[0054] On the other hand, in S107, when the user presses the power switch unit 70 of the mobile device 1 again (Figure 3: S107 YES), the power switch unit 70 turns OFF, the power switch unit 70 stops supplying power to the power supply unit 60, and the power supply unit 60 stops supplying power to each component. As a result, the operation of the mobile device 1 ends.
[0055] Here, for example, the LED display unit 80 may turn off in response to the power supply being cut off. Alternatively, the liquid crystal display unit 80 may temporarily display a message indicating the shutdown (for example, "Stop") in response to the power supply being cut off. This allows the user to be notified that the operation of the portable device 1 has ended.
[0056] Incidentally, the mobile device 1 can be used in various ways by registering a wide variety of wake words 50a. For example, as shown in Figure 6A, if the name of a person with hearing impairment (for example, "ABC") is stored in the wake word 50a, when someone nearby utters the name of the person with hearing impairment ("ABC"), the mobile device 1 will vibrate, alerting the person with hearing impairment.
[0057] Furthermore, for example, if the mobile device 1 vibrates in response to subsequent words, starting with the wake word 50a, as shown in Figure 6B, if someone nearby utters the name of the hearing-impaired person ("ABC") and then utters a word to warn of danger (for example, "danger"), the mobile device 1 will vibrate in response to the word to warn of danger ("danger"). This allows a specific person who knows the name of the hearing-impaired person ("ABC") to warn them of a dangerous situation, such as a car approaching. In this way, starting with the wake word 50a, people nearby can directly communicate various words to the hearing-impaired person.
[0058] By the way, the wake word 50a may be one or more, and may be configured to allow modification, addition, deletion, etc. Furthermore, there are no particular limitations on the configuration of the wake word 50a. For example, as described above, text was extracted from the collected audio, and the wake word 50a was converted into text and stored in the storage unit 50. In this case, regardless of who uttered the wake word 50a, the audio determination control unit 104 determines that the collected audio matches the wake word 50a.
[0059] On the other hand, in order to identify the person who made the sound, a wake word 50a may be constructed using speech recognition technology or voiceprint technology. For example, as shown in Figure 7A, the speech registration control unit 102 may be configured to register the characteristics of a person's voice in addition to the text of the wake word 50a in the memory unit 30. In this case, the characteristics can be, for example, Mel-frequency cepstrum coefficients (MFCC), linear predictive coding (LPC), pitch, or formant frequency. Furthermore, a Gaussian mixture model (GMM), i-vector, deep learning, or support vector machine (SVM) can be used as a model for recognizing a specific person. By training these models with the voice of a specific person, it is possible to identify that specific person.
[0060] Then, as shown in Figure 7B, the voice determination control unit 104 extracts text from the collected audio and extracts feature quantities from this audio, and determines whether the extracted text matches the text of the wake word 50a registered in the memory unit 30, and whether the extracted feature quantities match the feature quantities registered in the memory unit 30. If the extracted text matches the text of the wake word 50a registered in the memory unit 30, and the extracted feature quantities match the feature quantities registered in the memory unit 30, the voice determination control unit 104 determines that the collected audio matches the wake word 50a. If the extracted text does not match the text of the registered wake word 50a, or if the extracted feature quantities do not match the registered feature quantities, the voice determination control unit 104 determines that the collected audio does not match the wake word 50a. As a result, as shown in Figure 8, even in an environment where other people and ambient sounds are present, it is possible to react only to the voice of a specific person and alert the wearer.
[0061] Incidentally, since there are a wide variety of speech recognition and voiceprint technologies, as described below, you may use them as appropriate. For example, you can extract features from human speech, use an acoustic model to predict which sounds correspond to which phonemes or words based on the extracted features, correct the resulting phonemes and words using a language model, and finally obtain the most likely text by decoding. Examples of features include Mel-frequency cepstrum coefficients (MFCC) and spectral subband energy (SBE). MFCC divides the speech signal into short-time frames and analyzes the frequency components of each frame on a Mel scale. SBE extracts the energy for each frequency band of speech and quantifies the characteristics of the speech. Examples of acoustic models include recurrent neural networks (RNN), long-term short-term memory (LSTM), and convolutional neural networks (CNN). Examples of language models include n-gram models and neural network-based language models. Furthermore, decoding methods include the Viterbi algorithm and beam search. The Viterbi algorithm is used to find the optimal phoneme sequence, selecting the word sequence with the highest probability based on an acoustic model and transition probabilities. Beam search searches multiple candidates (beams) in parallel and selects the most promising one. In addition, end-to-end approaches such as CTC (Connectionist Temporal Classification) and Attention-based Models can also be used.
[0062] Furthermore, regarding speech recognition technology and voiceprint technology, if the mobile device 1 can connect to a network using the communication unit 90, it may utilize a speech recognition API provided on the network. Alternatively, speech recognition may be implemented by incorporating a speech recognition IC or the like into the control unit 40.
[0063] Furthermore, in the embodiment of the present invention, in S102, the portable device 1 is activated and the registration determination control unit 101 makes a determination (Figure 3: S102). If the wake word 50a is not registered (Figure 3: S102NO), the voice registration control unit 102 stores the voice as the wake word 50a in the storage unit 50 (Figure 3: S103). Here, regarding the storage (recording) of the wake word 50a, the voice registration control unit 102 stores it when the wake word 50a is not stored, but this is not limited to this. For example, the portable device 1 may be provided with a predetermined recording button to adjust the timing and recording time of the wake word 50a recording. For example, when a user presses the recording button, the voice registration control unit 102 uses the sound collection unit 20 to collect voice from the person and stores the wake word 50a. Furthermore, while the user continues to press the record button, the voice registration control unit 102 sets the duration of the press as the recording time and uses the sound collection unit 20 to collect voice from the person. When the user stops pressing the record button, the voice registration control unit 102 stores the collected voice as the wake word 50a. This makes it easy to rewrite the wake word 50a.
[0064] In this embodiment of the present invention, the portable device 1 is configured to include each control unit, but it is also possible to configure the device to store a program that implements each control unit on a storage medium and provide the storage medium. In this configuration, the program is read by the device, and the device implements each control unit. In this case, the program read from the recording medium itself performs the effects of the present invention. Furthermore, it is also possible to provide a method for storing the processes executed by each control unit on a hard disk.
[0065] As described above, the portable device and control method for the portable device according to the present invention are useful not only for pets and people with hearing impairments, but also for people and animals such as healthy individuals, people with hearing loss, and elderly people with reduced hearing ability. The portable device and control method for the portable device are effective in making the wearer aware of a specific person's voice even in environments with multiple people and ambient noise.
[0066] 1 Portable device 10 Mounting unit 20 Vibration unit 30 Sound collection unit 40 Control unit 50 Memory unit 101 Registration judgment control unit 102 Voice registration control unit 103 Voice sound collection control unit 104 Voice judgment control unit 105 Sound collection continuation control unit 106 Vibration control unit
Claims
1. A portable device comprising: an attachment part that can be attached to a part of the body of a person or animal; a sound collection unit attached to the attachment part for collecting sounds from people in the surrounding area; a vibration unit attached to the attachment part; a registration determination control unit that determines whether a predetermined wake word is registered in a predetermined storage unit when the portable device is activated; a voice registration control unit that, if the wake word is not registered in the storage unit, uses the sound collection unit to store a sound from a person as a wake word in the storage unit; a voice collection control unit that, if the wake word is registered in the storage unit, uses the sound collection unit to collect a sound from a person; a voice determination control unit that determines whether the collected sound matches the wake word; a sound collection continuation control unit that, if the sound does not match the wake word, continues to collect a sound from a person using the sound collection unit; and a vibration control unit that vibrates the vibration unit when the sound matches the wake word.
2. The portable device according to claim 1, wherein the vibration control unit activates the sound collection unit for a predetermined time to collect sound further emitted by the person, and vibrates the vibration unit with a vibration pattern corresponding to the collected sound.
3. The portable device according to claim 1, wherein the voice registration control unit registers the characteristic quantities of the person's voice in addition to the text of the wake word in the storage unit, and the voice determination control unit extracts text from the collected voice and extracts characteristic quantities from the voice, and determines that the collected voice matches the wake word if the extracted text matches the text of the wake word registered in the storage unit and the extracted characteristic quantities match the characteristic quantities registered in the storage unit.
4. A method for controlling a portable device comprising: an attachment step of attaching a portable device having an attachment part, a sound collection part, and a vibration part to a part of the body of a person or animal; a registration determination control step of determining whether a predetermined wake word is registered in a predetermined storage unit when the portable device is activated; a voice registration control step of using the sound collection unit to store a voice from a person as the wake word in the storage unit if the wake word is not registered in the storage unit; a voice collection control step of using the sound collection unit to collect a voice from a person if the wake word is registered in the storage unit; a voice determination control step of determining whether the collected voice matches the wake word; a sound collection continuation control step of using the sound collection unit to continue collecting a voice from a person if the voice does not match the wake word; and a vibration control step of vibrating the vibration part if the voice matches the wake word.