Information processing system and information processing method
The system automatically personalizes device settings by identifying users through biometric information, addressing the need for manual reconfiguration and separate authentication devices in existing systems, ensuring seamless user experience.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2024-11-08
- Publication Date
- 2026-05-20
AI Technical Summary
Existing information processing systems require manual reconfiguration of device settings, such as voice quality, each time a different user uses the device, and often necessitate the use of a separate personal authentication device for automatic configuration.
An information processing system that includes a user-worn detection device and an information processing device, utilizing sensors to detect biometric information from muscle movements, which automatically identifies the user and configures device settings based on pre-registered biometric data, enabling seamless personalization without additional hardware.
Enables easy personalization of device settings by automatically recognizing and configuring the system for the user, providing convenience and eliminating the need for manual reconfiguration when multiple users share the device.
Smart Images

Figure 2026083832000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing system that estimates character information, voice information, or a command for controlling a system from biological information.
Background Art
[0002] In recent years, it has been practiced to recognize the content of speech by using the voice information of a user. On the other hand, instead of voice information, biological information is acquired, and the expression and character information of the user are recognized from the biological information and output. Patent Document 1 discloses a technique for identifying the content of speech from biological information during voiceless operation acquired using various sensors. It is possible to output characters or voices based on the identified content of speech.
[0003] When using the device described in Patent Document 1, device settings such as the voice quality of voice output can be set by the connected PC. Further, by using it in combination with a personal authentication device as shown in Patent Document 2, the user can be recognized and used only by a specific user.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, in the case of Patent Document 1, when multiple people use the same device, device settings such as the voice quality of the audio output must be selected each time the device is used. It is also possible to configure the settings from a PC connected to the device, but even in that case, the settings must be configured each time the user changes. It is also possible to configure the settings automatically based on the authentication result by using it in conjunction with a personal authentication device as shown in Patent Document 2, but this requires the separate use of a personal authentication device.
[0006] The present invention aims to enable easy personalization of a device in an information processing system that estimates text information, voice information, or command information for controlling a system from biological information. [Means for solving the problem]
[0007] To achieve the objectives of the present invention, the information processing system of the present invention comprises: an acquisition unit that acquires a user's biometric information using a sensor; an estimation unit that estimates text information, voice information, or commands based on the biometric information acquired by the acquisition unit; a determination unit that determines the user based on the biometric information; and a control unit that configures the system based on the determination result of the determination unit. [Effects of the Invention]
[0008] According to the present invention, in an information processing system that estimates text information, voice information, or command information to control a system from biological information, it is possible to easily personalize the device settings. [Brief explanation of the drawing]
[0009] [Figure 1] A diagram showing the configuration of the information processing system according to the embodiment. [Figure 2] A flowchart illustrating the operation of the information processing system according to the embodiment. [Figure 3] A flowchart showing the initial setup of the information processing system according to the embodiment. [Figure 4] A diagram showing the configuration of the information processing system according to Example 1. [Figure 5] A schematic diagram of the detection device according to Example 1. [Figure 6] A diagram showing the configuration of the information processing system according to Example 2. [Figure 7] A schematic diagram of the detection device according to Example 2. [Modes for carrying out the invention]
[0010] Embodiments of the present invention will be described in detail below. A schematic diagram of the information processing system of the present invention is shown in Figure 1. The information processing system is a system that converts the content of a utterance from biological information such as muscle activity during utterance into textual information, voice information, or commands for controlling the system. In this disclosure, utterance includes utterances in which there is no or near-zero audible sound. For explanatory purposes, utterances in which there is no or near-zero audible sound will be specifically referred to as silent speech or voiceless utterance, and utterances in which there is audible sound at a normal level will be referred to as voiced utterance. In other words, in this disclosure, utterance is a term that includes both silent speech or voiceless utterance and voiced utterance.
[0011] The information processing system according to this embodiment is a system that estimates the content of a user's speech (including silent speech) from the user's biometric information. The information processing system according to this embodiment can also be considered as an information conversion system that converts the user's biometric information into text information or voice information representing the content of speech or commands for controlling the system.
[0012] The information processing system according to this embodiment mainly includes a user-worn detection device 100 and an information processing device 101 configured to communicate with the detection device 100.
[0013] The detection device 100 includes a processor, memory, and communication device. The detection device 100 has an acquisition unit 102, a first transmission unit 103, a first reception unit 104, and a device output unit 105. The acquisition unit 102 detects biometric information from one or more locations on the user. The first transmission unit 103 transmits the biometric information detected by the acquisition unit 102 to the information processing device 101. The first reception unit 104 receives the signal transmitted from the information processing device 101. The device output unit 105 provides information to the user based on the signal received by the first reception unit 104.
[0014] Figure 1 shows a configuration in which the detection device 100 has one acquisition unit, but the detection device 100 may have two or more acquisition units. Furthermore, the detection device 100 may have a mechanism for acquiring voice information in addition to biological information. Also, the detection device 100 can be rephrased as an acquisition unit that acquires biological information.
[0015] The acquisition unit 102 consists of sensors that detect biological information related to the user's muscle movements, the movements of various body parts such as skin and tongue. The biological information detected by the acquisition unit 102 may, for example, be related to the position of the sensors in accordance with the user's skeleton when there is no movement, i.e., when the detection device 100 is attached, or it may be biological information related to the movement of the user's vocal organs. The acquisition unit 102 may be, for example, an acceleration sensor and an angular velocity sensor that detect the movement of the user's mouth and tongue. Alternatively, the acquisition unit 102 may be an acceleration sensor and an angular velocity sensor that detect the movement of the user's mouth and tongue, and an electromyography sensor. Alternatively, the acquisition unit 102 may be an acceleration sensor and an angular velocity sensor that detect the movement of the user's mouth and tongue, and a tactile sensor.
[0016] An acceleration sensor is a sensor that detects acceleration and outputs data or a signal corresponding to the detected acceleration. An angular velocity sensor (gyro sensor) is a sensor that detects angular velocity and outputs data or a signal corresponding to the detected angular velocity. An electromyogram sensor is a sensor that detects weak electrical signals generated by muscle activity and outputs data or a signal. A tactile sensor is a sensor that outputs data or a signal corresponding to the force or the direction of the force transmitted to the sensor. The acceleration sensor, angular velocity sensor, electromyogram sensor, and tactile sensor in the acquisition unit 102 may be integrated.
[0017] The acquisition unit 102 is disposed at a position where it can detect the movements of the user's mouth and tongue. For example, the acquisition unit 102 is disposed around the user's mouth. Specifically, the acquisition unit 102 can be disposed at least at any one of the user's lower jaw, cheek, mouth periphery, throat, subauricular region, neck, and temple. As long as it is a position where it can detect biometric information related to the movements of the mouth and tongue, the acquisition unit 102 may be disposed at a site other than the above-mentioned sites.
[0018] The acquisition unit 102 may include an ultrasonic sensor, an optical sensor, a pressure sensor, etc. Specifically, the acquisition unit 102 may be one in which a pressure sensor is incorporated into an acceleration sensor and an angular velocity sensor. Also, the acquisition unit 102 may be one in which a geomagnetic sensor is incorporated. The acquisition unit 102 can also acquire information related to the movements of the skin and muscles at the same position.
[0019] The acceleration sensor and angular velocity sensor in the acquisition unit 102 can also detect movements other than the movements of the mouth and tongue such as the user's blinking and head movements.
[0020] The biometric information detected by the acquisition unit 102 is used for both the estimation process of corresponding character information, voice information, or a command for controlling the system or devices connected to the system, and the determination process of recognizing the user. In other words, the biometric information for converting into voice information, character information, or a command and the biometric information for identifying the user and selecting device settings are detected by the same sensor.
[0021] The first receiving unit 104 can receive signals from the information processing device 101. The device output unit 105 outputs sound, light, or vibration to the user in accordance with the signal received by the first receiving unit 104. The device output unit 105 may include earphones, light-emitting diodes, vibration elements, etc. If the device output unit 105 includes earphones, it outputs sound to the user in accordance with the signal received by the first receiving unit 104. Similarly, if the device output unit includes light-emitting diodes or vibration elements, it outputs light or vibration to the user in accordance with the signal received by the first receiving unit 104.
[0022] The information processing device 101 is a computer including a processor, memory, input / output devices, and communication devices. The information processing device 101 functions as a second receiving unit 106, a determination unit 107, a second transmitting unit 108, an estimation unit 109, a control unit 112, and a learning unit 113 by the processor executing a program stored in memory. The second receiving unit 106 receives biometric information transmitted from the detection device 100. The determination unit 107 determines, based on the biometric information, which of the pre-registered users the biometric information received by the second receiving unit 106 belongs to. The second transmitting unit 108 transmits information based on the determination result of the determination unit 107 to the detection device 100. The estimation unit 109 estimates character information, voice information, or commands corresponding to the biometric information based on the received biometric information. The learning unit generates a trained model for estimation by the estimation unit 109 based on the biometric information received by the second receiving unit 106. The control unit 112 controls the entire system based on the results from the determination unit 107 and the estimation unit 109.
[0023] The information processing device 101 may be a smartphone, personal computer (PC), tablet PC, etc., but is not limited to these. Furthermore, the information processing device 101 may be installed within some other device, such as a camera. If the information processing device 101 is, for example, a personal computer, the character information estimated by the estimation unit 109 is transmitted to the display unit 110, such as a display.
[0024] The display unit 110 displays character information. The display unit 110 may also display a user interface for controlling the detection device 100 and the information processing device 101. The information processing device 101 may have a display control unit (not shown) that controls the display format on the display unit 110.
[0025] The determination unit 107 uses a predetermined determination algorithm to determine whether a user is registered in advance from the biometric information detected by the acquisition unit 102.
[0026] The biological information used by the determination unit 107 is the biological information at the time the device is attached, and is biological information in a state without physical movement, but it may also be biological information resulting from a specific physical movement set in advance.
[0027] The determination unit 107 uses a predetermined determination method to determine whether a user is pre-registered based on the biometric information detected by the acquisition unit 102 when the detection device is attached. The determination method uses a trained model (trained model for determination) with an architecture composed of a neural network. The information processing device 101 has a storage unit (not shown) that stores the trained model. The determination unit 107 has a function to perform inference using the trained model. The trained model is a model generated using deep learning such as CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), or Transformer. The trained model may be a model derived from CNN, RNN, or Transformer, or it may be a model based on other machine learning techniques such as support vector machines, logistic regression, random forests, or hidden Markov models. In addition, the determination unit 107 may perform estimation using a rule-based method instead of or in addition to using the trained model.
[0028] The estimation unit 109 uses a predetermined estimation algorithm (conversion algorithm) to estimate corresponding character information, voice information, or a command for controlling the system from the biometric information detected by the acquisition unit 102. The estimation unit 109 can also be described as a conversion unit that converts the received biometric information into character information, voice information, or a command. The estimation unit 109 may also estimate voice information corresponding to the character information based on the estimated character information. Alternatively, the estimation unit 109 may directly estimate voice information based on the received biometric information.
[0029] The estimation unit 109 uses a predetermined estimation method when estimating character information, voice information, or commands from biometric information detected by the acquisition unit 102. The estimation method uses a trained model (trained estimation model) with an architecture composed of a neural network. The information processing device 101 has a storage unit (not shown) that stores the trained model. The estimation unit 109 has a function to perform inference using the trained model. The trained model is a model generated using deep learning such as CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), or Transformer. The trained model may be a model derived from CNN, RNN, or Transformer, or a model based on other machine learning techniques such as support vector machines, logistic regression, random forests, or hidden Markov models. In addition, the estimation unit 109 may perform estimation using a rule-based method instead of or in addition to using the trained model. Furthermore, the determination unit 107 and the estimation unit 109 may be integrated to perform determination and estimation through a single trained model. In other words, the trained model used by the determination unit 107 (trained model for determination) and the trained model used by the estimation unit 109 (trained model for estimation) may be the same trained model. In this case, the biometric information at the time of attachment is trained by the determination unit as a token (e.g., username) representing the user who attached the device.
[0030] In Figure 1, the determination unit 107 and the estimation unit 109 are configured in the information processing device 101. The determination unit 107 and the estimation unit 109 may also be configured on a detection device 100 or on the cloud.
[0031] The voice information estimated by the estimation unit 109 is transmitted to the sound information output unit 111. The sound information output unit 111 is a speaker and can reproduce the voice information. The text information and voice information estimated by the estimation unit 109 can also be transmitted to another receiver via a network or the like, and displayed and played back at the destination.
[0032] The control unit 112 controls the system based on the determination results of the determination unit 107 and the estimation unit 109. More specifically, the control unit 112 selects the voice quality for voice output set during initial setup based on the determination results of the registered user. It also selects a pre-trained estimation model specific to the determined user. Furthermore, in order to perform system control, if the estimation unit 109 determines that the user corresponds to one of the specific movements, it executes a command corresponding to that movement (system control).
[0033] The "specific movement for system control" (hereinafter also referred to as "specific movement") can be any movement that can be acquired by the detection device 100, and can be, for example, a movement different from the movement when speaking. An example of a "specific movement" is any of the following: movement of the mouth, movement of the tongue, movement of the cheeks, movement of the eyeballs or eyelids, movement of the face, or a combination thereof. The "specific movement" may also be a movement for uttering a predetermined keyword.
[0034] "System control" can be any control relating to the detection device 100 or the information processing device 101, or any control of a device connected to the system. Examples of "system control" include starting or ending the detection of biometric information and its conversion into text or audio information, and specifying the operating mode when detecting biometric information and converting it into text or audio information. For example, if the biometric information is estimated to correspond to a first specific movement when the text information estimation process is not being performed, the estimation unit starts estimating text information using the subsequently detected biometric information. Then, if the biometric information is determined to correspond to a second specific movement while the estimation process is being performed, the estimation by the estimation unit ends. The second specific movement may be a stationary movement from motion. Alternatively, the system may be controlled to switch to an operating mode corresponding to a movement or to launch an application corresponding to a movement, depending on which of a predetermined set of specific movements the biometric information corresponds to.
[0035] Furthermore, the information processing device 101 may be configured on the cloud. In this case, the first transmission unit 103 of the detection device 100 transmits the biometric information detected by the acquisition unit 102 to the cloud. The cloud performs user identification based on the biometric information and converts it into text information, voice information, or commands. If converted to text information, the text information is transmitted to the display unit 110. The display unit 110 displays the text information. Alternatively, the converted text information may be further converted into voice information and transmitted to the sound information output unit 111. Alternatively, the biometric information may be directly converted into voice information and transmitted to the sound information output unit 111. The converted command is then transmitted to the control unit 112, which executes the command corresponding to the movement.
[0036] Communication between the detection device 100 and the information processing device 101 may be wired or wireless. In the case of wired communication, the first transmitter 103 and first receiver 104 of the detection device 100 and the second receiver 106 and second transmitter 108 of the information processing device 101 are connected by a wire, such as a USB cable. In the case of wireless communication, each of the above parts is connected wirelessly by means of a wireless LAN such as WiFi, or short-range wireless communication such as Bluetooth®.
[0037] If the information processing device 101 is a smartphone or tablet PC, then the display unit 110 is a display. The sound information output unit 111 is a speaker located on the smartphone or tablet PC, or earphones connected to the smartphone or tablet PC.
[0038] Furthermore, the information processing device 101 can be considered an information conversion device that converts biometric information detected by the acquisition unit 102 into text information, voice information, or commands. The information processing device 101 can also be considered an information estimation device that performs text information, voice information, commands, or personal authentication from the biometric information detected by the acquisition unit 102. In addition, the information processing device 101 may include a processing unit that evaluates the biometric information detected by the acquisition unit 102 based on predetermined evaluation criteria and deletes text information or voice information corresponding to biometric information that does not meet the predetermined evaluation criteria.
[0039] Figure 2 shows a flowchart illustrating the flow of the information processing method executed by the information processing system according to this embodiment.
[0040] Step S100: The user attaches the detection device 100 to a location where biological information can be detected by the acquisition unit 102. For example, the user attaches the detection device 100 to the head, neck, or the area around the head and neck.
[0041] Step S101: During attachment, the acquisition unit 102 detects biological information that includes acceleration information and angular velocity information, or acceleration information and angular velocity information plus at least one of the following: electromyographic signals, tactile information, etc.
[0042] Step S102: The first transmission unit 103 transmits biological information to the information processing device 101. The first transmission unit 103 can transmit biological information and time information to the information processing device 101.
[0043] Step S103: The biological information detected by the detection device 100 is received by the second receiving unit 106 of the information processing device 101. The second receiving unit 106 transmits the received biological information to the determination unit 107.
[0044] Step S104: The determination unit 107 determines the user based on the initially set biometric information at the time of attachment using a predetermined determination algorithm. The biometric information acquired in advance at the time of attachment may be biometric information when the mouth is not moving, or it may be a specific movement that has been set in advance. By transmitting the determination result to the control unit 112, the control unit 112 selects a machine learning model optimized for the user, voice quality information to be output as voice, and correspondence information between movements and commands for controlling the system, based on the determined user.
[0045] Step S105: Based on the determination result, the determination unit 107 communicates to the user via the second transmission unit 108 and the first reception unit 104 from the device output unit 105 that the system startup is complete. If the user cannot be identified, the determination unit notifies the user to perform the initial setup.
[0046] Step S106: After the activation notification, the acquisition unit 102 detects the biometric information. The biometric information is accompanied by supplementary information regarding time information (time information).
[0047] Step S107: The first transmission unit 103 transmits biological information and time information to the information processing device 101.
[0048] Step S108: The biological information detected by the detection device 100 is received by the second receiving unit 106 of the information processing device 101. The second receiving unit 106 transmits the received biological information to the estimation unit 109.
[0049] Step S109: The estimation unit 109 estimates text information, voice information, or command information from the biological information using a predetermined estimation algorithm. It determines, using a predetermined determination algorithm, whether the biological information corresponds to a specific movement that has been set in advance (including not moving the mouth for a certain period of time). The estimation unit 109 may also estimate voice information from the biological information, or it may estimate text information and then perform speech synthesis to convert it into voice information. The estimation unit 109 outputs the information estimated from the biological information.
[0050] Step S110: If the control unit 112 determines that the biometric information corresponds to a movement that initiates conversion, it sends a signal as a command to the device output unit 105 via the first receiving unit 104 of the detection device 100 through the second transmitting unit 108. The device output unit 105 then notifies the user of the start of conversion with a sound. The control unit 112 also controls the estimation unit 109 to start estimating the biometric information to be converted, which will be acquired thereafter. After the estimation has started, if the control unit 112 determines that the biometric information has not moved the mouth for a certain period of time, it controls the estimation unit 109 to terminate the estimation. The display unit 110 displays the character information estimated by the estimation unit 109. The sound information output unit 111 outputs audio information obtained by further converting the character information estimated by the estimation unit 109.
[0051] Furthermore, the character information estimated by the estimation unit 109 can be stored in the storage unit of the information processing device 101. The character information can also be transferred via a network and displayed on an external terminal.
[0052] Furthermore, by converting text information into audio information, it is possible to play and record it on an external device, and the user can listen to the played audio through earphones. This allows the user to verify whether the estimation is correct. In addition, the audio information can be transferred over the network and played and recorded on other external devices.
[0053] Figure 3 shows a flowchart illustrating the flow of the initial information processing method according to this embodiment.
[0054] Step S200: The user enters user-related information, such as the username, using the application on the information processing device 101 for initial setup.
[0055] Step S201: The detection device 100 is attached multiple times to a position where biological information can be detected by the acquisition unit 102. During attachment, the acquisition unit 102 detects biological information that includes acceleration information and angular velocity information, or acceleration information and angular velocity information plus at least one of the following: electromyographic signals, tactile information, etc., and sends it to the learning unit 113 via the first transmission unit 103 and the second reception unit 106.
[0056] Step S202: Select a control item from the application as a command to control the system, acquire biometric information related to the movement to execute it using the acquisition unit 102, and send it to the learning unit 113 via the first transmission unit 103 and the second reception unit 106.
[0057] Step S203: The acquisition unit 102 acquires biometric information at the time of speech for multiple sentences presented by the application and sends it to the learning unit 113 via the first transmission unit 103 and the second receiving unit 106.
[0058] Step S204: Learn the information transmitted to the learning unit 113 and create a trained model for user recognition. Also, create a trained model for commands to control the system and a trained model for estimating textual or speech information from biometric information. Then, save the created trained models to the information processing device 101. Note that for learning, a pre-prepared large-scale trained model may be used as the initial model and then fine-tuned to create the trained model.
[0059] The information processing system according to this embodiment is highly convenient for users because it automatically identifies the user when the device is attached and configures the system accordingly. Furthermore, even when multiple people use the same device, the system identifies the user simply by attaching the device and automatically configures the system according to the user's initial settings, thus providing high user convenience.
[0060] (Example 1) Next, the information processing system according to Embodiment 1 will be described. The information processing system according to this embodiment converts sensing data (biometric information) obtained from a 3-axis acceleration sensor, a 3-axis angular velocity sensor, and a 3-axis tactile sensor attached to the user's cheek and under the chin into textual information or audio information.
[0061] Figure 4 shows an overview of the information processing system according to this embodiment, and Figure 5 shows a schematic diagram of the detection device 100 according to this embodiment. Here, the information processing device of the information processing system is shown to be a smartphone 300. The detection device 100 is powered by a battery (not shown).
[0062] As shown in Figure 5, the user wears the detection device 100 around their neck. The detection device 100 is a headset type worn on the user's ears. The detection device 100 consists of a first acquisition unit 401, a second acquisition unit 402, a transmitting / receiving unit 403, and a sound information output unit 404. The first acquisition unit 401 and the second acquisition unit 402 are each composed of a contact-type sensor unit that integrates a 3-axis acceleration sensor, a 3-axis angular velocity sensor, and a 3-axis tactile sensor, respectively. The user may wear the detection device 100 after powering it on, or after putting it on.
[0063] The transmitting / receiving unit 403 is connected to the first acquisition unit 401 and the second acquisition unit 402. The transmitting / receiving unit 403 can transmit biometric information detected by the first acquisition unit 401 and the second acquisition unit 402 to an external source. The transmitting / receiving unit 403 can also acquire information from the smartphone 300. The biometric information detected by the first acquisition unit 401 and the second acquisition unit 402 is wirelessly transferred to the smartphone 300 by the transmitting / receiving unit 403.
[0064] When a user puts on the detection device 100, the receiving unit 301 receives biometric information detected by the first acquisition unit 401 and the second acquisition unit 402, which have information about the wearing status. The biometric information is then sent to the determination unit 302, which determines the user from a trained model. After user determination, it selects a model for converting the biometric information into text information and voice information, and also selects the voice quality for voice output. Once all settings are complete, the transmitting unit 304 sends a signal to the detection device 100 indicating that the settings are complete. The detection device 100 then notifies the user from the output unit 404 that the device is ready for use.
[0065] Subsequently, the biometric information received by the receiving unit 301 is transmitted to the estimation unit 303, which determines, based on a trained model, whether the received biometric information corresponds to a specific action for controlling the smartphone. For example, if the estimation unit 303 determines that the biometric information corresponds to a nodding motion, the smartphone 300 transmits the determination result to the detection device 100 via the transmitting unit 304. The output unit 404 emits a sound to inform the user that biometric information estimation has begun, and the estimation unit 303 begins estimating voice information from the biometric information using the trained model. The estimation unit 303 continues estimation until it determines that speech has ended, based on the standard deviation of the signal acquired from the sensor over a 200-millisecond period being below a certain value. When it determines that speech has ended, based on the standard deviation of the signal acquired from the sensor over a 200-millisecond period being below a certain value, the estimation process ends.
[0066] The control unit 308 performs system control according to the movement when the estimation unit 303 estimates that the biometric information detected from one or more locations on the user corresponds to one of the specific movements required for system control.
[0067] In this embodiment, the generation of a trained model for identifying a user from biometric information was performed as follows. First, the biometric information detected by the first acquisition unit 401 and the second acquisition unit 402, which represent the biometric information when the user wears the detection device 100, is received. This information is then transferred from the data transfer unit 307 to the learning unit 501 in the cloud for training to extract the user's features. During training, biometric information acquired by having the user wear the detection device 100 five times is used. After training is complete, the trained model is transmitted to the smartphone 300 and stored on the smartphone.
[0068] The biometric information collected during wear primarily uses biometric data from a stationary state when the detection device 100 is worn, but biometric information related to specific movements determined by the user may also be used. Furthermore, learning to extract user characteristics is performed the first time the detection device 100 is used, but this learning can be performed again later. This data acquisition and learning during wear can be performed via an application on the smartphone 300.
[0069] Furthermore, the pre-trained models used for determining predetermined movements for system control and for estimating speech or text information from biometric information are identical. The pre-trained models were generated as follows: First, biometric information was acquired three times for each of 20 specific movements related to system control. Then, biometric information was acquired three times each for each of 10 specified sentences during silent speech movements. This data was used to fine-tune a large-scale pre-trained model prepared in advance, which was then used as the user's pre-trained model.
[0070] The audio information converted by the estimation unit 303 is transferred to the sound information output unit 306 and played back. If the conversion is to text information, the text information converted by the estimation unit 303 is displayed on the smartphone's display, which is the display unit 305. Furthermore, the simultaneously converted audio is transferred to the transmission / reception unit via the transmission unit 807. The transferred audio can be played back via the sound information output unit 704, such as earphones, allowing the conversion result to be confirmed.
[0071] The conversion from biometric data to text or audio data can be done by converting all biometric data up to the point where it is estimated that there has been no mouth movement for more than one second, or by converting and outputting sequentially from the start of estimation.
[0072] (Example 2) Next, an information processing system according to Example 2 will be described. The information processing system according to this example converts sensing data (biometric information) obtained from acceleration sensors and angular velocity sensors attached to the user's cheek and under the chin into textual information or voice information.
[0073] Figure 6 shows an overview of the information processing system according to this embodiment, and Figure 7 shows a schematic diagram of the detection device 100 according to this embodiment. Here, the information processing device of the information processing system is shown to be a smartphone 300.
[0074] As shown in Figure 7, the detection device 100 has two separate components. The user attaches the detection device 100 to the cheek and subchinenchyma, for example, with a self-adhesive gel.
[0075] The detection device 100 consists of a first acquisition unit 601, a second acquisition unit 602, a first transmitting / receiving unit 603, a second transmitting / receiving unit 604, and an output unit 605. The first acquisition unit 601 and the first transmitting / receiving unit 603, and the second acquisition unit 602 and the second transmitting / receiving unit 604 are each integrated and each has a built-in battery for power supply. The first transmitting / receiving unit 603 and the second transmitting / receiving unit 604 are wirelessly connected to each other and can communicate with each other, and can transmit biological information detected by the first acquisition unit 601 and the second acquisition unit 602 to the outside via the first transmitting / receiving unit 603.
[0076] Both the first acquisition unit 601 and the second acquisition unit 602 are 6-axis acceleration / angular velocity sensors. A 6-axis acceleration / angular velocity sensor is a sensor that can measure, for example, 3-axis translational acceleration and 3-axis angular acceleration. As a result, the first acquisition unit 601 and the second acquisition unit 602 can detect information regarding the movement of the user's body parts.
[0077] The biometric information detected by the first acquisition unit 601 and the second acquisition unit 602 is transmitted by the first transmission / reception unit 603 and the second transmission / reception unit 604 to, for example, a Bluetooth-connected smartphone 300. A transmission / reception unit 901 within the application of the smartphone 300 receives the data.
[0078] When a user puts on the detection device 100 and activates it, the estimation unit 902 estimates the user using a trained model based on the data received by the transmitting / receiving unit 901. Based on this information, the control unit 904 performs system settings, such as selecting a trained model that has been previously registered by the user. The data received at the time of attachment is information about the posture when the detection device 100 is attached. The system also sends a message to the first transmitting / receiving unit 603 via the transmitting / receiving unit 901 indicating that the system has started up. The output unit 605 notifies the user of the completion of the startup with a sound. Subsequently, the estimation unit 902 uses a trained model to determine whether the data received by the transmitting / receiving unit 901 corresponds to a specific movement or to the fact that the mouth has not moved for 1 second. An example of a specific movement is closing the eyes for 0.5 seconds or more. If the estimation unit 902 determines that the received data corresponds to the movement of closing the eyes for 0.5 seconds or more, the estimation unit 902 sends that information to the first transmitting / receiving unit 603 via the transmitting / receiving unit 901. The output unit 605 signals to the user with a sound that subsequent biometric information will be converted. After hearing the sound, the user can estimate the biometric information by silently performing the desired vocal movements. Furthermore, if the estimation unit 602 determines that the mouth has not moved for more than one second after the estimation has started, it determines that data acquisition for estimation is complete.
[0079] The estimation process in the estimation unit 602 is determined based on a trained model. In this embodiment, the trained model was generated based on a trained model that had been previously trained on a large dataset. When the user wore the detection device 100, biometric information was acquired five times for 0.5 seconds or longer for 20 pre-set movements, such as closing the eyes, with biometric information acquired five times for each movement. In addition, biometric information was acquired for 100 sentences during silent speech. Fine tuning was performed using this data. The neural network was trained by transferring data from the data transfer unit 903 to the learning unit 801 on the cloud.
[0080] The control unit 904 performs system control according to the movement when the estimation unit 902 estimates that the biometric information detected by the user corresponds to one of the specific movements required for system control.
[0081] The estimation unit 902 inputs biometric information received by the transmitting / receiving unit 901 during silent movement into a trained model to estimate corresponding character information or speech information. Specifically, character information or speech information is estimated by using sensor information from a total of 12 axes (6-axis x 2 acceleration / angular velocity sensors) as input data to the trained model. The estimation unit 902 has the function of estimating character information mixed with kanji and hiragana from the estimated phoneme sequence. It also has the function of synthesizing speech information from the phoneme sequence.
[0082] The audio information converted by the estimation unit 902 is transferred to the data transfer unit 903. The audio information converted by the estimation unit 902 is transferred to another smartphone 905 and played back. The user simultaneously receives the converted audio via the transmission / reception unit 901 to the detection device 100. The transferred audio is played back via the sound information output unit 605, such as earphones, and the user confirms the conversion result.
[0083] (Other examples) The above-described embodiment is merely one example of the information processing system relating to this disclosure, and can be modified as appropriate within the scope of the concept of this disclosure.
[0084] For example, the above description uses neckband-type, neckband-type and earphone-type combination, and headset-type detection devices as examples, but detection devices may also be glasses-type, mask-type, chin mask-type, choker-type, etc.
[0085] Furthermore, the biometric information used for inferring textual or audio information and for determining inputs for system control is not limited to the specific examples described above. The information processing related to this disclosure may use any biometric information that can be obtained from a part of the user. For example, biometric information related to the movement of the arms, legs, chest, or abdomen may be obtained. In addition, the detection of biometric information may be performed by sensors other than those permanently installed. For example, an image may be taken using a camera, and the movement of a part of the user's body obtained by image analysis may be used as biometric information.
[0086] Furthermore, while an information processing system consisting of a detection device and an information processing device has been given as an example, the information processing system may also include other devices, or it may be an integrated system in which the detection device and the information processing device are included in a single device. [Explanation of Symbols]
[0087] 102 Acquisition Department 107 Judgment section 109 Estimation part 112 Control Unit
Claims
1. An acquisition unit that acquires the user's biometric information using sensors, An estimation unit that estimates text information, voice information, or commands based on the biometric information acquired by the acquisition unit, A determination unit that determines the user based on the aforementioned biometric information, An information processing system comprising a control unit that configures the system based on the determination result of the determination unit.
2. The estimation unit includes a machine learning model for estimating text information, voice information, or commands based on the biometric information. The information processing system according to claim 1, characterized in that the aforementioned settings include the settings of the machine learning model.
3. The information processing system according to claim 1, characterized in that the setting includes setting the voice quality when converting the biometric information into voice information and outputting it.
4. The information processing system according to claim 2, characterized in that the machine learning model includes settings relating to the correspondence between the biometric information relating to the user's specific movements and commands for controlling the system.
5. The information processing system according to claim 1, characterized in that the acquisition unit acquires the biometric information for estimating voice information or text information and the biometric information for determining an individual's characteristics using the same sensor.
6. The information processing system according to claim 1, characterized in that the acquisition unit includes at least one of an acceleration sensor or an angular velocity sensor.
7. The information processing system according to claim 1, characterized in that the acquisition unit includes an acceleration sensor or an angular velocity sensor and an electromyography sensor.
8. The information processing system according to claim 1, characterized in that the acquisition unit includes an acceleration sensor or an angular velocity sensor and a tactile sensor.
9. The information processing system according to claim 1, characterized in that the acquisition unit acquires biological information from at least two or more locations of the user, including the sub-chin, cheek, throat, below the ear, neck, and temple.
10. The information processing system according to claim 9, characterized in that the acquisition unit is positioned on the user's cheek and under the chin.
11. The estimation unit estimates the character information or the voice information based on the biometric information using the first trained model. The information processing system according to claim 1, characterized in that the determination unit determines the characteristics of an individual based on the biometric information using a second trained model.
12. The information processing system according to claim 11, characterized in that the first trained model and the second trained model are the same trained model.
13. The information processing system according to claim 1, characterized in that the acquisition unit is included in a device that can be worn by the user, and the estimation unit, the determination unit, and the control unit are included in an information processing apparatus configured to communicate with the device.
14. The information processing system according to claim 1, characterized in that the acquisition unit and the determination unit are included in a device that can be worn by the user, and the estimation unit and the control unit are included in an information processing device configured to communicate with the device.
15. The information processing system according to claim 13 or 14, characterized in that the device comprises an output unit that outputs character information or voice information estimated by the estimation unit.
16. The information processing system according to claim 13 or 14, characterized in that the information processing device includes an output unit that outputs character information or voice information estimated by the estimation unit.
17. The acquisition step involves obtaining biometric information from the user, An information processing method comprising: a determination step of determining a user based on the biometric information acquired in the acquisition step; a setting step of configuring the system by a control unit based on the determination result of the determination step; and an estimation step of estimating text information, voice information, or a command for controlling the system based on the biometric information acquired in the acquisition step.
18. A program for causing a computer to perform each step of the information processing method described in claim 17.