Information processing system and information processing method

The information processing system enhances the accuracy of converting biometric information into text or voice by employing a two-tier inference model architecture with pre- and post-processing, addressing the limitations of existing methods with limited training data.

JP2026001897APending Publication Date: 2026-01-08CANON KK
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024099469
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-20
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Existing methods for converting biometric information, such as speech movements, into text or voice information face accuracy issues due to insufficient training data and reliance on simple threshold processing or machine learning models, which require large datasets to improve accuracy.

Method used

An information processing system utilizing a two-tier inference model architecture, where a first inference model processes biometric information from muscle movements, followed by a second model to enhance the accuracy of converting this information into text or voice, leveraging pre-processing and post-processing units to refine the output.

Benefits of technology

The system significantly improves the accuracy of converting biometric information into text or voice information by using trained inference models with pre- and post-processing, even with limited training data, ensuring high precision in character and audio outputs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026001897000001_ABST
    Figure 2026001897000001_ABST
Patent Text Reader

Abstract

To improve the accuracy of conversion from biological information on the movement of a part of a user such as a speech operation to character information and voice information.SOLUTION: An acquisition unit 210 that acquires first biological information related to a motion of a part of a user; A first inference unit 223A configured to output pseudo information based on the first biological information acquired by the acquiring unit 210, and a second inference unit 223A configured to output character information or voice information based on the pseudo information output from the first inference unit 223B by using a second inference model learned by using the pseudo information and the character information or the voice information corresponding to the pseudo information as teacher data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing system that converts biometric information relating to the movement of a user's body part, such as speech, into text information or voice information. [Background technology]

[0002] In recent years, the content of a user's speech has been estimated from the movement of the user's mouth. Patent Document 1 discloses a technology for calculating a phoneme score from a lip image and external sound to infer the content of a user's speech. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2021-162685 Summary of the Invention [Problem to be solved by the invention]

[0004] However, in Patent Document 1, the weighted sum of the lip phoneme score calculated based on the lip image and the speech phoneme score calculated based on the external sound is simply subjected to threshold processing, and the accuracy may not be sufficient. Meanwhile, in recent years, it has become known to make inferences using inference models trained by machine learning, but a large amount of training data is required to improve the accuracy of inference, and it has been difficult to improve the accuracy of inference with a small amount of training data.

[0005] The present invention aims to improve the accuracy of converting biometric information relating to the movement of a user's body parts, such as speech movements, into text information or voice information.

[0006] However, the problems to be solved by the embodiments disclosed in this specification and the drawings are not limited to the above problems. Problems corresponding to the effects of the configurations shown in the embodiments described below can also be positioned as other problems. [Means for solving the problem]

[0007] In order to achieve the above object, the information processing system according to the technology disclosed herein comprises: an acquisition unit that acquires first biological information related to a movement of a part of a user; a first inference unit that outputs pseudo information based on the first biometric information acquired by the acquisition unit, using a first inference model trained using second biometric information and pseudo information corresponding to the second biometric information as training data; a second inference unit that uses a second inference model trained using pseudo information and character information or voice information corresponding to the pseudo information as training data to output character information or voice information based on the pseudo information output from the first inference model; The present invention is characterized by having the following. [Effects of the Invention]

[0008] According to the technology disclosed herein, it is possible to improve the accuracy of converting biometric information relating to the movement of a user's body part, such as speech movements, into text information or voice information. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram showing an example of the functional configuration of an information processing system according to a first embodiment. [Figure 2] FIG. 2 is a diagram showing an example of the functional configuration of a conversion unit according to the first embodiment. [Figure 3] FIG. 2 is a diagram showing an example of the functional configuration of an inference unit according to the first embodiment. [Figure 4] 4 is a flowchart showing a conversion process of the information processing system according to the first embodiment. [Figure 5] FIG. 2 is a diagram showing an example of the functional configuration of a learning unit according to the first embodiment. [Figure 6] 5 is a flowchart showing a learning process in the first embodiment of the learning unit according to the first embodiment. [Figure 7] FIG. 1 is a schematic diagram of a detection device according to a first embodiment. [Figure 8] FIG. 3 is a diagram illustrating the operation of the conversion unit according to the first embodiment. [Figure 9] FIG. 10 is a schematic diagram of a detection device according to a second modification. [Figure 10] FIG. 10 is a diagram illustrating the operation of a conversion unit according to the second embodiment. [Figure 11] FIG. 10 is a diagram illustrating the operation of a conversion unit according to the third embodiment. [Figure 12] FIG. 10 is a diagram showing an example of the functional configuration of an inference unit according to the second embodiment. [Figure 13] 10 is a flowchart showing a learning process according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] The following describes the embodiments in detail. The present invention is not limited to the following embodiments, provided that the gist of the present invention is not exceeded. Furthermore, by adding publicly known information, such as audio data or video data, to other channels of input data to the inference model, it is possible to perform inference with higher accuracy.

[0011] [First embodiment] An information processing system 1 according to this embodiment will be described with reference to FIG. 1. The information conversion system 1000 includes a detection device 100 and an information processing device 200. The detection device 100 detects biometric information based on muscle movements caused by a user's speaking actions. The information processing device 200 acquires the biometric information detected by the detection device 100 and converts it into text information or audio information using an inference model. The detection device 100 and the information processing device 200 may each be configured as separate devices connected by a network, or may be configured as an information processing device having the functions of a detection device as an integrated device.

[0012] <Detection device 100> The detection device 100 that constitutes the information processing system 1000 includes a biometric information detection unit 101 that detects biometric information from one or more locations on the user, and a transmission unit 102 that transmits the biometric information detected by the biometric information detection unit 101 to the information processing device 200.

[0013] The detection device 100 is placed, for example, in contact with the user's skin in order to acquire biometric information from the user. The detection device 100 is preferably placed on a site such as the neck, lower jaw, around the mouth, or temples in order to detect biometric information based on muscle movements caused by the user's speaking actions. However, the detection device 100 may be placed in a location other than the above-mentioned locations as long as it can detect biometric information related to the movements of the mouth and tongue.

[0014] 1 shows a configuration in which the detection device 100 is realized by one device, but the biological information detection unit 101 and the transmission unit 102 may be configured by different devices. Each functional configuration of the detection device 100 will be described below.

[0015] The biometric information detection unit 101 acquires biometric information as first biometric information related to the movement of a user's body part. Specifically, the biometric information detection unit 101 is configured to include a sensor that detects biometric information based on the movement of the user's muscles. The biometric information detection unit 101 is configured to include at least one sensor such as a myoelectric potential sensor, an acceleration sensor, an ultrasonic sensor, a tactile sensor, an optical sensor, a pressure sensor, a strain sensor, or a magnetic sensor. Note that the biometric information detection unit 101 may be a sensor other than the above, such as a camera, as long as it detects biometric information, and may not be in direct contact with the skin.

[0016] The transmitting unit 102 transmits to the information processing device 200 the biological information based on muscle movements caused by the user's speaking movements, which is acquired from the biological information detecting unit 101.

[0017] <Information processing device 200> Here, the information processing device 200 constituting the information processing system 1000 includes a biometric information acquisition unit 210, a conversion unit 220, and an output unit 230. The biometric information acquisition unit 210 as an acquisition unit acquires biometric information (first biometric information) transmitted from the detection device 100. The conversion unit 220 converts the biometric information into text information or audio information by using the biometric information as input to an inference model and inferring text information or audio information corresponding to the biometric information. The output unit 230 outputs the text information or audio information converted by the conversion unit 220 to the display unit 110 or the audio information output unit 120 (audio output unit).

[0018] Here, the inference model is a trained inference model trained by the training unit 240 and stored in the storage unit 250. The training unit 240 and other functional components may be configured as independent devices. For example, the training unit 240 may train the inference model used for inference in the conversion unit 220 on the cloud.

[0019] Here, the information processing device 200 refers to, for example, a smartphone, a personal computer (PC), a tablet PC, etc. The information processing device 200 includes a CPU, a GPU, a RAM, a ROM, and a storage device, and is realized by connecting these via a system bus. If the information processing device 200 is, for example, a personal computer, the character information converted by the conversion unit 220 and output by the output unit 230 is output to a display unit 110 such as a display. Furthermore, if the information processing device 200 is converted into audio information, it is output to an audio information output unit 120 (audio output unit) such as a speaker. Note that the audio information output unit 120 may be incorporated into the device configuration of the information processing device 200.

[0020] Here, communication between the detection device 100 and the information processing device 200 may be wired or wireless. When communication between the detection device 100 and the information processing device 200 is realized by wire, the transmission unit 102 in the detection device 100 and the biometric information acquisition unit 210 in the information processing device 200 are connected by wire using a USB cable, HDMI (registered trademark), or the like.

[0021] When communication between the detection device 100 and the information processing device 200 is realized wirelessly, the transmitting unit 102 in the detection device 100 and the receiving unit 210 in the information processing device communicate wirelessly. For example, they are connected wirelessly by wireless LAN communication such as WiFi, or short-range wireless communication such as Bluetooth (registered trademark). If the information processing device 100 is a smartphone or a tablet PC, the display unit 110 is a display. The audio information output unit 120 is a speaker installed in the smartphone or tablet PC, or an earphone connected to the smartphone or tablet PC.

[0022] Hereinafter, each functional configuration of the information processing device 200 will be described.

[0023] The biometric information acquisition unit 210 as an acquisition unit acquires biometric information (first biometric information) based on muscle movements due to the user's speaking action detected by the detection device 100. Here, the detected biometric information is at least one of myoelectric potential information, acceleration information, and tactile information due to the user's speaking action. However, depending on the specifications of the inference model used by the conversion unit 220 to convert the biometric information, audio information and video information may be further acquired in addition to the above biometric information.

[0024] Furthermore, the biological information may be two or more types of information acquired from one location placed on the user. Details will be described later, but for example, the biological information detection unit 101 may be configured to include a myoelectric potential sensor and an acceleration sensor, and the biological information acquisition unit 210 acquires both myoelectric potential information and acceleration information. Both sensors may be placed at multiple locations relative to the user. When placed at multiple locations, each location may include multiple sensors, or each location may be configured with only a predetermined sensor.

[0025] 2, the conversion unit 220 is configured to include a signal acquisition unit 221, a pre-processing unit 222, an inference unit 223, and a post-processing unit 224. The signal acquisition unit 221 acquires biometric information acquired by the biometric information acquisition unit 210. The pre-processing unit 222 performs pre-processing of the signal. The inference unit 223 uses the signal pre-processed by the pre-processing unit 222 as input and infers character information or audio information corresponding to the biometric information using an inference model stored in the storage device 250. The post-processing unit 224 performs post-processing of the character information or audio information inferred by the inference unit 223.

[0026] 3, the inference unit 223 is composed of a first inference unit 223A and a second inference unit 223B. The first inference unit 223A receives a signal preprocessed by the preprocessing unit 222 as input, and infers pseudo information corresponding to the signal using a first inference model stored in the storage device 250. The second inference unit 223B receives pseudo information inferred and output by the first inference unit 223A as input, and infers text information or audio information using a second inference model stored in the storage device 250.

[0027] The output unit 230 outputs the character information or audio information converted by the conversion unit to the display unit 110 or the audio information output unit 120.

[0028] [Control Flow] An example of a control flow in the information processing device 200 will be described below with reference to FIG.

[0029] (Step S401: Obtaining Biometric Information) In step S401, the signal acquiring unit 221 in the conversion unit 220 acquires bio-information (first bio-information) such as a myoelectric potential signal transmitted from the detection device 100. The acquired bio-information is bio-information based on muscle movements caused by the user's speaking action, such as myoelectric potential information and acceleration information. After transmitting the bio-information to the pre-processing unit 222, the signal acquiring unit 221 proceeds to the next step.

[0030] (Step S402: Preprocessing) In step S402, the preprocessing unit 222 in the conversion unit 220 performs various preprocessing processes on the biometric information. If the biometric information is one-dimensional time-series data such as a myoelectric signal or an acceleration signal, the preprocessing processes include noise reduction processes such as elimination of outliers and signal smoothing, feature extraction processes such as octave analysis and spectrogram conversion, and standardization. If the biometric information is three-dimensional data such as a video of the user's face, the preprocessing processes include the above-mentioned noise reduction processes, background removal, grayscale conversion, image size conversion, and feature point extraction around the lips.

[0031] The preprocessing unit 222 preprocesses the biometric information, transmits the preprocessed data to the inference unit 223, and then proceeds to the next step. If preprocessing is not necessary, step S402 is not performed.

[0032] (Step S403: Inference) In step S403, the inference unit 223 uses the data transmitted from the preprocessing unit 222 as input to the inference model and infers the corresponding character information or audio information.

[0033] Here, the inference unit 223 uses a trained inference model with an architecture configured by a neural network, acquired from the storage device 240. The trained inference model is configured from a network of a known inference model. For example, an inference model having a network configuration of a convolutional neural network (CNN), a type of deep learning, a recurrent neural network (RNN), or a long short-term memory (LSTM) is used. An inference model is generated by applying a learning process to an inference model having such a network configuration. Note that the inference model used by the inference unit 223 for inference may be a model derived from a CNN, an RNN, or an LSTM, or may use other machine learning techniques such as a support vector machine, a logistic regression, or a random forest, or may use a rule-based method.

[0034] In the inference unit 223, examples of character information that the inference model outputs as an inference result include the following: syllables such as "a," "i," "u," "e," "o," "ka," and "ki," strings of syllables such as "hello" and "good night," and character strings including kanji, numbers, and katakana such as "call extension 5" and "turn off the lights."

[0035] The inference result by the inference unit 223 may be a phoneme string written in alphabetical order such as "katazukete" or "hajimemashite," or a mora string such as "de N maaku" or "by u cl feniiku." The language is not limited to Japanese.

[0036] The inference unit 223 transmits the character information that is the inference result to the output unit 230, and then proceeds to the next step. The conversion unit 220 may further convert the character information output by the inference unit 223. For example, the conversion unit 220 may convert "katazukete" to "tidy up" or convert "by u cl feniiku" to "go to the buffet," thereby making it easier for the user to understand the character information.

[0037] The conversion unit 220 may convert the information into audio information instead of text information. In the inference unit 223, examples of audio information output by the inference model as an inference result include sound pressure waveforms of monaural or stereo audio, and spectrograms. The audio information may be that of the user who made the speech, or may be audio information of another person. After sending the text information or audio information to the post-processing unit 224, the inference unit 223 proceeds to the next step.

[0038] (Step S404: Post-processing) In step S404, the post-processing unit 224 performs various post-processing operations on the text information or audio information transmitted by the inference unit 223. If the information transmitted by the inference unit 223 is text information, the post-processing may include further conversion of the text information. For example, if the text information is a phoneme string written in alphabetical notation, such as "hajimemashite," it may be converted into "nice to meet you" or "nice to meet you." In the above case, it may also be converted into another language, such as "How do you do?" Furthermore, if there is a possibility that the text information is mis-estimated, such as "hasimemashite," it may be corrected to "hajimemashite." For example, "Konnichiwa" may be corrected to "Konnichiwa." Furthermore, it may be converted into non-text information, such as emoticons or animations, that correspond to the text information, or the text information may be converted into audio information. If the information transmitted by the inference unit 223 is audio information, the post-processing may include further conversion of the audio information. For example, non-audio information and noise may be removed. It may also be converted into the voice tone of the user who made the speech or another person. Also, audio information may be converted into text information.

[0039] If post-processing is not required, step S404 may not be performed. Also, some or all of the processing included in the post-processing may be included in the inference model and performed by the inference unit 223.

[0040] (Step S405: Information output) In step S405, the output unit 230 outputs the text information or audio information acquired from the conversion unit 220 to an external device. In the case of text information, it is sent to the display unit 110 such as a monitor, and in the case of audio information, it is sent to the audio information output unit 120 such as a speaker or bone conduction earphones. The destination may be determined based on connection information with the external device, user setting information for the information processing device 200, etc.

[0041] With this configuration, the information processing system 1000 can convert biometric information based on muscle movements caused by a user's speaking actions into text information or audio information with high accuracy. Specifically, by performing inference using a first inference model that can input preprocessed biometric information and a second inference model that can input the output of the first inference model, the biometric information can be converted into more accurate text information or audio information.

[0042] [Learning Flow] The following describes the learning unit 240 for learning the inference model used in the conversion unit 220, and the storage device 250 for storing the learned inference model. Note that the learning unit 240 does not necessarily have to be provided by the same device or the same entity, and the processing of the learning unit 240 may be substituted by storing a learned inference model generated by a different device or a different entity in the storage device 250.

[0043] In the information processing device 200, the learning unit 240 generates a trained inference model by learning the inference model using training data. FIG. 5 shows an example of the configuration of the learning unit 240. The learning unit 240 includes a training data acquisition unit 241, an inference model acquisition unit 242, and a model learning unit 243. The training data acquisition unit 241 acquires training data from the storage device 250. The inference model acquisition unit 242 acquires the inference model to be learned from the storage device 250. The model learning unit 243 learns the inference model using the inference model and training data. The model architecture and weight parameters are collectively referred to as the inference model.

[0044] Here, the storage device 250 stores biometric information detected by the detection device 100 and text information or audio information corresponding to the biometric information. The storage device 250 also stores biometric information other than the biometric information detected by the detection device 100 and text information or audio information corresponding to the biometric information. Furthermore, the storage device 250 stores the architecture and weight parameters of a first inference model, the architecture and weight parameters of a second inference model, and parameters required for executing pre-processing and post-processing.

[0045] The learning flow of the inference model by the learning unit 240 will be explained below with reference to FIG.

[0046] (Step S601: Obtaining training data) In step S601, the teacher data acquisition unit 241 acquires, from the storage device 250, biometric information based on muscle movements during speech (second biometric information) and pseudo-information corresponding to the biometric information, and performs preprocessing on the biometric information. Here, the pseudo-information refers to text information or audio information inferred from the biometric information by an inference model. The teacher data acquisition unit 241 also acquires correct answer data by converting the pseudo-information into information suitable for learning as needed. For example, if one piece of pseudo-information is "Turn off the lights," it is converted into "Turn off the lights" or "raitooke shite." The teacher data acquisition unit 241 transmits the preprocessed biometric information data and correct answer data as teacher data to the model learning unit 243. The inference model acquisition unit 242 then acquires a first inference model from the storage device 250, transmits it to the model learning unit 243, and proceeds to the next step.

[0047] (Step S602: Learning the first inference model) The model learning unit 243 uses the teacher data transmitted from the teacher data acquisition unit 241 and the first inference model transmitted from the inference model acquisition unit 242 to perform a learning process on the first inference model and learn the weight parameters of the first inference model.

[0048] The first inference model learned by the model learning unit 243 is also used by the inference unit 223 of the conversion unit 220. The optimization function applied by the model learning unit 243 when learning the first inference model may be SGD (Stochastic Gradient Descent), Adam (Adaptive Moment Estimation), or the like. Furthermore, Cross-entropy loss or CTC (Connectionist Temporal Classification) loss may be used as the loss function. Note that the optimization function and loss function are not limited to these, and various functions may be used. Furthermore, the initial values ​​of the weight parameters may be random, or pre-learned values ​​may be used.

[0049] (Step S603: Storing the first inference model) In step S603, when the model learning unit 243 completes the learning process of the first inference model, it stores information about the learned first inference model in the storage device 250. In addition, the model learning unit 243 stores pseudo information obtained as a result of estimation by the learned first inference model using preprocessed biometric information as input.

[0050] (Step S604: Obtaining training data) In step S604, the teacher data acquisition unit 241 acquires the pseudo information and text information or audio information corresponding to the pseudo information from the storage device 250. The teacher data acquisition unit 241 may also acquire the pseudo information stored in step S603 and text information or audio information corresponding to the pseudo information from the storage device 250. The inference model acquisition unit 242 acquires a second inference model from the storage device 250. After the teacher data acquisition unit 241 transmits the pseudo information and text information or audio information corresponding to the pseudo information to the model learning unit 243 and the inference model acquisition unit 242 transmits the second inference model to the model learning unit 243, the process proceeds to the next step.

[0051] (Step S605: Learning the second inference model) The model learning unit 243 uses the training data transmitted from the training data acquisition unit 241 and the second inference model transmitted from the inference model acquisition unit 242 to perform a learning process on the second inference model and learn weight parameters of the second inference model. The second inference model learned by the model learning unit 243 is also used by the inference unit 223 of the conversion unit 220. The optimization function applied by the model learning unit 243 when learning the second inference model is, for example, Stochastic Gradient Descent (SGD) or Adaptive Moment Estimation (Adam). Furthermore, the loss function used is, for example, Cross-entropy loss or Connectionist Temporal Classification (CTC) loss. Note that the optimization function and loss function are not limited to these, and various functions can be used. Furthermore, the initial values ​​of the weight parameters may be random, or pre-learned values ​​may be used.

[0052] (Step S606: Storing the first inference model) In step S606, when the model learning unit 243 completes the learning process of the second inference model, it stores information about the learned second inference model in the storage device 250, and the learning of the inference model is completed.

[0053] The learning of the first inference model and the second inference model may be unsupervised learning without correct answer data or semi-supervised learning without correct answer data for some data. The learning unit 240 may be implemented as a function on a personal computer or may be configured on the cloud. Only some of the functions, such as the model learning unit 243, may be configured on the cloud. [Example]

[0054] An information processing system according to a first embodiment will be described with reference to Fig. 7. Fig. 7 shows a schematic diagram of a detection device 700 corresponding to the detection device 100 constituting the information conversion system 1. The detection device 700 is composed of a biometric information detection unit 701 and a transmission unit 702. The biometric information detection unit 701 is composed of a myoelectric potential sensor and an acceleration / angular velocity sensor, and can detect movements around the mouth by attaching it to the cheek, chin, neck, or temple of the user with tape or adhesive.

[0055] The myoelectric potential sensor can measure the myoelectric potential of the muscles at the attachment site, and the accelerometer / angular velocity sensor can measure movements around the mouth as three-axis translational acceleration and three-axis angular velocity.

[0056] The sampling rate of the myoelectric potential sensor and the acceleration / angular velocity sensor is set to 2 kHz.

[0057] The transmitter 702 is housed in a housing on the neckband. The myoelectric potential signal and acceleration / angular velocity signal detected by the biometric information detector 701 are sent to the transmitter 702 via a wired connection and then wirelessly transmitted to the information processing device 200. Specific examples include wireless LAN communication such as WiFi and short-range wireless communication such as Bluetooth (registered trademark). The detection device 700 is powered by a battery (not shown) in the neckband. While FIG. 7 illustrates a configuration in which biometric information is detected at two points on the face, biometric information may be detected at one point or three or more points. A different sensor may be provided at each point of contact with the user. The detected biometric information may be only myoelectric potential signals or only three-axis translational acceleration. The biometric information is not limited to myoelectric potential, translational acceleration, and angular velocity, and various other types of information may be detected, such as distortion, tactile sensation, magnetism, and ultrasound.

[0058] In this embodiment, the information processing device 200 is, for example, a smartphone, the display unit 110 is a display of the smartphone, and the audio information output unit 120 is a speaker or a wired or wireless earphone. The audio information output unit 120 may also be a component of the detection device 200, for example, the audio information output unit 120 may be a bone conduction earphone (not shown).

[0059] Next, the processing of the learning unit 240 included in the information processing device 200 will be described.

[0060] First, a large number of myoelectric potential signals, acceleration signals, and angular velocity signals measured by the detection device 100, as well as text information and audio information corresponding to each signal, are stored in the storage device 250 in advance. The architecture of a first inference model is stored in the storage device 250. Weight parameters of the first inference model may also be stored. The weight parameters of the first inference model may be random values ​​or values ​​learned in advance using arbitrary biometric information and text information. The architecture of a second inference model is also stored in the storage device 250. Weight parameters of the second inference model may also be stored. The weight parameters of the second inference model may be random values ​​or values ​​learned in advance using arbitrary information as training data.

[0061] [study] The pre-processing unit 222 performs a standardization process to remove noise from each signal, which is the biometric information (second biometric information) acquired from the storage device 250. The character information acquired from the storage device 250 is converted into a string of characters, such as "o N gakusaisee," but hiragana characters may also be used. The learning unit 240 first learns the weight parameters of the first inference model acquired from the storage device 250 using the signal and string of characters (pseudo information) converted by the pre-processing unit 222 as training data.

[0062] The model architecture includes multiple convolutional layers, BiGRU (Bidirectional Gated Recurrent Unit) layers, and linear combination layers. The optimization function is AdamW (Adaptive Moment Estimation with Weight Decay), and the loss function is CTC loss.

[0063] The model architecture, optimization function, and loss function are not limited to these, and various other functions are possible. The trained model obtained by the training unit 240 is stored in the storage device 250. In addition, the character information output by the trained model using the signal converted by the preprocessing unit 222 as input is stored in the storage device 250.

[0064] Next, the weight parameters of the second inference model are learned using the character information (pseudo information) acquired from the storage device 250 and the audio information corresponding to the character information as training data. The weight parameters of the second inference model are values ​​learned using arbitrary character information (pseudo information) and the audio information corresponding to the character information as training data. When the amount of data measured by the detection device 100 is small, the estimation accuracy can be improved by learning the weight parameters in advance using a large amount of available arbitrary character information and the corresponding audio information as training data.

[0065] As with the first inference model, various functions can be considered for the model architecture, optimization function, and loss function. The trained model obtained by the training unit 240 is stored in the storage device 250. Note that the training unit 240 may be configured as a separate device, or its function may be replaced by acquiring a trained inference model.

[0066] [conversion] Next, the processing of the conversion unit 220 constituting the information processing device 200 will be described with reference to Fig. 8. The biometric information (first biometric information) transmitted from the biometric information acquisition unit 210 and acquired by the signal acquisition unit 221 is converted into a waveform signal by the pre-processing unit 222. Next, the waveform signal is input to the trained first inference model acquired from the storage device 250, and character information is output as pseudo information corresponding to the biometric information.

[0067] Furthermore, the inference unit 223 inputs the output character information into a trained second inference model acquired from the storage device 250, and outputs corresponding audio information. Next, the post-processing unit 224 removes non-audio components and noise from the audio information output by the inference unit 223, and transmits the audio information to the output unit 230. Finally, the output unit 230 transmits the audio information to the audio information output unit 120, and audio is output from earphones as an example of the audio information output unit 120.

[0068] In this embodiment, an inference model combining the first inference model and the second inference model may be processed as a single inference model. That is, an inference model may be created in which the output of the first inference model is input to the second inference model, and a waveform signal may be input to this inference model to output audio information.

[0069] <Variation 1> An information processing system according to a first modification of the first embodiment will be described below, with overlapping descriptions omitted where appropriate.

[0070] In this modification, the teacher data acquisition unit 231 of the learning unit 240 and the preprocessing unit 222 of the conversion unit 220 convert signal data, which is biological information, into a spectrogram. Specifically, the signal is converted into a feature matrix in which multiple column vectors obtained by short-time Fourier transform are arranged. The converted feature matrix is ​​sent to the model learning unit 243 or the inference unit 223. Note that the feature matrix may be converted into various feature matrices other than a spectrogram, such as a mel spectrogram. In the first embodiment, the input to the first inference model is a waveform signal, i.e., one-dimensional data, whereas in this modification, it is a feature matrix, i.e., two-dimensional data. For this reason, the first layer of the architecture of the first inference model is changed from a one-dimensional convolutional layer to a two-dimensional convolutional layer. Furthermore, the architecture of layers other than the first layer may be changed.

[0071] <Variation 2> An information processing system according to a second modification of the first embodiment will be described. Note that descriptions of overlapping parts will be omitted where appropriate. In this modification, the detection device 900 has a biometric information detection unit 901 that is ear-hooked as shown in FIG. The biometric information detection unit 901, such as a myoelectric potential sensor or an acceleration / angular velocity sensor, is attached to a support made of resin, for example, and is in contact with the skin of the lower jaw, cheek, etc. with appropriate pressure. The biometric information detected by the biometric information detection unit 901 is sent via a wire to a transceiver unit 902 in the neckband, and then sent wirelessly to the information processing device 200. The detection device 900 is powered by a battery (not shown) in the neckband.

[0072] Furthermore, when the information processing device 200 outputs audio information, the transmitting / receiving unit 902 receives the audio information, and the audio information can be reproduced by an audio information output unit 903 such as a speaker or a bone conduction microphone in the ear hook portion.

[0073] <Variation 3> An information processing system according to a third modification of the first embodiment will be described.

[0074] In the above embodiment, the conversion unit 220 converts biometric information into character information using a trained model trained by the training unit 240.

[0075] In this modification, the learning unit 240 further performs additional learning of the inference model using training data that pairs correct answer data, which is text information based on the user's voice information, with learning data, which is waveform signals acquired from biometric information. Here, the conversion from voice information to text information may be realized by any known technology.

[0076] This configuration enables learning according to the characteristics of the user, and enables conversion to character information with even higher accuracy than at the time of distribution.

[0077] <Variation 4> A fourth modification of the first embodiment will be described. Note that the description of overlapping parts will be omitted as appropriate.

[0078] The processing of the learning unit 240 constituting the information processing device 200 in this modification will be described.

[0079] [study] The preprocessing unit 222 converts the biosignal (second bioinformation) acquired from the storage device 250 into a spectrogram. It also converts the audio information corresponding to the biosignal into a spectrogram. It also converts the character information corresponding to the biosignal into a phoneme string or a character string such as hiragana or kanji.

[0080] The learning unit 240 first learns weight parameters of the first inference model acquired from the storage device 250 using the spectrogram obtained by converting the biosignal in the preprocessing unit 222 and the spectrogram of the audio information as training data. As in the third embodiment, various functions can be considered for the model architecture, optimization function, and loss function. The weight parameters of the trained first inference model are stored in the storage device 250. In addition, the spectrogram that is the output result of the trained first inference model that uses the spectrogram obtained by converting the biosignal in the preprocessing unit 222 as input is also stored in the storage device 250. Here, the spectrogram that is the output result of the first inference model includes data that is different from the data used as training data. In other words, the spectrogram that is the output result of the first inference model includes erroneously estimated data.

[0081] Next, the learning unit 240 learns the weight parameters of the second inference model acquired from the storage device 250. During learning, the spectrogram data, which is the output result of the trained first inference model that uses as input a spectrogram obtained by converting the biological signal in the preprocessing unit 222, and character strings corresponding to the biological information, are used as training data. The weight parameters of the second inference model after learning are stored in the storage device 250. As with the third embodiment, various functions are possible for the model architecture, optimization function, and loss function. Furthermore, the weight parameters of the second inference model may use values ​​that have been trained in advance using a large amount of arbitrary spectrograms and corresponding character strings as training data. By training the inference model as described above, even if the amount of data measured by the detection device 100 is small, the weight parameters can be trained in advance using a large amount of arbitrary character information that is available as training data, thereby improving estimation accuracy. [Example]

[0082] An information processing system according to a second embodiment will be described below, and overlapping descriptions will be omitted as appropriate.

[0083] The processing of the learning unit 240 constituting the information processing device 200 in this embodiment will be described.

[0084] [study] The preprocessing unit 222 performs noise removal and standardization on the biometric signal as biometric information (second biometric information) acquired from the storage device 250. Also, character information acquired from the storage device 250 corresponding to the biometric information is converted into a phoneme string or a character string such as hiragana or kanji.

[0085] The learning unit 240 first learns the weight parameters of the first inference model acquired from the storage device 250 using the signal and character string (pseudo information) converted by the preprocessing unit 222 as training data. As in the first embodiment, various functions can be considered for the model architecture, optimization function, and loss function. The weight parameters of the trained first inference model are stored in the storage device 250. In addition, the character string that is the output result of the trained first inference model, which uses the signal converted by the preprocessing unit 222 as input, is also stored in the storage device 250. Here, the character string that is the output result of the first inference model includes both the same and different character strings from the character strings used as training data. In other words, the character string that is the output result of the first inference model includes erroneously estimated character strings.

[0086] Next, the learning unit 240 learns weight parameters of the second inference model acquired from the storage device 250 using the character string that is the output result of the trained first inference model, which inputs the signal converted by the preprocessing unit 222, and the character string corresponding to the biometric information as training data. The weight parameters of the trained second inference model are stored in the storage device 250. As with the first embodiment, various functions are possible for the model architecture, optimization function, and loss function. Furthermore, the weight parameters of the second inference model may be values ​​previously trained using a large number of arbitrary incorrect character strings and correct character strings as training data. Here, the character representations of the incorrect character strings and the correct character strings may be the same or different. For example, the incorrect character string may be "wata sh iwa" and the correct character string may be "watashi wa" (I am), and various combinations are possible. By training the inference model in the above manner, even when the data measured by the detection device 100 is small, the weight parameters can be previously trained using a large amount of arbitrary character information available as training data, thereby improving estimation accuracy.

[0087] [conversion] Next, the processing of the conversion unit 220 constituting the information processing device 200 will be described with reference to Fig. 10. The biometric information (first biometric information) transmitted from the biometric information acquisition unit 210 is acquired by the signal acquisition unit 221 and converted into a waveform signal by the pre-processing unit 222. Next, the waveform signal is input to a trained first inference model acquired from the storage device 250, and character information is output as pseudo information corresponding to the biometric information. Furthermore, the inference unit 223 inputs the output character information to a trained second inference model acquired from the storage device 250, and the corresponding character information is output. Next, the post-processing unit 224 transmits the character information output by the inference unit 223 to the output unit 230. Finally, the output unit 230 transmits the character information to the display unit 110, and the character information is displayed on a monitor as an example of the display unit 110.

[0088] In this embodiment, an inference model combining the first inference model and the second inference model may be processed as a single inference model. That is, an inference model may be created in which the output of the first inference model is input to the second inference model, and a waveform signal may be input to this inference model to output text information. [Example]

[0089] An information processing system according to a third embodiment will be described below. Note that overlapping descriptions will be omitted as appropriate.

[0090] The processing of the learning unit 240 constituting the information processing device 200 in this embodiment will be described.

[0091] [study] The preprocessing unit 222 performs noise removal and standardization on the biosignal as bioinformation (second bioinformation) acquired from the storage device 250. Also, audio information corresponding to the biosignal is converted into waveform data. Also, character information corresponding to the biosignal is converted into a phoneme string or a character string such as hiragana or kanji.

[0092] The learning unit 240 first learns the weight parameters of the first inference model acquired from the storage device 250 using the signal and waveform data (pseudo information) converted by the pre-processing unit 222 as training data. As in the first embodiment, various functions can be considered for the model architecture, optimization function, and loss function. The weight parameters of the trained first inference model are stored in the storage device 250. In addition, the voice signal that is the output result of the trained first inference model that uses the signal converted by the pre-processing unit 222 as input is also stored in the storage device 250. Here, the output result of the first inference model includes both the same data as the data used as training data and data that is different. In other words, the voice signal that is the output result of the first inference model includes erroneously estimated data.

[0093] Next, the learning unit 240 learns the weight parameters of the second inference model acquired from the storage device 250 using the waveform signal (pseudo information) that is the output result of the trained first inference model that uses the signal converted by the preprocessing unit 222 as input, and the character string corresponding to the biometric information as training data. The weight parameters of the second inference model after training are stored in the storage device 250. As with the first embodiment, various functions are possible for the model architecture, optimization function, and loss function. Furthermore, the weight parameters of the second inference model may use values ​​that have been trained in advance using a large amount of arbitrary erroneous waveform signals and correct character strings as training data. By training the inference model as described above, even if the data measured by the detection device 100 is small, the weight parameters can be trained in advance using a large amount of arbitrary character information that is available as training data, thereby improving estimation accuracy.

[0094] [conversion] Next, the processing of the conversion unit 220 constituting the information processing device 200 will be described with reference to FIG. 11. The biometric information (first biometric information) transmitted from the biometric information acquisition unit 210 is acquired by the signal acquisition unit 221 and converted into a waveform signal by the pre-processing unit 222. Next, the waveform signal is input to a trained first inference model acquired from the storage device 250, and audio information is output as pseudo information corresponding to the biometric information. Furthermore, the inference unit 223 inputs the output audio information to a trained second inference model acquired from the storage device 250, and outputs corresponding text information. Next, the post-processing unit 224 converts the text information output by the inference unit 223 into a character string that is easily recognizable by the user, for example, Chinese characters, and transmits it to the output unit 230. Finally, the output unit 230 transmits the text information to the display unit 110, and the text information is displayed on a monitor as an example of the display unit 110.

[0095] In this embodiment, an inference model combining the first inference model and the second inference model may be processed as a single inference model. That is, an inference model may be created in which the output of the first inference model is input to the second inference model, and a waveform signal may be input to this inference model to output text information. [Example]

[0096] An information processing system according to a fourth embodiment will be described below. Note that overlapping descriptions will be omitted as appropriate.

[0097] The processing of the learning unit 240 constituting the information processing device 200 in this embodiment will be described.

[0098] [study] The preprocessing unit 222 performs noise removal and standardization on the biosignal obtained from the storage device 250. Furthermore, audio information corresponding to the biosignal is converted into an audio waveform.

[0099] The learning unit 240 first learns the weight parameters of the first inference model acquired from the storage device 250 using the signal (second biometric information) and voice waveform (pseudo information) converted by the preprocessing unit 222 as training data. As with the first embodiment, various functions can be considered for the model architecture, optimization function, and loss function. The weight parameters of the trained first inference model are stored in the storage device 250. The voice waveform, which is the output result of the trained first inference model using the signal converted by the preprocessing unit 222 as input, is also stored in the storage device 250. Here, the output result of the first inference model includes data that differs from the data used as training data. In other words, the voice waveform, which is the output result of the first inference model, includes erroneously estimated data.

[0100] Next, the learning unit 240 learns weight parameters of the second inference model acquired from the storage device 250 using the speech waveform, which is the output result of the trained first inference model that uses the signal converted by the preprocessing unit 222 as input, and the speech waveform corresponding to the biometric information as training data. The weight parameters of the second inference model after training are stored in the storage device 250. As with the first embodiment, various functions are possible for the model architecture, optimization function, and loss function. Furthermore, the weight parameters of the second inference model may use values ​​that have been trained in advance using a large amount of arbitrary erroneous speech waveforms and correct speech waveforms as training data. By training the inference model as described above, even if the amount of data measured by the detection device 100 is small, the weight parameters can be trained in advance using a large amount of arbitrary available speech waveforms as training data, thereby improving estimation accuracy.

[0101] [conversion] Next, the processing of the conversion unit 220 constituting the information processing device 200 will be described. The preprocessing unit 222 performs preprocessing on the biometric information (first biometric information) transmitted from the biometric information acquisition unit 210 and acquired by the signal acquisition unit 221. Next, the preprocessed signal is input to a trained first inference model acquired from the storage device 250, and a voice waveform is output as pseudo information corresponding to the biometric information. Furthermore, the inference unit 223 inputs the output voice waveform to a trained second inference model acquired from the storage device 250, and outputs a corresponding voice waveform. Next, the postprocessing unit 224 removes non-voice components and noise from the voice waveform output by the inference unit 223, and transmits the voice waveform to the output unit 230. Finally, the output unit 230 transmits the voice information to the voice information output unit 120, and voice is output from earphones, for example.

[0102] In this embodiment, an inference model combining the first inference model and the second inference model may be processed as a single inference model. That is, an inference model may be created in which the output of the first inference model is input to the second inference model, and a waveform signal may be input to this inference model to output text information.

[0103] [Second embodiment] An information processing system according to a second embodiment will be described below, with explanations of parts that overlap with those of the first embodiment being omitted where appropriate.

[0104] The information processing system differs from the first embodiment in that the inference unit 223 constituting the conversion unit 220 is an inference unit. As shown in FIG. 12, the inference unit 223 of the information processing system is composed of a first inference unit 1223A, a second inference unit 1223B, and a third inference unit 1223C. The first inference unit 1223A receives a signal preprocessed by the preprocessing unit 222 as input and infers first pseudo information corresponding to the signal using a first inference model stored in the storage device 250. The second inference unit 1223B receives the first pseudo information as input and infers second pseudo information corresponding to the first pseudo information using a second inference model stored in the storage device 250. The third inference unit 1223C receives the second pseudo information as input and infers text information or audio information corresponding to the signal using a third inference model stored in the storage device 250.

[0105] [Learning Flow] Below, we will explain the learning unit 240 for learning the inference model used in the conversion unit 220 and the storage device 250 for storing the learned inference model.

[0106] In the information processing device 200, the storage device 250 stores biometric information detected by the detection device 100 and text information or audio information corresponding to the biometric information. The storage device 250 also stores biometric information other than the biometric information detected by the detection device 100 and text information or audio information corresponding to the biometric information. The storage device 250 also stores the architecture and weight parameters of a first inference model, the architecture and weight parameters of a second inference model, the architecture and weight parameters of a third inference model, and parameters required for executing pre-processing and post-processing.

[0107] Using FIG. 13, the learning flow of the inference model by the learning unit 240 will be explained.

[0108] (Step S1301: Obtaining training data) In step S1301, the teacher data acquisition unit 241 acquires, from the storage device 250, biometric information based on muscle movements during speech (second biometric information) and text information or audio information as pseudo-information corresponding to the biometric information, and performs preprocessing on the biometric information. The teacher data acquisition unit 241 also acquires correct answer data by converting the text information or audio information into information suitable for learning as needed. For example, if one piece of text information is "Turn off the lights," it converts it into "Rite wo keshite" or "raitooke sh ite." The teacher data acquisition unit 241 transmits the preprocessed biometric information data and correct answer data as teacher data to the model learning unit 243. The inference model acquisition unit 242 acquires a first inference model from the storage device 250, transmits it to the model learning unit 243, and proceeds to the next step.

[0109] (Step S1302: Learning the first inference model) The model learning unit 243 uses the teacher data transmitted from the teacher data acquisition unit 241 and the first inference model transmitted from the inference model acquisition unit 242 to perform a learning process on the first inference model and learn the weight parameters of the first inference model.

[0110] The first inference model learned by the model learning unit 243 is also used by the inference unit 223 of the conversion unit 220. The optimization function applied by the model learning unit 243 when learning the first inference model is, for example, SGD (Stochastic Gradient Descent) or Adam (Adaptive Moment Estimation). Furthermore, the loss function used is, for example, Cross-entropy loss or CTC (Connectionist Temporal Classification) loss. Note that the optimization function and loss function are not limited to these, and various functions can be used. Furthermore, the initial values ​​of the weight parameters may be random, or pre-learned values ​​may be used.

[0111] (Step S1303: Storing the first inference model) In step S1303, when the model learning unit 243 completes the learning process of the first inference model, it stores information about the learned first inference model in the storage device 250. In addition, the model learning unit 243 stores first pseudo information obtained as a result of estimation by the learned first inference model using preprocessed biometric information as input.

[0112] (Step S1304: Obtaining training data) In step S1304, the teacher data acquisition unit 241 acquires the first pseudo information and text information or audio information corresponding to the pseudo information from the storage device 250. The teacher data acquisition unit 241 may also acquire the pseudo information stored in step S1303 and text information or audio information corresponding to the pseudo information from the storage device 250. The inference model acquisition unit 242 acquires a second inference model from the storage device 250. After the teacher data acquisition unit 241 transmits the pseudo information and the text information or audio information corresponding to the other pseudo information to the model learning unit 243 and the inference model acquisition unit 242 transmits the second inference model to the model learning unit 243, the process proceeds to the next step.

[0113] (Step S1305: Learning the second inference model) The model learning unit 243 uses the training data transmitted from the training data acquisition unit 241 and the second inference model transmitted from the inference model acquisition unit 242 to perform a learning process on the second inference model and learn weight parameters of the second inference model. The second inference model learned by the model learning unit 243 is also used by the inference unit 223 of the conversion unit 220. The optimization function applied by the model learning unit 243 when learning the second inference model is, for example, Stochastic Gradient Descent (SGD) or Adaptive Moment Estimation (Adam). Furthermore, the loss function used is, for example, Cross-entropy loss or Connectionist Temporal Classification (CTC) loss. Note that the optimization function and loss function are not limited to these, and various functions can be used. Furthermore, the initial values ​​of the weight parameters may be random, or pre-learned values ​​may be used.

[0114] (Step S1306: Storing the second inference model) In step S1306, when the model learning unit 243 completes the learning process for the second inference model, it stores information about the learned second inference model in the storage device 250. In addition, the model learning unit 243 stores second pseudo information obtained as a result of estimation by the learned second inference model using the first pseudo information as input.

[0115] (Step S1307: Obtaining training data) In step S1307, the teacher data acquisition unit 241 acquires the second pseudo information and text information or audio information corresponding to the pseudo information from the storage device 250. The teacher data acquisition unit 241 may also acquire the pseudo information stored in step S1306 and text information or audio information corresponding to the pseudo information from the storage device 250. The inference model acquisition unit 242 acquires a third inference model from the storage device 250. After the teacher data acquisition unit 241 transmits the pseudo information and the text information or audio information corresponding to the other pseudo information to the model learning unit 243 and the inference model acquisition unit 242 transmits the third inference model to the model learning unit 243, the process proceeds to the next step.

[0116] (Step S1308: Learning the third inference model) The model learning unit 243 performs a learning process on the third inference model using the teacher data transmitted from the teacher data acquisition unit 241 and the third inference model transmitted from the inference model acquisition unit 242, and learns weight parameters of the third inference model. The third inference model learned by the model learning unit 243 is also used by the inference unit 223 of the conversion unit 220. The optimization function applied by the model learning unit 243 when learning the third inference model is, for example, SGD (Stochastic Gradient Descent) or Adam (Adaptive Moment Estimation). Furthermore, the loss function used is, for example, Cross-entropy loss or CTC (Connectionist Temporal Classification) loss. Note that the optimization function and loss function are not limited to these, and various functions can be used. Furthermore, the initial values ​​of the weight parameters may be random, or pre-learned values ​​may be used.

[0117] (Step S1309: Storing the third inference model) In step S1309, when the model learning unit 243 completes the learning process of the third inference model, it stores information about the learned third inference model in the storage device 250, and the learning of the inference model is completed.

[0118] The learning of the first inference model, the second inference model, and the third inference model may be unsupervised learning without correct answer data or semi-supervised learning without correct answer data for some data. The learning unit 240 may be implemented as a function on a personal computer or may be configured on the cloud. Only some of the functions, such as the model learning unit 243, may be configured on the cloud. [Example]

[0119] An information processing system according to a fifth embodiment will be described. Note that descriptions of parts that overlap with those of the first embodiment will be omitted as appropriate. The processing of the learning unit 240 included in the information processing device 200 in this embodiment will be described.

[0120] [study] The preprocessing unit 222 performs noise removal and standardization on the biometric signal obtained from the storage device 250. Furthermore, character information obtained from the storage device 250 and corresponding to the biometric information is converted into a phoneme string or a character string such as hiragana or kanji.

[0121] The learning unit 240 first learns weight parameters of the first inference model acquired from the storage device 250 using the signal (second biometric information) and character strings (pseudo information) converted by the preprocessing unit 222 as training data. As with the first embodiment, various functions can be considered for the model architecture, optimization function, and loss function. The weight parameters of the trained first inference model are stored in the storage device 250. The storage device 250 also stores first pseudo information, which is the output result of the trained first inference model that uses the signal converted by the preprocessing unit 222 as input. Here, the first pseudo information, which is the output result of the first inference model, i.e., the character strings, includes both the same and different character strings as those used as training data. In other words, the first pseudo information, which is the output result of the first inference model, includes erroneously estimated character strings.

[0122] Next, the learning unit 240 learns the weight parameters of the second inference model acquired from the storage device 250. During learning, first pseudo information, which is the output result of the trained first inference model using the signal converted by the preprocessing unit 222 as input, and speech waveforms, which are speech information corresponding to the biometric information, are used as training data. Here, the weight parameters of the second inference model may be initially set to values ​​trained in advance using a large amount of arbitrary character strings and correct speech waveforms as training data. The weight parameters of the trained second inference model are stored in the storage device 250. As in the first embodiment, various functions are possible for the model architecture, optimization function, and loss function. In addition, second pseudo information, which is the output result of the trained second inference model using the first pseudo information as input, is also stored in the storage device 250. Here, the second pseudo information, i.e., the speech waveform, which is the output result of the second inference model, includes speech waveforms that differ from those used as training data. In other words, the second pseudo information, which is the output result of the second inference model, includes erroneously estimated speech waveforms.

[0123] Next, the learning unit 240 learns the weight parameters of the third inference model obtained from the storage device 250 using the second pseudo information and the character string converted by the preprocessing unit 222 as training data. Here, the weight parameters of the third inference model may use values ​​learned in advance using a large amount of speech waveforms and character strings as training data as initial values. Also, various functions can be considered for the model architecture, optimization function, and loss function, as in Example 1. The trained weight parameters of the third inference model are stored in the storage device 250.

[0124] By training the inference model as described above, even if there is a small amount of data measured by the detection device 100, the weighting parameters can be trained in advance using a large amount of available arbitrary audio information and text information as training data, thereby improving estimation accuracy.

[0125] [conversion] Next, the processing of the conversion unit 220 constituting the information processing device 200 will be described. The biometric information (first biometric information) transmitted from the biometric information acquisition unit 210 by the signal acquisition unit 221 is converted into a waveform signal by the pre-processing unit 222. Next, the waveform signal is input to a trained first inference model acquired from the storage device 250, and character information is output as first pseudo information corresponding to the biometric information. Furthermore, the inference unit 223 inputs the output character information to a trained second inference model acquired from the storage device 250, and outputs audio information as corresponding second pseudo information. Furthermore, the inference unit 223 inputs the output audio information to a trained third inference model acquired from the storage device 250, and outputs corresponding character information.

[0126] Next, the post-processing unit 224 removes unnecessary information from the character information output by the inference unit 223 and transmits the character information to the output unit 230. Finally, the output unit 230 transmits the character information to the display unit 110, and the character information is displayed on a monitor as an example of the display unit 110.

[0127] In this embodiment, an inference model that connects the first inference model, the second inference model, and the third inference model may be processed as a single inference model. That is, an inference model may be created in which the output of the first inference model is used as the input of the second inference model, and the output of the second inference model is used as the input of the third inference model, and a waveform signal may be used as an input to this inference model to output text information.

[0128] (Other Examples) In addition to the inference models described in the above embodiments, multiple inference models with different training data may be stored in the storage device 250. The inference unit 223 may then be configured to use an inference model selected from the multiple inference models as the first inference model and the second inference model.

[0129] The present invention can also be realized by providing a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that realizes one or more functions. This program and a computer-readable storage medium storing the program are included in the embodiments.

[0130] It should be noted that the above-described embodiments are merely examples of specific embodiments for carrying out the present invention, and the technical scope of the present invention should not be construed as being limited by these embodiments. In other words, the present invention can be carried out in various forms without departing from its technical concept or main features. [Explanation of symbols]

[0131] 1000 Information Processing Systems 100 Detection Devices 200 Information processing device 220 Conversion Unit 223A First Inference Section 223B Second Inference Section

Claims

1. an acquisition unit that acquires first biological information related to a movement of a body part of the user; a first inference unit that outputs pseudo information based on the first biometric information acquired by the acquisition unit, using a first inference model trained using second biometric information and pseudo information corresponding to the second biometric information as training data; a second inference unit that uses a second inference model trained using pseudo information and character information or voice information corresponding to the pseudo information as training data to output character information or voice information based on the pseudo information output from the first inference unit; An information processing system comprising:

2. 2. The information processing system according to claim 1, wherein the first biological information is at least one of myoelectric potential information, acceleration information, and angular velocity information resulting from a user's speaking action.

3. 2. The information processing system according to claim 1, wherein the first biometric information includes two or more types of information acquired from one location of the user.

4. 2. The information processing system according to claim 1, wherein the first biological information includes two or more types of information acquired from a plurality of locations on the user.

5. 2. The information processing system according to claim 1, further comprising a display unit that displays the character information output from the second inference unit.

6. 2. The information processing system according to claim 1, further comprising a voice output unit that outputs the voice information output from the second inference unit.

7. An information processing system as described in claim 1, characterized in that it has a plurality of inference models with different training data, and an inference model selected from the plurality of inference models is used as the first inference model and the second inference model.

8. an acquisition unit that acquires first biological information related to a movement of a body part of the user; a first inference unit that outputs first pseudo information based on the first biometric information acquired by the acquisition unit, using a first inference model trained using second biometric information and pseudo information corresponding to the second biometric information as training data; a second inference unit that uses a second inference model trained using the pseudo information output from the first inference unit and pseudo information corresponding to the pseudo information output from the first inference unit as training data, and outputs second pseudo information based on the first pseudo information output from the first inference model; a third inference unit that uses a third inference model trained using the pseudo information output from the second inference unit and character information or voice information corresponding to the pseudo information as training data to output character information or voice information based on the second pseudo information output from the second inference unit; An information processing system comprising:

9. an acquisition unit that acquires first biological information related to a movement of a body part of the user; a first inference unit that outputs pseudo information based on the first biometric information acquired by the acquisition unit, using a first inference model trained using second biometric information and pseudo information corresponding to the second biometric information as training data; a second inference unit that uses a second inference model trained using pseudo information and character information or voice information corresponding to the pseudo information as training data to output character information or voice information based on the pseudo information output from the first inference unit; An information processing device comprising:

10. an acquisition unit that acquires first biological information related to a movement of a body part of the user; a first inference unit that outputs first pseudo information based on the first biometric information acquired by the acquisition unit, using a first inference model trained using second biometric information and pseudo information corresponding to the second biometric information as training data; a second inference unit that outputs second pseudo information based on the first pseudo information output from the first inference unit, using a second inference model trained using the pseudo information output from the first inference unit and pseudo information corresponding to the pseudo information output from the first inference unit as training data; a third inference unit that uses a third inference model trained using the pseudo information output from the second inference unit and character information or voice information corresponding to the pseudo information as training data to output character information or voice information based on the second pseudo information output from the second inference unit; An information processing device comprising:

11. outputting pseudo information based on first biometric information of a user's body part acquired from the user, using a first inference model trained using second biometric information and pseudo information corresponding to the second biometric information as training data; a step of outputting character information or audio information based on the pseudo information output using the first inference model, using a second inference model trained using pseudo information and character information or audio information corresponding to the pseudo information as training data; 10. A method for controlling an information processing device, comprising:

12. outputting first pseudo information based on first biometric information relating to the movement of a part of the user, using a first inference model trained using second biometric information and pseudo information corresponding to the second biometric information as training data; a step of outputting second pseudo information based on the first pseudo information output using the first inference model, using a second inference model trained using the pseudo information output from the first inference unit and pseudo information corresponding to the pseudo information output from the first inference unit as training data; a step of outputting character information or audio information based on the second pseudo information output from the second inference unit using a third inference model trained using the pseudo information output using the second inference model and character information or audio information corresponding to the pseudo information as training data; An information processing method comprising:

13. A program for executing the steps of the control method according to claim 11 or 12.

Citation Information

Patent Citations

  • Method and device for speech recognition

    JP2006267664A

  • Speech recognition device using myoelectric potential signal

    JP2008233438A

  • Voice recognition device and program

    JP2015141253A

  • Information conversion system, information processing device, information processing method and program

    JP2023154894A

  • Utterance section detection device, voice recognition device, utterance section detection system, utterance section detection method, and utterance section detection program

    JP2021162685A