Information processing system and information processing method

The information processing system enhances the accuracy of converting biometric information into text or voice information by employing a two-tier inference model approach, addressing the limitations of existing technologies with limited training data.

WO2025263458A1PCT designated stage Publication Date: 2025-12-26CANON KK
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/021564
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-20
Filing Date
2025-06-16
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing technologies for converting biometric information, such as speech movements, into text or voice information face accuracy issues, particularly when limited training data is available, and the use of machine learning inference models does not sufficiently improve accuracy.

Method used

An information processing system utilizing a two-tier inference model approach, where a first inference model processes biometric information using a trained model, and a second model processes the output of the first to enhance accuracy, supported by pseudo information and text/audio training data.

Benefits of technology

The system significantly improves the accuracy of converting biometric information into text or voice information, even with limited training data, by leveraging a dual inference model structure and additional training data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025021564_26122025_PF_FP_ABST
    Figure JP2025021564_26122025_PF_FP_ABST
Patent Text Reader

Abstract

This information processing system comprises: an acquisition unit 210 that acquires first biological information relating to movement of a part of a user; a first inference unit 223A that outputs pseudo information on the basis of the first biological information acquired by the acquisition unit 210 using a first inference model trained with second biological information and pseudo information corresponding to the second biological information as teaching data; and a second inference unit 223B that outputs character information or voice information on the basis of the pseudo information outputted from the first inference unit 223A using a second inference model trained with pseudo information and character information or voice information corresponding to the pseudo information as teaching data.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing system and information processing method

[0001] The present invention relates to an information processing system that converts biometric information relating to the movement of a user's body part, such as speech, into text information or voice information.

[0002] In recent years, the content of a user's speech has been estimated from the movement of the user's mouth. Patent Document 1 discloses a technology for calculating a phoneme score from a lip image and an external sound to infer the content of a user's speech.

[0003] Japanese Patent Application Laid-Open No. 2021-162685

[0004] However, in Patent Document 1, the weighted sum of the lip phoneme score calculated based on the lip image and the speech phoneme score calculated based on the external sound is simply subjected to threshold processing, and the accuracy may not be sufficient. Meanwhile, in recent years, it has become known to make inferences using inference models trained by machine learning, but a large amount of training data is required to improve the accuracy of inference, and it has been difficult to improve the accuracy of inference with a small amount of training data.

[0005] The present invention aims to improve the accuracy of converting biometric information relating to the movement of a user's body parts, such as speech movements, into text information or voice information.

[0006] However, the problems to be solved by the embodiments disclosed in this specification and the drawings are not limited to the above problems. Problems corresponding to the effects of the configurations shown in the embodiments described below can also be positioned as other problems.

[0007] In order to achieve the above-mentioned object, the information processing system relating to the technology disclosed herein is characterized by having an acquisition unit that acquires first biometric information regarding the movement of a user's body part; a first inference unit that uses a first inference model trained with second biometric information and pseudo information corresponding to the second biometric information as training data to output pseudo information based on the first biometric information acquired by the acquisition unit; and a second inference unit that uses a second inference model trained with the pseudo information and text information or audio information corresponding to the pseudo information as training data to output text information or audio information based on the pseudo information output from the first inference model.

[0008] According to the technology disclosed herein, it is possible to improve the accuracy of converting biometric information relating to the movement of a user's body part, such as speech movements, into text information or voice information.

[0009] 1 is a diagram showing an example of the functional configuration of an information processing system according to a first embodiment. 2 is a diagram showing an example of the functional configuration of a conversion unit according to the first embodiment. 3 is a diagram showing an example of the functional configuration of an inference unit according to the first embodiment. 4 is a flowchart showing a conversion process of the information processing system according to the first embodiment. 5 is a diagram showing an example of the functional configuration of a learning unit according to the first embodiment. 6 is a flowchart showing a learning process in a first embodiment of the learning unit according to the first embodiment. 7 is a schematic diagram of a detection device according to Example 1. 8 is an explanatory diagram of the operation of the conversion unit according to Example 1. 9 is a schematic diagram of a detection device of Modification Example 2. 10 is an explanatory diagram of the operation of the conversion unit according to Example 2. 11 is an explanatory diagram of the operation of the conversion unit according to Example 3. 12 is a diagram showing an example of the functional configuration of an inference unit according to a second embodiment. 13 is a flowchart showing a learning process according to the second embodiment.

[0010] The following describes the embodiments in detail. The present invention is not limited to the following embodiments, provided that the gist of the present invention is not exceeded. Furthermore, by adding publicly known information, such as audio data or video data, to other channels of input data to the inference model, it is possible to perform inference with higher accuracy.

[0011] [First Embodiment] An information processing system 1 according to this embodiment will be described with reference to FIG. 1 . The information conversion system 1000 includes a detection device 100 and an information processing device 200. The detection device 100 detects biometric information based on muscle movements caused by a user's speaking actions. The information processing device 200 acquires the biometric information detected by the detection device 100 and converts it into text information or audio information using an inference model. The detection device 100 and the information processing device 200 may each be configured as separate devices connected by a network, or may be configured as an information processing device having the functions of a detection device as an integrated device.

[0012] <Detection device 100> The detection device 100 that constitutes the information processing system 1000 is configured to include a biometric information detection unit 101 that detects biometric information from one or more locations on the user, and a transmission unit 102 that transmits the biometric information detected by the biometric information detection unit 101 to the information processing device 200.

[0013] The detection device 100 is placed, for example, in contact with the user's skin to acquire biometric information from the user. The detection device 100 is preferably placed on a site such as the neck, lower jaw, around the mouth, or temples to detect biometric information based on muscle movements caused by the user's speaking actions. However, the detection device 100 may be placed in a location other than the above-mentioned locations as long as it can detect biometric information related to the movements of the mouth and tongue.

[0014] 1 shows a configuration in which the detection device 100 is realized by a single device, but the biological information detection unit 101 and the transmission unit 102 may be configured by different devices. Each functional configuration of the detection device 100 will be described below.

[0015] The biometric information detection unit 101 acquires biometric information as first biometric information related to the movement of a user's body part. Specifically, the biometric information detection unit 101 includes a sensor that detects biometric information based on the movement of the user's muscles. The biometric information detection unit 101 includes at least one sensor such as a myoelectric potential sensor, an acceleration sensor, an ultrasonic sensor, a tactile sensor, an optical sensor, a pressure sensor, a strain sensor, or a magnetic sensor. Note that the biometric information detection unit 101 may be a sensor other than the above, such as a camera, as long as it detects biometric information, and may not be in direct contact with the skin.

[0016] The transmitting unit 102 transmits the biological information based on muscle movements caused by the user's speaking action, which is acquired from the biological information detecting unit 101, to the information processing device 200.

[0017] <Information Processing Device 200> Here, the information processing device 200 constituting the information processing system 1000 is configured to include a biometric information acquisition unit 210, a conversion unit 220, and an output unit 230. The biometric information acquisition unit 210 as an acquisition unit acquires biometric information (first biometric information) transmitted from the detection device 100. The conversion unit 220 converts the biometric information into text information or audio information by using the biometric information as input to an inference model and inferring text information or audio information corresponding to the biometric information. The output unit 230 outputs the text information or audio information converted by the conversion unit 220 to the display unit 110 or the audio information output unit 120 (audio output unit).

[0018] Here, the inference model is a trained inference model trained by the training unit 240 and stored in the storage unit 250. The training unit 240 and other functional components may be configured as independent devices. For example, the training unit 240 may train the inference model used for inference by the conversion unit 220 on the cloud.

[0019] Here, the information processing device 200 refers to, for example, a smartphone, a personal computer (PC), a tablet PC, etc. The information processing device 200 includes a CPU, a GPU, a RAM, a ROM, and a storage device, which are connected via a system bus. If the information processing device 200 is, for example, a personal computer, the character information converted by the conversion unit 220 and output by the output unit 230 is output to a display unit 110 such as a display. Furthermore, if the information processing device 200 is converted into audio information, it is output to an audio information output unit 120 (audio output unit) such as a speaker. Note that the audio information output unit 120 may be incorporated into the device configuration of the information processing device 200.

[0020] Here, communication between the detection device 100 and the information processing device 200 may be wired or wireless. When communication between the detection device 100 and the information processing device 200 is realized by wire, the transmitting unit 102 in the detection device 100 and the biometric information acquiring unit 210 in the information processing device 200 are connected by wire such as a USB cable, HDMI (registered trademark), or the like.

[0021] When communication between the detection device 100 and the information processing device 200 is realized wirelessly, the transmitter 102 of the detection device 100 and the receiver 210 of the information processing device communicate wirelessly. For example, they are connected wirelessly by wireless LAN communication such as Wi-Fi or short-range wireless communication such as Bluetooth (registered trademark). If the information processing device 100 is a smartphone or tablet PC, the display unit 110 is a display. The audio information output unit 120 is a speaker installed in the smartphone or tablet PC, or an earphone connected to the smartphone or tablet PC.

[0022] Hereinafter, each functional configuration of the information processing device 200 will be described.

[0023] The biometric information acquisition unit 210 as an acquisition unit acquires biometric information (first biometric information) based on muscle movements caused by the user's speaking action, detected by the detection device 100. Here, the detected biometric information is at least one of myoelectric potential information caused by the user's speaking action, acceleration information, and tactile information. However, depending on the specifications of the inference model used by the conversion unit 220 to convert the biometric information, audio information and video information may also be acquired in addition to the above biometric information.

[0024] Furthermore, the biometric information may be two or more types of information acquired from one location placed on the user. Details will be described later, but for example, the biometric information detection unit 101 may be configured to include a myoelectric potential sensor and an acceleration sensor, and the biometric information acquisition unit 210 may acquire both myoelectric potential information and acceleration information. Both sensors may be placed in multiple locations relative to the user. When placed in multiple locations, each location may include multiple sensors, or each location may be configured with only a predetermined sensor.

[0025] As shown in Fig. 2, the conversion unit 220 includes a signal acquisition unit 221, a pre-processing unit 222, an inference unit 223, and a post-processing unit 224. The signal acquisition unit 221 acquires biometric information acquired by the biometric information acquisition unit 210. The pre-processing unit 222 performs pre-processing of the signal. The inference unit 223 uses the signal pre-processed by the pre-processing unit 222 as input and infers character information or audio information corresponding to the biometric information using an inference model stored in the storage device 250. The post-processing unit 224 performs post-processing of the character information or audio information inferred by the inference unit 223.

[0026] 3, the inference unit 223 is composed of a first inference unit 223A and a second inference unit 223B. The first inference unit 223A receives a signal preprocessed by the preprocessing unit 222 as input, and infers pseudo information corresponding to the signal using a first inference model stored in the storage device 250. The second inference unit 223B receives pseudo information inferred and output by the first inference unit 223A as input, and infers text information or audio information using a second inference model stored in the storage device 250.

[0027] The output unit 230 outputs the character information or audio information converted by the conversion unit to the display unit 110 or the audio information output unit 120 .

[0028] [Control Flow] An example of a control flow in the information processing device 200 will be described below with reference to FIG.

[0029] (Step S401: Acquiring Biometric Information) In step S401, the signal acquiring unit 221 in the conversion unit 220 acquires biometric information (first biometric information) such as a myoelectric potential signal transmitted from the detection device 100. The acquired biometric information is biometric information based on muscle movements caused by the user's speaking actions, such as myoelectric potential information and acceleration information. After transmitting the biometric information to the preprocessing unit 222, the signal acquiring unit 221 proceeds to the next step.

[0030] (Step S402: Preprocessing) In step S402, the preprocessing unit 222 in the conversion unit 220 performs various preprocessing processes on the biometric information. If the biometric information is one-dimensional time-series data such as an electromyography signal or an acceleration signal, preprocessing may include noise reduction processes such as elimination of outliers and signal smoothing, feature extraction processes such as octave analysis and spectrogram conversion, and standardization. If the biometric information is three-dimensional data such as a video of the user's face, preprocessing may include the above-mentioned noise reduction processes, background removal, grayscale conversion, image size conversion, and feature point extraction around the lips.

[0031] The preprocessing unit 222 preprocesses the biometric information, transmits the preprocessed data to the inference unit 223, and then proceeds to the next step. If preprocessing is not necessary, step S402 is not performed.

[0032] (Step S403: Inference) In step S403, the inference unit 223 uses the data transmitted from the preprocessing unit 222 as input to the inference model and infers corresponding character information or audio information.

[0033] Here, the inference unit 223 uses a trained inference model with an architecture configured by a neural network acquired from the storage device 240 for inference. The trained inference model is configured from a network of a known inference model. For example, an inference model having a network configuration such as a convolutional neural network (CNN), a recurrent neural network (RNN), or a long short term memory (LSTM), which is a type of deep learning, is used. An inference model is generated by applying a learning process to an inference model having such a network configuration. Note that the inference model used by the inference unit 223 for inference may be a model derived from a CNN, an RNN, or an LSTM, or may be another machine learning technique such as a support vector machine, a logistic regression, or a random forest, or may be a rule-based method.

[0034] Examples of character information that the inference model outputs as an inference result in the inference unit 223 include the following: syllables such as "a," "i," "u," "e," "o," "ka," and "ki," strings of syllables such as "hello" and "good night," and character strings including kanji, numbers, and katakana such as "call extension 5" and "turn off the lights."

[0035] The inference result by the inference unit 223 may be a phoneme string written in the alphabet such as "k a t a zu ke te e" or "h a j i m e m a s h i t e", or a mora string such as "d e N m a ak u" or "b y u cl f e n i i ku u". The language is not limited to Japanese.

[0036] The inference unit 223 transmits the character information that is the inference result to the output unit 230, and then proceeds to the next step. The conversion unit 220 may further convert the character information output by the inference unit 223. For example, the conversion unit 220 may convert "k a t a z u k e t e" to "tidy up" or convert "b y u cl f e n i i ku" to "go to the buffet," thereby making it easier for the user to understand the character information.

[0037] The conversion unit 220 may convert the information into audio information instead of text information. In the inference unit 223, examples of audio information output by the inference model as an inference result include sound pressure waveforms of monaural or stereo audio, and spectrograms. The audio information may be that of the user who made the speech, or may be audio information of another person. After sending the text information or audio information to the post-processing unit 224, the inference unit 223 proceeds to the next step.

[0038] (Step S404: Post-processing) In step S404, the post-processing unit 224 performs various post-processing operations on the text information or audio information transmitted by the inference unit 223. When the information transmitted by the inference unit 223 is text information, the post-processing operations include further conversion of the text information. For example, when the text information is a phoneme string written in the alphabet, such as "h a j i m e m a s h i t e," this may be converted into "nice to meet you" or "nice to meet you." In addition, in the above case, it may be converted into another language, such as "How do you do?". Furthermore, when there is a possibility of erroneous estimation of the text information, such as "h a s i m e m a s h i t e," it may be converted to "h a j i m e m a s h i t e." For example, "h a s i m e m a s h i t e" may be corrected to "h a ji m e m a s h i t e." Furthermore, the text information may be converted into non-text information such as emoticons or animations that correspond to the text information, or the text information may be converted into audio information. If the information transmitted by the inference unit 223 is audio information, the post-processing may involve further conversion of the audio information. For example, non-audio information and noise may be removed. The audio may also be converted into the voice tone of the user who made the speech or another person. Audio information may also be converted into text information.

[0039] If post-processing is not required, step S404 may not be performed. Also, some or all of the processing included in the post-processing may be included in the inference model and performed by the inference unit 223.

[0040] (Step S405: Information Output) In step S405, the output unit 230 outputs the text information or audio information acquired from the conversion unit 220 to an external device. In the case of text information, it is sent to the display unit 110 such as a monitor, and in the case of audio information, it is sent to the audio information output unit 120 such as a speaker or bone conduction earphones. In addition, the destination of transmission may be determined based on connection information with the external device, user setting information for the information processing device 200, etc.

[0041] With this configuration, the information processing system 1000 can convert biometric information based on muscle movements caused by a user's speaking actions into text information or audio information with high accuracy. Specifically, by performing inference using a first inference model that can input preprocessed biometric information and a second inference model that can input the output of the first inference model, the biometric information can be converted into text information or audio information with higher accuracy.

[0042] [Learning Flow] The following describes the learning unit 240 for learning an inference model to be used in the conversion unit 220, and the storage device 250 for storing the learned inference model. Note that the learning unit 240 does not necessarily have to be provided by the same device or the same entity, and the processing of the learning unit 240 may be replaced by storing a learned inference model generated by a different device or a different entity in the storage device 250.

[0043] In the information processing device 200, the learning unit 240 generates a trained inference model by learning the inference model using teacher data. FIG. 5 shows an example of the configuration of the learning unit 240. The learning unit 240 includes a teacher data acquisition unit 241, an inference model acquisition unit 242, and a model learning unit 243. The teacher data acquisition unit 241 acquires teacher data from the storage device 250. The inference model acquisition unit 242 acquires the inference model to be learned from the storage device 250. The model learning unit 243 learns the inference model using the inference model and teacher data. The model architecture and weight parameters are collectively referred to as the inference model.

[0044] Here, the storage device 250 stores biometric information detected by the detection device 100 and text information or audio information corresponding to the biometric information. The storage device 250 also stores biometric information other than the biometric information detected by the detection device 100 and text information or audio information corresponding to the biometric information. The storage device 250 also stores the architecture and weight parameters of a first inference model, the architecture and weight parameters of a second inference model, and parameters required for executing pre-processing and post-processing.

[0045] The learning flow of the inference model by the learning unit 240 will be explained below with reference to FIG.

[0046] (Step S601: Acquisition of Teacher Data) In step S601, the teacher data acquisition unit 241 acquires, from the storage device 250, biometric information based on muscle movements caused by speaking (second biometric information) and pseudo information corresponding to the biometric information, and performs preprocessing on the biometric information. Here, the pseudo information is text information or audio information inferred based on the biometric information using an inference model. The teacher data acquisition unit 241 also acquires correct answer data by converting the pseudo information into information suitable for learning as needed. For example, if one piece of pseudo information is "Turn off the lights," it converts it into "Turn off the lights" or "R a i to o o k e s h i t e." The teacher data acquisition unit 241 transmits the preprocessed biometric information data and correct answer data to the model learning unit 243 as teacher data, and the inference model acquisition unit 242 acquires a first inference model from the storage device 250 and transmits it to the model learning unit 243, and proceeds to the next step.

[0047] (Step S602: Learning the first inference model) The model learning unit 243 uses the teacher data transmitted from the teacher data acquisition unit 241 and the first inference model transmitted from the inference model acquisition unit 242 to perform a learning process on the first inference model and learn the weight parameters of the first inference model.

[0048] The first inference model learned by the model learning unit 243 is also used by the inference unit 223 of the conversion unit 220. The optimization function applied by the model learning unit 243 when learning the first inference model may be SGD (Stochastic Gradient Decent), Adam (Adaptive Moment Estimation), or the like. Furthermore, Cross-entropy loss or CTC (Connectionist Temporal Classification) loss is used as the loss function. The optimization function and loss function are not limited to these, and various functions may be used. Furthermore, the initial values ​​of the weight parameters may be random, or pre-learned values ​​may be used.

[0049] (Step S603: Storing the first inference model) In step S603, when the model learning unit 243 completes the learning process of the first inference model, it stores information about the learned first inference model in the storage device 250. In addition, the model learning unit 243 stores pseudo-information obtained as a result of estimation by the learned first inference model using preprocessed biometric information as input.

[0050] (Step S604: Acquiring Teacher Data) In step S604, the teacher data acquisition unit 241 acquires pseudo information and text information or audio information corresponding to the pseudo information from the storage device 250. The teacher data acquisition unit 241 may also acquire the pseudo information stored in step S603 and text information or audio information corresponding to the pseudo information from the storage device 250. The inference model acquisition unit 242 acquires a second inference model from the storage device 250. After the teacher data acquisition unit 241 transmits the pseudo information and text information or audio information corresponding to the pseudo information to the model learning unit 243 and the inference model acquisition unit 242 transmits the second inference model to the model learning unit 243, the process proceeds to the next step.

[0051] (Step S605: Learning the second inference model) The model learning unit 243 performs a learning process on the second inference model using the teacher data transmitted from the teacher data acquisition unit 241 and the second inference model transmitted from the inference model acquisition unit 242, and learns weight parameters of the second inference model. The second inference model learned by the model learning unit 243 is also used by the inference unit 223 of the conversion unit 220. The optimization function applied by the model learning unit 243 when learning the second inference model may be SGD (Stochastic Gradient Decent), Adam (Adaptive Moment Estimation), or the like. Furthermore, cross-entropy loss and CTC (Connectionist Temporal Classification) loss are used as loss functions. Note that the optimization function and loss function are not limited to these, and various functions can be used. Furthermore, the initial values ​​of the weight parameters may be random, or pre-learned values ​​may be used.

[0052] (Step S606: Storing the first inference model) In step S606, when the model learning unit 243 completes the learning process of the second inference model, it stores information about the learned second inference model in the memory device 250, and the learning of the inference model is completed.

[0053] The learning of the first inference model and the second inference model may be unsupervised learning without correct answer data or semi-supervised learning without correct answer data for some data. The learning unit 240 may be implemented as a function on a personal computer or may be configured on the cloud. Only some of the functions, such as the model learning unit 243, may be configured on the cloud.

[0054] An information processing system according to a first embodiment will be described with reference to Fig. 7. Fig. 7 shows a schematic diagram of a detection device 700 corresponding to the detection device 100 constituting the information conversion system 1. The detection device 700 is composed of a biometric information detection unit 701 and a transmission unit 702. The biometric information detection unit 701 is composed of a myoelectric potential sensor and an acceleration / angular velocity sensor, and can detect movements around the mouth by attaching it to the cheek, chin, neck, or temple of the user with tape or adhesive.

[0055] The myoelectric potential sensor can measure the myoelectric potential of the muscles at the attachment site, and the acceleration / angular velocity sensor can measure movements around the mouth as three-axis translational acceleration and three-axis angular velocity.

[0056] The sampling rate of the myoelectric potential sensor and the acceleration / angular velocity sensor is set to 2 kHz.

[0057] The transmitter 702 is housed in a housing on the neckband. The myoelectric potential signal and acceleration / angular velocity signal detected by the biometric information detector 701 are sent to the transmitter 702 via a wired connection and then wirelessly transmitted to the information processing device 200. Specifically, this can be achieved via wireless LAN communication such as Wi-Fi or short-range wireless communication such as Bluetooth (registered trademark). The detection device 700 is powered by a battery (not shown) in the neckband. While FIG. 7 illustrates a configuration in which biometric information is detected at two locations on the face, biometric information may be detected at one location or three or more locations. A different sensor may be provided at each contact point with the user. The detected biometric information may be only myoelectric potential signals or only three-axis translational acceleration. The biometric information is not limited to myoelectric potential, translational acceleration, and angular velocity; various other types of information may also be detected, such as distortion, tactile sensation, magnetism, and ultrasound.

[0058] In this embodiment, the information processing device 200 is, for example, a smartphone, the display unit 110 is a display of the smartphone, and the audio information output unit 120 is a speaker or a wired or wireless earphone. The audio information output unit 120 may also be a component of the detection device 200, and for example, the audio information output unit 120 may be a bone conduction earphone (not shown).

[0059] Next, the processing of the learning unit 240 that constitutes the information processing device 200 will be described.

[0060] First, a large number of myoelectric potential signals, acceleration signals, and angular velocity signals measured by the detection device 100, as well as text information and audio information corresponding to each signal, are stored in the storage device 250 in advance. The architecture of a first inference model is stored in the storage device 250. Weight parameters of the first inference model may also be stored. The weight parameters of the first inference model may be random values ​​or values ​​learned in advance using arbitrary biometric information and text information. The architecture of a second inference model is also stored in the storage device 250. Weight parameters of the second inference model may also be stored. The weight parameters of the second inference model may be random values ​​or values ​​learned in advance using arbitrary information as training data.

[0061] [Learning] The preprocessing unit 222 performs a standardization process to remove noise from each signal that is biometric information (second biometric information) acquired from the storage device 250. Furthermore, the character information acquired from the storage device 250 is converted into a character string such as "o N g a ku u s a i s e e", but hiragana characters may also be used. The learning unit 240 first learns weight parameters of the first inference model acquired from the storage device 250 using the signal and character string (pseudo information) converted by the preprocessing unit 222 as training data.

[0062] The model architecture includes multiple convolutional layers, BiGRU (Bidirectional Gated Recurrent Unit) layers, and linear combination layers. The optimization function is AdamW (Adaptive Moment Estimation with Weight decay), and the loss function is CTC loss.

[0063] The model architecture, optimization function, and loss function are not limited to these, and various other functions are possible. The trained model obtained by the training unit 240 is stored in the storage device 250. In addition, the character information output by the trained model using the signal converted by the preprocessing unit 222 as input is stored in the storage device 250.

[0064] Next, weight parameters of the second inference model are learned using the character information (pseudo information) acquired from the storage device 250 and the audio information corresponding to the character information as training data. The weight parameters of the second inference model are values ​​learned using arbitrary character information (pseudo information) and the audio information corresponding to the character information as training data. When the amount of data measured by the detection device 100 is small, the estimation accuracy can be improved by learning the weight parameters in advance using a large amount of available arbitrary character information and the corresponding audio information as training data.

[0065] As with the first inference model, various functions can be considered for the model architecture, optimization function, and loss function. The trained model obtained by the training unit 240 is stored in the storage device 250. Note that the training unit 240 may be configured as a separate device, or its function may be replaced by acquiring a trained inference model.

[0066] [Conversion] Next, the processing of the conversion unit 220 constituting the information processing device 200 will be described with reference to Fig. 8. The biometric information (first biometric information) transmitted from the biometric information acquisition unit 210 and acquired by the signal acquisition unit 221 is converted into a waveform signal by the pre-processing unit 222. Next, the waveform signal is input to the trained first inference model acquired from the storage device 250, and character information is output as pseudo information corresponding to the biometric information.

[0067] Furthermore, the inference unit 223 inputs the output character information into a trained second inference model obtained from the storage device 250, and outputs corresponding audio information. Next, the post-processing unit 224 removes non-audio components and noise from the audio information output by the inference unit 223, and transmits the audio information to the output unit 230. Finally, the output unit 230 transmits the audio information to the audio information output unit 120, and audio is output from earphones as an example of the audio information output unit 120.

[0068] In this embodiment, an inference model formed by linking the first inference model and the second inference model may be processed as a single inference model. That is, an inference model may be created in which the output of the first inference model is input to the second inference model, and a waveform signal may be input to this inference model to output audio information.

[0069] <Modification 1> An information processing system according to Modification 1 of Example 1 will be described below. Note that descriptions of overlapping parts will be omitted as appropriate.

[0070] In this modification, the teacher data acquisition unit 231 of the learning unit 240 and the preprocessing unit 222 of the conversion unit 220 convert signal data, which is biological information, into a spectrogram. Specifically, the signal is converted into a feature matrix in which multiple column vectors obtained by short-time Fourier transform are arranged. The converted feature matrix is ​​sent to the model learning unit 243 or the inference unit 223. Note that the feature matrix may be converted into various feature matrices other than a spectrogram, such as a mel spectrogram. In the first embodiment, the input to the first inference model is a waveform signal, i.e., one-dimensional data, whereas in this modification, the input is a feature matrix, i.e., two-dimensional data. Therefore, the first layer of the architecture of the first inference model is changed from a one-dimensional convolutional layer to a two-dimensional convolutional layer. Furthermore, the architecture other than the first layer may be changed.

[0071] <Modification 2> An information processing system according to Modification 2 of Example 1 will be described. Note that descriptions of overlapping parts will be omitted as appropriate. In this modification, the detection device 900 has a biometric information detection unit 901 that is ear-hooked as shown in FIG. The biometric information detection unit 901, such as a myoelectric potential sensor or an acceleration / angular velocity sensor, is attached to a support made of, for example, a resin, and is in contact with the skin of the lower jaw, cheek, etc. with appropriate pressure. The biometric information detected by the biometric information detection unit 901 is sent via a wire to a transceiver unit 902 in the neckband, and then sent wirelessly to the information processing device 200. The detection device 900 is powered by a battery (not shown) in the neckband.

[0072] Furthermore, when the information processing device 200 outputs audio information, the transmitting / receiving unit 902 receives the audio information and the audio information can be reproduced by an audio information output unit 903 such as a speaker or a bone conduction microphone in the ear hook portion.

[0073] <Modification 3> An information processing system according to Modification 3 of the first embodiment will be described.

[0074] In the above embodiment, the conversion unit 220 converts biometric information into character information using a trained model trained by the training unit 240.

[0075] In this modification, the learning unit 240 further performs additional learning of the inference model using training data that pairs correct answer data, which is text information based on the user's voice information, with learning data, which is waveform signals acquired from biometric information. Here, the conversion from voice information to text information may be achieved by any known technology.

[0076] This configuration enables learning according to the characteristics of the user, and enables conversion to character information with even higher accuracy than at the time of distribution.

[0077] <Modification 4> Modification 4 of Example 1 will be described. Note that descriptions of overlapping parts will be omitted as appropriate.

[0078] The processing of the learning unit 240 constituting the information processing device 200 in this modification will be described.

[0079] [Learning] The preprocessing unit 222 converts the biosignal (second bioinformation) acquired from the storage device 250 into a spectrogram. It also converts the audio information corresponding to the biosignal into a spectrogram. It also converts the text information corresponding to the biosignal into a phoneme sequence or a character string such as hiragana or kanji.

[0080] The learning unit 240 first learns weight parameters of the first inference model acquired from the storage device 250 using the spectrogram obtained by converting the biosignal in the preprocessing unit 222 and the spectrogram of the audio information as training data. As with Example 3, various functions can be considered for the model architecture, optimization function, and loss function. The weight parameters of the trained first inference model are stored in the storage device 250. In addition, a spectrogram that is the output result of the trained first inference model that uses the spectrogram obtained by converting the biosignal in the preprocessing unit 222 as input is also stored in the storage device 250. Here, the spectrogram that is the output result of the first inference model includes data that is different from the data used as training data. In other words, the spectrogram that is the output result of the first inference model includes erroneously estimated data.

[0081] Next, the learning unit 240 learns the weight parameters of the second inference model acquired from the storage device 250. During learning, the spectrogram data, which is the output result of the trained first inference model that uses as input the spectrogram obtained by converting the biological signal in the preprocessing unit 222, and the character string corresponding to the biological information, are used as training data. The weight parameters of the second inference model after learning are stored in the storage device 250. As with Example 3, various functions are possible for the model architecture, optimization function, and loss function. Furthermore, the weight parameters of the second inference model may use values ​​that have been trained in advance using a large amount of arbitrary spectrograms and corresponding character strings as training data. By training the inference model as described above, even if the data measured by the detection device 100 is small, the weight parameters can be trained in advance using a large amount of arbitrary character information that is available as training data, thereby improving estimation accuracy.

[0082] An information processing system according to a second embodiment will be described below. Note that overlapping descriptions will be omitted as appropriate.

[0083] The processing of the learning unit 240 constituting the information processing device 200 in this embodiment will be described.

[0084] [Learning] The preprocessing unit 222 performs noise removal and standardization on the biometric signal as biometric information (second biometric information) acquired from the storage device 250. In addition, character information acquired from the storage device 250 corresponding to the biometric information is converted into a phoneme string or a character string such as hiragana or kanji.

[0085] The learning unit 240 first learns weight parameters of the first inference model acquired from the storage device 250 using the signal and character string (pseudo information) converted by the preprocessing unit 222 as training data. As in Example 1, various functions can be considered for the model architecture, optimization function, and loss function. The weight parameters of the trained first inference model are stored in the storage device 250. In addition, character strings that are the output result of the trained first inference model, which uses the signal converted by the preprocessing unit 222 as input, are also stored in the storage device 250. Here, the character strings that are the output result of the first inference model include both the same and different character strings from those used as training data. In other words, the character strings that are the output result of the first inference model include erroneously estimated character strings.

[0086] Next, the learning unit 240 learns weight parameters of the second inference model acquired from the storage device 250 using a character string that is the output result of the trained first inference model that uses the signal converted by the preprocessing unit 222 as input and a character string corresponding to the biometric information as training data. The weight parameters of the trained second inference model are stored in the storage device 250. As with the first embodiment, various functions are possible for the model architecture, optimization function, and loss function. Furthermore, the weight parameters of the second inference model may be values ​​that have been trained in advance using a large number of arbitrary incorrect character strings and correct character strings as training data. Here, the character representation methods for the incorrect character strings and the correct character strings may be the same or different. For example, the incorrect character string may be "w a t a s h i wa" and the correct character string may be "watashi wa," and various combinations are possible. By training the inference model in the above manner, even when there is little data measured by the detection device 100, the weight parameters can be trained in advance using a large amount of arbitrary character information that is available as training data, thereby improving estimation accuracy.

[0087] [Conversion] Next, the processing of the conversion unit 220 constituting the information processing device 200 will be described with reference to FIG. 10 . The biometric information (first biometric information) transmitted from the biometric information acquisition unit 210 is acquired by the signal acquisition unit 221 and converted into a waveform signal by the pre-processing unit 222. Next, the waveform signal is input to a trained first inference model acquired from the storage device 250, and character information is output as pseudo information corresponding to the biometric information. Furthermore, the inference unit 223 inputs the output character information into a trained second inference model acquired from the storage device 250, and the corresponding character information is output. Next, the post-processing unit 224 transmits the character information output by the inference unit 223 to the output unit 230. Finally, the output unit 230 transmits the character information to the display unit 110, and the character information is displayed on a monitor, an example of the display unit 110.

[0088] In this embodiment, an inference model formed by linking the first inference model and the second inference model may be processed as a single inference model. That is, an inference model may be created in which the output of the first inference model is input to the second inference model, and a waveform signal may be input to this inference model to output character information.

[0089] An information processing system according to a third embodiment will be described below. Note that overlapping descriptions will be omitted as appropriate.

[0090] The processing of the learning unit 240 constituting the information processing device 200 in this embodiment will be described.

[0091] [Learning] The preprocessing unit 222 performs noise removal and standardization on the biometric signal as biometric information (second biometric information) acquired from the storage device 250. Furthermore, audio information corresponding to the biometric signal is converted into waveform data. Furthermore, character information corresponding to the biometric signal is converted into a phoneme sequence or a character string such as hiragana or kanji.

[0092] The learning unit 240 first learns the weight parameters of the first inference model obtained from the storage device 250 using the signal converted by the preprocessing unit 222 and waveform data (pseudo information) as training data. As in Example 1, various functions can be considered for the model architecture, optimization function, and loss function. The weight parameters of the trained first inference model are stored in the storage device 250. In addition, the audio signal that is the output result of the trained first inference model that uses the signal converted by the preprocessing unit 222 as input is also stored in the storage device 250. Here, the output result of the first inference model includes both the same data as the data used as training data and data that is different. In other words, the audio signal that is the output result of the first inference model includes erroneously estimated data.

[0093] Next, the learning unit 240 learns weight parameters of the second inference model acquired from the storage device 250 using the waveform signal (pseudo information) that is the output result of the trained first inference model that uses the signal converted by the preprocessing unit 222 as input, and character strings corresponding to the biometric information as training data. The weight parameters of the second inference model after training are stored in the storage device 250. As with the first embodiment, various functions are possible for the model architecture, optimization function, and loss function. Furthermore, the weight parameters of the second inference model may use values ​​that have been trained in advance using a large amount of arbitrary erroneous waveform signals and correct character strings as training data. By training the inference model as described above, even if the amount of data measured by the detection device 100 is small, the weight parameters can be trained in advance using a large amount of arbitrary character information that is available as training data, thereby improving estimation accuracy.

[0094] [Conversion] Next, the processing of the conversion unit 220 constituting the information processing device 200 will be described with reference to FIG. 11 . The biometric information (first biometric information) transmitted from the biometric information acquisition unit 210 is acquired by the signal acquisition unit 221 and converted into a waveform signal by the pre-processing unit 222. Next, the waveform signal is input to a trained first inference model acquired from the storage device 250, and audio information is output as pseudo-information corresponding to the biometric information. Furthermore, the inference unit 223 inputs the output audio information into a trained second inference model acquired from the storage device 250, and outputs corresponding text information. Next, the post-processing unit 224 converts the text information output by the inference unit 223 into a character string that is easily recognizable by the user, for example, kanji, and transmits it to the output unit 230. Finally, the output unit 230 transmits the text information to the display unit 110, and the text information is displayed on a monitor, an example of the display unit 110.

[0095] In this embodiment, an inference model formed by linking the first inference model and the second inference model may be processed as a single inference model. That is, an inference model may be created in which the output of the first inference model is input to the second inference model, and a waveform signal may be input to this inference model to output character information.

[0096] An information processing system according to a fourth embodiment will be described below. Note that overlapping descriptions will be omitted as appropriate.

[0097] The processing of the learning unit 240 constituting the information processing device 200 in this embodiment will be described.

[0098] [Learning] The preprocessing unit 222 performs noise removal and standardization on the biosignal acquired from the storage device 250. Furthermore, audio information corresponding to the biosignal is converted into an audio waveform.

[0099] The learning unit 240 first learns weight parameters of the first inference model acquired from the storage device 250 using the signal (second biometric information) and voice waveform (pseudo information) converted by the preprocessing unit 222 as training data. As with Example 1, various functions can be considered for the model architecture, optimization function, and loss function. The weight parameters of the trained first inference model are stored in the storage device 250. The voice waveform, which is the output result of the trained first inference model using the signal converted by the preprocessing unit 222 as input, is also stored in the storage device 250. Here, the output result of the first inference model includes data that differs from the data used as training data. In other words, the voice waveform, which is the output result of the first inference model, includes erroneously estimated data.

[0100] Next, the learning unit 240 learns weight parameters of the second inference model acquired from the storage device 250 using the speech waveform that is the output result of the trained first inference model that uses the signal converted by the preprocessing unit 222 as input and the speech waveform corresponding to the biometric information as training data. The weight parameters of the second inference model after training are stored in the storage device 250. As with the first embodiment, various functions are possible for the model architecture, optimization function, and loss function. Furthermore, the weight parameters of the second inference model may use values ​​that have been trained in advance using a large amount of arbitrary erroneous speech waveforms and correct speech waveforms as training data. By training the inference model as described above, even if the amount of data measured by the detection device 100 is small, the weight parameters can be trained in advance using a large amount of arbitrary available speech waveforms as training data, thereby improving estimation accuracy.

[0101] [Conversion] Next, the processing of the conversion unit 220 constituting the information processing device 200 will be described. The preprocessing unit 222 performs preprocessing on the biometric information (first biometric information) transmitted from the biometric information acquisition unit 210 and acquired by the signal acquisition unit 221. Next, the preprocessed signal is input to a trained first inference model acquired from the storage device 250, and a voice waveform is output as pseudo information corresponding to the biometric information. Furthermore, the inference unit 223 inputs the output voice waveform to a trained second inference model acquired from the storage device 250, and outputs a corresponding voice waveform. Next, the postprocessing unit 224 removes non-voice components and noise from the voice waveform output by the inference unit 223, and transmits the voice waveform to the output unit 230. Finally, the output unit 230 transmits the voice information to the voice information output unit 120, and voice is output from earphones, for example.

[0102] In this embodiment, an inference model formed by linking the first inference model and the second inference model may be processed as a single inference model. That is, an inference model may be created in which the output of the first inference model is input to the second inference model, and a waveform signal may be input to this inference model to output character information.

[0103] Second Embodiment An information processing system according to a second embodiment will be described below, with the description of parts that overlap with the first embodiment being omitted as appropriate.

[0104] The information processing system differs from the first embodiment in that the inference unit 223 constituting the conversion unit 220 is an inference unit. As shown in FIG. 12 , the inference unit 223 of the information processing system is composed of a first inference unit 1223A, a second inference unit 1223B, and a third inference unit 1223C. The first inference unit 1223A receives a signal preprocessed by the preprocessing unit 222 as input and infers first pseudo information corresponding to the signal using a first inference model stored in the storage device 250. The second inference unit 1223B receives the first pseudo information as input and infers second pseudo information corresponding to the first pseudo information using a second inference model stored in the storage device 250. The third inference unit 1223C receives the second pseudo information as input and infers character information or audio information corresponding to the signal using a third inference model stored in the storage device 250.

[0105] [Learning Flow] Below, we will explain the learning unit 240 for learning the inference model used in the conversion unit 220 and the storage device 250 for storing the learned inference model.

[0106] In the information processing device 200, the storage device 250 stores biometric information detected by the detection device 100 and character information or audio information corresponding to the biometric information. The storage device 250 also stores biometric information other than the biometric information detected by the detection device 100 and character information or audio information corresponding to the biometric information. The storage device 250 also stores the architecture and weight parameters of a first inference model, the architecture and weight parameters of a second inference model, the architecture and weight parameters of a third inference model, and parameters required for executing pre-processing and post-processing.

[0107] Using FIG. 13, the learning flow of the inference model by the learning unit 240 will be explained.

[0108] (Step S1301: Acquisition of Teacher Data) In step S1301, the teacher data acquisition unit 241 acquires, from the storage device 250, biometric information based on muscle movements caused by speaking and text information or audio information as pseudo-information corresponding to the biometric information, and preprocesses the biometric information. The teacher data acquisition unit 241 also acquires correct answer data by converting the text information or audio information into information suitable for learning as needed. For example, if one piece of text information is "Turn off the lights," it converts it into "Turn off the lights" or "R a i to o o k e s h i t e." The teacher data acquisition unit 241 transmits the preprocessed biometric information data and correct answer data to the model learning unit 243 as teacher data. The inference model acquisition unit 242 acquires a first inference model from the storage device 250 and transmits it to the model learning unit 243, and the process proceeds to the next step.

[0109] (Step S1302: Learning the first inference model) The model learning unit 243 uses the teacher data transmitted from the teacher data acquisition unit 241 and the first inference model transmitted from the inference model acquisition unit 242 to perform a learning process on the first inference model and learn the weight parameters of the first inference model.

[0110] The first inference model learned by the model learning unit 243 is also used by the inference unit 223 of the conversion unit 220. The optimization function applied by the model learning unit 243 when learning the first inference model may be SGD (Stochastic Gradient Decent), Adam (Adaptive Moment Estimation), or the like. Furthermore, Cross-entropy loss or CTC (Connectionist Temporal Classification) loss is used as the loss function. The optimization function and loss function are not limited to these, and various functions may be used. Furthermore, the initial values ​​of the weight parameters may be random, or pre-learned values ​​may be used.

[0111] (Step S1303: Storing the first inference model) In step S1303, when the model learning unit 243 completes the learning process of the first inference model, it stores information about the learned first inference model in the storage device 250. In addition, the model learning unit 243 stores the first pseudo information obtained as a result of estimation by the learned first inference model using preprocessed biometric information as input.

[0112] (Step S1304: Acquiring Teacher Data) In step S1304, the teacher data acquisition unit 241 acquires the first pseudo information and text information or audio information corresponding to the pseudo information from the storage device 250. The teacher data acquisition unit 241 may also acquire the pseudo information stored in step S1303 and text information or audio information corresponding to the pseudo information from the storage device 250. The inference model acquisition unit 242 acquires a second inference model from the storage device 250. After the teacher data acquisition unit 241 transmits the pseudo information and the text information or audio information corresponding to the other pseudo information to the model learning unit 243 and the inference model acquisition unit 242 transmits the second inference model to the model learning unit 243, the process proceeds to the next step.

[0113] (Step S1305: Learning the second inference model) The model learning unit 243 performs a learning process on the second inference model using the teacher data transmitted from the teacher data acquisition unit 241 and the second inference model transmitted from the inference model acquisition unit 242, and learns weight parameters of the second inference model. The second inference model learned by the model learning unit 243 is also used by the inference unit 223 of the conversion unit 220. The optimization function applied by the model learning unit 243 when learning the second inference model may be SGD (Stochastic Gradient Decent), Adam (Adaptive Moment Estimation), or the like. Furthermore, cross-entropy loss and CTC (Connectionist Temporal Classification) loss are used as loss functions. Note that the optimization function and loss function are not limited to these, and various functions can be used. Furthermore, the initial values ​​of the weight parameters may be random, or pre-learned values ​​may be used.

[0114] (Step S1306: Storing the second inference model) In step S1306, when the model learning unit 243 completes the learning process for the second inference model, it stores information about the learned second inference model in the storage device 250. In addition, the model learning unit 243 stores the second pseudo information obtained as a result of estimation by the learned second inference model using the first pseudo information as input.

[0115] (Step S1307: Acquiring Teacher Data) In step S1307, the teacher data acquisition unit 241 acquires the second pseudo information and text information or audio information corresponding to the pseudo information from the storage device 250. The teacher data acquisition unit 241 may also acquire the pseudo information stored in step S1306 and text information or audio information corresponding to the pseudo information from the storage device 250. The inference model acquisition unit 242 acquires a third inference model from the storage device 250. After the teacher data acquisition unit 241 transmits the pseudo information and the text information or audio information corresponding to the other pseudo information to the model learning unit 243 and the inference model acquisition unit 242 transmits the third inference model to the model learning unit 243, the process proceeds to the next step.

[0116] (Step S1308: Learning the third inference model) The model learning unit 243 performs a learning process on the third inference model using the teacher data transmitted from the teacher data acquisition unit 241 and the third inference model transmitted from the inference model acquisition unit 242, and learns weight parameters of the third inference model. The third inference model learned by the model learning unit 243 is also used by the inference unit 223 of the conversion unit 220. The optimization function applied by the model learning unit 243 when learning the third inference model may be SGD (Stochastic Gradient Decent), Adam (Adaptive Moment Estimation), or the like. Furthermore, cross-entropy loss and CTC (Connectionist Temporal Classification) loss are used as loss functions. Note that the optimization function and loss function are not limited to these, and various functions can be used. Furthermore, the initial values ​​of the weight parameters may be random, or pre-learned values ​​may be used.

[0117] (Step S1309: Storing the third inference model) In step S1309, when the model learning unit 243 completes the learning process of the third inference model, it stores information about the learned third inference model in the memory device 250, and the learning of the inference model is completed.

[0118] The learning of the first inference model, the second inference model, and the third inference model may be unsupervised learning without correct answer data or semi-supervised learning without correct answer data for some data. The learning unit 240 may be implemented as a function on a personal computer or may be configured on the cloud. Only some of the functions, such as the model learning unit 243, may be configured on the cloud.

[0119] An information processing system according to a fifth embodiment will be described. Note that descriptions of parts that overlap with those of the first embodiment will be omitted as appropriate. The processing of a learning unit 240 constituting an information processing device 200 in this embodiment will be described.

[0120] [Learning] The preprocessing unit 222 performs noise removal and standardization on the biometric signal acquired from the storage device 250. Furthermore, character information acquired from the storage device 250 corresponding to the biometric information is converted into a phoneme string or a character string such as hiragana or kanji.

[0121] The learning unit 240 first learns weight parameters of the first inference model acquired from the storage device 250 using the signal (second biometric information) and character string (pseudo information) converted by the preprocessing unit 222 as training data. As with Example 1, various functions can be considered for the model architecture, optimization function, and loss function. The weight parameters of the trained first inference model are stored in the storage device 250. In addition, the storage device 250 also stores first pseudo information, which is the output result of the trained first inference model that uses the signal converted by the preprocessing unit 222 as input. Here, the first pseudo information, which is the output result of the first inference model, i.e., the character string, includes both the same and different character strings as those used as training data. In other words, the first pseudo information, which is the output result of the first inference model, includes erroneously estimated character strings.

[0122] Next, the learning unit 240 learns the weight parameters of the second inference model acquired from the storage device 250. During learning, first pseudo information, which is the output result of the trained first inference model using the signal converted by the preprocessing unit 222 as input, and speech waveforms, which are speech information corresponding to the biometric information, are used as training data. Here, the weight parameters of the second inference model may be initially set to values ​​trained in advance using a large amount of arbitrary character strings and correct speech waveforms as training data. The weight parameters of the trained second inference model are stored in the storage device 250. As in Example 1, various functions are possible for the model architecture, optimization function, and loss function. In addition, second pseudo information, which is the output result of the trained second inference model using the first pseudo information as input, is also stored in the storage device 250. Here, the second pseudo information, i.e., the speech waveform, which is the output result of the second inference model, includes speech waveforms that differ from the speech waveforms used as training data. In other words, the second pseudo information, which is the output result of the second inference model, includes incorrectly estimated speech waveforms.

[0123] Next, the learning unit 240 learns the weight parameters of the third inference model obtained from the storage device 250 using the second pseudo information and the character string converted by the preprocessing unit 222 as training data. Here, the weight parameters of the third inference model may use values ​​learned in advance using a large amount of speech waveforms and character strings as training data as initial values. Also, various functions can be considered for the model architecture, optimization function, and loss function, as in Example 1. The trained weight parameters of the third inference model are stored in the storage device 250.

[0124] By training the inference model as described above, even if there is a small amount of data measured by the detection device 100, the weighting parameters can be trained in advance using a large amount of available arbitrary audio information and character information as training data, thereby improving estimation accuracy.

[0125] [Conversion] Next, the processing of the conversion unit 220 constituting the information processing device 200 will be described. The biometric information (first biometric information) transmitted from the biometric information acquisition unit 210 by the signal acquisition unit 221 is converted into a waveform signal by the pre-processing unit 222. Next, the waveform signal is input to a trained first inference model acquired from the storage device 250, and character information is output as first pseudo information corresponding to the biometric information. Furthermore, the inference unit 223 inputs the output character information to a trained second inference model acquired from the storage device 250, and outputs audio information as corresponding second pseudo information. Furthermore, the inference unit 223 inputs the output audio information to a trained third inference model acquired from the storage device 250, and outputs corresponding character information.

[0126] Next, the post-processing unit 224 removes unnecessary information from the character information output by the inference unit 223 and transmits the character information to the output unit 230. Finally, the output unit 230 transmits the character information to the display unit 110, and the character information is displayed on a monitor as an example of the display unit 110.

[0127] In this embodiment, an inference model that connects the first inference model, the second inference model, and the third inference model may be processed as a single inference model. That is, an inference model may be created in which the output of the first inference model is used as the input of the second inference model and the output of the second inference model is used as the input of the third inference model, and a waveform signal may be used as an input to this inference model to output character information.

[0128] (Other Examples) In addition to the inference models described in the above embodiments, multiple inference models with different training data may be stored in the storage device 250. The inference unit 223 may then be configured to use an inference model selected from the multiple inference models as the first inference model and the second inference model.

[0129] The present invention can also be realized by providing a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that realizes one or more functions. This program and a computer-readable storage medium storing the program are included in the embodiments.

[0130] It should be noted that the above-described embodiments are merely examples of specific embodiments for carrying out the present invention, and the technical scope of the present invention should not be construed as being limited by these embodiments. In other words, the present invention can be carried out in various forms without departing from its technical concept or main features.

[0131] The present invention is not limited to the above-described embodiments, and various modifications and variations can be made without departing from the spirit and scope of the present invention. Therefore, the following claims are appended to apprise the public of the scope of the present invention.

[0132] This application claims priority based on Japanese Patent Application No. 2024-099469, filed June 20, 2024, the entire contents of which are incorporated herein by reference.

[0133] 1000 Information processing system 100 Detection device 200 Information processing device 220 Conversion unit 223A First inference unit 223B Second inference unit

Claims

1. An information processing system comprising: an acquisition unit that acquires first biometric information relating to the movement of a user's body part; a first inference unit that uses a first inference model trained using second biometric information and pseudo-information corresponding to the second biometric information as training data to output pseudo-information based on the first biometric information acquired by the acquisition unit; and a second inference unit that uses a second inference model trained using the pseudo-information and text information or audio information corresponding to the pseudo-information as training data to output text information or audio information based on the pseudo-information output from the first inference unit.

2. The information processing system according to claim 1, wherein the first biological information is at least one of myoelectric potential information, acceleration information, and angular velocity information resulting from the user's speaking movements.

3. The information processing system according to claim 1, wherein the first biometric information includes two or more types of information acquired from one location of the user.

4. The information processing system according to claim 1, wherein the first biometric information includes two or more types of information acquired from multiple locations on the user.

5. The information processing system according to claim 1, further comprising a display unit for displaying the character information output from the second inference unit.

6. An information processing system according to claim 1, further comprising an audio output unit that outputs the audio information output from the second inference unit.

7. An information processing system as described in claim 1, characterized in that it has multiple inference models with different training data, and an inference model selected from the multiple inference models is used as the first inference model and the second inference model.

8. An information processing system comprising: an acquisition unit that acquires first biometric information related to the movement of a user's body part; a first inference unit that outputs first pseudo information based on the first biometric information acquired by the acquisition unit, using a first inference model trained using second biometric information and pseudo information corresponding to the second biometric information as teacher data; a second inference unit that outputs second pseudo information based on the first pseudo information output from the first inference model, using a second inference model trained using the pseudo information output from the first inference unit and pseudo information corresponding to the pseudo information output from the first inference unit as teacher data; and a third inference unit that outputs text information or audio information based on the second pseudo information output from the second inference unit, using a third inference model trained using the pseudo information output from the second inference unit and text information or audio information corresponding to the pseudo information as teacher data.

9. An information processing device comprising: an acquisition unit that acquires first biometric information relating to the movement of a user's body part; a first inference unit that outputs pseudo information based on the first biometric information acquired by the acquisition unit using a first inference model trained using second biometric information and pseudo information corresponding to the second biometric information as training data; and a second inference unit that outputs text information or audio information based on the pseudo information output from the first inference unit using a second inference model trained using the pseudo information and text information or audio information corresponding to the pseudo information as training data.

10. An information processing device comprising: an acquisition unit that acquires first biometric information related to the movement of a user's body part; a first inference unit that outputs first pseudo information based on the first biometric information acquired by the acquisition unit, using a first inference model trained using second biometric information and pseudo information corresponding to the second biometric information as teacher data; a second inference unit that outputs second pseudo information based on the first pseudo information output from the first inference unit, using a second inference model trained using the pseudo information output from the first inference unit and pseudo information corresponding to the pseudo information output from the first inference unit as teacher data; and a third inference unit that outputs text information or audio information based on the second pseudo information output from the second inference unit, using a third inference model trained using the pseudo information output from the second inference unit and text information or audio information corresponding to the pseudo information as teacher data.

11. A method for controlling an information processing device, comprising: a step of outputting pseudo information based on first biometric information of a user's body part obtained from the user using a first inference model trained using second biometric information and pseudo information corresponding to the second biometric information as training data; and a step of outputting text information or audio information based on the pseudo information output using a second inference model trained using pseudo information and text information or audio information corresponding to the pseudo information as training data.

12. An information processing method comprising: a step of outputting first pseudo information based on first biometric information relating to movement of a user's body part using a first inference model trained using second biometric information and pseudo information corresponding to the second biometric information as training data; a step of outputting second pseudo information based on the first pseudo information output using a second inference model trained using the pseudo information output from the first inference unit and pseudo information corresponding to the pseudo information output from the first inference unit as training data; and a step of outputting text information or audio information based on the second pseudo information output from the second inference unit using a third inference model trained using the pseudo information output using the second inference model and text information or audio information corresponding to the pseudo information as training data.

13. A program for executing each step of the control method according to claim 11 or 12.

Citation Information

Patent Citations

  • Method and device for speech recognition

    JP2006267664A

  • Speech recognition device, method, and program

    WO2021144901A1

  • Information conversion system, information processing device, information processing method, and program

    WO2023195323A1