Machine learning-based input device, input system, and method for processing input

The machine learning-based input device processes sound signals to generate control commands, addressing installation and environmental limitations of conventional input devices, ensuring robust and compatible operation.

WO2026095582A1PCT designated stage Publication Date: 2026-05-07COMMONLINK INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
COMMONLINK INC
Filing Date
2025-10-28
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Conventional input devices are constrained by their installation requirements, exposed to environmental factors, and lack compatibility with other devices, leading to potential malfunctions and design compromises.

Method used

A machine learning-based input device that acquires sound signals through an object, processes them using a learning model, and generates control signals to operate external devices, allowing for robust and compatible input without external exposure.

Benefits of technology

Enables precise and clear command recognition, enhances device compatibility, and reduces failure risks due to environmental factors by internal installation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025017343_07052026_PF_FP_ABST
    Figure KR2025017343_07052026_PF_FP_ABST
Patent Text Reader

Abstract

In one embodiment, provided may be a machine learning-based input device which is coupled to or in contact with an object that serves as a medium for transmitting sound, the input device comprising: a sound reception unit for acquiring the sound as an input signal via the object, and electrically converting the input signal; a processing unit for receiving the input signal from the sound reception unit, and inputting sound data corresponding to the input signal into a pre-trained learning model so as to output a command corresponding to the sound data; and a signal generation unit for generating a control signal for executing the command.
Need to check novelty before this filing date? Find Prior Art

Description

Machine learning-based input device, input system, and method for processing input

[0001] The present embodiment relates to an input device for processing sound signals, and to a technology for processing input based on machine learning.

[0002] Generally, mechanical and / or electronic devices are equipped with input devices. Simply put, switches or buttons are primarily used, while keyboards, touchscreens, touchpads, trackballs, or mice are used to implement complex functions.

[0003] However, conventional input devices have various limitations. For example, since input devices must be indispensably installed somewhere within an electronic device or a product containing it, they impose constraints such as compromising the design or requiring compatibility with other internal components. Furthermore, because conventional input devices must be operated by a user, they are exposed to the outside of the product; in this case, they may malfunction or break down due to exposure to moisture, temperature, or foreign substances. To solve these problems, there is an urgent need to develop a new concept of input device that is robust against external environments, provides a simple input method, and is compatible with any other device.

[0004] Accordingly, the inventor of the present invention has completed the present invention after a long period of research and trial and error.

[0005] Against this backdrop, one objective of the present embodiment is to provide an input device coupled to or in contact with an object that serves as a medium for transmitting sound, wherein the input device acquires sound as an input signal through said object, inputs sound data corresponding to the input signal into a pre-learned learning model to output a command corresponding to said sound data, and generates a control signal for executing said command.

[0006] Meanwhile, other unspecified objectives of the present invention will be further considered to the extent that they can be easily inferred from the following detailed description and effects.

[0007] To achieve the aforementioned objective, one embodiment provides a machine learning-based input device comprising: a sound receiving unit that acquires the sound as an input signal through the object and electrically converts the input signal, wherein the input device is coupled to or contacted with an object that is a medium for transmitting sound; a processing unit that receives the input signal from the sound receiving unit and inputs sound data corresponding to the input signal into a pre-learned learning model to output a command corresponding to the sound data; and a signal generating unit that generates a control signal for executing the command.

[0008] In the above device, the sound is generated by stimulating the surface of the object and can be transmitted to the sound receiving unit through the object.

[0009] In the above device, the stimulus may be electrical or non-electric, and the object may be made of a conductive material or a non-conductive material.

[0010] In the above device, the learning model may be composed of an artificial neural network including an RNN (recurrent neural network) or a CNN (convolutional neural network).

[0011] In the above device, the learning model may be composed of a combination of an RNN artificial neural network and a CNN artificial neural network.

[0012] In the above device, the learning model includes a plurality of learning models, and the processing unit can output the command based on the output results of the plurality of learning models for the sound data.

[0013] In the above device, the sound receiving unit includes a plurality of sound receiving units that each acquire the sound as an input signal through the object and electrically convert the plurality of input signals, and receives the plurality of input signals from the plurality of sound receiving units and inputs a plurality of sound data corresponding to the plurality of input signals into the learning model to output a single command corresponding to the plurality of sound data.

[0014] The above device includes a communication unit that communicates with an electronic device via a wired or wireless network and transmits the control signal to the electronic device, and the control signal can control the electronic device so that the electronic device executes the command.

[0015] Another embodiment provides a machine learning-based input system comprising: an input device; and an object to which the input device is coupled on one surface, wherein the sound is generated by stimulating the surface of the object and transmitted to the sound receiver through the object.

[0016] Another embodiment provides a method for processing input based on machine learning, comprising the steps of: acquiring the sound as an input signal through the object as a medium for transmitting sound; electrically converting the input signal; generating a command corresponding to the input signal from the input signal, wherein the command is generated by receiving the input signal and inputting sound data corresponding to the input signal into a pre-learned learning model to output a command corresponding to the sound data; and generating a control signal for executing the command.

[0017] In the above method, the sound is generated by stimulating the surface of the object and can be transmitted to the sound receiver through the object.

[0018] Another embodiment may provide a computer-readable medium that records a program for performing a method of processing input.

[0019] As explained above, according to the present embodiment, by deriving commands for sounds through a machine learning-based learning model, it is possible to recognize various sounds and generate commands accordingly in a precise and clear manner.

[0020] In addition, according to the present embodiment, since processing of sound-type input is possible, the user can easily issue commands simply by generating sound.

[0021] In addition, according to the present embodiment, various types of user interfaces can be constructed by providing sound-type input.

[0022] In addition, according to the present embodiment, an input device linked with a surrounding electronic device can process sound-type input and issue commands to the electronic device, thereby increasing the expandability of the input device and increasing compatibility with surrounding electronic devices.

[0023] In addition, according to the present embodiment, the input device is installed inside the product without being exposed on the surface of the electronic device or the product including it, thereby enabling secret input by the user.

[0024] In addition, according to the present embodiment, the input device is not exposed to the outside but is contained within the product, so the design of the product is not compromised and the occurrence of failures (malfunctions caused by humidity, temperature, and foreign substances) due to the influence of the external environment can be reduced.

[0025] Meanwhile, it should be added that even if an effect is not explicitly mentioned here, the effects described in the following specification and the provisional effects expected by the technical features of the present invention are treated as described in the specification of the present invention.

[0026] FIG. 1 is a configuration diagram of an input device according to one embodiment.

[0027] FIG. 2 is an illustrative diagram for explaining the use of an input device according to one embodiment.

[0028] FIG. 3 is a drawing showing an example of an input device according to one embodiment.

[0029] FIG. 4 is a drawing showing another example of an input device according to one embodiment.

[0030] FIG. 5 is an example of sound data input to an input device according to one embodiment.

[0031] FIG. 6 is an example of preprocessed sound data of sound data input to an input device according to one embodiment.

[0032] FIG. 7 is an example diagram of a learning model of an input device according to one embodiment.

[0033] FIG. 8 is another example of a learning model of an input device according to one embodiment.

[0034] FIG. 9 is another example of a learning model of an input device according to one embodiment.

[0035] FIG. 10 is a flowchart of an operation in which an input device according to one embodiment processes input through a learning model.

[0036] FIG. 11 is a flowchart of an operation in which an input device according to one embodiment processes input through a plurality of learning models.

[0037] FIG. 12 is a diagram illustrating the operation of an input device according to one embodiment processing input through a plurality of learning models.

[0038] FIG. 13 is a flowchart of an operation in which an input device according to one embodiment processes a plurality of inputs through a learning model.

[0039] FIG. 14 is an example of a plurality of sound data input to an input device according to one embodiment.

[0040] It should be noted that the attached drawings are provided as examples for reference to help understand the technical concept of the present invention, and the scope of the rights of the present invention is not limited by them.

[0041] In describing the present invention, detailed descriptions of related known functions are omitted if they are deemed obvious to a person skilled in the art and could unnecessarily obscure the essence of the invention.

[0042] The terms used in this application are used merely to describe specific embodiments and are not intended to limit the invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, terms such as "comprising" or "having" are intended to specify the presence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0043] Terms such as "first," "second," etc. are merely identifiers for distinguishing identical or corresponding components, and identical or corresponding components are not limited by terms such as "first," "second," etc.

[0044] Hereinafter, embodiments according to the present invention will be described in detail with reference to the accompanying drawings. In describing with reference to the accompanying drawings, identical or corresponding components are given the same reference numerals, and redundant descriptions thereof will be omitted.

[0045] FIG. 1 is a configuration diagram of an input device according to one embodiment, and FIG. 2 is an example diagram for explaining the use of an input device according to one embodiment.

[0046] An input device (100) according to one embodiment has the primary purpose of processing input and may provide some operation or result as an auxiliary function after processing the input. This is because the input device (100) is specialized in processing input provided by a user in the form of sound. In addition, for the utility of using the input device (100), after processing the input in the special form of sound, a command may be issued as a result of the input. This command may be generated to operate another electronic device rather than to operate the input device (100). This command may be transmitted to another electronic device via wired or wireless means.

[0047] Additionally, the input device (100) may be manufactured as a small device or module that processes sound-shaped input on its own and used independently. Alternatively, the input device (100) may be included in an electronic device and used in combination with other objects as a component of the electronic device.

[0048] Referring to FIG. 1, an input device (100) according to one embodiment may include a sound receiving unit (110), a processing unit (120), a signal generating unit (130), a communication unit (140), a preprocessing unit (150), and a storage unit (not shown). The input device (100) may be designed to be dedicated to processing input provided by a user. In particular, the input device (100) may receive and process sound generated from the outside. Sound or tone can be understood as a wave generated by the vibration (movement) of an object and transmitted through air as a medium. This wave is transmitted to a person's ear and causes the eardrum to vibrate, thereby allowing the person to perceive the sound. Sound may have a meaning that a person can understand, but it may also have no meaning, such as noise.

[0049] The sound receiver (110) can detect sound as an input signal. The sound may originate from the outside surrounding the input device (100). First, the sound may originate from the object (D). The sound may be transmitted to the input device (100) using the object (D) as a medium as the object (D) is stimulated by the user (UE). Alternatively, the sound may originate from the surrounding environment rather than the object (D). Preferably, such sound may be generated by the user operating the input device (100). When a sound having waves, i.e., a sound wave, is input to the sound receiver (110), the sound receiver (110) can electrically convert the input signal to generate an analog signal. The electrically converted analog input signal is represented as a wave in which the intensity or pressure of the sound changes over time, and the intensity or pressure of the sound of the electrical input signal can be converted into a voltage level or a current level. For example, the sound receiving unit (110) may include a microphone capable of detecting sound.

[0050] Here, the waveform of an analog input signal can be determined by sound waves; in particular, the amplitude of the waveform can correspond proportionally to sound intensity or sound pressure. Externally generated sound waves have an amplitude that varies over time, and this amplitude can be expressed as 'intensity.' Sound intensity can refer to the instantaneous amount of sound energy transmitted per unit area by sound waves. Furthermore, sound intensity can be defined as the sound energy (E) transmitted vertically per unit time (t) and unit area (A). Therefore, sound intensity has units of W / m² and can be represented by sound energy. Meanwhile, sound pressure is the pressure of the medium with units of N / m², and since sound intensity is determined by sound pressure, sound pressure can represent sound intensity. Since sound intensity represents sound energy, sound pressure can also correspond to sound energy. Therefore, sound intensity or sound pressure can be regarded as 'sound energy,' and hereinafter, sound energy may be used as a term representing the magnitude, intensity, or pressure of sound.

[0051] The processing unit (120) receives an analog input signal and can convert it into a digital input signal. The processing unit (120) can perform computational processing on digital data corresponding to the sound input signal. Specifically, the processing unit (120) can generate digital sound data corresponding to the analog input signal. The processing unit (120) can generate commands corresponding to the sound data through a pre-machine-learned learning model (122). The processing unit (120) can input sound data into the learning model (122) and output commands corresponding thereto.

[0052] The processing unit (120) may include a learning unit (121) that machine learns the learning model (122). The learning unit (121) may input sound data regarding an input signal received from the sound receiving unit (110) into the learning model (122) and train the learning model (122) to output a command corresponding to the input signal from the learning model (122).

[0053] For example, a user (UE) can draw a specific shape or form on the surface of an object (D) to generate sound according to that shape or form. The sound can be transmitted along the object (D) and enter the input device (100). The input device (100) can output a command corresponding to the sound through a learning model (122), and furthermore, can generate a control signal for that command to operate an external electronic device. By drawing a '3' on the surface of the object (D), the user (UE) can generate a sound corresponding to the shape of '3', identify the sound represented by this '3', and output a command to the electronic device to 'Power OFF' in response. Alternatively, by drawing a '4' on the surface of the object (D), the user (UE) can generate a sound corresponding to the shape of '4', identify the sound represented by this '4', and output a command to the electronic device to 'Power ON' in response. Various sounds are generated as the user (UE) stimulates the object (D), and these various sounds act as individual inputs through the input device (100), and the input device (100) can derive each command according to the input—that sound. The user (UE) can generate sounds by stimulating the object (D) not only with numbers but also with the shape of letters or symbols, and derive commands accordingly. Furthermore, the user (UE) can draw shapes or forms created by themselves, and the input device (100) can derive commands corresponding to them. The learning unit (121) can train the learning model (122) to output a corresponding command when the learning model (122) receives a sound according to the shape drawn on the object (D). Therefore, there may actually be no limit to the number of input signals—the number of cases of sounds acting as inputs—that the user (UE) can command through the input device (100).

[0054] The learning model (122) can learn the sound signal to output a command for the sound input signal. The learning model (122) can perform machine learning in a supervised or unsupervised manner. The learning model (122) may include various machine learning algorithms, and preferably may include a recurrent neural network (RNN) or a convolutional neural network (CNN). An RNN may be a model specialized for processing time-series data. Since the sound signal is time-series—a change in sound energy over time—sound data can also be considered time-series data. Therefore, an RNN may be advantageous for learning sound data as time-series data. Additionally, a CNN may be a model specialized for processing image data. The sound signal may have a waveform and, essentially, represent a change in sound energy over time. Since this waveform itself is represented as image data, sound data can also be considered as image data. Therefore, a CNN may be advantageous for learning sound data as image data. Alternatively, the sound signal may be converted into a spectrogram. A spectrogram can represent a sound signal by plotting time on the horizontal axis, frequency on the vertical axis, and pixel values ​​as sound energy at a corresponding frequency. Since a sound signal converted into a spectrogram is represented as image data, sound data can also be considered as image data. Therefore, CNNs can be advantageous for learning sound data as image data.

[0055] The signal generation unit (140) can generate a signal corresponding to a command to execute the command output from the processing unit (120) on an external electronic device connected to the input device (100). Since the command controls the operation of the electronic device, this signal can be defined as a control signal. The signal generation unit (140) can generate a control signal by processing command data so that the command can be transmitted to the electronic device and executed by the electronic device. For example, since the communication protocol or data processing protocol differs for each electronic device, the signal generation unit (130) can process the command data according to the protocol. The signal generation unit (130) can change the voltage or current level representing the value of the command data or modulate the command data to suit the communication environment. Modulation can also be optionally performed by the communication unit (140).

[0056] The communication unit (140) can perform communication with an external device. Specifically, the communication unit (140) can communicate with an external electronic device by being connected to a network via wired or wireless means. Here, the external electronic device may be a server, a smartphone, a tablet, a PC, etc. The communication unit (140) may include a communication module that supports one of various wired or wireless communication methods. For example, the communication module may be in the form of a chipset, or it may be a sticker / barcode containing information necessary for communication. Additionally, the communication module may be a short-range communication module or a wired communication module. For example, the communication unit (140) can support at least one of Wireless LAN, Wi-Fi (Wireless Fidelity), WFD (Wi-Fi Direct), Bluetooth, BLE (Bluetooth Low Energy), Wired LAN, NFC (Near Field Communication), Zigbee, infrared (IrDA, infrared Data Association), 3G, 4G, and 5G.

[0057] Accordingly, the communication unit (140) is connected to the electronic device via a wireless network to communicate and can transmit a control signal generated by the processing unit (120) to the electronic device. Here, the control signal can control the electronic device to execute a command.

[0058] And the storage unit (not shown) may store data or information required for the input device (100) to process the input signal. For example, the storage unit (not shown) may store learning data that the learning model (122) learns. The learning data may be a data set containing sound data and a corresponding command. Depending on the case, the learning unit (121) may train the learning model (122) with the learning data stored in the storage unit (not shown). When new learning data is stored in the storage unit (not shown) or is periodically updated, the learning unit (121) may train the learning model (122) with the new learning data. When the learning model (122) trained with the learning data receives sound data according to the input signal, it may output a corresponding command.

[0059] The preprocessing unit (150) receives an input signal—a sound signal—from the sound receiving unit (110) and can preprocess the input signal to generate a preprocessed signal (preprocessed data). When the input signal is transmitted to the processing unit (120), sound data can be obtained from the input signal. The processing unit (120) can input this sound data into a learning model (122) to derive a command. The sound data transmitted directly from the sound receiving unit (110) to the processing unit (120) can be defined as 'first sound data'. Meanwhile, the input signal is transmitted to the preprocessing unit (150), and the preprocessing unit (150) can obtain a preprocessed signal (preprocessed data) from the input signal. This input signal may be identical to the input signal of the first sound data. However, the preprocessing unit (150) can preprocess the input signal to generate a preprocessed input signal. Sound data can be obtained from this preprocessed input signal. The processing unit (120) can input this sound data into a learning model (122) to derive a command. The sound data transmitted from the sound receiving unit (110) to the processing unit (120) via the preprocessing unit (150) can be defined as 'second sound data'. Here, the first sound data may be time-series data representing sound energy over time. Additionally, the first sound data may be image data representing sound energy over time. The second sound data may represent image data representing sound energy over time and frequency. For example, the second sound data may include an image-spectrogram- in which the horizontal axis and the vertical axis are divided into time and frequency, respectively, and sound energy is represented as pixel values.

[0060] Referring to FIG. 2, an example of the use of an input device (100) according to one embodiment may be illustrated. The input device (100) may be coupled to or in contact with an object (D) and operated by a user (UE). The user (UE) may combine or contact the input device (100) with one side of the object (D) and generate sound through the object (D) on the other side. The user (UE) may stimulate the other side of the object (D)—by rubbing or scraping it with a hand—to produce sound in a first direction (D1). The user (UE) may use the object (D) to easily generate sound. This is because it may be difficult to continuously produce sound indicating a certain direction without a medium such as the object (D). Whenever the user (UE) wishes to apply an input signal to the input device (100), the user may combine with the object (D) to input sound in the first direction (D1) into the input device (100). As described above, the sound varies depending on the shape drawn by the user (UE) on the object (D), and as the shape varies, the commands output by the input device (100) can also vary.

[0061] Additionally, the input device (100) can be simply configured as a miniaturized device or module. The input device (100) may be equipped with at least one sound receiving unit (110). This sound receiving unit (110) can detect sound generated through the target object (D). The input device (100) can determine a command corresponding to sound data of the sound generated according to the first direction (D1).

[0062] Here, the object (D) is exemplified as a desk, but it is not limited to this; any means capable of connecting or contacting the input device (100) and allowing the user (UE) to generate sound may be possible. Since the input of the input device (100) is sound, the source generating the sound does not need to have electrical characteristics. In particular, the difference between the input device (100) according to one embodiment and a touch input is that in touch input, the touch panel where the user (UE)'s stimulation occurs and the electronic device are electrically connected so that the user's stimulation is transmitted electrically, whereas in sound input according to one embodiment, the object (D) where the user's (UE) stimulation occurs and the input device (100) are not electrically connected at all, and the user's (UE) stimulation is also transmitted non-electrically. Therefore, the user's (UE) stimulation is non-electric, and the object (D) can be composed of a non-conductive material—e.g., wood, plastic, etc. As long as the object (D) functions as a medium for transmitting sound, any material can be used as the object (D). Therefore, even if the input device (100) is positioned on an object (D), the input device (100) can receive sound as input according to the shape drawn by the user (UE) on the object (D) and output a corresponding command. The user (UE) can control an external electronic device from anywhere through the input device (100).

[0063] Additionally, the input device (100) can receive sound input signals transmitted from the object (D) while placed on the surface of the object (D). Therefore, when the input device (100) is placed on the upper surface of the object (D), it comes into contact with the upper surface. Conversely, when the input device (100) is located on the lower surface of the object (D), it is coupled to the lower surface. That is, the input device (100) can be fixed to the object (D) by coupling to the lower surface of the object (D)—coupling in various ways such as vacuum or attachment. In this drawing, an example is shown in which the input device (100) is coupled to the lower surface of the object (D).

[0064] FIG. 3 is a drawing showing an example of an input device according to one embodiment.

[0065] Referring to FIG. 3, an example of a form in which an input device (100) according to one embodiment is implemented may be illustrated. The input device (100)

[0066] The sound receiving unit, processing unit, signal generation unit, communication unit, preprocessing unit, and storage unit can all be included and implemented in a single form. A form in which all components of the input device (100) are included in a single module or device in this manner can be referred to as a fixed type.

[0067] And the sound receiver may be positioned to be suitable for detecting sound. The input device (100) may include a single sound receiver positioned at any location. The input device (100) may include a plurality of sound receivers (110-1, 110-2) to detect a plurality of input signals—a plurality of sounds—coming from a plurality of directions.

[0068] FIG. 4 is a drawing showing another example of an input device according to one embodiment.

[0069] Referring to FIG. 4, another example of an implementation form of an input device according to one embodiment may be illustrated. The sound receiving unit may be located outside the input device, and components other than the sound receiving unit—such as a communication unit, a processing unit, a signal generation unit, a preprocessing unit, and a storage unit—may be located inside the input device. This form in which the sound receiving unit is located outside the input device may be referred to as a variable type. In this drawing, two sound receiving units (110-1, 110-2) will be described as examples.

[0070] Here, the externally located sound receiving unit (110-1, 110-2) can communicate with the communication unit (140) via wired or wireless means and exchange data. The sound receiving unit (110-1, 110-2) may include a communication module itself and can communicate with the communication unit (140) of the input device (100) through this communication module.

[0071] For example, according to the first method, the sound receiving unit (110-1, 110-2) can communicate with the communication unit (140) via a wire (see I in the drawing). The sound receiving unit (110-1, 110-2) can detect the input signal of sound and transmit the electrically converted input signal to the processing unit (120) via a wire through the communication unit (140). The processing unit (120) can determine a command through a learning model from the sound data of the input signal. The input device (100) can generate a control signal (SIG_CON) containing a command and transmit it to the electronic device (ED) via a wire. The electronic device (ED) can operate according to the control signal (SIG_CON).

[0072] According to the second method as another example, the sound receiving unit (110-1, 110-2) can communicate wirelessly with the communication unit (140) (see II in the drawing). The sound receiving unit (110-1, 110-2) can detect the input signal of sound and transmit the electrically converted input signal to the processing unit (120) through the communication unit (140) via a wireless network. The processing unit (120) can determine a command through a learning model from the sound data of the input signal. The processing unit (120) can generate a control signal (SIG_CON) containing the command and transmit it to an electronic device (ED) via a wireless network. The electronic device (ED) can operate according to the control signal (SIG_CON).

[0073] According to the third method as another example, the sound receiving unit (110-1, 110-2) can communicate wirelessly with the communication unit (140) (see III in this drawing). The sound receiving unit (110-1, 110-2) can detect the input signal of sound and transmit the electrically converted input signal to the processing unit (120) via the communication unit (140) using short-range communication—e.g., Bluetooth, Zigbee, etc. The processing unit (120) can determine a command through a learning model from the sound data of the input signal. The processing unit (120) can generate a control signal (SIG_CON) containing the command and transmit it to an electronic device (ED) via short-range communication—e.g., Bluetooth, Zigbee, etc. The electronic device (ED) can operate according to the control signal (SIG_CON).

[0074] FIG. 5 is an example of sound data input to an input device according to one embodiment.

[0075] Referring to FIG. 5, an example of sound data input to an input device according to one embodiment may be illustrated. When a processing unit obtains an input signal from a sound receiving unit, sound data such as this example can be extracted. This sound data may correspond to first sound data representing sound energy over time. The first sound data is a change in sound energy over time, i.e., time series data, and can be applied to a learning model, mainly including an RNN model.

[0076] In addition, the first sound data is image data representing sound energy over time, and can be applied to a learning model, mainly including a CNN model.

[0077] FIG. 6 is an example of preprocessed sound data of sound data input to an input device according to one embodiment.

[0078] Referring to FIG. 6, an example of preprocessed sound data input to an input device according to one embodiment may be illustrated. The sound receiving unit may transmit the input signal to a preprocessing unit in addition to the processing unit. The preprocessing unit may preprocess the input signal to generate a preprocessed input signal. The preprocessing unit may convert the input signal in the time domain into the frequency domain—for example, using STFT (Short-Time Fourier Transform)—to obtain a spectrogram for the first sound data. The processing unit may convert the preprocessed input signal into digital data to extract or generate second sound data. The second sound data may represent sound energy according to time and frequency. Sound energy according to frequency over time is represented as pixel values, and this form may become image data. The second sound data is image data and can be applied to a learning model, mainly including a CNN model.

[0079] FIG. 7 is an example diagram of a learning model of an input device according to one embodiment.

[0080] Referring to FIG. 7, an example of a learning model of an input device (100) according to one embodiment may be illustrated. As an example, the learning model may include an RNN. An RNN is a model that performs deep learning among machine learning and may have the structure of an artificial neural network. A learning model that performs deep learning may generally be composed of an input layer, a hidden layer, and an output layer, and in this figure, a case in which it is composed of one input layer, two hidden layers, and one output layer may be illustrated.

[0081] Each layer of the input, hidden, and output layers contains neurons (or nodes) marked with circles, where each circle acts as a nerve cell (neuron) and these are connected to form a neural network. Neurons in adjacent layers are connected by lines, and as data flows from the input layer to the output layer through these lines, the value of each node is calculated and the final value—the command—is output.

[0082] FIG. 8 is another example of a learning model of an input device according to one embodiment.

[0083] Referring to FIG. 8, another example of a learning model of an input device (100) according to one embodiment may be illustrated. As another example, the learning model may include a CNN. A CNN is also a model that performs deep learning among machine learning, and basically has the structure of an artificial neural network such as an RNN. The learning model may include additional configurations for learning image data. The learning model is a convolutional neural network with a pyramid structure and can output or extract commands corresponding to an input signal—sound data in the form of an image. Furthermore, the learning model may include a plurality of convolution layers that perform filtering composed of 3×3 and 1×1 convolutions and a plurality of max pooling layers that perform downsampling, and may apply batch normalization to the input of each convolution layer and apply Leaky ReLU (rectified linear unit) as an activation function. Subsequently, a fully connected layer may be included. Once image features are derived from the convolution layer and pooling layer, they can be passed to the fully connected layer for classification. Finally, a command can be derived from the fully connected layer.

[0084] FIG. 9 is another example of a learning model of an input device according to one embodiment.

[0085] Referring to FIG. 9, another example of a learning model of an input device (100) according to one embodiment may be illustrated. As another example, the learning model may be composed of an artificial neural network including an RNN and an artificial neural network including a CNN. The learning model may be composed of a combination of at least one RNN artificial neural network and at least one CNN artificial neural network. For example, the learning model may be composed of one RNN artificial neural network and a plurality of CNN artificial neural networks. The first CNN artificial neural network, the RNN artificial neural network, and the second CNN artificial neural network may be connected sequentially, and the RNN artificial neural network and the second CNN artificial neural network may be connected in parallel with each other. The learning model may receive first sound data in the form of an image representing sound energy over time, or receive second sound data in the form of an image representing sound energy over time and frequency. The first or second sound data, which is image data, is input to the first CNN artificial neural network, and the first CNN artificial neural network may output intermediate data. And intermediate data is input into an RNN artificial neural network, and the RNN artificial neural network can output first result data. At the same time, intermediate data is input into a second CNN artificial neural network, and the second CNN artificial neural network can output second result data. The learning model can combine the first and second result data to output a single command corresponding to the input data—image data—that is input into the learning model. Here, the method of outputting a single command from the first and second result data may be identical or similar to the process of generating a single command from the results output from the multiple learning models described below.

[0086] FIG. 10 is a flowchart of an operation in which an input device according to one embodiment processes input through a learning model.

[0087] Referring to FIG. 10, the flow of operation in which an input device according to one embodiment processes input through a learning model can be illustrated.

[0088] The sound receiving unit of the input device can acquire sound as an input signal through the object (D) (Step S1001). The sound receiving unit can also electrically convert the sound to generate an analog input signal (Step S1003).

[0089] The processing unit of the input device receives an analog input signal and can generate a command corresponding to the input signal through a learning model (step S1005). The processing unit can extract digital sound data from the analog input signal and input the sound data into the learning model to output a command corresponding to the sound data. The learning model is a pre-trained model and may include an RNN or a CNN. The learning unit can input sound data generated according to various stimuli to the user's object into the learning model and train the learning model to output a command corresponding to it. Here, the learning data may include time-series data of sound data, or image data in the form of an image of sound data—for example, a waveform image (see FIG. 5) or a spectrogram image (see FIG. 6).

[0090] And the signal generation unit (140) of the input device can generate a control signal to control an external device in response to the output command (step S1007).

[0091] FIG. 11 is a flowchart of an operation in which an input device according to one embodiment processes input through a plurality of learning models.

[0092] Referring to FIG. 11, the flow of operation in which an input device according to one embodiment processes input through a plurality of learning models may be illustrated. An input device according to one embodiment may include two or more learning models and generate a signal to control an external electronic device by outputting a single command from two or more learning models.

[0093] The sound receiving unit of the input device can acquire sound as an input signal through the object (D) (step S1101). The sound receiving unit can also electrically convert the sound to generate an analog input signal (step S1103).

[0094] The processing unit of the input device can extract digital sound data from an analog input signal (Step S1105). The analog input signal can be converted into digital data to correspond to the first sound data. At the same time, the preprocessing unit can preprocess the analog input signal (Step S1107). The preprocessed input signal can be converted into digital data to correspond to the second sound data.

[0095] The processing unit of the input device can input the first sound data of the input signal and the second sound data of the preprocessed input signal to a plurality of learning models (step S1109). The first sound data can be input as time-series data or as image data to the first learning model, which is an RNN model among the plurality of learning models (see FIG. 5). The second sound data can be input as image data—spectrogram image—to the second learning model, which is a CNN model among the plurality of learning models (see FIG. 6).

[0096] The processing unit of the input device can output a single command based on the output results of multiple learning models (step S1111). Since the first and second sound data are input to the first and second learning models, output results can be produced from each learning model. The processing unit can generate a single command by analyzing the output results of each learning model. This will be explained later.

[0097] And the signal generation unit (140) of the input device can generate a control signal to control an external device in response to the output command (step S1113).

[0098] FIG. 12 is a diagram illustrating the operation of an input device according to one embodiment processing input through a plurality of learning models.

[0099] Referring to FIG. 12, a process in which a single command is generated from the output of a plurality of learning models illustrated in FIG. 11 can be illustrated. When an input device according to one embodiment uses a plurality of learning models, it can learn and process different sound data. In this drawing, it will be described as including two learning models.

[0100] When the sound receiving unit (110) receives sound as an input signal, it can electrically convert it to generate an analog input signal. The analog input signal can be converted into digital data in the processing unit (120) to generate the first sound data. Meanwhile, the analog input signal can be preprocessed—spectrographically converted—in the preprocessing unit (150) and converted into digital data to generate the second sound data. The first sound data can be input to the first learning model (122-1), and the second sound data can be input to the second learning model (122-2). Here, the first learning model (122-1) can be trained to output commands from time-series data such as the first sound data, and the second learning model (122-2) can be trained to output commands from image data such as the second sound data. Accordingly, the first learning model (122-1) can include an RNN model, and the second learning model (122-2) can include a CNN model.

[0101] The first learning model (122-1) can output a first command from the first sound data, and the second learning model (122-2) can output a second command from the second sound data. The processing unit (120) can generate a single command based on the output results—the first and second commands—of the first and second learning models (122-1, 122-2). For example, the processing unit (120) can derive the first or second command as the final command only when the output results of the first and second learning models (122-1, 122-2) are identical. For instance, if both the first and second commands indicate power ON, the final command can be determined as power ON. If both the first and second commands indicate power OFF, the final command can be determined as power OFF. In addition to the result of the first learning model (122-1) learned with time series data in an input device according to one embodiment, the result of the second learning model (122-2) learned with image data of a different type from time series data is used, thereby increasing the accuracy and clarity of the process of deriving commands from input signals.

[0102] FIG. 13 is a flowchart of an operation in which an input device according to one embodiment processes a plurality of inputs through a learning model, and FIG. 14 is an example diagram of a plurality of sound data input to an input device according to one embodiment.

[0103] Referring to FIG. 13, the flow of operation in which an input device according to one embodiment processes multiple inputs through a learning model may be illustrated. An input device according to one embodiment may include a single learning model and output a single command corresponding to multiple input signals.

[0104] The sound receiving unit of the input device can acquire a plurality of sounds as input signals through the target object (D) (Step S1301). The sound receiving unit can electrically convert the plurality of sounds to generate a plurality of input signals in an analog form (Step S1303). Referring to FIG. 14, an example of a plurality of input signals in an analog form generated by the sound receiving unit of the input device according to one embodiment may be illustrated. A first input signal (top) and a second input signal (bottom) are illustrated side by side, and each may represent a sound transmitted through the target object by the first and second sound receiving units. A single stimulus from a user (UE) can generate a plurality of input signals as shown in the figure by the plurality of sound receiving units.

[0105] Returning to Fig. 13, the processing unit of the input device can extract multiple digital sound data from multiple analog input signals (step S1305).

[0106] The processing unit of the input device can input multiple sound data into the learning model (step S1307). The learning model receives a pair of multiple sound data as a single input data and can output a single command corresponding to this pair of multiple sound data (step S1309). The learning model can be pre-trained to output a single command corresponding to this pair of multiple sound data. Then, the signal generation unit can generate a control signal to execute the single command and send it to an external electronic device (step S1311).

[0107] Aspects of the subject described herein may be described in the context of computer-executable instructions, such as program modules, executed on a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., which perform specific tasks or specific abstract data types. Aspects of the subject described herein may also be executed in distributed computing environments where tasks are performed by remote processing devices linked via a communication network. In a distributed computing environment, program modules may be located on both local and remote computer storage media, including memory storage devices.

[0108] Alternatively or additionally, the functions described herein may be performed at least partially by one or more hardware logic components. As examples, not limitations, exemplary types of hardware logic components that may be used include field-programmable gate arrays (FPGAs), program-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), etc.

[0109] Meanwhile, the disclosed embodiments may be implemented in the form of a recording medium that stores a program and / or instructions executable by a computer. The instructions may be stored in the form of program code and, when executed by a processor, may generate a program module to perform the operation of the disclosed embodiments. The recording medium may be implemented as a computer-readable recording medium.

[0110] Computer-readable recording media include all types of recording media that store instructions that can be decoded by a computer. Examples include ROM (read-only memory), RAM (random access memory), magnetic tape, magnetic disk, flash memory, optical data storage devices, etc.

[0111] The scope of protection of the present invention is not limited to the description and expression of the embodiments explicitly described above. Furthermore, it is added once again that the scope of protection of the present invention cannot be limited by obvious changes or substitutions in the technical field to which the present invention belongs.

Claims

1. An input device coupled to or in contact with an object that is a medium for transmitting sound, A sound receiving unit that acquires the above sound as an input signal through the above object and electrically converts the input signal; A processing unit that receives the input signal from the sound receiving unit, inputs sound data corresponding to the input signal into a pre-trained learning model, and outputs a command corresponding to the sound data; and A signal generation unit comprising a signal generating unit for generating a control signal to execute the above command Machine learning-based input device.

2. In Paragraph 1, The above sound is generated when the surface of the object is stimulated and is transmitted to the sound receiving unit through the object. Machine learning-based input device.

3. In Paragraph 2, The above stimulus is electrical or non-electric, and The above object is composed of a conductive material or a non-conductive material. Machine learning-based input device.

4. In Paragraph 1, The above learning model is composed of an artificial neural network including an RNN (recurrent neural network) or a CNN (convolutional neural network). Machine learning-based input device.

5. In Paragraph 4, The above learning model is composed of a combination of an RNN artificial neural network and a CNN artificial neural network. Machine learning-based input device.

6. In Paragraph 1, The above learning model includes a plurality of learning models, and The processing unit outputs the command based on the output results of the plurality of learning models for the sound data. Machine learning-based input device.

7. In Paragraph 1, The above sound receiving unit includes a plurality of sound receiving units that each acquire the sound as an input signal through the object and electrically convert the plurality of input signals. A plurality of input signals are received from the plurality of sound receivers, and a plurality of sound data corresponding to the plurality of input signals are input into the learning model to output a single command corresponding to the plurality of sound data. Machine learning-based input device.

8. In Paragraph 1, It includes a communication unit that communicates by being connected to an electronic device via a wired or wireless network and transmits the control signal to the electronic device. The above control signal controls the electronic device so that the electronic device executes the above command. Machine learning-based input device.

9. Input device according to paragraph 1; and It includes an object to which the above-mentioned input device is coupled on one side, and The above sound is generated when the surface of the object is stimulated and is transmitted to the sound receiving unit through the object. Machine learning-based input system.

10. A method for processing input in an input device coupled to an object as a medium for transmitting sound, A step of acquiring the above sound as an input signal through the above object; A step of electrically converting the above input signal; A step of generating a command corresponding to the input signal from the input signal, wherein the input signal is received, sound data corresponding to the input signal is input into a pre-trained learning model to output a command corresponding to the sound data; and A step comprising generating a control signal for executing the above command A method for processing input based on machine learning.

11. In Paragraph 10, The above sound is generated when the surface of the object is stimulated and is transmitted to the sound receiving unit through the object. A method for processing input based on machine learning.

12. A computer-readable medium storing a program that performs the method of paragraph 10.