Electronic device and method for identifying speech signals for invocation in various languages, and non-transitory computer readable storage medium

The neural network-based system in electronic devices effectively identifies voice signals in multiple languages by segmenting and correlating voice and text data, improving speech recognition and human-machine interaction.

WO2026116542A1PCT designated stage Publication Date: 2026-06-04NCSOFT CORP

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
NCSOFT CORP
Filing Date
2024-11-28
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Existing electronic devices struggle to accurately identify voice signals in multiple languages due to limitations in speech recognition systems, particularly in training without labeled voice information.

Method used

The electronic device employs a neural network-based system that includes multiple neural networks to process voice signals, segmenting them into phonetic symbols and determining the correspondence between voice and text, allowing for the identification of predicted phonemes without requiring explicit voice labels.

Benefits of technology

This approach enables effective voice signal recognition across various languages, enhancing human-machine interaction capabilities by accurately identifying voice commands and controlling devices through trained neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024019211_04062026_PF_FP_ABST
    Figure KR2024019211_04062026_PF_FP_ABST
Patent Text Reader

Abstract

A processor of an electronic device, according to one embodiment, may acquire: from a first neural network into which a speech signal received through a microphone has been input, a first sequence of portions of the speech signal corresponding to designated frame units; from a second neural network into which designated text has been input, a second sequence of one or more phonetic symbols for the designated text; from a third neural network into which the first sequence and the second sequence have been input, a first dataset indicating the degree to which each of the one or more phonetic symbols corresponds to each of the portions of the speech signal; and, from a pattern discriminator into which the first dataset has been input, predicted phonemes of the portions of the speech signal.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device, method, and non-transient computer-readable storage medium for identifying voice signals called in various languages

[0001] The present disclosure relates to an electronic device, a method, and a non-transient computer-readable storage medium for identifying voice signals called in various languages.

[0002] Recently, the proliferation of various types of electronic devices, such as smartphones, tablet PCs, wireless earphones, and / or smartwatches, has been expanding. These electronic devices can provide functions for interacting with users based on a human-machine interface (HMI).

[0003] The electronic device can interact with the user through a speech recognition service that uses an artificial intelligence system to provide a response to the user's voice or to control the electronic device based on the user's voice.

[0004] An electronic device according to one embodiment can identify predicted phonemes represented by parts of a voice signal received through a microphone using a neural network.

[0005] An electronic device according to one embodiment can train an audio encoder based on predicted phonemes without voice information (or label information) regarding parts of a voice signal received through a microphone.

[0006] The technical problems to be solved in this document are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this invention belongs from the description below.

[0007] An electronic device according to one embodiment may include a microphone, a memory, and a processor. The processor may be configured to obtain a first sequence of parts of a voice signal corresponding to a designated frame unit from a first neural network into which a voice signal received through the microphone is input. The processor may be configured to obtain a second sequence of one or more phonetic symbols for a designated text from a second neural network into which a designated text is input. The processor may be configured to obtain a first data set indicating the degree to which each of the one or more phonetic symbols corresponds to each of the parts of the voice signal from a third neural network into which the first sequence and the second sequence are input. The processor may be configured to obtain prediction phonemes of the parts of the voice signal from a pattern discriminator into which the first data set is input.

[0008] A method performed by an electronic device according to one embodiment may include an operation of obtaining a first sequence of parts of a voice signal corresponding to a designated frame unit from a first neural network into which a voice signal received through the microphone is input. The method may include an operation of obtaining a second sequence of one or more phonetic symbols for a designated text from a second neural network into which a designated text is input. The method may include an operation of obtaining a first data set from a third neural network into which the first sequence and the second sequence are input, wherein each of the one or more phonetic symbols corresponds to each of the parts of the voice signal. The method may include an operation of obtaining predicted phonemes of the parts of the voice signal from a pattern discriminator into which the first data set is input.

[0009] A non-transient computer-readable storage medium according to one embodiment may store one or more programs. The one or more programs may include instructions that cause the electronic device to obtain a first sequence of parts of the voice signal corresponding to a designated frame unit from a first neural network to which a voice signal received through the microphone is input when executed by the processor of the electronic device. The one or more programs may include instructions that cause the electronic device to obtain a second sequence of one or more phonetic symbols for a designated text from a second neural network to which a designated text is input when executed by the processor of the electronic device. The one or more programs may include instructions that cause the electronic device to obtain a first data set indicating the degree to which each of the one or more phonetic symbols corresponds to each of the parts of the voice signal from a third neural network to which the first sequence and the second sequence are input when executed by the processor of the electronic device. The above one or more programs may include instructions that cause the electronic device to acquire predicted phonemes of the parts of the voice signal from a pattern discriminator into which the first data set is input when executed by a processor of the electronic device.

[0010] An electronic device according to one embodiment can identify predicted phonemes represented by parts of a voice signal received through a microphone using a neural network.

[0011] An electronic device according to one embodiment can train an audio encoder based on predicted phonemes without voice information (or label information) regarding parts of a voice signal received through a microphone.

[0012] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs from the description below.

[0013] FIG. 1 illustrates an example of a block diagram of an electronic device according to one embodiment.

[0014] FIG. 2 illustrates an example for explaining a neural network obtained from a set of parameters stored in memory by an electronic device according to one embodiment.

[0015] FIG. 3a illustrates an example of an operation in which an electronic device according to one embodiment identifies a voice signal corresponding to a specified text through a neural network.

[0016] FIG. 3b illustrates an example of a block diagram of a second neural network of an electronic device according to one embodiment.

[0017] FIG. 3c illustrates an example of a block diagram of a tripon module of a second neural network of an electronic device according to one embodiment.

[0018] FIG. 3d illustrates an example of a block diagram of a third neural network of an electronic device according to one embodiment.

[0019] FIG. 4 illustrates an example of an operation in which an electronic device according to one embodiment learns a neural network to identify a voice signal corresponding to a specified text.

[0020] FIG. 5 illustrates an example of an operation in which an electronic device according to one embodiment infers a voice signal corresponding to a specified text.

[0021] FIG. 6 illustrates an example of a user interface for setting a specified text in an electronic device according to one embodiment.

[0022] FIG. 7 illustrates an example of a user interface for indicating whether a voice signal according to one embodiment corresponds to a specified text.

[0023] FIG. 8 illustrates an example of a state in which an electronic device according to one embodiment receives a voice signal to perform a specified function.

[0024] FIG. 9 illustrates an example of a flowchart showing the operation of an electronic device according to one embodiment.

[0025] The electronic device according to the various embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, an electronic device, or a home appliance. The electronic device according to the embodiments of this document is not limited to the devices described above.

[0026] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, each of phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish a component from another component and do not limit the components in any other aspect (e.g., importance or order). Where any component (e.g., the first) is referred to as "coupled" or "connected" to another component (e.g., the second), with or without the terms "functionally" or "communicationally," it means that said component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.

[0027] The term "module" as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0028] Various embodiments of this document may be implemented as software (e.g., a program) comprising one or more instructions stored in a storage medium (e.g., internal memory or external memory) readable by a machine. For example, the processor of the machine may call at least one of the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, "non-transitory" simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily in the storage medium.

[0029] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0030] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations among the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

[0031] FIG. 1 illustrates an example of a block diagram of an electronic device according to one embodiment. Referring to FIG. 1, an electronic device (101) according to one embodiment may include a terminal owned by a user. The terminal may include a personal computer (PC), such as a laptop and a desktop, a smartphone, a smartpad, a tablet PC, a smartwatch, and a smart accessory such as a head-mounted device (HMD).

[0032] Referring to FIG. 1, an electronic device (101) according to one embodiment may include at least one of a processor (110), a memory (120), a microphone (160), or a display (170). The processor (110), the memory (120), the microphone (160), and the display (170) may be electrically and / or operably coupled with each other by an electronic component such as a communication bus. The type and / or number of hardware components included in the electronic device (101) are not limited to those shown in FIG. 1. For example, the electronic device (101) may include only some of the hardware components shown in FIG. 1. The elements within the memory described below (e.g., layers and / or a plurality of neural networks (130)) may be in a logically separated state. However, they are not limited thereto.

[0033] A processor (110) of an electronic device (101) according to one embodiment may include a hardware component for processing data based on one or more instructions. The hardware component for processing data may include, for example, an arithmetic and logic unit (ALU), a field programmable gate array (FPGA), and / or a central processing unit (CPU). The number of processors (110) may be one or more. For example, the processor (110) may have the structure of a multi-core processor such as a dual core, a quad core, or a hexa core.

[0034] A memory (120) of an electronic device (101) according to one embodiment may include a hardware component for storing data and / or instructions that are input and / or output to a processor (110). The memory (120) may include, for example, volatile memory such as random-access memory (RAM) and / or non-volatile memory such as read-only memory (ROM). Volatile memory may include, for example, at least one of dynamic RAM (DRAM), static RAM (SRAM), cache RAM, and pseudo SRAM (PSRAM). Non-volatile memory may include, for example, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, hard disk, compact disk, and embedded multi-media card (eMMC).

[0035] A microphone (160) of an electronic device (101) according to one embodiment can receive a voice signal (e.g., a user's voice). The microphone (160) may be placed in a part of the housing of the electronic device (101). The microphone (160) may be referred to as a feedback microphone in a side positioned adjacent to a speaker (not shown). The microphone (160) may be placed in a part of the housing containing a sensor (not shown) of the electronic device (101). The microphone (240) may be referred to as a feedforward microphone in a side positioned facing the outside of the electronic device (101), but is not limited thereto.

[0036] In one embodiment, a display (170) of an electronic device (101) can output visualized information to a user of the electronic device (101). For example, the display (170) can be controlled by a processor (110) including a circuit such as a GPU (graphic processing unit) to output visualized information to a user. The display (170) may include a flat panel display (FPD) and / or electronic paper. The FPD may include a liquid crystal display (LCD), a plasma display panel (PDP), and / or one or more light emitting diodes (LEDs). The LED may include an organic LED (OLED).

[0037] According to one embodiment, within the memory (120) of the electronic device (101), one or more instructions (or commands) representing operations and / or operations to be performed on data by the processor (110) of the electronic device (101) may be stored. A set of one or more instructions may be referred to as firmware, an operating system, a process, a routine, a sub-routine, and / or an application. For example, the electronic device (101), and / or the processor (110), may perform at least one of the operations of FIG. 9 when a set of a plurality of instructions distributed in the form of an operating system, firmware, a driver, and / or an application is executed. In the following, the statement that an application is installed within an electronic device (101) may mean that one or more instructions provided in the form of an application are stored in memory (120), and that said one or more applications are stored in an executable format (e.g., a file having an extension specified by the operating system of the electronic device (101)) by the processor (110). For example, an application may include a program and / or library related to a service provided to a user.

[0038] A set of parameters associated with a plurality of neural networks (130) may be stored in the memory (120) of an electronic device (101) according to one embodiment. The plurality of neural networks (130) are a recognition model implemented in software or hardware that mimics the computational power of a biological system using a large number of artificial neurons (or nodes). The plurality of neural networks (130) can perform human cognitive actions or learning processes through artificial neurons. The parameters associated with the plurality of neural networks (130) may represent, for example, weights assigned to a plurality of nodes included in the plurality of neural networks (130) and / or connections between said plurality of nodes. The structure of the plurality of neural networks (130) represented by the set of parameters stored in the memory (120) of the electronic device (101) according to one embodiment will be described later through FIG. 2. The number of multiple neural networks (130) stored in memory (120) is not limited to that shown in FIG. 1, and sets of parameters corresponding to each of the multiple neural networks may be stored in memory (120).

[0039] For example, an electronic device (101) can use a first neural network (131) to divide a voice signal received through a microphone based on a specified period. The electronic device (101) can obtain a first sequence of feature information for parts of a voice signal corresponding to a specified frame unit from the first neural network into which the voice signal is input. The first neural network (131) may be an example of a pre-trained neural network for obtaining parts of a voice signal from a voice signal based on a specified period. The first neural network (131) may be referred to as an audio encoder in that it obtains a first sequence of feature information for parts of a voice signal.

[0040] For example, the electronic device (101) can obtain a second sequence of one or more phonetic symbols for a specified text included in text information (150) by using the second neural network (132). For example, the electronic device (101) can obtain a second sequence of a specified number of consecutive phonemes among one or more phonemes corresponding to each of the one or more phonetic symbols by using the second neural network (132). In one embodiment, the phonetic symbols may be the International Phonetic Alphabet (IPA). The second neural network (132) may be an example of an IPA-based fully trainable neural network for isolating the specified text into phoneme units using the phonetic symbols corresponding to the specified text. The second neural network (132) may be referred to as a text encoder (or phoneme encoder) in that it obtains a second sequence of one or more phonetic symbols for the specified text.

[0041] For example, the electronic device (101) can obtain data sets indicating the degree to which a voice signal and a designated text correspond from a third neural network (133) into which a first sequence and a second sequence are input. The electronic device (101) can obtain data sets indicating the degree to which each of one or more phonetic symbols included in the second sequence corresponds to a voice signal from the third neural network (133) into which a first sequence and a second sequence are input. The third neural network (133) may be referred to as an integrated model or a multimodal processing model in that it obtains data sets through the voice signal and the designated text. The third neural network (133) may be referred to as a pattern extractor in that it obtains data sets through the voice signal and the designated text.

[0042] For example, the electronic device (101) may obtain a first parameter indicating whether it is a speech signal corresponding to a specified text from a fourth neural network (134) into which one or more data sets are input to indicate the relationship between a first sequence and a second sequence. The fourth neural network (134) may be referred to as a discriminator and / or a discriminating model in terms of determining whether it is a speech signal corresponding to a specified text.

[0043] For example, the electronic device (101) may obtain a second parameter for identifying (or determining) whether each part of a speech signal is matched to each of one or more phonetic symbols based on the number of one or more phonetic symbols corresponding to a specified text from a fifth neural network (135) into which one or more data sets are input. The fifth neural network (135) may be referred to as a discriminator and / or a discriminative model in terms of determining whether each part of a speech signal is matched to each of one or more phonetic symbols.

[0044] For example, the electronic device (101) may obtain a third parameter for identifying (or, discriminating) one or more prediction phonemes corresponding to a speech signal from a sixth neural network (136) into which one or more data sets are input. The sixth neural network (136) may be referred to as a discriminator and / or a discriminative model in terms of discriminating one or more prediction phonemes corresponding to a speech signal.

[0045] For example, the electronic device (101) can train a plurality of neural networks (130) using a first parameter, a second parameter, and a third parameter. For example, the electronic device (101) can train a first neural network (131), a second neural network (132), and / or a third neural network (133) using a first parameter, a second parameter, and a third parameter. The electronic device (101) can tune the first neural network (131) using the third parameter. The electronic device (101) can tune the first neural network (131) using the loss between the third parameter and the phonemes represented by parts of the speech signal. An operation in which the electronic device (101) trains at least a portion of the plurality of neural networks (130) is described later in FIGS. 3a to 4.

[0046] A calling software application (140) according to one embodiment may be an example of a software application for displaying to a user whether a voice signal corresponding to a designated text has been received. The calling software application (140) may be used to execute at least one function by utilizing a plurality of neural networks (130). The electronic device (101) may perform at least one operation based on the reception of a voice signal calling a designated text referring to the electronic device (101) by utilizing the calling software application (140). An example of an operation in which the electronic device (101) executes at least one function based on the execution of the calling software application (140) is described later in FIGS. 6 to 8.

[0047] For example, text information (150) may include a training data set for training multiple neural networks (130). Text information (150) may include information about one or more phonetic symbols (or, the International Phonetic Alphabet, IPA). Phonetic symbols may refer to codes recorded by representing voice in a form visible to the user. For example, text information (150) may include information about a call word for a user of the electronic device (101) to use the functions of the electronic device (101) based on the execution of a call software application (140). After identifying a voice signal corresponding to the call word, the electronic device (101) may wait to receive another voice signal indicating the initiation of the functions of the electronic device (101). However, it is not limited thereto. The electronic device (101) may provide a user interface (e.g., the user interface (610) of FIG. 6) for changing the call word to the user.

[0048] An electronic device (101) according to one embodiment as described above can identify a call word not included in a training data set for training by using the trained neural networks (130) after training a plurality of neural networks (130). The electronic device (101) can provide functions related to human-machine interaction (HMI) to a user by using the plurality of neural networks (130) capable of identifying a call word not included in the training data set.

[0049] FIG. 2 illustrates an example for explaining a neural network obtained from a set of parameters stored in memory by an electronic device according to one embodiment. Referring to FIG. 2, at least a portion of a plurality of neural networks (130) may include a plurality of layers. For example, a plurality of neural networks (130) may include an input layer (210), one or more hidden layers (220), and an output layer (230). The input layer (210) may receive a vector representing input data (e.g., a vector having elements corresponding to the number of nodes included in the input layer (210)). Signals generated at each of the nodes within the input layer (210), which are generated by the input data, may be transmitted from the input layer (210) to the hidden layers (220). The output layer (230) may generate output data of the plurality of neural networks (130) based on one or more signals received from the hidden layers (220). The above output data may include, for example, a vector having elements corresponding to the number of nodes included in the output layer (230).

[0050] Referring to FIG. 2, one or more hidden layers (220) may be located between the input layer (210) and the output layer (230) and may convert input data transmitted through the input layer (210) into a predictable value. The input layer (210), one or more hidden layers (220), and the output layer (230) may include multiple nodes. The one or more hidden layers (220) are not limited to the illustrated feedforward-based topology and may be, for example, convolution filters or fully connected layers in a convolutional neural network (CNN), or various types of filters or layers grouped based on special functions or features. In one embodiment, the one or more hidden layers (220) may be layers based on a recurrent neural network (RNN) in which the output value is input back into the hidden layer at the current time. For example, an input layer (210), one or more hidden layers (220) and / or an output layer (230) may be some layers of a transformer model. A plurality of neural networks (130) according to one embodiment may include a plurality of hidden layers (220) to form a deep neural network. Training a deep neural network is called deep learning. Among the nodes of the plurality of neural networks (130), a node included in the hidden layers (220) is referred to as a hidden node.

[0051] Nodes included in the input layer (210) and one or more hidden layers (220) may be connected to each other through connecting lines having connection weights, and nodes included in the hidden layer and output layer may also be connected to each other through connecting lines having connection weights. Tuning and / or training a plurality of neural networks (130) may mean changing the connection weights between nodes included in each of the layers (e.g., input layer (210), one or more hidden layers (220), and output layer (230)) included in the plurality of neural networks (130). Tuning of the plurality of neural networks (130) may be performed, for example, based on supervised learning and / or unsupervised learning.

[0052] An electronic device according to one embodiment may tune a plurality of neural networks (130) based on reinforcement learning in unsupervised learning. For example, the electronic device may change policy information used by the plurality of neural networks (130) to control an agent based on the interaction between the agent and the environment. The policy information is a rule by which the electronic device determines the agent's actions within the environment using the neural network, and the electronic device may change the policy information of the neural network by training the neural network based on the interaction between the agent and the environment. For example, the policy information may be changed to determine the optimal action and / or sequence of actions to achieve the reward and / or goal obtainable by the agent. An electronic device according to one embodiment may cause the change of the policy information by the plurality of neural networks (130) to maximize the agent's goal and / or reward due to the interaction.

[0053] FIG. 3a illustrates an example of an operation in which an electronic device according to one embodiment identifies a voice signal corresponding to a specified text through a neural network. FIG. 3b illustrates an example of a block diagram of a second neural network of an electronic device according to one embodiment. FIG. 3c illustrates an example of a block diagram of a tripon module of the second neural network of an electronic device according to one embodiment. FIG. 3d illustrates an example of a block diagram of a third neural network of an electronic device according to one embodiment.

[0054] The electronic device (101) of FIG. 3a may include the electronic device (101) of FIG. 1.

[0055] Referring to FIG. 3a, an electronic device (101) according to one embodiment can receive a voice signal (310) from a user using a microphone (160). For example, the voice signal (310) may be 'Hallow'.

[0056] For example, the electronic device (101) can phonemicize the voice signal (310) using the first neural network (131). The electronic device (101) can divide the voice signal (310) based on a specified period. The electronic device (101) can acquire parts (310-1, 310-2, 310-3, 310-4, 310-5) of the voice signal (310) based on a specified frame unit. For example, the specified frame unit may be identified based on phonemes corresponding to the voice signal (310). For example, the specified frame unit may be changed based on the length of the specified text, but is not limited thereto. For example, if the voice signal (310) is 'Hallow', one or more phonemes corresponding to parts (310-1, 310-2, 310-3, 310-4, 310-5) of the voice signal (310) are, It may include, however, not be limited to.

[0057] For example, the electronic device (101) can convert the voice signal (310) into first voice data (311) based on frequency. The electronic device (101) can convert the voice signal (310) into second voice data (312) based on time. The electronic device (101) can input the first voice data (311) and the second voice data (312) into a first neural network (131) to obtain a first sequence (320). The order and / or number of the first sequence (320) may vary according to the embodiment based on parts (310-1, 310-2, 310-3, 310-4, 310-5) of the voice signal (310) separated based on a specified frame unit. The first sequence (320) may be an example of a matrix based on N dimensions. The first sequence (320) may include feature information based on multiple dimensions. The first sequence (320) may include vector parameters embedded from a voice signal (310).

[0058] For example, the electronic device (101) may process the first voice data (311) using a first operator (131-1) and / or a second operator (131-2) included in the first neural network (131). The first operator (131-1) may be an example of an operator used to identify frequency-based changes in the first voice data (311) over time.

[0059] In one embodiment, the first operator (131-1) may include a short-time Fourier transform (STFT). In one embodiment, the electronic device (101) may convert the first voice data (311) through the STFT within the first operator (131-1).

[0060] In one embodiment, the second operator (131-2) may include a one-dimensional convolution operator (e.g., Conv1D), batch normalization (BN), and / or an activation function. The activation function may include a non-linear function such as a ReLu (rectified linear unit) function or a sigmoid operation. The number of second operators (131-2) may be one or more.

[0061] In one embodiment, the electronic device (101) may sequentially apply a one-dimensional convolution operator (e.g., Conv1D), batch normalization (BN), and / or activation function in a second operator (131-2) to the output of a first operator (131-1) (or, first voice data (311) converted via STFT).

[0062] For example, if there are two second operators (131-2), the kernel size of the convolution operator based on the one dimension of the first second operator (131-2) (e.g., Conv1D) may be 3 and the stride may be 2. For example, if there are two second operators (131-2), the kernel size of the convolution operator based on the one dimension of the second second operator (131-2) (e.g., Conv1D) may be 3 and the stride may be 1.

[0063] For example, the electronic device (101) can process the second voice data (312) using the third operator (131-3) and / or the fourth operator (131-4) included in the first neural network (131).

[0064] For example, the third operator (131-3) may include a pre-trained embedder to vectorize the second voice data (312). For example, the third operator (131-3) may have a specified time window (e.g., 775 milliseconds). The electronic device (101) may compute a feature vector of a specified dimension (e.g., 96 dimensions) for the second voice data (312) through the third operator (131-3) at a specified time interval (e.g., 80 milliseconds).

[0065] For example, the fourth operator (131-4) may include a time-based one-dimensional convolution operator (e.g., TConv1D). The fourth operator (131-4) may include an activation function. The fourth operator (131-4) may be used to identify one-dimensional sequence data, but is not limited thereto.

[0066] For example, the electronic device (101) can obtain first sequence data by processing first voice data (311) using a first operator (131-1) and a second operator (131-2). The electronic device (101) can obtain second sequence data by processing second voice data (312) using a third operator (131-3) and a fourth operator (131-4).

[0067] For example, an electronic device (101) can obtain a first sequence (320) using first sequence data and second sequence data. The electronic device (101) can obtain the first sequence (320) based on calculating the first sequence data and second sequence data through a specified operator (e.g., element-wise sum). For example, the electronic device (101) can obtain the first sequence (320) by summing the first sequence data and second sequence data element-wise. In one embodiment, the first sequence (320) is T a XH dimensional audio embeddings (E a It can be. Here, T a can represent the length of the voice signal (310) (or, the number of parts (310-1, 310-2, 310-3, 310-4, 310-5)) of the voice signal (310). H can represent the embedding dimension.

[0068] According to one embodiment, an electronic device (101) can obtain a second sequence (330) from a second neural network (132) into which a designated text (315) included in text information (e.g., text information (150) of FIG. 1) is input. For example, the electronic device (101) can obtain the second sequence (330) using one or more phonetic symbols corresponding to the designated text (315). The electronic device (101) can segment the designated text (315) based on one or more phonetic symbols. For example, if the designated text (315) is 'Hello', the one or more phonetic symbols corresponding to the designated text (315) are It may include, however, not be limited to.

[0069] For example, the electronic device (101) can change a designated text (315) (or one or more phonetic symbols corresponding to the designated text (315)) into a second sequence (330) using a fifth operator (132-1) and / or a sixth operator (132-2) included in the second neural network (132).

[0070] For example, the fifth operator (132-1) may include an embedding for vectorizing (or data-izing) a specified text (315). For example, referring to FIG. 3b, the fifth operator (132-1) of the second neural network (132) may include an embedding layer (381), fully connected layers (FC), and / or an activation function (e.g., LeakyReLU (leaky rectified linear unit)). For example, the fifth operator (132-1) may be referred to as an embedding module in terms of vectorizing (or data-izing) the specified text (315).

[0071] For example, the data of the specified text (315) processed through the fifth operator (132-1) may be an example of vector data based on multiple dimensions (e.g., two dimensions). For example, the data of the specified text (315) processed through the fifth operator (132-1) is T t XH dimensional text embedding (E t It can be. Here, T t can indicate the length of the specified text (315) (or, the number of one or more phonetic symbols corresponding to the specified text (315)).

[0072] For example, the electronic device (101) can text embedding (E) through the sixth operator (132-2). tBy processing ), data indicating the association between consecutive phonemes can be obtained. For example, referring to FIG. 3b, the sixth operator (132-2) may include a triphone module (384) and / or an element-wise sum operation (385).

[0073] For example, referring to FIG. 3c, the tripon module (384) of the sixth operator (132-2) has a text embedding (E t It may include a batch normalization (BN) layer (384-1) for processing, a two-dimensional convolution operator (e.g., Conv2D) (384-2), and / or an activation function (e.g., GeLU (Gaussian error linear unit)) (384-3). For example, the number of tripon modules (384) of the sixth operator (132-2) may be one or more (e.g., three).

[0074] For example, referring to FIG. 3c, the electronic device (101) through the tripon module (384) of the sixth operator (132-2) text embedding (E t Three consecutive phonemes based on the i-th phoneme among the phonemes of ) (e.g., E t i-1 , E t i , E t i+1 For ), a batch normalization (BN) layer (384-1), a 2D-based convolution operator (e.g., Conv2D) (384-2), and / or an activation function (e.g., GeLU) (384-3) can be applied sequentially.

[0075] For example, the electronic device (101) through the tripon module (384) of the sixth operator (132-2) three consecutive phonemes based on the i-th phoneme (e.g., E t i-1 , E t i , E t i+1For text embeddings concatenated with ), batch normalization (BN) layers (384-1), a 2D-based convolution operator (e.g., Conv2D) (384-2), and / or an activation function (e.g., GeLU) (384-3) may be applied sequentially. Three consecutive phonemes based on the i-th phoneme may include the i-1th phoneme, the i-th phoneme, and the i+1th phoneme. Three consecutive phonemes based on the i-th phoneme (e.g., E t i-1 , E t i , E t i+1 The text embedding formed by concatenating ) is T t It can have three dimensions of XHX. For example, text embeddings (E) processed through the tripon module (384) of the sixth operator (132-2). t ) is T t XH dimensional Triphon embedding (E tri It can be.

[0076] Referring again to FIG. 3b, the electronic device (101) performs an elemental summation operation (385) of the sixth operator (132-2) on the tripon embedding (E tri ) and text embeddings (E t By summing ), a second sequence (330) can be obtained. The second sequence (330) is T t Final text embedding in XH dimensions (E t' It can be ). For example, the final text embedding (E t' ) may include vector values ​​(330-1, 330-2, 330-3, 330-4, 330-5, 330-6) for each of the phonemes.

[0077] An electronic device (101) according to one embodiment may input a first sequence (320) and a second sequence (330) into a third neural network (133) to obtain one or more data sets (340). For example, the electronic device (101) may obtain one or more data sets (340) representing the relationship between the first sequence (320) and the second sequence (330) from the third neural network (133) into which the first sequence (320) and the second sequence (330) are input. The one or more data sets (340) may include information representing the association (or relationship) between each of the first sequence data sets included in the first sequence (320) and each of the second sequence data sets included in the second sequence (330). The one or more data sets (340) may include pattern information between the first sequence (320) and the second sequence (330). In terms of obtaining pattern information between the first sequence (320) and the second sequence (330), the third neural network (133) may be referred to as a pattern extractor.

[0078] In one embodiment, the electronic device (101) comprises a third neural network (133), a first sequence (320) (or, audio embedding (E a )) and the second sequence (330) (or, the final text embedding (E t' A sequence connecting )) (or, connected embeddings (E c Patterns can be extracted for )). For example, linked embeddings (E c ) is (T a + T t ) can have XH dimensions.

[0079] Referring to FIG. 3d, the third neural network (133) may include a self-attention layer (391), a summation and batch normalization layer (392), a fully connected layer (FC) (393), and / or a summation and batch normalization layer (394). For example, the electronic device (101) may have a final connected embedding (E) through Equation 1 below, which represents the self-attention layer (391) and the summation and batch normalization layer (392) of the third neural network (133). c' ) can be operated on.

[0080]

[0081] In Equation 1, BN(·) can represent a batch normalization operation. In Equation 1, Self-Attention(·) can represent a self-attention operation.

[0082] For example, the electronic device (101) represents the joint embedding (E) through the following mathematical formula 2, which represents the fully connected layer (FC) (393) and the summation and batch normalization layer (394) of the third neural network (133). j )(or, one or more data sets (340)) can be computed.

[0083]

[0084] In mathematical equation 2, W c' can represent the weights of operations for a fully connected layer (FC). W c' can be a lower triangular mask (or lower triangular matrix). W c' may be referred to as an attention mask. Wc' may be a learnable weight that ensures the monotonicity constraint of the speech signal (310) while preserving all information about the specified text (315). For example, joint embeddings (E j )(or, one or more data sets (340)) is (Ta + T t ) can have XH dimensions.

[0085] Among one or more data sets (340) according to one embodiment, the first data set (340-1) may indicate the degree of association of the second sequence (330) with the first sequence (320). For example, the first data set (340-1) may include information indicating the degree to which each of one or more phonetic symbols corresponds to each of the parts (310-1, 310-2, 310-3, 310-4, 310-5) of the voice signal (310). For example, the first data set (340-1) may include information indicating the degree to which the entire designated text (315) corresponds to the entire voice signal (310). For example, among one or more data sets (340), the second data set (340-2) may indicate the degree of association of the first sequence (320) with the second sequence (330). The second data set (340-1) may include information indicating the relationship between one or more phonetic symbols and parts (310-1, 310-2, 310-3, 310-4, 310-5) of the speech signal (310). For example, the second data set (340-2) may include information indicating the degree to which each of the parts (310-1, 310-2, 310-3, 310-4, 310-5) of the speech signal (310) corresponds to each of the one or more phonetic symbols. However, it is not limited thereto.

[0086] An electronic device (101) according to one embodiment can obtain a first parameter (350) from a fourth neural network (134) into which one or more data sets (340) are input. The first parameter (350) (or, P utt) may include a first probability (350-1) indicating the degree to which a voice signal (310) and a designated text (315) correspond. The first probability (350-1) may indicate whether the voice signal (310) matches the designated text (315). The first probability (350-1) may indicate the degree of matching of the utterance level between the voice signal (310) and the designated text (315). For example, the first probability (350-1) may indicate the degree of matching between 'Hallow' of the voice signal (310) and 'Hello' of the designated text (315). For example, referring to FIG. 3a, the first probability (350-1) may indicate that 'Hallow' of the voice signal (310) and 'Hello' of the designated text (315) do not match each other (e.g., false).

[0087] For example, the fourth neural network (134) may include a seventh operator (134-1). The seventh operator (134-1) may include one or more additional operators. The seventh operator (134-1) may include a deep learning framework (e.g., LSTM (long short term memory) or GRU (gated recurrent unit)), a fully connected layer, and / or an activation function (e.g., a sigmoid function).

[0088] For example, the electronic device (101) can identify whether the voice signal (310) matches the specified text (315) by using the first parameter (350) obtained through the seventh operator (134-1).

[0089] For example, a plurality of neural networks (130) can be trained based on a result (380-1) indicating whether a voice signal (310) corresponds to a specified text (315).

[0090] An electronic device (101) according to one embodiment receives from a fifth neural network (135) in which at least some of one or more data sets (340) (e.g., a second data set (340-2)) is input, a second parameter (360) (or, P phon You can obtain ).

[0091] The second parameter (360) may include one or more second probabilities (360-1, 360-2, 360-3, 360-4, 360-5) indicating the degree to which parts of the speech signal (310) correspond to each of one or more phonetic symbols. One or more second probabilities (360-1, 360-2, 360-3, 360-4, 360-5) may indicate the degree of phoneme level matching between the speech signal (310) and the specified text (315).

[0092]

[0093]

[0094] For example, the fifth neural network (135) may include an eighth operator (135-1). The eighth operator (135-1) may include a fully connected layer and / or an activation function (e.g., a sigmoid function). The number of one or more second probabilities (360-1, 360-2, 360-3, 360-4, 360-5) in the second parameter (360) may correspond to the number of one or more phonetic symbols corresponding to the specified text (315).

[0095] For example, a plurality of neural networks (130) can be trained based on a result (380-2) indicating whether one or more second probabilities (360-1, 360-2, 360-3, 360-4, 360-5) correspond to one or more phonetic symbols corresponding to the specified text (315).

[0096] For example, the electronic device (101) can identify whether parts of the voice signal (310) (310-1, 310-2, 310-3, 310-4, 310-5) are matched to each of one or more phonetic symbols by using the second parameter (360) obtained through the eighth operator (135-1).

[0097] An electronic device (101) according to one embodiment receives from a sixth neural network (136) in which at least some of one or more data sets (340) (e.g., a first data set (340-01)) is input, a third parameter (370) (or, P ctc ) can be obtained. The third parameter (370) may represent predicted phonemes (370-1, 370-2, 370-3, 370-4, 370-5) corresponding to parts (310-1, 310-2, 310-3, 310-4, 310-5) of the speech signal (310).

[0098] For example, referring to Fig. 3a, the predicted phonemes (370-1, 370-2, 370-3, 370-4, 370-5) can be compared with the ground truth phoneme data (380-3).

[0099]

[0100]

[0101] The sixth neural network (136) may include a ninth operator (136-1). The ninth operator (136-1) may include a fully connected layer and / or an activation function (e.g., a sigmoid function). The number of one or more predicted phonemes (370-1, 370-2, 370-3, 370-4, 370-5) in the third parameter (370) may correspond to the number of parts (310-1, 310-2, 310-3, 310-4, 310-5) of the speech signal (310).

[0102] For example, the electronic device (101) can identify predicted phonemes of parts (310-1, 310-2, 310-3, 310-4, 310-5) of the voice signal (310) using the third parameter (370) obtained through the ninth operator (136-1).

[0103] For example, a plurality of neural networks (130) can be trained based on a result (380-3) indicating whether the predicted phonemes of parts (310-1, 310-2, 310-3, 310-4, 310-5) of the speech signal (310) correspond to the phonemes labeled in the speech signal (310).

[0104] Hereinafter, with reference to FIG. 4, an example of an operation in which an electronic device (101) according to one embodiment trains a plurality of neural networks (130) using a first parameter (350), a second parameter (360), and / or a third parameter (370) is described below.

[0105] FIG. 4 illustrates an example of an operation in which an electronic device according to one embodiment learns a neural network for identifying a voice signal corresponding to a specified text. The electronic device (101) of FIG. 4 may include the electronic device (101) of FIG. 1. A plurality of neural networks (130) of FIG. 4 may include a first neural network (131), a second neural network (132), a third neural network (133), a fourth neural network (134), a fifth neural network (135), and / or a sixth neural network (136).

[0106] An electronic device (101) according to one embodiment can identify a voice signal (310) (e.g., 'NCYA') in general speech data (401). The electronic device (101) can identify a designated text (315) (e.g., 'NCYA') included in text information. The designated text (315) can correspond to a call word for calling the electronic device (101).

[0107] For example, an electronic device (101) can obtain predicted values ​​(440) from a plurality of neural networks (130) into which a voice signal (310) and a specified text (315) are input.

[0108] The predicted values ​​(440) are a first predicted value (440-1) (or, P) indicating the degree of matching of the utterance level between the voice signal (310) and the specified text (315). utt It may include ). The first predicted value (440-1) may have at least one value between 0 and 1. For example, the first predicted value (440-1) may be 0.95. The first predicted value (440-1) may correspond to the first parameter (350) of FIG. 3A.

[0109] The predicted values ​​(440) are a second predicted value (440-2) (or, P) indicating the degree of phoneme level matching between each of one or more phonetic symbols of the specified text (315) and each of the parts (310-1, 310-2, 310-3, 310-4, 310-5) of the speech signal (310). phon It may include ). The second predicted value (440-2) may include partial predicted values ​​corresponding to the number of parts (310-1, 310-2, 310-3, 310-4, 310-5). Each of the partial predicted values ​​in the second predicted value (440-2) may have at least one value between 0 and 1. For example, each of the partial predicted values ​​in the second predicted value (440-2) may be 0.88, ... 0.74. The second predicted value (440-2) may correspond to the second parameter (360) of FIG. 3A.

[0110]

[0111] An electronic device (101) according to one embodiment can set labeling values ​​(450) (or truth phoneme data) for learning a plurality of neural networks (130) based on supervised learning.

[0112] Since the electronic device (101) matches the text corresponding to the voice signal (310) with the designated text (315), the first labeling value (450-1) indicating the degree of matching of the utterance level among the labeling values ​​(450) can be set to a value (e.g., 1) indicating that the voice signal (310) and the designated text (315) are matched.

[0113] Since each of the one or more phonetic symbols of the specified text (315) and each of the parts (310-1, 310-2, 310-3, 310-4, 310-5) of the voice signal (310) are matched, the electronic device (101) can set a second labeling value (450-2) indicating the degree of phoneme-level matching among the labeling values ​​(450) to a value (e.g., 1) indicating that each of the one or more phonetic symbols of the specified text (315) and each of the parts (310-1, 310-2, 310-3, 310-4, 310-5) of the voice signal (310) are matched. For example, partial labeling values ​​within the second labeling value (450-2) indicating the degree of matching at the phoneme level can be set to values ​​(e.g., 1) indicating that each of one or more phonetic symbols of the specified text (315) and each of the parts (310-1, 310-2, 310-3, 310-4, 310-5) of the speech signal (310) are matched.

[0114]

[0115] For example, the electronic device (101) can train a plurality of neural networks (130) so that the predicted values ​​(440) and the labeling values ​​(450) match. For example, the electronic device (101) can train a plurality of neural networks (130) so that the predicted values ​​(440) converge to the labeling values ​​(450). For example, the electronic device (101) can train a plurality of neural networks (130) so that a first predicted value (440-1) (e.g., 0.95) corresponding to a first parameter (e.g., the first parameter (350) in FIG. 3a) matches the first labeling value (450-1). For example, the electronic device (101) may train a plurality of neural networks (130) based on supervised learning so that a second predicted value (440-2) corresponding to a second parameter (e.g., the second parameter (360) in FIG. 3A) matches a second labeled value (450-2). For example, the electronic device (101) may train a plurality of neural networks (130) so that a third predicted value (440-3) corresponding to a third parameter (e.g., the third parameter (370) in FIG. 3A) matches a third labeled value (450-3). For example, the electronic device (101) may obtain parameters for training a plurality of neural networks (130) based on binary cross entropy (BCE) error. For example, an electronic device (101) can obtain parameters for training multiple neural networks (130) based on a single connectionist temporal classification (CTC) error. For example, the BCE error may represent an error at the utterance level (e.g., the error between a first predicted value (440-1) and a first labeled value (450-1)). For example, the BCE error may represent an error at the phoneme level (e.g., the error between a second predicted value (440-2) and a second labeled value (450-2)).For example, the CTC error may represent a phoneme-level audio recognition error (e.g., the error between the third predicted value (440-3) and the third labeled value (450-3)).

[0116] For example, an electronic device (101) can train multiple neural networks (130) using a loss value (470) (or a loss parameter representing the error rate) that represents the difference between the predicted values ​​(440) and the labeled values ​​(450).

[0117] For example, loss value (470) (or, L total ) is the loss between the first predicted value (440-1) and the first labeled value (450-1) (or, L utt May include ). L utt It can be used to identify similarity in the utterance level between the voice signal (310) and the specified text (315).

[0118] For example, loss value (470) (or, L total ) is the loss between the second predicted value (440-2) and the second labeled value (450-2) (or, L phon May include ). L phon It can be used to identify phoneme-level similarity between a voice signal (310) and a specified text (315).

[0119] For example, loss value (470) (or, L total ) is the loss between the third predicted value (440-3) and the third labeled value (450-3) (or, L ctc May include ). L ctc It can be used to improve the performance of the audio encoder (or, the first neural network (131)) without voice information for parts (310-1, 310-2, 310-3, 310-4, 310-5) of the voice signal (310). ctc It can represent a loss of audio recognition at the phoneme level.

[0120] For example, the electronic device (101) can train at least a portion of the first neural network (131) (e.g., second operator (131-2)), at least a portion of the second neural network (132) (e.g., sixth operator (132-6)), the third neural network (133), the fourth neural network (134), and / or the fifth neural network (135) based on a loss value (470). For example, the electronic device (101) can train L of the loss value (470). ctc Based on this, the first neural network (131) can be trained.

[0121] An electronic device (101) according to one embodiment can train a plurality of neural networks (130) using speech data used in everyday life, which is distinct from a training data set for training neural networks. By training a plurality of neural networks (130) using the speech data, the electronic device (101) can accumulate more data. Based on training a plurality of neural networks (130) using speech data, the electronic device (101) can improve the performance of a plurality of neural networks (130). Based on training a plurality of neural networks (130) using speech data, the electronic device (101) can be trained to be more suitable for the user.

[0122] Hereinafter, an example of an electronic device (101) that infers a call word using a plurality of learned neural networks (130) is described with reference to FIG. 5.

[0123] FIG. 5 illustrates an example of an operation in which an electronic device according to one embodiment infers a voice signal corresponding to a designated text. The electronic device (101) of FIG. 5 may include the electronic device (101) of FIG. 1 to FIG. 4. Referring to FIG. 5, a plurality of neural networks (130) according to one embodiment may be an example of a neural network trained to identify whether a voice signal received through a microphone matches a designated text. The plurality of neural networks (130) of FIG. 4 may include a first neural network (131), a second neural network (132), a third neural network (133), and / or a fourth neural network (134).

[0124] An electronic device (101) according to one embodiment may receive a voice signal (510) using a microphone (e.g., the microphone (160) of FIG. 1). The voice signal (510) may correspond to information representing a word, a sentence, and / or utterance.

[0125] For example, based on receiving a voice signal (510), the electronic device (101) can infer (or identify) whether at least a portion of the voice signal (510) corresponds to a designated text (515) using a plurality of neural networks (130). For example, the electronic device (101) can set the designated text (515). The electronic device (101) can receive input from a user to change the designated text (515) using a user interface for setting the designated text (515). Based on receiving the input, the electronic device (101) can set the designated text (515). By changing the designated text (515), the electronic device (101) can at least temporarily refrain from training the plurality of neural networks (130) to identify the voice signal corresponding to the designated text (515). Since the electronic device (101) has trained multiple neural networks (130) using a speech level data set, it can infer whether the received voice signal (510) corresponds to the specified text (515) independently of the training of the multiple neural networks (130) corresponding to the changed specified text (515).

[0126] For example, an electronic device (101) can input a specified text (515) and a voice signal (510) into a plurality of neural networks (130) to obtain a first parameter (520). The electronic device (101) can obtain parts of the voice signal (510) based on a specified frame unit. The electronic device (101) can identify one or more International Phonetic Symbols for the specified text (515).

[0127] The electronic device (101) can obtain a first parameter (520) indicating whether the entire voice signal (510) and the entire specified text (515) correspond. For example, the first parameter (520) can be obtained based on a binary indicator (e.g., a value having 0 or 1). The electronic device (101) can infer the value of the first parameter (520) based on a specified probability distribution for identifying the first parameter (520). The electronic device (101) can obtain a first parameter (520) having a value closer to 1 as the degree to which the entire voice signal (510) and the entire specified text (515) correspond is relatively high.

[0128] The electronic device (101) can determine that it has identified a voice signal (510) corresponding to a set call word based on the first parameter (520) indicating a specified binary indicator (e.g., a value having 1). The electronic device (101) can determine that it has identified a voice signal (510) corresponding to a set call word based on the first parameter (520) indicating a specified probability value greater than or equal to (e.g., 0.9). For example, the electronic device (101) can provide the user with a function corresponding to another voice signal that follows the voice signal (510) based on identifying the voice signal (510) corresponding to the set call word.

[0129] The electronic device (101) may determine that the voice signal (510) corresponding to the set call word is not identified based on the first parameter (520) indicating a different binary indicator (e.g., a value having 0). The electronic device (101) may determine that the voice signal (510) corresponding to the set call word is not identified based on the first parameter (520) indicating less than a specified probability value. For example, the electronic device (101) may ignore the voice signal (510) and other voice signals consecutive to the voice signal (510) based on not identifying the voice signal (510) corresponding to the set call word.

[0130] As described above, an electronic device (101) according to one embodiment can infer whether it has identified a voice signal corresponding to at least some randomly set call words by using a plurality of learned neural networks (130).

[0131] Referring to FIG. 6 below, an example of a user interface for setting a call word according to one embodiment is described below.

[0132] FIG. 6 illustrates an example of a user interface for setting a specified text in an electronic device according to one embodiment. The electronic device (101) of FIG. 6 may include the electronic device (101) of FIG. 1 through FIG. 5. Referring to FIG. 6, in state (600), the electronic device (101) according to one embodiment may display on a display (170) a user interface (610) for setting a call word for calling the electronic device (101) based on the execution of the call software application (140) of FIG. 1. For example, the user interface (610) may include a visual object (620) for obtaining a call word from the user to initiate the execution of the call software application (140). The electronic device (101) may receive a text object (e.g., Envya) representing the call word from the user using the visual object (620). For example, the electronic device (101) may set the text object as a call word based on receiving an input indicating that an icon (630) has been selected. The electronic device (101) may initiate the execution of a call software application (140) based on identifying a voice signal corresponding to the call word through a plurality of neural networks (130) using a microphone. The electronic device (101) may perform at least one function based on the execution of the call software application (140). The electronic device (101) may initiate the execution of at least one function based on receiving another voice signal indicating the name of at least one function based on the execution of the call software application (140). However, it is not limited thereto.

[0133] FIG. 7 illustrates an example of a user interface for indicating whether a voice signal according to one embodiment corresponds to a specified text. The electronic device (101) of FIG. 7 may include the electronic device (101) of FIG. 1 to 6. The electronic device (101) may provide the user with whether a voice signal corresponding to a specified text (720) has been received in a state (700) based on the execution of a calling software application (140).

[0134] An electronic device (101) according to one embodiment may display a user interface (710) on a display to indicate whether a voice signal corresponding to a designated text (720) has been received. For example, the electronic device (101) may guide the user using a visual object (750) whether a voice signal received using a microphone corresponds to the designated text (720). For example, the electronic device (101) may guide the user using visual objects (760) whether each part of the voice signal corresponds to each of one or more phonetic symbols (730). For example, a visual object (750) may correspond to a first parameter (e.g., the first parameter (520) of FIG. 5). For example, when the first parameter is 1, the electronic device (101) may display a visual object (750) indicating that the voice signal and the designated text are matched. For example, visual objects (760) may correspond to one or more second parameters (e.g., one or more second parameters (530) of FIG. 5). For example, the electronic device (101) may display a visual object indicating that when at least one of the one or more second parameters is 0, at least one of the parts of the voice signal corresponding to said at least one and at least one of the one or more phonetic symbols (730) corresponding to said at least one do not match. However, it is not limited thereto.

[0135] An electronic device (101) according to one embodiment may receive an input indicating that a visual object (740) is selected. Based on receiving the input, the electronic device (101) may provide a notification requesting the user to speak using a microphone, but is not limited thereto. The designated text (720) may be included in the text information (150) of FIG. 1.

[0136] As described above, an electronic device (101) according to one embodiment may display a user interface (710) for correcting a user's pronunciation on a display. The electronic device (101) may guide the user on the accuracy of speech in a second language received from a user based on a first language through a plurality of neural networks (130).

[0137] FIG. 8 illustrates an example of a state in which an electronic device according to one embodiment receives a voice signal to perform a specified function. The electronic device (101) of FIG. 8 may include the electronic device (101) of FIG. 1 to FIG. 7. For example, the electronic device (101) may display a user interface (810) representing a game on a display in state (800).

[0138] An electronic device (101) according to one embodiment can identify text information contained within a user interface (810). For example, the electronic device (101) can identify a first nickname (813) (e.g., User A) of a first character (815) representing a user of the electronic device (101), and a second nickname (833) (e.g., User B) of a second character (835) representing another user. For example, the electronic device (101) can receive a voice signal (805) using a microphone. The electronic device (101) can identify text corresponding to the voice signal (805) through a plurality of neural networks (130) of FIG. 1. If the text matches the second nickname (833), the electronic device (101) can transmit data representing the voice signal (805) to an external electronic device corresponding to the second character (835).

[0139] For example, the electronic device (101) can control the first character (815) to identify text indicating an available skill (820) within the user interface (810). The electronic device (101) can control the first character (815) to use the skill (820) when a voice signal (805) received using a microphone matches the text indicating the skill (820). However, it is not limited thereto. Based on receiving a voice signal to control the first character (815) within the user interface (810), the electronic device (101) can perform a function corresponding to the voice signal using the first character (815).

[0140] For example, an electronic device (101) may receive a voice signal (805) representing a sentence through a microphone. The electronic device (101) may identify whether the voice signal (805) representing the sentence corresponds to a specified text through a plurality of neural networks (130). For example, the electronic device (101) may control the first character (815) to use the skill (820) on the second character (835) based on identifying the voice signal (805) representing a sentence (e.g., attack the second character (835) using the skill (820). However, it is not limited thereto.

[0141] As described above, an electronic device (101) according to one embodiment can acquire a voice signal (805) using a microphone while displaying a user interface (810) for providing a game service. The voice signal (805) may correspond to a command for controlling a first character (815) corresponding to a user of the electronic device (101) within the user interface (810). Based on acquiring the voice signal (805), the electronic device (101) can identify a designated text that matches the voice signal (805) and control the first character (815) to perform a function corresponding to the designated text. The electronic device (101) can provide a service that can control the first character (815) based on the user's voice signal. The electronic device (101) can provide a service that can control the first character (815) based on hands-free operation.

[0142] FIG. 9 illustrates an example of a flowchart showing the operation of an electronic device according to one embodiment. The electronic device of FIG. 9 may include the electronic device (101) of FIG. 1. At least one of the operations of FIG. 9 may be performed by the electronic device (101) of FIG. 1. At least one of the operations of FIG. 9 may be controlled by the processor (110) of FIG. 1. Each of the operations of FIG. 9 may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each of the operations may be changed, and at least two operations may be performed in parallel.

[0143] Referring to FIG. 9, in operation 910, a processor according to one embodiment can obtain a first sequence of parts of a voice signal corresponding to a specified frame unit from a first neural network into which a voice signal is input.

[0144] The processor can acquire parts of a voice signal based on a specified time from the voice signal. The processor can acquire parts of the voice signal by dividing the voice signal based on a specified time. The processor can change parts of the voice signal into first voice data (311) of FIG. 3A based on frequency. The processor can change parts of the voice signal into second voice data (312) of FIG. 3A based on time. The processor can acquire a first sequence (e.g., first sequence (320) of FIG. 3A) through a first neural network (e.g., first neural network (131) of FIG. 1) into which the first voice data and the second voice data are input.

[0145] Referring to FIG. 9, in operation 920, a processor according to one embodiment can obtain a second sequence of one or more phonetic symbols for a specified text from a second neural network into which a specified text is input.

[0146] The processor can obtain a second sequence (e.g., the second sequence (330) of FIG. 3a) by using one or more phonetic symbols corresponding to a specified text. The processor can obtain the second sequence based on the number of one or more phonetic symbols.

[0147] Referring to FIG. 9, in operation 930, a processor according to one embodiment can obtain a first data set indicating the degree to which one or more phonetic symbols correspond to each of the parts of a speech signal from a third neural network to which a first sequence and a second sequence are input.

[0148] For example, the processor can identify a first data set indicating the degree to which each of one or more phonetic symbols corresponds to each of the parts of the speech signal among one or more data sets. The processor can input the first data set into a fourth neural network to obtain a first parameter (e.g., the first parameter (350) of FIG. 3a) indicating whether it is a speech signal corresponding to a specified text.

[0149] For example, the processor may identify a second data set indicating the degree to which each part of the speech signal corresponds to each of one or more phonetic symbols among one or more data sets. The processor may input the second data set into a fifth neural network to obtain one or more second parameters (e.g., one or more second parameters (350) of FIG. 3a) having a number equal to the number of one or more phonetic symbols. The one or more second parameters may indicate whether the parts of the speech signal correspond to each of the one or more phonetic symbols.

[0150] Referring to FIG. 9, in operation 940, a processor according to one embodiment can obtain prediction phonemes of parts of a speech signal from a pattern discriminator to which a first data set is input.

[0151] For example, the processor can obtain a first parameter by inputting a first data set and a second data set into a fourth neural network. The first parameter (or, P utt ) may include a first probability indicating the degree to which a voice signal and a specified text correspond. The first probability may indicate the degree of matching of utterance levels between the voice signal and the specified text.

[0152] For example, the processor may input a second data set into a fifth neural network to obtain a second parameter. The second parameter may include one or more second probabilities representing the degree to which parts of the speech signal correspond to each of one or more phonetic symbols. One or more second probabilities may represent the degree of phoneme-level matching between the speech signal and a specified text.

[0153] Subsequently, the processor can train multiple neural networks (or the first neural network (131)) using the first parameter. For example, the processor can train multiple neural networks such that the third parameter matches the third label value (e.g., the third label value (450-3) in FIG. 4). For example, the processor can train multiple neural networks based on a single connectionist temporal classification (CTC) error. For example, the CTC error may represent a phoneme-level audio recognition error (e.g., the error between the third parameter and the third label value (e.g., the third label value (450-3) in FIG. 4).

[0154] As described above, the electronic device (101) may include a microphone (160), a memory (120), and a processor (110). The processor (110) may be configured to obtain a first sequence (320) of parts of the voice signal (310) corresponding to a designated frame unit from a first neural network (131) into which the voice signal (310) received through the microphone (160) is input. The processor (110) may be configured to obtain a second sequence (330) of one or more phonetic symbols for the designated text (315) from a second neural network (132) into which the designated text (315) is input. The processor (110) may be configured to obtain a first data set (340-1) indicating the degree to which each of the one or more phonetic symbols corresponds to each of the parts of the speech signal (310) from a third neural network (133) into which the first sequence (320) and the second sequence (330) are input. The processor (110) may be configured to obtain prediction phonemes (370-1, 370-2, 370-3, 370-4, 370-5) of the parts of the speech signal (310) from a pattern discriminator (e.g., a sixth neural network (136)) into which the first data set (340-1) is input.

[0155] The processor (110) may be configured to obtain connectionist temporal classification (CTC) errors for the predicted phonemes (370-1, 370-2, 370-3, 370-4, 370-5) of the parts. The processor (110) may be configured to train the first neural network (131) based on the CTC errors.

[0156] The second neural network (132) above is a text embedding (E) for the specified text (315).t An embedding module for obtaining ) (e.g., the 5th operator (132-1)), and the text embedding (E t Triphone embeddings (E) representing the association between consecutive phonemes through ) tri It may be configured to include a Triphon module (e.g., a sixth operator (132-2)) for obtaining the Triphon embedding (E tri ) and the above text embedding (E t It can be configured to obtain the second sequence (330) by summing the elements of ).

[0157] The above tripon module (e.g., 6th operator (132-2)) is the text embedding (E t It may be configured to include a batch normalization layer (384-1) for batch normalizing the tripons within ), a convolution operator (384-2) for performing a convolution operation on the batch normalized tripons, and an activation function (384-3) for obtaining the tripon embedding by performing an activation operation on the convolutionally operated tripons. The tripons are the text embedding (E t It includes the i-th phoneme among the phonemes within ), and two phonemes adjacent to the i-th phoneme, wherein i may be an integer greater than or equal to 1 and less than or equal to the number of phonemes.

[0158] The above activation function (384-3) may include a GeLU (Gaussian error linear unit).

[0159] The above embedding module (e.g., fifth operator (132-1)) comprises an embedding layer (381) for embedding one or more phonetic symbols for the specified text (315), a fully connected layer (382) for fully connecting the one or more embedded phonetic symbols, and by performing an activation operation on the one or more fully connected phonetic symbols, the text embedding (E t It can be configured to include an activation function (383) for obtaining ).

[0160] The above activation function (383) may include LeakyReLU (leaky rectified linear unit).

[0161] The third neural network (133) is a connected embedding (E) that connects the first sequence (320) and the second sequence (330). c Final connected embedding (E) for ) c' A self-focusing layer (391) for obtaining ) and the final connected embedding (E c' A joint embedding (E) including the first data set (340-1) for ) j It may be configured to include a fully connected layer (393) for obtaining ).

[0162] The processor (110) may be configured to obtain a second data set (340-2) from the third neural network (133) which indicates the degree to which each of the parts of the speech signal (310) corresponds to each of the first data set (340-1) and each of the one or more phonetic symbols. The processor (110) may be configured to obtain a first parameter (350) from the fourth neural network (134) into which the first data set (340-1) is input, which indicates whether the speech signal (310) corresponds to the designated text (315). The processor (110) may be configured to obtain a second parameter (360) from the fifth neural network (135) into which the second data set (340-2) is input, which indicates whether the parts of the speech signal (310) correspond to each of the one or more phonetic symbols.

[0163] The above third neural network (133) can use a lower triangular matrix as an attention mask.

[0164] The method described above may be performed by an electronic device (101). The method may include the operation of obtaining a first sequence (320) of parts of the voice signal (310) corresponding to a designated frame unit from a first neural network (131) into which the voice signal (310) received through a microphone (160) is input. The method may include the operation of obtaining a second sequence (330) of one or more phonetic symbols for the designated text (315) from a second neural network (132) into which the designated text (315) is input. The method may include the operation of obtaining a first data set (340-1) (340-1) indicating the degree to which each of the one or more phonetic symbols corresponds to each of the parts of the voice signal (310) from a third neural network (133) into which the first sequence (320) and the second sequence (330) are input. The above method may include the operation of obtaining prediction phonemes (370-1, 370-2, 370-3, 370-4, 370-5) of the parts of the speech signal (310) from a pattern discriminator (e.g., the sixth neural network (136)) into which the first data set (340-1) (340-1) is input.

[0165] The above method may include an operation of obtaining a connectionist temporal classification (CTC) error for the predicted phonemes (370-1, 370-2, 370-3, 370-4, 370-5) of the above parts. The above method may include an operation of training the first neural network (131) based on the CTC error.

[0166] The operation of acquiring the second sequence (330) is to obtain a text embedding (E) for the specified text (315) through an embedding module (e.g., fifth operator (132-1)) within the second neural network (132). t The operation of obtaining the second sequence (330) may include obtaining the text embedding (E) through a tripon module (e.g., the sixth operator (132-2)) within the second neural network (132). t Triphone embeddings (E) representing the association between consecutive phonemes through ) tri It may include an operation to obtain ). The operation to obtain the second sequence (330) includes the tripon embedding (E tri ) and the above text embedding (E t The operation of obtaining the second sequence (330) by summing the elements of ) may be included.

[0167] The above Triphon embedding (E tri The operation of obtaining ) is through the placement normalization layer (384-1) within the tripon module (e.g., the 6th operator (132-2)), the text embedding (E t It may include an operation to batch normalize the tripons within ). The tripons are the text embedding (E t It includes the i-th phoneme among the phonemes within ), and two phonemes adjacent to the i-th phoneme, wherein i may be an integer greater than or equal to 1 and less than or equal to the number of phonemes. The tryphone embedding (E tri The operation of obtaining ) may include the operation of performing a convolution operation on the batch-normalized tripons through a convolution operator (384-2) within the tripon module (e.g., the 6th operator (132-2)). The tripon embedding (E triThe operation of obtaining ) may include the operation of obtaining the second sequence (330) by performing an activation operation on the convolutional tripons through an activation function (384-3) in the tripon module (e.g., the sixth operator (132-2)).

[0168] The above method may include the operation of embedding one or more phonetic symbols for the specified text (315) through an embedding layer (381) within the embedding module (e.g., fifth operator (132-1)). The above method may include the operation of fully connecting the one or more embedded phonetic symbols through a fully connected layer (382) within the embedding module (e.g., fifth operator (132-1)). The above method may perform an activation operation on the one or more fully connected phonetic symbols through an activation function (383) within the embedding module (e.g., fifth operator (132-1)), thereby [translating] the text embedding (E t It may include an action to acquire ).

[0169] The above method is a connected embedding (E) that connects the first sequence (320) and the second sequence (330) through the self-focusing layer (391) in the third neural network (133). c Final connected embedding (E) for ) c' The method may include an operation to obtain ). The method may include obtaining the final connected embedding (E) through the fully connected layer (393) in the third neural network (133). c' A joint embedding (E) including the first data set (340-1) for ) j It may include an action to acquire ).

[0170] The above third neural network (133) can use a lower triangular matrix as an attention mask.

[0171] A non-transient computer-readable storage medium as described above may store one or more programs. The one or more programs may include instructions that cause the electronic device (101) to obtain a first sequence (320) of parts of the voice signal (310) corresponding to a designated frame unit from a first neural network (131) into which a voice signal (310) received through a microphone (160) is input when executed by the processor (110) of the electronic device (101). The one or more programs may include instructions that cause the electronic device (101) to obtain a second sequence (330) of one or more phonetic symbols for the designated text (315) from a second neural network (132) into which a designated text (315) is input when executed by the processor (110) of the electronic device (101). The above one or more programs may include instructions that cause the electronic device (101) to obtain a first data set (340-1) (340-1) indicating the degree to which each of the one or more phonetic symbols corresponds to each of the parts of the voice signal (310), from a third neural network (133) into which the first sequence (320) and the second sequence (330) are input when executed by the processor (110) of the electronic device (101). The above one or more programs may include instructions that cause the electronic device (101) to obtain prediction phonemes (370-1, 370-2, 370-3, 370-4, 370-5) of the parts of the speech signal (310) from a pattern discriminator (e.g., a sixth neural network (136)) into which the first data set (340-1) is input when executed by the processor (110) of the electronic device (101).

[0172] The above one or more programs may include instructions that cause the electronic device (101) to obtain a connectionist temporal classification (CTC) error for the predicted phonemes (370-1, 370-2, 370-3, 370-4, 370-5) of the parts when executed by the processor (110) of the electronic device (101). The above one or more programs may include instructions that cause the electronic device (101) to train the first neural network (131) based on the CTC error when executed by the processor (110) of the electronic device (101).

[0173] When the above one or more programs are executed by the processor (110) of the electronic device (101), text embedding (E) for the specified text (315) is obtained through an embedding module (e.g., fifth operator (132-1)) within the second neural network (132). t The electronic device (101) may include instructions that cause the one or more programs to obtain the text embedding (E) when executed by the processor (110) of the electronic device (101), through a trypon module (e.g., a sixth operator (132-2)) within the second neural network (132). t Triphone embeddings (E) representing the association between consecutive phonemes through ) tri It may include instructions that cause the electronic device (101) to obtain ). When the one or more programs are executed by the processor (110) of the electronic device (101), the trypon embedding (E tri ) and the above text embedding (E t The electronic device (101) may include instructions that cause the second sequence (330) to be obtained by summing the elements of the electronic device (101).

[0174] The device described above may be implemented as a hardware component, a software component, and / or a combination of a hardware component and a software component. For example, the device and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. In addition, other processing configurations, such as parallel processors, are also possible.

[0175] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or instruct the processing unit independently or collectively. Software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and stored or executed in a distributed manner. Software and data may be stored on one or more non-transient computer-readable storage media.

[0176] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. In this case, the medium may continuously store a computer-executable program, or temporarily store it for execution or download. Furthermore, the medium may be various recording or storage means in the form of a single or several hardware combined, and is not limited to a medium directly connected to a computer system, but may also exist distributed over a network. Examples of media may include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and media configured to store program instructions, including ROM, RAM, and flash memory. Additionally, other examples of media may include storage media managed by app stores that distribute applications or sites and servers that supply or distribute various other software.

[0177] Although the embodiments have been described above with reference to limited embodiments and drawings, those skilled in the art can make various modifications and variations from the description above. For example, appropriate results can be achieved even if the described techniques are performed in a different order than described, and / or the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.

[0178] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims set forth below.

Claims

1. In an electronic device, mike, Memory, and Includes a processor, The above processor is, From a first neural network into which a voice signal received through the microphone is input, a first sequence of parts of the voice signal corresponding to a designated frame unit is obtained, and From a second neural network into which a specified text is input, a second sequence of one or more phonetic symbols for said specified text is obtained, and From a third neural network to which the first sequence and the second sequence are input, a first data set is obtained that indicates the degree to which each of the one or more phonetic symbols corresponds to each of the parts of the speech signal, and A pattern discriminator configured to acquire prediction phonemes of the parts of the speech signal from the pattern discriminator into which the first data set is input, Electronic device.

2. In paragraph 1, the processor, Obtain connectionist temporal classification (CTC) errors for the predicted phonemes of the above parts, and Configured to train the first neural network based on the above CTC error, Electronic device.

3. In Paragraph 1, The above second neural network is, An embedding module for obtaining a text embedding for the above-mentioned specified text, and It is configured to include a tripon module for obtaining a tripon embedding that indicates the association between consecutive phonemes through the above text embedding, and The above second neural network is, Configured to obtain the second sequence by summing the elements of the above tripon embedding and the above text embedding, Electronic device.

4. In paragraph 3, the above-mentioned tryphon module is, A batch normalization layer for batch normalizing tripons within the text embedding, wherein the tripons include the i-th phoneme among the phonemes within the text embedding, and two phonemes adjacent to the i-th phoneme, wherein i is an integer greater than or equal to 1 and less than or equal to the number of phonemes, and A convolution operator for performing a convolution operation on the above batch-normalized tripons, and A configuration comprising an activation function for obtaining the tripon embedding by performing an activation operation on the tripons convolutionally operated above, Electronic device.

5. In Paragraph 4, The above activation function includes GeLU (Gaussian error linear unit), Electronic device.

6. In paragraph 3, the embedding module is, An embedding layer for embedding one or more phonetic symbols for the above-specified text, A fully connected layer for fully connecting the one or more of the embedded phonetic symbols, and A method configured to include an activation function for obtaining the text embedding by performing an activation operation on one or more fully connected phonetic symbols. Electronic device.

7. In Paragraph 6, The above activation function includes LeakyReLU (leaky rectified linear unit), Electronic device.

8. In Paragraph 1, The above third neural network is, A self-focusing layer for obtaining a final connected embedding for a connected embedding concatenating the first sequence and the second sequence, and Configured to include a fully connected layer for obtaining a joint embedding comprising the first data set for the final connected embedding, Electronic device.

9. In paragraph 1, the processor, From the third neural network, a second data set is obtained that indicates the degree to which each of the parts of the speech signal corresponds to each of the first data set and each of the one or more phonetic symbols, and From the fourth neural network into which the first data set is input, a first parameter indicating whether it is a voice signal corresponding to the specified text is obtained, and A fifth neural network to which the second data set is input is configured to obtain a second parameter indicating whether the parts of the speech signal correspond to each of the one or more phonetic symbols, Electronic device.

10. In Paragraph 1, The above third neural network uses a lower triangular matrix as an attention mask, Electronic device.

11. A method performed by an electronic device, wherein the method comprises: The operation of obtaining a first sequence of parts of the voice signal corresponding to a designated frame unit from a first neural network into which the voice signal received through a microphone is input, The operation of obtaining a second sequence of one or more phonetic symbols for a specified text from a second neural network into which a specified text is input, The operation of obtaining a first data set indicating the degree to which each of the one or more phonetic symbols corresponds to each of the parts of the speech signal from a third neural network to which the first sequence and the second sequence are input, and The operation of obtaining predicted phonemes of the parts of the speech signal from a pattern discriminator into which the first data set is input method.

12. In paragraph 11, the above method is, The operation of obtaining connectionist temporal classification (CTC) errors for the predicted phonemes of the above parts, and The operation of training the first neural network based on the above CTC error method.

13. In paragraph 11, the operation of acquiring the second sequence is, The operation of obtaining a text embedding for the specified text through an embedding module within the second neural network, The operation of obtaining a tripon embedding representing the association between consecutive phonemes through the text embedding via the tripon module in the second neural network, and The operation of obtaining the second sequence by summing the elements of the above tripon embedding and the above text embedding method.

14. In paragraph 13, the operation of acquiring the above-mentioned tryphone embedding is, An operation of batch normalizing tripons in the text embedding through a batch normalization layer in the tripon module, wherein the tripons include the i-th phoneme among the phonemes in the text embedding and two phonemes adjacent to the i-th phoneme, and i is an integer greater than or equal to 1 and less than or equal to the number of phonemes. The operation of performing a convolution operation on the batch-normalized tripons through a convolution operator within the tripon module, and The operation of obtaining the second sequence by performing an activation operation on the convolutionally operated tripons through an activation function within the tripon module. method.

15. In Paragraph 13, the above method is, The operation of embedding one or more phonetic symbols for the specified text through an embedding layer within the embedding module, The operation of fully connecting the one or more embedded phonetic symbols through a fully connected layer within the embedding module, and The operation of obtaining the text embedding by performing an activation operation on one or more fully connected phonetic symbols through an activation function within the embedding module, method.

16. In Paragraph 11, the above method is, The operation of obtaining a final connected embedding for a connected embedding that concatenates the first sequence and the second sequence through a self-focusing layer in the third neural network, The operation of obtaining a joint embedding including the first data set for the final connected embedding through a fully connected layer in the third neural network, method.

17. In Paragraph 11, The above third neural network uses a lower triangular matrix as an attention mask, method.

18. In a non-transient computer-readable storage medium storing one or more programs, said one or more programs when executed by a processor of an electronic device, From a first neural network into which a voice signal received through a microphone is input, a first sequence of parts of the voice signal corresponding to a designated frame unit is obtained, and From a second neural network into which a specified text is input, a second sequence of one or more phonetic symbols for said specified text is obtained, and From a third neural network to which the first sequence and the second sequence are input, a first data set is obtained that indicates the degree to which each of the one or more phonetic symbols corresponds to each of the parts of the speech signal, and The electronic device comprises instructions that cause the electronic device to obtain prediction phonemes of the parts of the voice signal from a pattern discriminator into which the first data set is input. Non-transient computer-readable storage media.

19. In Paragraph 18, When the above one or more programs are executed by the processor of the electronic device, Obtain connectionist temporal classification (CTC) errors for the predicted phonemes of the above parts, and Instructions that cause the electronic device to train the first neural network based on the above CTC error, Non-transient computer-readable storage media.

20. In Paragraph 18, When the above one or more programs are executed by the processor of the electronic device, Text embeddings for the specified text are obtained through the embedding module within the second neural network, and Through the tripon module in the second neural network, a tripon embedding representing the association between consecutive phonemes is obtained through the text embedding, and The electronic device includes instructions that cause the second sequence to be obtained by summing the elements of the tripon embedding and the text embedding, Non-transient computer-readable storage media.