Artificial intelligence apparatus for training acoustic model
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2019-12-20
- Publication Date
- 2026-08-12
Smart Images

Figure 112019132075608-PAT00006_ABST
Abstract
Description
Technology Field
[0001] The present disclosure discloses an artificial intelligence device for acquiring voice data corresponding to a user utterance and training an acoustic model including a shared network and a branch network using a phoneme set corresponding to the acquired voice data. Background Technology
[0002] Speech recognition is a technology that receives a user's voice input and automatically converts it into text for recognition. Recently, speech recognition has been used as an interface technology to replace keyboard input in devices such as smartphones and TVs.
[0003] Generally, speech recognition systems can perform speech recognition using acoustic models, grammar models, and pronunciation dictionaries. In order to recognize a specific word from a speech signal in a speech recognition system, it is necessary to pre-construct a grammar model and a pronunciation dictionary for that specific word.
[0004] Meanwhile, the speech produced by users varies from person to person. Specifically, differences in pronunciation accuracy occur depending on the shape of the mouth and the position of the tongue. These characteristics are distinguished in acoustic models, and in the case of globally used languages such as English, various variations in speech data are produced. These variations hinder the consistency of acoustic signals, which has the problem of lowering the accuracy of speech recognition. The problem to be solved
[0005] The purpose of the present disclosure is to improve the accuracy of speech recognition without the process of unifying the different phoneme sets of native and non-native speakers into one by using a shared network and a branch network connected to the shared network.
[0006] The present disclosure discloses an artificial intelligence device that trains an acoustic model comprising a shared network and a branch network connected to the shared network, using a phoneme set corresponding to voice data.
[0007] In addition, the present disclosure discloses an artificial intelligence device that trains an acoustic model so that the shared network can output a hidden representation by reflecting common acoustic information of the received voice data, and trains an acoustic model so that the branch network can output a phoneme set corresponding to the received voice data using the hidden representation output by the shared network. Effects of the invention
[0008] The present disclosure can improve performance on non-native speaker voice data while maintaining performance on native speaker voice data by training an acoustic model using a shared network and a branch network that output a hidden expression reflecting common acoustic information of native speaker voice data and non-native speaker voice data, and by acquiring a native speaker phoneme set using the trained acoustic model. Brief explanation of the drawing
[0009] FIG. 1 shows an artificial intelligence device (100) according to one embodiment of the present disclosure. FIG. 2 shows an artificial intelligence server (200) according to one embodiment of the present disclosure. FIG. 3 shows an artificial intelligence system (1) according to one embodiment of the present disclosure. FIG. 4 shows an AI device according to one embodiment of the present disclosure. Figure 5 shows a general speech recognition process. FIG. 6 is a flowchart of the present disclosure. FIG. 7 is a learning example of the present disclosure. FIG. 8 is a learning example of the present disclosure. FIG. 9 is an example of use of the present disclosure. Specific details for implementing the invention
[0010] The embodiments described below are merely examples of the present invention, and the present invention may be modified in various forms. Accordingly, the specific configurations and functions disclosed below do not limit the scope of the claims.
[0011] Hereinafter, embodiments disclosed in this specification will be described in detail with reference to the attached drawings. Identical or similar components, regardless of drawing symbols, are assigned the same reference number, and redundant descriptions thereof will be omitted. The suffixes "module" and "part" used for components in the following description are assigned or used interchangeably solely for the ease of drafting the specification and do not inherently possess distinct meanings or roles. Furthermore, in describing embodiments disclosed in this specification, if it is determined that a detailed description of related prior art could obscure the essence of the embodiments disclosed in this specification, such detailed description will be omitted. Additionally, the attached drawings are intended only to facilitate understanding of the embodiments disclosed in this specification; the technical concept disclosed in this specification is not limited by the attached drawings, and it should be understood that they include all modifications, equivalents, and substitutions that fall within the concept and technical scope of this disclosure.
[0012] Terms including ordinal numbers, such as first, second, etc., may be used to describe various components, but said components are not limited by said terms. These terms are used solely for the purpose of distinguishing one component from another.
[0013] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.
[0014] Artificial Intelligence (AI)
[0015] Artificial intelligence refers to the field of researching artificial intelligence or the methodologies to create it, while machine learning refers to the field of researching methodologies to define and solve various problems addressed within the field of artificial intelligence. Machine learning is also defined as an algorithm that improves performance on a task through continuous experience.
[0016] An Artificial Neural Network (ANN) is a model used in machine learning that can refer to any model capable of problem-solving, composed of artificial neurons (nodes) that form a network through the connection of synapses. An artificial neural network can be defined by connection patterns between neurons in different layers, a learning process that updates model parameters, and an activation function that generates output values.
[0017] An artificial neural network may include an input layer, an output layer, and optionally one or more hidden layers. Each layer may include one or more neurons, and the artificial neural network may include synapses connecting the neurons. In an artificial neural network, each neuron may output a function value of an activation function for input signals, weights, and biases input through the synapses.
[0018] Model parameters refer to parameters determined through learning, including synaptic connection weights and neuron biases. Hyperparameters, on the other hand, refer to parameters that must be set prior to training in a machine learning algorithm, including the learning rate, number of iterations, mini-batch size, and initialization function.
[0019] The objective of training an artificial neural network can be viewed as determining model parameters that minimize the loss function. The loss function can be used as an indicator to determine optimal model parameters during the training process of an artificial neural network.
[0020] Machine learning can be classified into supervised learning, unsupervised learning, and reinforcement learning depending on the learning method.
[0021] Supervised learning refers to a method of training an artificial neural network with labels provided for the training data; a label can refer to the correct answer (or result) that the neural network must infer when the training data is input. Unsupervised learning refers to a method of training an artificial neural network without labels provided for the training data. Reinforcement learning refers to a learning method in which an agent defined within an environment is trained to select an action or sequence of actions that maximizes the cumulative reward in each state.
[0022] Machine learning implemented using a Deep Neural Network (DNN) that includes multiple hidden layers among artificial neural networks is also called Deep Learning, and Deep Learning is a part of Machine Learning. Hereinafter, Machine Learning is used in a sense that includes Deep Learning.
[0023] Robot
[0024] A robot can refer to a machine that automatically processes or operates a given task based on its own capabilities. In particular, a robot that has the ability to perceive its environment, make decisions on its own, and perform actions can be called an intelligent robot.
[0025] Robots can be classified into industrial, medical, household, military, etc., depending on their purpose of use or field.
[0026] A robot is equipped with a drive unit including actuators or motors to perform various physical movements, such as moving robot joints. Additionally, a mobile robot may include wheels, brakes, propellers, etc., in the drive unit, enabling it to drive on the ground or fly in the air.
[0027] Self-Driving
[0028] Autonomous driving refers to technology that drives itself, and an autonomous vehicle refers to a vehicle that drives without user intervention or with minimal user intervention.
[0029] For example, autonomous driving can include technologies such as maintaining the driving lane, automatically adjusting speed like adaptive cruise control, automatically driving along a set route, and automatically setting a route and driving once a destination is set.
[0030] Vehicles encompass vehicles equipped solely with internal combustion engines, hybrid vehicles equipped with both internal combustion engines and electric motors, and electric vehicles equipped solely with electric motors, and may include not only automobiles but also trains, motorcycles, etc.
[0031] In this case, an autonomous vehicle can be viewed as a robot with autonomous driving capabilities.
[0032] Extended Reality (XR)
[0033] Extended Reality is a collective term for Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR). VR technology provides real-world objects or backgrounds solely as CG images, AR technology provides virtual CG images superimposed on real-world images, and MR technology is a computer graphics technology that mixes and combines virtual objects with the real world.
[0034] MR technology is similar to AR technology in that it displays real-world objects and virtual objects together. However, there is a difference in that while virtual objects in AR technology are used to complement real-world objects, virtual objects and real-world objects are used as equals in MR technology.
[0035] XR technology can be applied to HMDs (Head-Mount Displays), HUDs (Head-Up Displays), mobile phones, tablet PCs, laptops, desktops, TVs, digital signage, etc., and devices to which XR technology is applied can be called XR devices.
[0037] FIG. 1 shows an artificial intelligence device (100) according to one embodiment of the present disclosure.
[0038] The AI device (100) can be implemented as a stationary device or a mobile device, such as a TV, projector, mobile phone, smartphone, desktop computer, laptop, digital broadcasting terminal, PDA (personal digital assistants), PMP (portable multimedia player), navigation, tablet PC, wearable device, set-top box (STB), DMB receiver, radio, washing machine, refrigerator, desktop computer, digital signage, robot, vehicle, etc.
[0039] Referring to FIG. 1, the terminal (100) may include a communication unit (110), an input unit (120), a learning processor (130), a sensing unit (140), an output unit (150), a memory (170), and a processor (180), etc.
[0040] The communication unit (110) can transmit and receive data with external devices, such as other AI devices (100a to 100e) or an AI server (200), using wired or wireless communication technology. For example, the communication unit (110) can transmit and receive sensor information, user input, learning models, control signals, etc., with external devices.
[0041] At this time, the communication technologies used by the communication unit (110) include GSM (Global System for Mobile communication), CDMA (Code Division Multi Access), LTE (Long Term Evolution), 5G, WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Bluetooth (Bluetooth), RFID (Radio Frequency Identification), Infrared Data Association (IrDA), ZigBee, NFC (Near Field Communication), etc.
[0042] The input unit (120) can acquire various types of data.
[0043] At this time, the input unit (120) may include a camera for inputting a video signal, a microphone for receiving an audio signal, a user input unit for receiving information from a user, etc. Here, the camera or microphone may be treated as a sensor, and the signal obtained from the camera or microphone may be referred to as sensing data or sensor information.
[0044] The input unit (120) can obtain training data for model training and input data to be used when obtaining an output using a training model. The input unit (120) may also obtain unprocessed input data, in which case the processor (180) or the learning processor (130) can extract input feature points as a preprocessing step for the input data.
[0045] The learning processor (130) can train a model composed of an artificial neural network using training data. Here, the trained artificial neural network may be referred to as a learning model. The learning model can be used to infer a result value for new input data other than the training data, and the inferred value can be used as a basis for judgment to perform an action.
[0046] At this time, the learning processor (130) can perform AI processing together with the learning processor (240) of the AI server (200).
[0047] At this time, the learning processor (130) may include memory integrated into or implemented in the AI device (100). Alternatively, the learning processor (130) may be implemented using memory (170), external memory directly coupled to the AI device (100), or memory maintained in an external device.
[0048] The sensing unit (140) can obtain at least one of internal information of the AI device (100), surrounding environment information of the AI device (100), and user information using various sensors.
[0049] At this time, the sensors included in the sensing unit (140) include a proximity sensor, an illuminance sensor, an accelerometer, a magnetic sensor, a gyroscope, an inertial sensor, an RGB sensor, an IR sensor, a fingerprint recognition sensor, an ultrasonic sensor, a light sensor, a microphone, a lidar, a radar, etc.
[0050] The output unit (150) can generate output related to sight, hearing, or touch.
[0051] At this time, the output unit (150) may include a display unit that outputs visual information, a speaker that outputs auditory information, a haptic module that outputs tactile information, etc.
[0052] The memory (170) can store data that supports various functions of the AI device (100). For example, the memory (170) can store input data, training data, training models, training history, etc. obtained from the input unit (120).
[0053] The processor (180) can determine at least one executable action of the AI device (100) based on information determined or generated using a data analysis algorithm or a machine learning algorithm. The processor (180) can perform the determined action by controlling the components of the AI device (100).
[0054] To this end, the processor (180) can request, search, receive, or utilize data from the learning processor (130) or memory (170), and can control the components of the AI device (100) to execute a predicted operation or a preferred operation among the at least one executable operation.
[0055] At this time, if the processor (180) requires the connection of an external device to perform a determined operation, it can generate a control signal to control the external device and transmit the generated control signal to the external device.
[0056] The processor (180) can obtain intent information regarding user input and determine the user's requirements based on the obtained intent information.
[0057] At this time, the processor (180) can obtain intent information corresponding to the user input by using at least one of a Speech To Text (STT) engine for converting voice input into a string or a Natural Language Processing (NLP) engine for obtaining intent information of natural language.
[0058] At this time, at least one of the STT engine or NLP engine may be composed of an artificial neural network in which at least a portion is learned according to a machine learning algorithm. Also, at least one of the STT engine or NLP engine may be learned by a learning processor (130), learned by a learning processor (240) of an AI server (200), or learned through distributed processing thereof.
[0059] The processor (180) may collect history information, including the operation details of the AI device (100) or user feedback regarding the operation, and store it in memory (170) or a learning processor (130), or transmit it to an external device such as an AI server (200). The collected history information may be used to update a learning model.
[0060] The processor (180) can control at least some of the components of the AI device (100) to run an application stored in memory (170). Furthermore, the processor (180) can operate two or more of the components included in the AI device (100) in combination with each other to run the application.
[0062] FIG. 2 shows an artificial intelligence server (200) according to one embodiment of the present disclosure.
[0063] Referring to FIG. 2, the AI server (200) may refer to a device that trains an artificial neural network using a machine learning algorithm or uses a trained artificial neural network. Here, the AI server (200) may be composed of multiple servers to perform distributed processing and may be defined as a 5G network. At this time, the AI server (200) may be included as part of the configuration of the AI device (100) and may perform at least part of the AI processing together.
[0064] The AI server (200) may include a communication unit (210), memory (230), a learning processor (240), and a processor (260), etc.
[0065] The communication unit (210) can transmit and receive data with external devices such as AI devices (100).
[0066] The memory (230) may include a model storage unit (231). The model storage unit (231) may store a model (or artificial neural network, 231a) that is being learned or has been learned through a learning processor (240).
[0067] The learning processor (240) can train the artificial neural network (231a) using training data. The training model may be used while mounted on the AI server (200) of the artificial neural network, or it may be used while mounted on an external device such as an AI device (100).
[0068] The learning model may be implemented in hardware, software, or a combination of hardware and software. If part or all of the learning model is implemented in software, one or more instructions constituting the learning model may be stored in memory (230).
[0069] The processor (260) can use a learning model to infer a result value for new input data and generate a response or control command based on the inferred result value.
[0070] FIG. 3 shows an artificial intelligence system (1) according to one embodiment of the present disclosure.
[0071] Referring to FIG. 3, the AI system (1) is connected to a cloud network (10) at least one of an AI server (200), a robot (100a), an autonomous vehicle (100b), an XR device (100c), a smartphone (100d), or a home appliance (100e). Here, the robot (100a), the autonomous vehicle (100b), the XR device (100c), the smartphone (100d), or the home appliance (100e) to which AI technology is applied may be referred to as AI devices (100a to 100e).
[0072] A cloud network (10) may mean a network that constitutes part of a cloud computing infrastructure or exists within a cloud computing infrastructure. Here, the cloud network (10) may be configured using a 3G network, a 4G or LTE (Long Term Evolution) network or a 5G network, etc.
[0073] That is, each device (100a to 100e, 200) constituting the AI system (1) can be connected to each other through a cloud network (10). In particular, each device (100a to 100e, 200) may communicate with each other through a base station, but may also communicate directly with each other without going through a base station.
[0074] The AI server (200) may include a server that performs AI processing and a server that performs operations on big data.
[0075] The AI server (200) is connected via a cloud network (10) to at least one of the AI devices constituting the AI system (1), such as a robot (100a), an autonomous vehicle (100b), an XR device (100c), a smartphone (100d), or a home appliance (100e), and can assist in at least some of the AI processing of the connected AI devices (100a to 100e).
[0076] At this time, the AI server (200) can train an artificial neural network according to a machine learning algorithm on behalf of the AI devices (100a to 100e), and can directly store the training model or transmit it to the AI devices (100a to 100e).
[0077] At this time, the AI server (200) receives input data from the AI devices (100a to 100e), infers a result value for the received input data using a learning model, and generates a response or control command based on the inferred result value and transmits it to the AI devices (100a to 100e).
[0078] Alternatively, the AI device (100a to 100e) may use a direct learning model to infer a result value for input data and generate a response or control command based on the inferred result value.
[0079] Hereinafter, various embodiments of AI devices (100a to 100e) to which the above-described technology is applied will be described. Here, the AI devices (100a to 100e) illustrated in FIG. 3 can be seen as specific embodiments of the AI device (100) illustrated in FIG. 1.
[0080] <AI+로봇>
[0081] The robot (100a) can be implemented as a guide robot, transport robot, cleaning robot, wearable robot, entertainment robot, pet robot, unmanned flying robot, etc. by applying AI technology.
[0082] The robot (100a) may include a robot control module for controlling operation, and the robot control module may mean a software module or a chip that implements the same in hardware.
[0083] The robot (100a) can use sensor information obtained from various types of sensors to obtain state information of the robot (100a), detect (recognize) surrounding environment and objects, generate map data, determine movement path and driving plan, determine response to user interaction, or determine action.
[0084] Here, the robot (100a) can use sensor information obtained from at least one sensor among lidar, radar, and camera to determine a movement path and driving plan.
[0085] The robot (100a) can perform the above-mentioned actions using a learning model composed of at least one artificial neural network. For example, the robot (100a) can recognize the surrounding environment and objects using the learning model, and can determine actions using the recognized surrounding environment information or object information. Here, the learning model may be learned directly by the robot (100a) or learned from an external device such as an AI server (200).
[0086] At this time, the robot (100a) may perform an operation by generating a result using a direct learning model, but it may also perform an operation by transmitting sensor information to an external device such as an AI server (200) and receiving the result generated accordingly.
[0087] The robot (100a) can determine a movement path and a driving plan using at least one of map data, object information detected from sensor information, or object information obtained from an external device, and control a driving unit to drive the robot (100a) according to the determined movement path and driving plan.
[0088] Map data may include object identification information for various objects placed in the space where the robot (100a) moves. For example, map data may include object identification information for fixed objects such as walls and doors, and movable objects such as flowerpots and desks. In addition, the object identification information may include names, types, distances, locations, etc.
[0089] Additionally, the robot (100a) can perform actions or drive by controlling the drive unit based on the user's control / interaction. At this time, the robot (100a) can acquire intention information of interaction based on the user's actions or voice utterances, and can perform actions by determining a response based on the acquired intention information.
[0090] <AI+자율주행>
[0091] The autonomous vehicle (100b) can be implemented as a mobile robot, vehicle, unmanned aerial vehicle, etc. by applying AI technology.
[0092] The autonomous vehicle (100b) may include an autonomous driving control module for controlling autonomous driving functions, and the autonomous driving control module may refer to a software module or a chip that implements the same in hardware. The autonomous driving control module may be included internally as a component of the autonomous vehicle (100b), but may also be configured and connected as separate hardware externally to the autonomous vehicle (100b).
[0093] The autonomous vehicle (100b) can use sensor information obtained from various types of sensors to obtain state information of the autonomous vehicle (100b), detect (recognize) surrounding environment and objects, generate map data, determine a travel path and driving plan, or determine an action.
[0094] Here, the autonomous vehicle (100b) can use sensor information obtained from at least one sensor among lidar, radar, and camera, just like the robot (100a), to determine a travel path and a driving plan.
[0095] In particular, the autonomous vehicle (100b) can recognize environments or objects in areas where the field of view is obscured or in areas beyond a certain distance by receiving sensor information from external devices, or by receiving information directly recognized from external devices.
[0096] The autonomous vehicle (100b) can perform the above-mentioned operations using a learning model composed of at least one artificial neural network. For example, the autonomous vehicle (100b) can recognize surrounding environments and objects using the learning model, and can determine a driving path using the recognized surrounding environment information or object information. Here, the learning model may be learned directly in the autonomous vehicle (100b) or learned from an external device such as an AI server (200).
[0097] At this time, the autonomous vehicle (100b) may perform operations by generating results using a direct learning model, but may also perform operations by transmitting sensor information to an external device such as an AI server (200) and receiving the results generated accordingly.
[0098] The autonomous vehicle (100b) can determine a movement path and a driving plan using at least one of map data, object information detected from sensor information or object information obtained from an external device, and control a driving unit to drive the autonomous vehicle (100b) according to the determined movement path and driving plan.
[0099] Map data may include object identification information for various objects placed in a space (e.g., a road) where the autonomous vehicle (100b) is driving. For example, the map data may include object identification information for fixed objects such as streetlights, rocks, and buildings, and movable objects such as vehicles and pedestrians. In addition, the object identification information may include names, types, distances, locations, etc.
[0100] Additionally, the autonomous vehicle (100b) can perform operations or drive by controlling the drive unit based on the user's control / interaction. At this time, the autonomous vehicle (100b) can acquire intention information of the interaction based on the user's actions or voice utterances, and can perform operations by determining a response based on the acquired intention information.
[0101] <AI+XR>
[0102] The XR device (100c) can be implemented as a Head-Mount Display (HMD), a Head-Up Display (HUD) equipped in a vehicle, a television, a mobile phone, a smartphone, a computer, a wearable device, a home appliance, digital signage, a vehicle, a stationary robot, or a mobile robot by applying AI technology.
[0103] The XR device (100c) can obtain information about surrounding space or real objects by analyzing 3D point cloud data or image data obtained through various sensors or from an external device to generate position data and attribute data for 3D points, and can render and output an XR object to be output. For example, the XR device (100c) can output an XR object containing additional information about a recognized object by associating it with the recognized object.
[0104] The XR device (100c) can perform the above-mentioned operations using a learning model composed of at least one artificial neural network. For example, the XR device (100c) can recognize real-world objects in 3D point cloud data or image data using the learning model and can provide information corresponding to the recognized real-world objects. Here, the learning model may be learned directly by the XR device (100c) or learned from an external device such as an AI server (200).
[0105] At this time, the XR device (100c) may perform an operation by generating a result using a direct learning model, but it may also perform an operation by transmitting sensor information to an external device such as an AI server (200) and receiving the result generated accordingly.
[0106] <AI+로봇+자율주행>
[0107] The robot (100a) can be implemented as a guide robot, transport robot, cleaning robot, wearable robot, entertainment robot, pet robot, unmanned flying robot, etc. by applying AI technology and autonomous driving technology.
[0108] A robot (100a) equipped with AI technology and autonomous driving technology may refer to the robot itself having autonomous driving capabilities, or a robot (100a) that interacts with an autonomous driving vehicle (100b).
[0109] A robot (100a) with autonomous driving capabilities can be collectively referred to as a device that moves on its own along a given path without user control, or moves by determining its own path.
[0110] A robot (100a) and an autonomous vehicle (100b) having autonomous driving capabilities may use a common sensing method to determine one or more of a travel path or a driving plan. For example, a robot (100a) and an autonomous vehicle (100b) having autonomous driving capabilities may determine one or more of a travel path or a driving plan by using information sensed through a lidar, radar, or camera.
[0111] A robot (100a) interacting with an autonomous vehicle (100b) exists separately from the autonomous vehicle (100b) and can perform actions linked to the autonomous driving function inside the autonomous vehicle (100b) or linked to a user riding in the autonomous vehicle (100b).
[0112] At this time, the robot (100a) interacting with the autonomous vehicle (100b) can control or assist the autonomous driving function of the autonomous vehicle (100b) by acquiring sensor information on behalf of the autonomous vehicle (100b) and providing it to the autonomous vehicle (100b), or by acquiring sensor information and generating surrounding environment information or object information and providing it to the autonomous vehicle (100b).
[0113] Alternatively, a robot (100a) interacting with an autonomous vehicle (100b) may monitor a user riding in the autonomous vehicle (100b) or control the functions of the autonomous vehicle (100b) through interaction with the user. For example, if the robot (100a) determines that the driver is drowsy, it may activate the autonomous driving function of the autonomous vehicle (100b) or assist in controlling the drive unit of the autonomous vehicle (100b). Here, the functions of the autonomous vehicle (100b) controlled by the robot (100a) may include not only the autonomous driving function but also functions provided by a navigation system or an audio system equipped inside the autonomous vehicle (100b).
[0114] Alternatively, a robot (100a) interacting with an autonomous vehicle (100b) may provide information to or assist functions to the autonomous vehicle (100b) from outside the autonomous vehicle (100b). For example, the robot (100a) may provide traffic information, such as signal information, to the autonomous vehicle (100b), such as a smart traffic light, or may interact with the autonomous vehicle (100b) to automatically connect an electric charger to the charging port, such as an automatic electric charger for an electric vehicle.
[0115] <AI+로봇+XR>
[0116] The robot (100a) can be implemented as a guide robot, transport robot, cleaning robot, wearable robot, entertainment robot, pet robot, unmanned flying robot, drone, etc. by applying AI technology and XR technology.
[0117] A robot (100a) to which XR technology is applied may refer to a robot that is the subject of control / interaction within an XR image. In this case, the robot (100a) is distinguished from the XR device (100c) and can be interconnected with it.
[0118] When a robot (100a) that is the target of control / interaction within an XR image acquires sensor information from sensors including a camera, the robot (100a) or the XR device (100c) can generate an XR image based on the sensor information, and the XR device (100c) can output the generated XR image. Furthermore, the robot (100a) can operate based on a control signal input through the XR device (100c) or user interaction.
[0119] For example, the user can view an XR image corresponding to the viewpoint of the remotely linked robot (100a) through an external device such as an XR device (100c), and through interaction, can adjust the autonomous driving path of the robot (100a), control its movement or driving, or check information about surrounding objects.
[0120] <AI+자율주행+XR>
[0121] The autonomous vehicle (100b) can be implemented as a mobile robot, vehicle, unmanned aerial vehicle, etc. by applying AI technology and XR technology.
[0122] An autonomous vehicle (100b) equipped with XR technology may refer to an autonomous vehicle equipped with means for providing XR images, or an autonomous vehicle that is the subject of control / interaction within the XR images. In particular, an autonomous vehicle (100b) that is the subject of control / interaction within the XR images may be distinguished from an XR device (100c) and may be interconnected with it.
[0123] An autonomous vehicle (100b) equipped with means for providing XR images can acquire sensor information from sensors including cameras and output an XR image generated based on the acquired sensor information. For example, the autonomous vehicle (100b) can provide an XR object corresponding to a real object or an object in the screen to the occupant by providing an XR image by outputting an XR image with a HUD.
[0124] At this time, when the XR object is displayed on the HUD, at least a portion of the XR object may be displayed so as to overlap with the actual object to which the occupant's gaze is directed. On the other hand, when the XR object is displayed on a display provided inside the autonomous vehicle (100b), at least a portion of the XR object may be displayed so as to overlap with an object on the screen. For example, the autonomous vehicle (100b) may display XR objects corresponding to objects such as lanes, other vehicles, traffic lights, traffic signs, motorcycles, pedestrians, buildings, etc.
[0125] When an autonomous vehicle (100b) that is the subject of control / interaction within an XR image acquires sensor information from sensors including a camera, the autonomous vehicle (100b) or the XR device (100c) can generate an XR image based on the sensor information, and the XR device (100c) can output the generated XR image. Furthermore, the autonomous vehicle (100b) can operate based on control signals input through an external device such as the XR device (100c) or user interaction.
[0127] FIG. 4 shows an AI device (100) according to one embodiment of the present disclosure.
[0128] Descriptions that overlap with Figure 1 are omitted.
[0129] In the present disclosure, the artificial intelligence device (100) includes an edge device.
[0130] Referring to FIG. 4, the input unit (120) may include a camera (Camera, 121) for inputting a video signal, a microphone (Microphone, 122) for receiving an audio signal, and a user input unit (User Input Unit, 123) for receiving information from a user.
[0131] Voice data or image data collected from the input unit (120) can be analyzed and processed into user control commands.
[0132] The input unit (120) is for inputting video information (or signal), audio information (or signal), data, or information input from a user, and for inputting video information, the AI device (100) may be equipped with one or more cameras (121).
[0133] The camera (121) processes image frames, such as still images or video, obtained by an image sensor in video call mode or shooting mode. The processed image frames may be displayed on a display unit (Display Unit, 151) or stored in memory (170).
[0134] The microphone (122) processes external acoustic signals into electrical voice data. The processed voice data can be utilized in various ways depending on the function (or application running) being performed on the AI device (100). Meanwhile, various noise removal algorithms can be applied to the microphone (122) to remove noise generated during the process of receiving external acoustic signals.
[0135] The user input unit (123) is for receiving information from a user, and when information is input through the user input unit (123), the processor (180) can control the operation of the AI device (100) to correspond to the input information.
[0136] The user input unit (123) may include mechanical input means (or mechanical keys, such as buttons, dome switches, jog wheels, jog switches, etc. located on the front / rear or side of the AI device (100)) and touch input means. As an example, the touch input means may consist of a virtual key, soft key, or visual key displayed on a touchscreen through software processing, or a touch key placed on a part other than the touchscreen.
[0137] The sensing unit (140) can be called a sensor unit.
[0138] The output unit (150) may include at least one of a display unit (Display Unit, 151), a sound output unit (Sound Output Unit, 152), a haptic module (Haptic Module, 153), and an optical output unit (Optical Output Unit, 154).
[0139] The display unit (151) displays (outputs) information processed by the AI device (100). For example, the display unit (151) can display information on the execution screen of an application running on the AI device (100), or UI (User Interface) and GUI (Graphic User Interface) information based on such execution screen information.
[0140] The display unit (151) can implement a touch screen by forming a layered structure with the touch sensor or by being formed as an integral unit. This touch screen functions as a user input unit (123) that provides an input interface between the AI device (100) and the user, and at the same time, can provide an output interface between the terminal (100) and the user.
[0141] The sound output unit (152) can output audio data received from the communication unit (110) or stored in the memory (170) in call signal reception, call mode or recording mode, voice recognition mode, broadcast reception mode, etc.
[0142] The sound output unit (152) may include at least one of a receiver, a speaker, and a buzzer.
[0143] The haptic module (153) generates various tactile effects that the user can feel. A typical example of the tactile effect generated by the haptic module (153) can be vibration.
[0144] The light output unit (154) outputs a signal to indicate the occurrence of an event using the light from the light source of the AI device (100). Examples of events occurring in the AI device (100) may include receiving a message, receiving a call signal, a missed call, an alarm, a schedule notification, receiving an email, receiving information through an application, etc.
[0146] Figure 5 shows a general speech recognition process.
[0147] In a typical speech recognition model, when speech data is received, text data is derived through a preprocessing stage, an acoustic model, and a language model. Specifically, a speech recognition model can be composed of a feature vector extraction model, an acoustic model, a pronunciation model, and a language model.
[0148] Specifically, when voice data is input into a voice recognition model (S510), a feature vector can be extracted from the input voice data through a feature vector extraction module (S520). At this time, the process of extracting the feature vector can be used in combination with the preprocessing process. Generally, voice recognition models use the 12th-order Mel-Cepstrum (MFCC), log energy, and the first and second derivatives thereof as algorithms for extracting feature vectors.
[0149] The feature data obtained from feature extraction undergoes a similarity measurement and recognition process. For similarity measurement and recognition, an acoustic model that models and compares the signal characteristics of speech and a language model that models the linguistic sequential relationships, such as words or syllables corresponding to the recognized vocabulary, are used (S530, S540). Then, when the language model undergoes a process of correcting the converted text into corrected text, the final speech recognition result can be obtained (S550).
[0150] FIG. 6 is a flowchart of the present disclosure.
[0151] Before explaining the flowchart of the present disclosure, when a feature vector corresponding to the input speech data is extracted in a general speech recognition process, the feature vector is processed through an acoustic model, a pronunciation model, and a language model during the pattern matching process for speech recognition. Therefore, methods to improve the performance of a speech recognition model can be classified into an acoustic model perspective, a language model perspective, and a pronunciation model perspective.
[0152] The present invention takes into account an acoustic model-theoretic perspective and aims to improve the performance of acoustic models. The present disclosure discloses an artificial intelligence device for improving speech recognition performance for non-native speakers while maintaining the performance of speech recognition for native speakers.
[0153] Figure 6 is explained below.
[0154] Referring to FIG. 6, the input unit (120) of the artificial intelligence device (100) can receive voice data (S610). At this time, the input unit (120) may include a microphone and may be implemented with at least one sensor. Additionally, the voice data may include voice data corresponding to a voice signal spoken by a user.
[0155] When voice data is received, the processor (180) of the artificial intelligence device (100) can extract features of the voice data (S620). Specifically, feature extraction may include a preprocessing step to remove noise from the voice data. The noise can be removed using a filter used in signal processing.
[0156] The processor (180) can extract actual valid sound features from noise and background sounds of the input voice data. Specifically, Linear Prediction Coefficients (LPC) and Linear Prediction Cepstral Coefficient (LPCC) using an MFCC or HMM Classifier may be used.
[0157] Additionally, the process of extracting features of voice data (S620) may include a process of performing language embeddings corresponding to feature vectors, which is a method of quantifying input voice data.
[0158] The processor (180) of the artificial intelligence device (100) according to the present disclosure can train an acoustic model using the voice data and phonemes corresponding to the voice data (S630 to S650).
[0159] In this context, an acoustic model can be used to represent the relationship between an audio signal and phonemes or other linguistic units that constitute speech. Additionally, an acoustic model can model the relationship between the audio signal and the phonemes of a language.
[0160] Specifically, the acoustic model of the present disclosure may include a shared network and a branch network. Additionally, the acoustic model of the present disclosure may include a shared network and a branch network connected to said shared network. A processor (180) may input voice data into the acoustic model and train the acoustic model by setting a phoneme set corresponding to the voice data as a result. The internal configuration of the acoustic model will be described below.
[0161] The acoustic model of the present disclosure may include a shared network.
[0162] The shared network of the present disclosure may include a first artificial intelligence model that outputs a hidden representation by reflecting common acoustic information of a plurality of voice data used during acoustic model training. The shared network may also transmit the hidden representation to a layer of the branch network.
[0163] In this case, the shared network may include an artificial intelligence model that combines various artificial neural networks, such as a Convolutional Neural Network (CNN) and a Time Delay Neural Network (TDNN) that learns the time sequence of limited data. For example, the shared network may be composed of a single artificial neural network by combining a CNN model and a TDNN model. Meanwhile, the internal configuration of the shared network is not limited to the above example.
[0164] The acoustic model of the present disclosure may include a branch network.
[0165] The branch network of the present disclosure can be connected to the shared network.
[0166] The branch network of the present disclosure can classify phonemes for the language of input voice data.
[0167] And the branch network may include a second artificial intelligence model that outputs a phoneme set corresponding to the speech data based on a hidden representation transmitted by a layer of the shared network.
[0168] At this time, the branch network may include an artificial intelligence model that combines various artificial neural networks, such as a Recurrent Neural Network (RNN) that uses sequence data, a Long-Short Term Memory (LSTM) that introduces the concept of gates to the RNN to reflect temporal characteristics well, and a Time Delay Neural Network (TDNN) that learns the time sequence of limited data.
[0169] For example, a branch network can be composed of a single artificial neural network by combining an LSTM model and a TDNN. Meanwhile, the internal configuration of the branch network is not limited to the above example.
[0170] The acoustic model of the present disclosure may include a shared network and a branch network connected to said shared network. Additionally, the acoustic model of the present disclosure may include a plurality of branch networks. The plurality of branch networks may include a native speaker branch network that outputs a phonemes of a native language when voice data is input, and a non-native speaker branch network that outputs a phonemes of a foreign language when voice data is input. Also,
[0171] For convenience of explanation, in this disclosure, voice data spoken by American speakers is defined as native speaker voice data, and voice data spoken by British English, other English-speaking speakers, or non-native speakers is defined as non-native speaker voice data. However, the above description is merely an example for convenience of explanation, and the criteria for a native speaker may change.
[0172] The processor (180) of the artificial intelligence device (100) can set feature data corresponding to the input voice data as the input value of the acoustic model (S630). At this time, the acoustic model includes a branch network connected to a shared network, so that the input value of the acoustic model can be used in combination with the input value of the shared network.
[0173] The processor (180) of the artificial intelligence device (100) can set a phoneme set corresponding to the input voice data as the correct answer value of the acoustic model (S650).
[0174] Specifically, the processor (180) can train an acoustic model using voice data and a phoneme set corresponding to the voice data. During the training process, the shared network included in the acoustic model and the branch network connected to the shared network can be trained such that the error between the result value derived from the weights that fully connect each layer and the correct value is minimized (S630, S640).
[0175] In addition, the learning method of the present disclosure may include supervised learning, which sets a phoneme set corresponding to speech data as the correct answer value; unsupervised learning, which clusters similar data without setting a correct answer value; or reinforcement learning, which enables learning by setting a reward.
[0176] A learning example is described below in Fig. 7.
[0177] FIG. 7 is a learning example of the present disclosure.
[0178] First, with reference to Fig. 7, we assume the process of an acoustic model learning speech data "APPLE" spoken by native speakers and non-native speakers.
[0179] The processor (180) of the artificial intelligence device (100) can acquire voice data corresponding to the speaker's utterance (S710). The processor (180) can extract features of the acquired voice data and extract feature vectors corresponding to the received voice data through word embedding (S720, S730).
[0180] For example, the processor (180) of the artificial intelligence device (100) may receive voice data called "APPLE" spoken by a speaker and extract voice features of "APPLE". At this time, the result of extracting the features of "APPLE" may vary depending on features such as the position of the tongue, the shape of the mouth, the intensity of the sound, and the frequency when the speaker pronounces it. In addition, the voice features obtained may be generated differently depending on the nationality of the speaker or whether they are a native speaker or a non-native speaker.
[0181] In this specification, voice features may be used interchangeably with terms corresponding to received voice data, such as voice data and feature data.
[0182] The processor (180) of the artificial intelligence device (100) can train an acoustic model including a shared network and a branch network connected to the shared network using acquired voice data and a phoneme set corresponding to the voice data.
[0183] At this time, the acoustic model can be trained by setting voice data or feature data corresponding to voice data as input values for the acoustic model, and setting a phoneme set corresponding to the voice data as output values.
[0184] In addition, voice data obtained from the input unit (120) of the present disclosure may include native speaker voice data or non-native speaker voice data.
[0185] In an embodiment of the present disclosure, when the native speaker voice data is received, the processor (180) can train the shared network and the native speaker branch network connected to the shared network. Specifically, when the voice data is voice data spoken by a native speaker, the processor (180) can train the shared network and the native speaker branch network connected to the shared network using the native speaker voice data and the native speaker phoneme set corresponding to the native speaker voice data.
[0186] For example, the processor (180) can use voice data corresponding to "APPLE," a voice signal spoken by an American, as an input value for a first acoustic model including a shared network and a native speaker branch network connected to the shared network.
[0187] And, the processor (180) can set an American English phoneme set (e.g., CMU pronunciation sequence “AE1 P AH0 L”) corresponding to the input “APPLE” as the result value of the first acoustic model.
[0188] The processor (180) can train the first acoustic model such that the output value of the first acoustic model and the value of the error L1 of the correct answer value "AE1 P AH0 L" are minimized.
[0189] In another embodiment, when the processor (180) receives the non-native speaker voice data, it can train the shared network and the non-native speaker branch network connected to the shared network. Specifically, if the voice data is voice data spoken by a non-native speaker, the processor (180) can train the shared network and the non-native speaker branch network connected to the shared network using the non-native speaker voice data and the non-native speaker phonemes of a foreign language corresponding to the non-native speaker voice data.
[0190] For example, the processor (180) can use voice data corresponding to "APPLE," a voice signal spoken by a British person, as an input value for a second acoustic model including a shared network and a non-native speaker branch network connected to the shared network.
[0191] And, the processor (180) can set a British English phoneme set (e.g., the BEEP pronunciation sequence “ae pl”) corresponding to the input “APPLE” as the result value of the second acoustic model.
[0192] The processor (180) can train the second acoustic model such that the output value of the second acoustic model and the value of the error L2 of the correct value "ae pl" are minimized.
[0193] As another example, the processor (180) can use voice data corresponding to "APPLE," a voice signal spoken by an Indian person, as an input value for a third acoustic model including a shared network and a non-native speaker branch network connected to the shared network.
[0194] And, the processor (180) can set an Indian English phoneme set (e.g., X-SAMPA phoneme sequence “{ p @ l”) corresponding to the input “APPLE” as the result of the second acoustic model.
[0195] The processor (180) can train the third acoustic model such that the output value of the third acoustic model and the value of the error L3 of the correct value "{ p @ l" are minimized.
[0196] The present disclosure allows the use of a shared network regardless of whether the speaker is a native or non-native speaker during learning. Therefore, when the processor (180) trains the shared network, it can be trained so that common features of the received voice data are reflected in the shared network regardless of whether the speaker is a native or non-native speaker.
[0197] In the above process, by training the shared network regardless of whether the speaker is a native or non-native speaker, when native speaker speech data or non-native speaker speech data is subsequently received by the acoustic model, the shared network can extract common speech features.
[0198] In other words, the shared network can prevent the updated acoustic model from overfitting to a specific country's language by conveying common features of native / non-native speech data as hidden expressions. Additionally, during the training process, it is possible to achieve a regularization effect of the acoustic model that does not lean toward speech features of specific individuals among native and non-native speakers.
[0199] In addition, the present disclosure allows the branch network connected to the shared network to be trained to perform the role of classifying phonemes appropriate for each language by setting the branch network connected to the shared network differently depending on whether the learner is a native speaker or a non-native speaker during learning.
[0200] Specifically, the number of phonemes varies for each language and phoneme set. For example, the CMU phoneme set corresponding to American English may consist of 39, while the BEEP phoneme set corresponding to British English may consist of 50.
[0201] In other words, since the number of corresponding phoneme sets differs for native and non-native speech data, the final output of the acoustic model may vary. Therefore, a separate branch network may be required for each language.
[0202] The artificial intelligence device (100) of the present disclosure can use a branch network individually for each language, thereby integrating the phoneme sets of non-native speech data and native speech data into a single phoneme set for use in learning, or can train an acoustic model without the effort of establishing rules for one-to-one mapping of phonemes and collecting data whenever a non-native country is added.
[0203] The following describes a process in which, when multiple voice data are acquired, the artificial intelligence device (100) processor (180) classifies the acquired voice data by country and, based on the classified voice data, changes the branch network connected to the shared network to train an acoustic model. Duplicate explanations are omitted.
[0204] According to the present disclosure, the input unit of the artificial intelligence device (100) can acquire a plurality of voice data (S710). The processor (180) can extract features of the input plurality of voice data (S720, S730).
[0205] The processor (180) can determine whether the input voice data is a speech by a native speaker or a speech by a non-native speaker. Specifically, the processor (180) can separate the nationality of the speaker by using the relationship between the embedding information for the voice data and a database storing the linguistic features of native / non-native speaker speech, and can determine whether the input voice data is native or non-native speaker voice data.
[0206] The processor (180) of the artificial intelligence device (100) can obtain a native speaker phoneme set stored in memory (170) if the voice data is determined to be native speaker voice data. Additionally, if the voice data is determined to be non-native speaker voice data, it can obtain a non-native speaker phoneme set stored in memory (170).
[0207] The processor (180) of the present disclosure can train the shared network and the native speaker branch network connected to the shared network when the native speaker voice data is received, and train the shared network and the non-native speaker branch network connected to the shared network when the non-native speaker voice data is received.
[0208] Specifically, when native speaker voice data is received, the processor (180) can construct a first acoustic model by connecting a shared network within the acoustic model and a native speaker branch network. The processor (180) can train the first acoustic model by inputting native speaker voice data as an input value for the first acoustic model and setting a native speaker phoneme set corresponding to the native speaker voice data as an output value.
[0209] Additionally, when non-native speaker voice data is received, the processor (180) can construct a second acoustic model by connecting the shared network within the acoustic model with the non-native speaker branch network. The processor (180) can train the second acoustic model by inputting the non-native speaker voice data as an input value for the second acoustic model and setting a non-native speaker phoneme set corresponding to the non-native speaker voice data as an output value.
[0210] The processor (180) of the present disclosure can design a Cost Function such that the sum of L1, which is the error between the output value of the first acoustic model and the correct value, and L2, which is the error between the output value of the second acoustic model and the correct value, is minimized, and thus train the acoustic model so that the error is minimized.
[0211] FIG. 8 is a learning example of the present disclosure.
[0212] Figure 8 is an example of training an acoustic model in a conversation in which languages from various countries are mixed in the utterance.
[0213] The processor (180) of the present disclosure can receive voice data consisting of a conversation containing languages of several countries (S810). The processor (180) can extract feature data of the acquired voice data (S820, S830). The processor (180) can segment the languages of different countries included in the conversation using the voice features of the voice data.
[0214] The processor (180) of the present disclosure can train an acoustic model using a phoneme set corresponding to a plurality of segmented voice data.
[0215] The processor (180) can change the branch network connected to the shared network according to the country corresponding to the segmented multiple voice data, and train the acoustic model for each country using the changed acoustic model.
[0216] And, the processor (180) can set the Cost Function of the acoustic model such that the sum of the errors between each phoneme set output by the acoustic model and the correct answer value corresponding to each phoneme set is minimized based on the input of multiple segmented voice data to the acoustic model. Through the above process, the performance of the acoustic model can be improved regardless of whether the speaker is a native or non-native speaker.
[0217] Additionally, the processor (180) may assign weights to the errors of each phoneme set output by the acoustic model and the correct answer value corresponding to each phoneme set based on the input of a plurality of segmented voice data to the acoustic model, and set the Cost Function of the acoustic model such that the sum of the weighted errors is minimized. Through the above process, the performance of the acoustic model of the language that the user wants to focus on can be improved.
[0218] For example, let's assume that the processor (180) receives voice data consisting of a conversation containing the languages of several countries, such as "(British speech) Hello", "(Indian speech) I'm mohhamad", and "(American speech) Nice to meet you".
[0219] The processor (180) can segment the voice data into "Hello", "I'M MOHHAMAD" and "Nice to meet you" from the voice data.
[0220] The processor (180) can obtain a phoneme set corresponding to each segmented speech data. Specifically, the processor (180) can obtain a British English phoneme set (e.g., BEEP) "hh ax l ow" corresponding to "hello". And it can obtain an Indo-English phoneme set (e.g., X-SAMPA) "aI mm @ h { mm I d" corresponding to "I'm mohhamad" and an American English phoneme set (e.g., CMU) "N AY1 ST UW1 M IY1 TY UW1" corresponding to "nice to meet you".
[0221] The processor (180) of the present disclosure can train an acoustic model using a phoneme set corresponding to a plurality of segmented voice data.
[0222] Specifically, the processor (180) can change the branch network connected to the shared network according to the country of the segmented multiple voice data. And, using the changed acoustic model, it can train an acoustic model corresponding to each language.
[0223] And the processor (180) of the present disclosure can train the acoustic model by applying weights to the loss (L1) for the US voice data “nice to meet you”, the loss (L2) for the UK voice data “hello”, and the loss (L3) for the Indian voice data “I'm mohammad” while inputting the cost function of the acoustic model, so that the total cost of all said weights is minimized.
[0224] FIG. 9 is an example of use of the present disclosure.
[0225] Figure 9 is a diagram illustrating the use of an acoustic model after training.
[0226] The processor (180) of the artificial intelligence device (100) can receive voice data and extract features of the voice data (S910, S920).
[0227] The processor (180) of the present disclosure inputs the voice data into an updated acoustic model including a shared network and a native speaker branch network updated according to the learning, and can perform voice recognition using a native speaker phoneme set output by the updated acoustic model.
[0228] Specifically, the processor (180) of the present disclosure may configure an acoustic model such that a shared network and a native speaker branch network are connected in order to perform speech recognition. The processor (180) may input acquired speech data into the shared network included in the acoustic model that has been updated according to learning. (S930)
[0229] And the shared network of the present disclosure outputs a hidden expression that reflects the acoustic information of the input voice data and can transmit the hidden expression to a layer of the native speaker branch network (S940).
[0230] The native speaker branch network of the present disclosure can output a native speaker phoneme set that reflects the features of the input speech data based on the hidden representation transmitted by the layer of the shared network (S950).
[0231] The present disclosure uses branch networks connected to a shared network differently for each language during learning, but uses an updated acoustic model by connecting a native speaker branch network and a shared network regardless of the input language during use.
[0232] Therefore, when used, the result output by the updated acoustic model is a native speaker phoneme set, and when decoding the speech recognition model, a speech recognition result corresponding to the native speaker phoneme set can be derived regardless of language (S970).
[0233] In another embodiment of the present disclosure, the input unit (120) of the artificial intelligence device (100) can receive user input selecting a native speaker or a non-native speaker.
[0234] The processor (180) can modify the updated acoustic model based on the user's selection so that the shared network and the branch network corresponding to the user's selection are connected.
[0235] The processor (180) can input voice data into a modified acoustic model and perform voice recognition using the phoneme set output by the modified acoustic model.
[0236] For example, a British person (referred to as a non-native speaker for convenience in this specification) can change a native speaker branch network included in an acoustic model into a British language branch network by using the input unit (120) of the artificial intelligence device (100). The British language branch network can be connected to a shared network to construct an acoustic model.
[0237] At this time, when voice data corresponding to the user's speech is received, the processor (180) of the artificial intelligence device (100) can input the voice data into an acoustic model. The acoustic model can output a phoneme set corresponding to British English (e.g., BEPP).
[0238] Through the above process, the artificial intelligence device (100) can derive a voice recognition result using a phoneme set familiar to the user when performing voice recognition.
[0240] The above operations may be performed simultaneously and are not bound by the order in which they are performed. Furthermore, the present disclosure may consist of software, firmware, or a combination of software or firmware. Differences relating to such variations and applications should be interpreted as being included within the scope of the invention as defined in the appended claims.
[0241] The above-described disclosure can be implemented as computer-readable code on a medium on which a program is recorded. A computer-readable medium includes all types of recording devices in which data that can be read by a computer system is stored. Examples of computer-readable media include a Hard Disk Drive (HDD), a Solid State Disk (SSD), a Silicon Disk Drive (SDD), ROM, RAM, a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc. Additionally, the computer may include a processor (180) of a terminal.
[0242] Furthermore, although the above description has focused on the services and embodiments, this is merely illustrative and does not limit the invention. Those skilled in the art will understand that various modifications and applications not exemplified above are possible within the scope of the essential characteristics of the services and embodiments. For example, each component specifically shown in the embodiments may be modified and implemented. Differences related to such modifications and applications should be interpreted as being included within the scope of the invention as defined in the appended claims.
Claims
Claim 1 An artificial intelligence device for recognizing speech, comprising: an input unit for acquiring speech data corresponding to a user utterance; and a processor for training an acoustic model including a shared network and a branch network connected to the shared network using the speech data and phonemes corresponding to the speech data, wherein the shared network outputs a hidden representation by reflecting common acoustic information of a plurality of speech data used during training, and includes a first artificial intelligence model that transmits the hidden representation to a layer of the branch network. Claim 2 delete Claim 3 An artificial intelligence device according to claim 1, wherein the plurality of voice data includes native speaker voice data or non-native speaker voice data, and the processor trains the shared network and the native speaker branch network connected to the shared network when the native speaker voice data is received, and trains the shared network and the non-native speaker branch network connected to the shared network when the non-native speaker voice data is received. Claim 4 An artificial intelligence device according to claim 1, wherein the branch network includes a second artificial intelligence model that outputs a phoneme set corresponding to the voice data based on a hidden representation transmitted by a layer of the shared network. Claim 5 ◈Claim 5 was abandoned upon payment of the registration fee.◈ In claim 1, the processor is an artificial intelligence device that, when the voice data is voice data spoken by a native speaker, trains the shared network and the native speaker branch network connected to the shared network using the native speaker voice data and the native speaker phoneme set corresponding to the native speaker voice data. Claim 6 ◈Claim 6 was abandoned upon payment of the registration fee.◈ In Claim 1, the processor is an artificial intelligence device that, when the voice data is voice data spoken by a non-native speaker, trains the shared network and the non-native speaker branch network connected to the shared network using the non-native speaker voice data and the non-native speaker phoneme set (phonemes of a foreign language) corresponding to the non-native speaker voice data. Claim 7 An artificial intelligence device for recognizing speech, comprising: an input unit for acquiring speech data corresponding to a user utterance; and a processor for training an acoustic model including a shared network and a branch network connected to the shared network using the speech data and phonemes corresponding to the speech data, wherein the input unit receives second speech data, and the processor inputs the second speech data into an updated acoustic model including a shared network and a native speaker branch network updated according to the training, and performs speech recognition using a native speaker phoneme set output by the updated acoustic model. Claim 8 ◈Claim 8 was abandoned upon payment of the registration fee.◈ An artificial intelligence device according to claim 7, wherein the second voice data includes non-native speaker voice data, and the processor inputs the second voice data into the updated acoustic model and performs voice recognition using a native speaker phoneme set output by the updated acoustic model. Claim 9 An artificial intelligence device according to claim 7, wherein the input unit receives user input selecting a native speaker or a non-native speaker, and the processor modifies the updated acoustic model so that the shared network and the branch network corresponding to the user's selection are connected based on the user's selection, inputs the second voice data into the modified acoustic model, and performs voice recognition using the phoneme set output by the modified acoustic model. Claim 10 A method of operation of an artificial intelligence device for recognizing speech, comprising: a step of acquiring speech data corresponding to a user utterance; and a step of training an acoustic model including a shared network and a branch network connected to the shared network using the speech data and phonemes corresponding to the speech data, wherein the step of training the acoustic model includes a step of outputting a hidden representation by reflecting common acoustic information of a plurality of speech data used during training of the shared network, and a step of transmitting the hidden representation to a layer of the branch network. Claim 11 delete Claim 12 A method of operation of an artificial intelligence device according to claim 10, wherein the plurality of voice data includes native speaker voice data or non-native speaker voice data, and the step of training the acoustic model comprises: a step of training the shared network and the native speaker branch network connected to the shared network when the native speaker voice data is received; and a step of training the shared network and the non-native speaker branch network connected to the shared network when the non-native speaker voice data is received. Claim 13 A method of operation of an artificial intelligence device, wherein the step of training an acoustic model comprises the step of the branch network outputting a phoneme set corresponding to the speech data based on a hidden representation transmitted by a layer of the shared network. Claim 14 ◈Claim 14 was abandoned upon payment of the registration fee.◈ In claim 10, the step of training the acoustic model comprises, when the voice data is voice data spoken by a native speaker, the step of training the shared network and the native speaker branch network connected to the shared network using the native speaker voice data and the native speaker phoneme set corresponding to the native speaker voice data, a method of operation of an artificial intelligence device. Claim 15 ◈Claim 15 was abandoned upon payment of the registration fee.◈ In claim 10, the step of training the acoustic model comprises, when the voice data is voice data spoken by a non-native speaker, training the shared network and the non-native branch network connected to the shared network using the non-native voice data and the non-native phoneme set (phonemes of a foreign language) corresponding to the non-native voice data, a method of operation of an artificial intelligence device. Claim 16 In claim 10, the method of operation of the artificial intelligence device further comprises the steps of: receiving second voice data; inputting the second voice data into an updated acoustic model including a shared network and a native speaker branch network updated according to the learning; and performing voice recognition using a native speaker phoneme set output by the updated acoustic model. Claim 17 ◈Claim 17 was abandoned upon payment of the registration fee.◈ Method of operation of an artificial intelligence device, wherein the step of performing the voice recognition according to claim 16 comprises, when the second voice data is voice data spoken by a non-native speaker, the step of performing voice recognition using a native speaker phoneme set output by the updated acoustic model. Claim 18 A method of operation of an artificial intelligence device according to claim 16, further comprising: receiving user input selecting a native speaker or a non-native speaker; modifying the updated acoustic model so that the shared network and the branch network corresponding to the user's selection are connected based on the user's selection; the step of inputting the second voice data includes the step of inputting the second voice data into the modified acoustic model; and the step of performing voice recognition includes the step of performing voice recognition using the phoneme set output by the modified acoustic model.
Citation Information
Patent Citations
Method and apparatus for generating acoustic models for speaker independent speech recognition of foreign words uttered by non-native speakers
US20050197835A1
Improving Automatic Speech Recognition of Multilingual Named Entities
US20170287474A1