Minimizing Computational Requirements in Model-Agnostic Cross-Lingual Transfer Using Neural Task Representations as Weak Supervision
By sharing high-level model architecture in the cross-language transfer framework and training neural models with labeled and unlabeled loss functions, the high-cost and time-consuming problems in the existing technology are solved, and low-cost cross-language model transfer and accurate prediction are achieved.
Patent Information
- Application Number
- CN201980067534.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-10-18
- Filing Date
- 2019-10-11
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2039-10-11
AI Technical Summary
The prior art transfer of neural models from one language to another is expensive and time-consuming, especially due to the dependence on high-quality annotated data and the dependence on machine translation or bilingual dictionaries.
Using a cross-language transfer framework, the neural models of the second language are trained by training the neural models on the labeled data of the first language and using labeled and unlabeled loss functions on parallel data between the first and second languages, sharing a high-level model architecture, relying only on parallel data and defined loss functions, avoiding dependence on translation systems and dictionaries.
Effectively reduces the currency and computational costs of transferring neural models from one language to another, while maintaining prediction accuracy, suitable for various neural architectures without the need for training data for the target language.
Smart Images

Figure CN112840344B_ABST
Abstract
Description
Technical Field
[0001] The present subject matter generally relates to transferring a neural model from one language to a second language. More specifically, the present subject matter relates to transferring a neural model from one language to a second language using representation projection as weak supervision. Background Art
[0002] Currently, natural language processing is largely centered around English, and the need for models that work in languages other than English is greater than ever. However, the task of transferring a model from one language to another can be costly in terms of factors such as annotation costs, engineering time, and effort.
[0003] Current research in natural language processing (NLP) and deep learning has produced systems that can achieve human parity in several key research areas such as speech recognition and machine translation. That is, these systems perform at the same or higher level as humans. However, much of this research has centered around English-centric models, methods, and data sets.
[0004] It is estimated that only about 350 million people are native English speakers, while another 500 million to 1 billion people speak English as a second language. This accounts for at most 20% of the world's population. As language technology enters people's digital lives, there is a need for NLP applications that can understand the other 80% of the world's languages. However, building such systems from scratch can be expensive, time-consuming, and technically challenging. Summary of the Invention
[0005] According to one aspect of the present technology, a method for cross-lingual neural model transfer may include: training a first neural model of a first language having multiple layers on annotated data of the first language based on a token-based loss function, wherein training the first neural model includes defining and updating parameters of each layer in the layers of the first neural model; and training a second neural model of a second language having multiple layers on parallel data between the first language and the second language based on an unlabeled loss function, wherein training the second neural model includes copying all layers of the first neural model except the lowest layer, and defining and updating parameters of the lowest layer of the second neural model.
[0006] The training may be a two-stage training process, wherein the first model is fully trained before training the second model, or alternatively, in a joint training process, the first model and the second model may be co-trained after an initial training of the first model.
[0007] The following description and the drawings set forth certain illustrative aspects of the claimed subject matter. However, these aspects merely indicate some of the various ways in which the principles of the present invention may be employed, and the claimed subject matter is intended to encompass all such aspects and their equivalents. Other advantages and novel features of the claimed subject matter will become apparent from the following detailed description of the invention when considered in conjunction with the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Non-limiting and non-exhaustive examples are described with reference to the following drawings.
[0009] Figure 1 A framework for cross-lingual neural model transfer according to an embodiment is shown;
[0010] Figure 2 A neural model architecture according to an embodiment is shown;
[0011] Figure 3 A flowchart depicting a method for cross-lingual neural model transfer according to an embodiment is shown;
[0012] Figure 4 A flowchart depicting a method for cross-lingual neural model transfer according to another embodiment is shown;
[0013] Figure 5 An exemplary block diagram of a computer system in which an embodiment may be implemented is shown. DETAILED DESCRIPTION
[0014] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof, and in which are shown by way of illustration specific embodiments or examples. These aspects may be combined, other aspects may be utilized, and structural changes may be made without departing from the present disclosure. Embodiments may be practiced as a method, system, or apparatus. Thus, embodiments may take the form of a hardware implementation, a fully software implementation, or an implementation combining software and hardware aspects. Accordingly, the following detailed description should not be construed as limiting, and the scope of the present disclosure is defined by the appended claims and their equivalents.
[0015] One reason that building an NLP system from scratch is expensive, time-consuming, and technically challenging is that high-performance NLP models typically rely on large amounts of high-quality annotated data, which comes at the cost of annotator time, effort, and money. The data being annotated is some linguistic artifact (e.g., any text) that is annotated with some additional artifact. For example, the text can be checked against a criterion, and a label or annotation can be added to the text based on that criterion. By way of example, the criterion can be sentiment, and the label or annotation can include positive sentiment or negative sentiment.
[0016] Other exemplary guidelines include style classification, where the label can include whether the artifact is formal or informal; intent understanding, where the label can include a prediction of the intent of the artifact selected from a plurality of predefined intents (such as scheduling an event, requesting information, or providing an update); message routing, where the label can include a prediction of the primary recipient among a plurality of recipients; task duration, where the label can include a prediction of the event duration; or structured content recognition, where the label can include a prediction of the artifact category (such as classifying an email into categories such as flight itinerary, shipping notice, or hotel reservation).
[0017] Given the huge cost of building a system from scratch, many efforts in the research community to build tools for other languages rely on transferring existing English models to other languages.
[0018] Previous efforts to transfer English models to other languages relied on machine translation (MT) to translate training data or test data from English to the target language. Other efforts also considered leveraging bilingual dictionaries to directly transfer features.
[0019] Building a state-of-the-art MT system requires expertise and a large amount of training data, which is expensive. At the same time, building a bilingual dictionary can be equally expensive if done manually and contains significant noise if introduced automatically.
[0020] Other research includes the study of the transferability of neural network components in the context of image recognition. This study illustrates the technical problems in traditional techniques, namely that the higher layers of the network tend to be more specialized and domain-specific and thus less generalizable.
[0021] However, the technical solution according to an embodiment of the present disclosure includes a framework for cross-language transfer in the opposite direction: specifically, the higher layers of the network are shared between models in different languages while maintaining separate language-specific embeddings (i.e., the parameters of the lower layers of the network). By sharing the higher layers of the network, accurate models can be generated in multiple languages without relying on MT, bilingual dictionaries, or data annotated in the model language.
[0022] Cross-domain sharing of information is also related to multi-task learning. The work in this area can be roughly divided into two methods: hard parameter sharing and soft parameter sharing. In hard parameter sharing, the model shares a common architecture with some task-specific layers, while in soft parameter sharing, the tasks have their own parameter sets, which are constrained by some sharing costs.
[0023] Previous studies including label projection, feature projection, and weak supervision are different from the embodiments of the present disclosure, which are attracted to a neural framework that integrates task characterization, model learning, and cross-language transfer in a joint scheme but at the same time has sufficient flexibility to adapt to various target applications.
[0024] When solving the technical problems faced by traditional technologies, embodiments of the general framework of the present disclosure can easily and effectively transfer a neural model from one language to other languages. On the one hand, the framework relies on task representation as a form of weak supervision and is model- and task-agnostic. Generally, a neural network includes a series of nodes that are arranged in layers including an input layer and a prediction layer. The part of the neural network between the input layer and the prediction layer may include one or more layers that transform the input into a representation. Each layer after the input layer is trained on the previous layer, so the feature complexity and abstraction of each layer increase. The task representation captures an abstract description of the prediction problem and is embodied as a layer before the prediction layer in the neural network model. By leveraging the disclosed framework, many existing neural architectures can be transplanted into other languages with minimal effort.
[0025] The only requirements for transferring a neural model according to embodiments of the present disclosure are parallel data and a loss defined on the task representation.
[0026] The framework according to embodiments of the present disclosure can reduce monetary and computational costs by the following: abandoning the reliance on machine translation or bilingual dictionaries while accurately capturing semantically rich and meaningful representations across various languages. By eliminating any reliance on or interaction with translation means, the framework can reduce the number of instructions processed by the processor, thereby increasing system speed, saving memory, and reducing power consumption.
[0027] Regarding these and other general considerations, embodiments of the present disclosure are described below. Additionally, although relatively specific problems have been discussed, it should be understood that the embodiments should not be limited to solving the specific problems identified above.
[0028] Hereinafter, a framework according to an embodiment is described, which can transfer an existing neural model in a first language to a second language with minimal cost and effort.
[0029] Specifically, the framework: (i) is model- and task-agnostic and thus applicable to various new and existing neural architectures; (ii) only requires a parallel corpus and does not require target language training data, a translation system, or a bilingual dictionary; (iii) has a unique modeling requirement of defining a loss on the task representation, thereby greatly reducing the engineering effort, monetary cost, and computational cost involved in transferring a model from one language to another.
[0030] The embodiments are particularly useful when a high-quality MT system is not applicable to the target language or a specialized domain. Traditionally, an MT system, a bilingual dictionary, or a pivot lexicon is required to transfer a model from one language to another; however, according to the embodiments, none of these are required to accurately predict results at a rate comparable to or even exceeding that of traditional solutions.
[0031] A framework for transferring a neural model from a first language to a second language according to an embodiment is described in more detail. For example, an embodiment in which the first language is English and the second language is French is shown and described. Of course, the present technology is not limited to this, and it should be understood that the only limitation on the first language and the second language is that they are not the same dialect of the same language.
[0032] Figure 1 An exemplary framework 100 for transferring an English neural model 200 to a French neural model 300 is shown. As Figure 1 shown, the framework includes a training part 101 or module and a testing part 102 or module. Figure 1 Implementations of joint training and two-stage training are depicted and will be discussed in detail below.
[0033] The training part 101 depicts the English neural model 200 and the French neural model 300. As Figure 1 shown, the training part 101 depicts how the English neural model 200 is trained and how the English neural model 200 is transferred to the French neural model 300. The training part 101 of the framework 100 utilizes the labeled English data D L and the unlabeled parallel data D P , and the unlabeled parallel data D P includes English parallel data PE and French parallel data PF.
[0034] Labeled data is data that is typically directly supplemented with context information by humans and can also be referred to as annotated data.
[0035] As long as the parallel data is aligned between languages, it can be aligned at any level, including character level, word level, sentence level, paragraph level, or other levels.
[0036] According to the example embodiments as Figure 1 and 2 shown, the labeled English data D L is provided to the English neural model 200.
[0037] The English neural model 200 can be a neural NLP model and can include three different components: an embedding layer 201, a task-appropriate model architecture 202, and a prediction layer 203.
[0038] More specifically, the English neural NLP model 200 includes a first layer, namely an embedding layer 201, which transforms a language unit w (character, word, sentence, paragraph, pseudo-paragraph, etc.) into a mathematical representation of the language unit w. The mathematical representation may preferably be a dense representation of a vector mainly including non-zero values, or alternatively may be a sparse representation of a vector including many zero values.
[0039] The third layer is a prediction layer 203, which is used to generate a probability distribution over the space of output labels. According to an exemplary embodiment, the prediction layer 203 may include a softmax function.
[0040] Between the prediction layer 203 and the embedding layer 201 is a task-appropriate model architecture 202.
[0041] Since the framework 100 is model- and task-agnostic, the structure of the task-appropriate model architecture 202 may include any number of layers and any number of parameters. That is, the task-appropriate model architecture 202 is what makes the model suitable for a specific task or application, and the configuration and number of layers of the network do not affect the application of the general framework.
[0042] Therefore, for simplicity, the task-appropriate model architecture 202 is depicted as including an x-layer network 202a (where x is a non-zero integer of layers), and a task representation layer 202b that is the layer immediately preceding the prediction layer 203.
[0043] As Figure 1 shown, the test part 102 includes a French model 300, a French embedding layer 301, a task-appropriate model architecture 302, and a prediction layer 303. According to an embodiment of the framework 100, the test part 102 represents classifying unlabeled French data D F using the French model 300.
[0044] Figure 2 Shows an example of the model architectures of the neural models 200 and 300 according to one embodiment. The neural models 200 and 300 may be configured as a hierarchical recurrent neural network (RNN) 400, but it should be understood that this is only exemplary and the architecture is not limited thereto.
[0045] As Figure 2 shown, the data set is embedded by the embedding layer 401 into a sequence of language units w 11 -w nm in. The sequence of language units w 11 -w nmIs converted by the sentence RNN 402 into a sentence representation 403, and the sequence of sentence representations 403 is converted by the review-level RNN 404 into a task representation 405 by the review RNN 404. The task representation 405 is then converted into a prediction layer 406, which is used to generate a probability distribution over the space of output labels 407. The number of output labels 407 is equal to the number of results of the prediction task.
[0046] According to one embodiment, the RNN may include, for example, a gated recurrent unit (GRU). However, it should be understood that the present disclosure is not limited thereto, and the RNN may also be a long short-term memory network (LSTM) or other network.
[0047] The model transfer according to the embodiment depends on two features. First, the architecture and prediction layer suitable for the task are shared across languages. Second, all the information required for successful prediction is included in the task representation layer.
[0048] As Figure 1 shown, in the case where the English model 200 is transferred to the French model 300, the only difference between the English model 200 and the French model 300 is the language-specific embeddings included in the embedding layers 201, 301 of the English model 200 and the French model 300, as shown by the contrasting shaded lines of the embeddings. Second, the task representation layers 204, 304 of the English model 200 and the French model 300 contain all the information required for successful prediction.
[0049] An indication of successful model transfer is that when considering parallel data, the French model and the British model will predict the same thing. That is, the content of the prediction is irrelevant, but the success of the model transfer is based on the similarity of the predictions of the French model and the English model. When the scenario is label projection, the content of the prediction can be the actual label. Alternatively, in the case where the goal is to generate the same task representation in two languages, representation projection can be utilized. Compared with label projection-based supervision, representation projection is a softer form of weak supervision and is the preferred projection according to the embodiment.
[0050] To better illustrate the framework according to the embodiment, consider the task T and the labeled data D L ={(x i , y i )|0≤i≤N}, where x i is the English input, y i is the output with K possible values, such that each x i is annotated with the value y i , and N is the number of language units included in the labeled data D L . Without loss of generality, assume the input x i ={e il,..., e il} is a sequence of English words. In addition, the parallel data set D P = {(e j , f j ) | 0 ≤ j ≤ M}, where e j = {e jl ,..., e jl} and f j = {f jl ,..., f jl} are parallel English and French language units respectively, and M is the number of language unit pairs included in the parallel data D P .
[0051] The English embeddings included in the English embedding layer 201 can be expressed as such that each word in the English vocabulary V E has a vector The English vocabulary includes all the words found in the input x i . The French embeddings included in the French embedding layer 301 can be expressed as such that each word in the French vocabulary V F has a vector
[0052] In the case of a shared model architecture, the dimensions d of the vectors and must be the same. The mapping of the English sequence e j = {e j1 ,…, e jm} to a sequence of vectors is expressed as and the mapping of the French sequence f j = {f j1 ,…, f jn} to a sequence of vectors is expressed as The x-layer model 202b is expressed as μ with parameters θ μ , which takes the embedding sequence as input and produces a task representation. Specifically, for the English input x i , the task representation is expressed as:
[0053]
[0054] Finally, the prediction layer 203 is expressed as π with parameters θ π , which produces a probability distribution over K output variables:
[0055]
[0056] where π k is the k-th neuron of this layer, and the shorthand is used to represent Then, two losses are optimized according to the framework of an embodiment.
[0057] Loss of the tokens: Assume that the model takes as input the tokenized English data D L Then the following loss is optimized for the combined network:
[0058]
[0059] where ΔL is a loss function defined between and the variable y i For example, in the binary case, Δ L might be the cross-entropy loss, although it should be understood that this is merely exemplary and the framework is not limited to this.
[0060] Loss of the untokens: The English task representations generated by the model are used as weak supervision for the parallel data on the French side. Specifically:
[0061]
[0062] where ΔP is a loss function between the task representations generated on the parallel inputs. Since the task representations are vectors, the mean squared error between them might be an appropriate loss, for example, although the framework is not limited to this.
[0063] Then, the final optimization is given by where α is a hyperparameter that controls the mixing strength between the two loss components.
[0064] Contrary to conventional frameworks, in the framework according to the embodiment, there is no requirement for MT because neither the training nor the test data has been translated. Nor are any other resources used, such as a pivot dictionary or a bilingual dictionary. The only requirements are the parallel data and the definition of the loss function The model architecture μ and the loss of the tokens are properties defined for the English-only model.
[0065] Using the well-defined loss functions Δ L and Δ P , the training consists of backpropagating the error through the network and updating the parameters of the model.
[0066] Figure 3 and Figure 4 show two methods for transferring a neural model from a first language to a second language according to an embodiment. Specifically, Figure 3 shows a two-stage training method, Figure 4 shows a joint training method.
[0067] As Figure 3As shown, in two-stage training, the model architecture is defined in step S301. Since the framework is model-agnostic, the model can be defined as shown in Figure 2 , but it should be understood that the framework is not limited to this way.
[0068] The labeled loss is defined in step S302 and in step S303, by finding the first model 200 is trained on the labeled data D of the first language L . Here, "*" represents the optimized value for the arg max function in step S303.
[0069] After the first model 200 is trained, in step S304, the embedding U of the first model and the shared model parameters θ μ and θ π are frozen.
[0070] The unlabeled loss is defined in step S305 and in step S306, by optimizing the unlabeled loss is trained on the parallel data D P . That is, in the second stage of two-stage training, only the second embedding V of the second model is updated on the parallel data.
[0071] In step S307, the first embedding U of the embedding layer 201 of the first model 200 is replaced with the second embedding V of the embedding layer 301 of the second model 300. The combined model is the updated second model 300. Therefore, the updated second model 300 includes the parameters V*, θ μ , θ π .
[0072] As Figure 4 shown, in joint training, the model architecture is defined in step S401. Since the framework is model-agnostic, the model can be defined as shown in Figure 2 , but it should be understood that the framework is not limited to this way.
[0073] The labeled loss is defined in step S402 and in step S403, by finding the labeled loss is trained on the labeled data D L .
[0074] The unlabeled loss is defined in step S404 and in step S405, by optimizing to train the unlabeled loss on the parallel data D P . L is a weighted combination of the labeled loss and the unlabeled loss, and is given by is given, where α is a hyperparameter that controls the mixing strength between the two loss components.
[0075] In joint training, when processing parallel data D P the parameters of both the first model 200 and the second model 300 are updated in step S404.
[0076] In step S406, the first embedding U of the embedding layer 201 of the first model 200 is replaced with the second embedding V of the embedding layer 301 of the second model 300. This combined model is the updated second model 300. Thus, the updated second model 300 includes the parameters V*, θ* μ , θ* π .
[0077] Example model transfer: Sentiment classification
[0078] To better illustrate the general framework according to the embodiments, in the following illustrative example, a sentiment classifier is transferred from one language to another.
[0079] In this example, the sentiment classifier predicts whether a language artifact is positive or negative. According to the embodiments, the only necessary steps are to define the model architecture μ and two loss functions and
[0080] Given the binary nature of the prediction task, the prediction layer can be given as a sigmoid layer with one output neuron that computes the probability of the positive label: The labeled loss may be the cross-entropy loss:
[0081]
[0082] On the parallel side, the unlabeled loss may be the mean squared error loss:
[0083]
[0084] where d T is the dimension of the task representation R T , and R T (i) represents its i-th dimension.
[0085] Although the above example defines loss functions for a binary system, it should be understood that other loss functions can be defined for other systems, and the system can have any number of possible outputs.
[0086] Cross-lingual word associations
[0087] To demonstrate that the task representation is weakly supervised, Table 1 shows several English words with sentiment in the joint model according to one embodiment, and their closest French neighbors (by vector cosine distance on their respective embeddings).
[0088]
[0089] Table 1
[0090] As can be seen from Table 1 above, the definitions of positive (or negative) sentiment terms in English are similar to the positive (or negative) terms of their closest neighbors in French. Although the closest neighbor terms in French are not necessarily direct translations, or even synonyms, the sentiment prediction task does not require translation; it is sufficient to identify words that respond to the same sentiment. Thus, the framework for model transfer according to the embodiment is able to identify cross - language sentiment similarities without direct supervision and using only weak fuzzy signals from the representation projection.
[0091] Machine translation is utilized
[0092] Although the framework does not require MT, according to one embodiment, MT can be utilized.
[0093] For example, training - time translation (TrnT) can be utilized, which translates the training data from the first language into another language and then trains the sentiment model in that language. Test - time translation (TstT) can be utilized, which trains the sentiment model in the first language and uses the trained sentiment model to classify language artifacts that are translated into the first language at test time.
[0094] Thus, the framework according to the embodiment can optionally be used in combination with a translator even if there may not even be a translation engine.
[0095] Multimodal model transfer
[0096] The framework can be applied to multimodal (rather than multilingual) transfer. That is, the model can be transferred between different modalities including language, image, video, audio clip, etc. For example, sentiment understanding can be transferred to images without explicit image annotation. In such multimodal transfer, the annotated data can include labeled sentiment data in the first language. The parallel data can include images with captions in the first language. Once the framework is trained on the annotated parallel data, the framework can predict the sentiment of images without captions.
[0097] Figure 5FIG. shows a schematic diagram of an exemplary computer or processing system that can implement any of the systems, methods, and computer program products described in the embodiments of the present disclosure herein, such as the English neural model 200 and the French neural model 300. This computer system is only an example of a suitable processing system and is not intended to impose any limitation on the scope of use or functionality of the embodiments of the methods described herein. The illustrated processing system can operate with many other general-purpose or special-purpose computing system environments or configurations. Well-known examples of computing systems, environments, and / or configurations that may be suitable for use with Figure 5 the processing system shown in FIG. may include, but are not limited to, personal computer systems, server computer systems, thin clients, fat clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments including any of the above systems or devices, etc.
[0098] The computer system may be described in the general context of computer system-executable instructions, such as program modules, executed by the computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types. The computer system may be practiced in a distributed cloud computing environment where tasks are performed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules may be located in both local and remote computer system storage media including memory storage devices.
[0099] The components of the computer system may include, but are not limited to, a server 500, one or more processors or processing units 510, and a system memory 520. The processing unit 510 may include software modules that execute the methods described herein. The modules may be programmed into the integrated circuit of the processing unit 510, or may be loaded from the memory 520 or a network (not shown), or a combination thereof.
[0100] The computer system may include various computer system-readable media. Such media may be any available media accessible by the computer system and may include volatile and non-volatile media, removable and non-removable media.
[0101] Volatile memory may include random access memory (RAM) and / or cache memory or other memory. Other removable / non-removable, volatile / non-volatile computer system storage media may include a disk drive for reading from and writing to a removable non-volatile disk (such as a "floppy disk"), and an optical disk drive that can provide a means for reading or writing a removable non-volatile optical disk such as a CD-ROM, DVD-ROM, or other optical media.
[0102] As will be understood by those skilled in the art, various aspects of a framework can be embodied as a system, a method, or a computer program product. Accordingly, various aspects of the disclosed technology may take the form of: a full hardware embodiment, a full software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, which may generally be collectively referred to herein as "circuitry", "module", or "system". In addition, aspects of the disclosed technology may take the form of a computer program product embodied in one or more computer-readable media having embodied thereon computer-readable program code.
[0103] Any combination of one or more computer-readable media may be utilized. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium would include the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium may be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0104] A computer-readable signal medium may include, for example, a propagated data signal embodied in baseband or as part of a carrier wave, the propagated data signal having computer-readable program code embodied therein. Such a propagated signal may take any of a variety of forms, including but not limited to electromagnetic, optical, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium that is not a computer-readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0105] Any suitable medium may be used to send the program code embodied on the computer-readable medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0106] Computer program code for performing operations in aspects of the disclosed technology can be written in any combination of one or more programming languages, including: object-oriented programming languages such as Java, Smalltalk, C++; and conventional procedural programming languages such as the "C" programming language or similar programming languages; scripting languages such as Perl, VBS or similar languages; and / or functional languages such as Lisp and ML; and logic-oriented languages such as Prolog. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can establish a connection with an external computer (e.g., through the Internet using an Internet service provider).
[0107] Aspects of the disclosed technology are described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create means for implementing the functions / actions specified in the flowchart and / or block Figure 1 diagram block or blocks.
[0108] These computer program instructions can also be stored in a computer-readable medium that can direct a computer, other programmable data processing device, or other device to operate in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture including instructions for implementing the functions / actions specified in the flowchart and / or block Figure 1 diagram block or blocks.
[0109] The computer program instructions can also be loaded onto a computer, other programmable data processing device, or other device to cause a series of operational steps to be performed on the computer, other programmable device, or other device to produce a computer-implemented process, such that the instructions executed on the computer or other programmable device provide a process for implementing the functions / actions specified in the flowchart and / or block Figure 1 diagram block or blocks.
[0110] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of code, which includes one or more executable instructions for implementing the specified (multiple) logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, depending on the functionality involved, two consecutive blocks shown may actually be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by a system based on dedicated hardware that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.
[0111] A computer program product may include all the respective features that are capable of implementing the implementation of the methods described herein and, when loaded into a computer system, are capable of executing the methods. In this context, a computer program, software program, program, or software refers to any expression of a set of instructions in any language, code, or notation, which is intended to cause a system with information processing capabilities to perform a specific function either directly or after any one or both of the following: (a) being converted into another language, code, or notation; and / or (b) being reproduced in a different material form.
[0112] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the disclosure. As used herein, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising", when used in this specification, specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0113] All parts or steps in the following claims, plus the corresponding structures, materials, acts, and equivalents of the functional elements, if any, are intended to include any structure, material, or act for performing the function in combination with other claimed elements as expressly claimed. The description of the disclosed technology has been presented for purposes of illustration and description, but is not intended to be exhaustive or limiting. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the disclosure. The embodiments were chosen and described in order to best explain the principles of the disclosure and its practical application, and to enable others of ordinary skill in the art to understand the disclosure in various embodiments with various modifications that are suited to the particular use contemplated.
[0114] Aspects of the present disclosure may be embodied as a program, software, or computer instructions embodied in a computer or machine-usable or readable medium, which, when executed on a computer, processor, and / or machine, cause the computer or machine to perform the steps of a method. A machine-readable program storage device is also provided, which tangibly embodies an instruction program executable by the machine to perform the various functions and methods described in the present disclosure.
[0115] The systems and methods of the present disclosure may be implemented and run on a general-purpose or special-purpose computer system. The terms "computer system" and "computer network" that may be used in this application may include various combinations of fixed and / or portable computer hardware, software, peripherals, and storage devices. The computer system may include multiple individual components that are networked or otherwise linked to cooperate in execution, or may include one or more stand-alone components. The hardware and software components of the computer system of this application may include and may be included in fixed and portable devices such as desktop computers, laptop computers, and / or servers. A module may be a component of a device, software, program, or system that implements certain "functions", which may be embodied as software, hardware, firmware, electronic circuits, etc.
[0116] Although specific embodiments have been described, those skilled in the art will understand that there are other embodiments equivalent to the described embodiments. Therefore, it should be understood that the present disclosure is not limited to the specific embodiments shown, but is only limited by the scope of the appended claims.
[0117] Concept
[0118] Concept 1: A system for transferring cross-lingual neural models, comprising: a processor and a memory, wherein a first neural model and a second neural model are stored in the memory, wherein the language or dialect of the first neural model is different from the language or dialect of the second neural model; and an operating environment that uses the processor to execute commands to train the first neural model on annotated data based on a labeled loss function to define and update the parameters of each layer in multiple layers of the first neural model; and train the first neural model and the second neural model on parallel data between the first language or dialect and the second language or dialect based on an unlabeled loss function to update each layer in multiple layers of the first neural model and to define and update the parameters of each layer in multiple layers of the second neural model, wherein all layers except the lowest layer of the first neural model are copied to the second neural model.
[0119] Concept 2. A system according to any (multiple) preceding or subsequent concepts, wherein the first neural model comprises: a first embedding layer that converts language units of the first language or dialect into a vector representation; a first task-suited model architecture having a predetermined network configuration including one or more layers; and a first prediction layer, wherein one of the layers included in the first task-suited model architecture is a first task representation layer, and wherein the first task representation layer is immediately before the first prediction layer
[0120] Concept 3. A system according to any (multiple) preceding or subsequent concepts, wherein the second neural model comprises: a second embedding layer that converts language units of the second language or dialect into a vector representation; a second task-suited model architecture having a predetermined network configuration including one or more layers; and a second prediction layer.
[0121] Concept 4. A system according to any (multiple) preceding or subsequent concepts, wherein the tasks of the task-suited model architecture include one of the following: sentiment classification, style classification, intent understanding, message routing, duration prediction, or structured content recognition.
[0122] Concept 5. A system according to any (multiple) preceding or subsequent concepts, wherein the second neural model is trained without data annotated in the second language or dialect.
[0123] Concept 6. A system according to any (multiple) preceding or subsequent concepts, wherein the second neural model is trained without a translation system, a dictionary, or a pivot dictionary.
[0124] Concept 7. A system according to any (multiple) preceding or subsequent concepts, wherein the training resources include data annotated in the first language or dialect and unannotated parallel data in both the first language or dialect and the second language or dialect.
[0125] Concept 8. A computer-implemented method for cross-lingual neural model transfer, comprising: supplying annotated data in a first language to a first neural model in the first language; training the first neural model in the first language on the annotated data based on a token-based loss function to define and update parameters of the first neural model in the first language; supplying unannotated parallel data between the first language and a second language to the first neural model in the first language and a second neural model in the second language; training the first neural model in the first language and the second neural model in the second language on the parallel data to update the parameters of the first neural model in the first language and to define and update parameters of the second neural model in the second language; and merging a portion of the parameters of the first neural model in the first language into the second neural model in the second language.
[0126] Concept 9. A system according to any preceding or subsequent concept, wherein the task of the neural model comprises one of the following: sentiment classification, style classification, intent understanding, message routing, duration prediction, or structured content recognition.
[0127] Concept 10. A system according to any preceding or subsequent concept, wherein the second neural model in the second language is trained without annotated data in the second language, a translation system, a dictionary, and a pivot dictionary.
[0128] Concept 11. A system according to any preceding or subsequent concept, wherein the training resources comprise annotated data in the first language and unannotated parallel data in both the first language and the second language.
[0129] Concept 12. A system according to any preceding or subsequent concept, wherein training the first neural model in the first language on the annotated data to define and update the parameters of the first neural model comprises optimizing the token-based loss function of the first neural model in the first language.
[0130] Concept 13. A system according to any preceding or subsequent concept, wherein training the first neural model in the first language and the second neural model in the second language on the parallel data to update the parameters of the first neural model in the first language and to define and update the parameters of the second neural model in the second language comprises: optimizing a loss function between task representations generated by the first neural model in the first language and the second neural model in the second language on the parallel data.
[0131] Concept 14: A computer-implemented method for cross-lingual neural model transfer, comprising: supplying annotated data in a first language to a first neural model in the first language, training the first neural model in the first language on the annotated data based on a token-based loss function to define and update parameters of the first neural model in the first language; freezing the parameters of the first neural model in the first language; supplying unannotated parallel data between the first language and the second language to the first neural model in the first language and a second neural model in the second language; training the second neural model in the second language on the unannotated parallel data to define and update parameters of the second neural model in the second language; and incorporating a portion of the parameters of the first neural model in the first language into the second neural model in the second language.
[0132] Concept 15. A system according to any preceding or subsequent concept, wherein the task of the neural model comprises one of the following: sentiment classification, style classification, intent understanding, message routing, duration prediction, or structured content recognition.
[0133] Concept 16. A system according to any preceding or subsequent concept, wherein the second neural model in the second language is trained without annotated data in the second language.
[0134] Concept 17. A system according to any preceding or subsequent concept, wherein the second neural model in the second language is trained without a translation system, dictionary, or pivot dictionary.
[0135] Concept 18. A system according to any preceding or subsequent concept, wherein the training resources include annotated data in the first language and unannotated parallel data in both the first language and the second language.
[0136] Concept 19. A system according to any preceding or subsequent concept, wherein training the first neural model in the first language on the annotated data to define and update parameters of the first neural model in the first language comprises: optimizing the token-based loss function of the first neural model in the first language.
[0137] Concept 20. A system according to any preceding or subsequent concept, wherein training the second neural model in the second language on the unannotated parallel data to define and update parameters of the second neural model in the second language comprises: optimizing an unlabeled loss function between task representations, the task representations being generated by the first neural model in the first language and the second neural model in the second language on the unannotated parallel data.
Claims
1. A system for transferring cross - language neural models, comprising: A processor and a memory, wherein the memory includes instructions that, when executed by the processor, cause the processor to perform actions, the actions including: Training a first neural model on labeled data in a first language or dialect such that the first neural model, when trained, performs a classification task on a first text received in the first language or dialect, wherein training the first neural model includes learning first parameters of a first embedding layer that receives the first text in the first language or dialect and converts the first text into a first embedding, wherein the first embedding layer is the lowest layer in the first neural model; and Training a second neural model on parallel data between the first language or dialect and a second language or dialect based on an unlabeled loss function such that, when trained, the second neural model performs the classification task on a second text received in the second language or dialect, wherein training the second neural model includes learning second parameters of a second embedding layer of the second neural network that receives the second text in the second language or dialect and converts the second text into a second embedding, wherein the second embedding layer is the lowest layer in the second neural model; and Updating the first neural model by replacing the first embedding layer with the second embedding layer such that, when updated, the first neural model performs the classification task on the second text received in the second language or dialect.
2. The system according to claim 1, wherein the first embedding layer converts language units in the first language or dialect into a vector representation, and wherein the first neural model further includes: A first task - suitable model architecture having a predetermined network configuration including one or more layers; And A first prediction layer, wherein one of the one or more layers included in the first task - suitable model architecture is a first task representation layer, and wherein the first task representation layer is immediately before the first prediction layer.
3. The system according to claim 2, wherein the second embedding layer converts language units in the second language or dialect into a vector representation, and wherein the second neural model further includes: A second task - suitable model architecture having a predetermined network configuration including one or more layers; And A second prediction layer.
4. The system according to claim 3, wherein the tasks of the task - suitable model architecture include one of the following: sentiment classification, style classification, intent understanding, message routing, duration prediction, and structured content recognition.
5. The system according to claim 1, wherein the second neural model is trained without data annotated in the second language or dialect.
6. The system according to claim 1, wherein the second neural model is trained without a translation system, a dictionary, or a pivot dictionary.
7. The system according to claim 1, wherein the training resources include annotated data in the first language or dialect and unannotated parallel data in both the first language or dialect and the second language or dialect.
8. A method performed by a computing system, the method comprising: training a first neural model on tokenized data in a first language or dialect such that the first neural model, when trained, performs a classification task on a first text received in the first language or dialect, wherein training the first neural model includes learning first parameters of a first embedding layer that receives the first text in the first language or dialect and converts the first text into a first embedding, and further wherein the first embedding layer is the lowest layer in the first neural model; training a second neural model on parallel data between the first language or dialect and a second language or dialect based on an unlabeled loss function such that, when trained, the second neural model performs the classification task on a second text received in the second language or dialect, wherein training the second neural model includes learning second parameters of a second embedding layer of the second neural network that receives the second text in the second language or dialect and converts the second text into a second embedding, and further wherein the second embedding layer is the lowest layer in the second neural model; and updating the first neural model by replacing the first embedding layer with the second embedding layer such that, when updated, the first neural model performs the classification task on the second text received in the second language or dialect.
9. The method according to claim 8, wherein the first embedding layer converts language units in the first language or dialect into a vector representation, and wherein the first neural model further comprises: a first task-suited model architecture having a predetermined network configuration including one or more layers; and a first prediction layer, wherein one of the one or more layers included in the first task-suited model architecture is a first task representation layer, and wherein the first task representation layer is immediately before the first prediction layer.
10. The method according to claim 9, wherein the second embedding layer converts language units in the second language or dialect into a vector representation, and wherein the second model further comprises: a second task-suited model architecture having a predetermined network configuration including one or more layers; and a second prediction layer.
11. The method according to claim 10, wherein the tasks of the task-suited model architecture include one of the following: sentiment classification, style classification, intent understanding, message routing, duration prediction, and structured content recognition.
12. The method according to claim 8, wherein the second neural model is trained without annotated data in the second language or dialect.
13. The method according to claim 8, wherein the second neural model is trained without a translation system, a dictionary, or a pivot dictionary.
14. The method according to claim 8, wherein the training resources include annotated data in the first language or dialect and unannotated parallel data in both the first language or dialect and the second language or dialect.
15. A non-transitory computer-readable medium comprising instructions that, when executed by a processor, cause the processor to perform actions, the actions including: training a first neural model on tokenized data in a first language or dialect such that the first neural model, when trained, performs a classification task on a first text received in the first language or dialect, wherein training the first neural model includes learning first parameters of a first embedding layer that receives the first text received in the first language or dialect and converts the first text into a first embedding, and further wherein the first embedding layer is the lowest layer in the first neural model; and training a second neural model on parallel data between the first language or dialect and a second language or dialect based on an unlabeled loss function such that, when trained, the second neural model performs the classification task on a second text received in the second language or dialect, wherein training the second neural model includes learning second parameters of a second embedding layer of the second neural network that receives the second text in the second language or dialect and converts the second text into a second embedding, and further wherein the second embedding layer is the lowest layer in the second neural model; and updating the first neural model by replacing the first embedding layer with the second embedding layer such that, when updated, the first neural model performs the classification task on the second text received in the second language or dialect.
16. The non-transitory computer-readable medium according to claim 15, wherein the first embedding layer converts a language unit in the first language or dialect into a vector representation, and wherein the first neural model further includes: a first task-suited model architecture having a predetermined network configuration including one or more layers; and a first prediction layer, wherein one of the one or more layers included in the first task-suited model architecture is a first task representation layer, and wherein the first task representation layer is immediately before the first prediction layer.
17. The non-transitory computer-readable medium according to claim 16, wherein the second embedding layer converts a language unit in the second language or dialect into a vector representation, and wherein the second neural model further includes: a second task-suited model architecture having a predetermined network configuration including one or more layers; and a second prediction layer.
18. The non-transitory computer-readable medium according to claim 17, wherein the tasks of the task-suited model architecture include one of the following: sentiment classification, style classification, intent understanding, message routing, duration prediction, and structured content recognition.
19. The non-transitory computer-readable medium according to claim 15, wherein the second neural model is trained without data annotated with the second language or dialect.
20. The non-transitory computer-readable medium according to claim 15, wherein the second neural model is trained without a translation system, a dictionary, or a pivot dictionary.
Citation Information
Patent Citations
Multilingual, acoustic deep neural networks
US9460711B1