Zero-shot cross-lingual transfer learning
By determining instance weights for source language text units based on target language text units and updating network parameters accordingly, the method facilitates zero-shot cross-lingual transfer learning in NLP models, addressing the challenge of limited labeled data in target languages and enhancing cross-lingual accuracy.
Patent Information
- Application Number
- JP2023510436
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-09-16
- Filing Date
- 2021-09-14
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2041-09-14
AI Technical Summary
Existing natural language processing (NLP) models face challenges in cross-lingual transfer learning, particularly in zero-shot scenarios where labeled data for the target language is unavailable, and most labeled data is in a few languages with English being dominant.
The method involves determining an instance weight for labeled source language text units based on unlabeled target language text units, scaling errors between predicted and correct labels using these weights, and updating network parameters of a prediction neural network model for the target language to provide a zero-shot cross-lingual transfer learning solution.
This approach enables the development of NLP models that can perform text classification and sequence labeling tasks in target languages without labeled data, improving cross-lingual accuracy and expanding language support beyond English.
Smart Images

Figure 0007695017000003 
Figure 0007695017000004 
Figure 0007695017000005
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to zero-shot cross-lingual transfer learning, and more particularly to zero-shot cross-lingual transfer learning using instance weighting.
Background Art
[0002] Most methods for training natural language processing (NLP) models utilize labeled data in a multi-shot or few-shot training process. Most of the labeled data available for training NLP models is available in only a few languages, with English being the most common. Cross-lingual NLP methods attempt to transfer the learning of a model trained in a source language to a target language.
Summary of the Invention
[0003] To facilitate a basic understanding of one or more embodiments of the present disclosure, a summary is presented below. This summary is not intended to identify key or critical elements, nor is it intended to delineate the scope of particular embodiments or the scope of the claims. The sole purpose of this summary is to present concepts in a simplified form as a prelude to the more detailed description that is presented later. In one or more embodiments described herein, a device, system, computer-implemented method, apparatus, or computer program product or a combination thereof is provided that enables zero-shot cross-lingual transfer learning of a natural language processing model.
[0004] Aspects of the present invention relate to providing a prediction model for a target language by determining an example weight of a labeled source language text unit according to a set of unlabeled target language text units, scaling an error between a predicted label for the source language text unit and a correct label for the source language text unit according to the example weight, updating network parameters of a prediction neural network model for the target language according to the error, and providing the prediction neural network model for the target language.
[0005] Aspects of the present invention relate to providing a prediction model for a target language by inputting source language text units and target language text units into a language neural network, generating a source language vector from the source language text units, generating a target language vector from the target language text units, measuring a similarity between a source language vector and a target language vector, forming a scalar example weight for each source language vector, generating a scaled error between a predicted label and a correct label for a source language text unit, calculating an update to network parameters of a downstream prediction neural network model according to the scaled error, calculating an update to network parameters of the language neural network according to the scaled error, and providing a prediction neural network model for the target language including the downstream prediction neural network and the pre-trained language neural network.
[0006] Aspects of the present invention relate to providing a prediction model for a target language by: inputting, by one or more computer processors, a labeled source language text unit and an unlabeled target language text unit into a language neural network; using the language neural network to determine a scalar instance weight for one source language text unit; scaling an error between a label predicted by a downstream neural network for the one source language text unit and a correct label for the one source language text unit; calculating updates to network parameters of the downstream neural network and network parameters of the language neural network; and providing a prediction model for the target language that includes the language neural network and the downstream neural network. Disclosed are methods, systems, and computer-readable media related to providing a prediction model for a target language.
[0007] According to one aspect of the present invention, there is provided a method for providing a prediction model for a target language, the method comprising: determining, by one or more computer processors, an instance weight for one labeled source language text unit according to a set of unlabeled target language text units; scaling, by the one or more computer processors, an error between a predicted label for the one source language text unit and a correct label for the one source language text unit according to the instance weight; updating, by the one or more computer processors, network parameters of the prediction neural network model for the target language according to the error; and providing, by the one or more computer processors, the prediction neural network model for the target language to a user.
[0008] According to another aspect of the present invention, a method for providing a prediction model for a target language, comprising: inputting, by one or more computer processors, a source language text unit and a target language text unit into a pre-trained language neural network; generating, by the one or more computer processors, a source language vector from the source language text unit; generating, by the one or more computer processors, a target language vector from the target language text unit; measuring, by the one or more computer processors, a similarity between one source language vector among the source language vectors and a set of target language vectors among the target language vectors; determining, by the one or more computer processors, a scalar instance weight for each source language vector according to the similarity; generating, by the one or more computer processors, a scaled error between a predicted label for one source language text unit and a correct label for the one source language text unit by using the scalar instance weight for the source language vector; calculating, by the one or more computer processors, an update to network parameters of a downstream prediction neural network model according to the scaled error; calculating, by the one or more computer processors, an update to network parameters of the language neural network according to the scaled error; and providing, by the one or more computer processors, the prediction neural network model for the target language, including the downstream prediction neural network and the language neural network.
[0009] According to another aspect of the present invention, there is provided a method for providing a prediction model for a target language, the method comprising: inputting, by one or more computer processors, a labeled source language text unit and an unlabeled target language text unit into a language neural network; determining, by the one or more computer processors, a scalar case weight of one source language text unit using the language neural network; scaling, by the one or more computer processors, an error between a label predicted by a downstream neural network for the one source language text unit and a correct label for the one source language text unit; calculating, by the one or more computer processors, updates to network parameters of the downstream neural network and network parameters of the language neural network; and providing, by the one or more computer processors, the prediction model for the target language, including the language neural network and the downstream neural network.
[0010] According to another aspect of the present invention, there is provided a computer program product for providing a prediction model for a target language, the computer program product comprising: one or more computer-readable storage devices; and program instructions collectively stored on the one or more computer-readable storage devices, the stored program instructions including: program instructions for determining a case weight of one labeled source language text unit according to a set of unlabeled target language text units; program instructions for scaling an error between a predicted label for the one source language text unit and a correct label for the one source language text unit according to the case weight; program instructions for updating network parameters of the prediction neural network model for the target language according to the error; and program instructions for providing the prediction neural network model for the target language to a user. According to another aspect of the present invention, there is provided a computer system for providing a prediction model for a target language, including one or more computer processors, one or more computer-readable storage devices, and program instructions stored in the one or more computer-readable storage devices for execution by the one or more computer processors. The stored program instructions include program instructions for determining an example weight of one labeled source language text unit according to a set of unlabeled target language text units, program instructions for scaling an error between a predicted label for the one source language text unit and a correct label for the one source language text unit according to the example weight, program instructions for updating network parameters of the prediction neural network model for the target language according to the error, and program instructions for providing the prediction neural network model for the target language to a user. A computer system is provided.
Brief Description of the Drawings
[0011] Through the more detailed description of some embodiments of the present disclosure in the accompanying drawings, the above and other objects, features and advantages of the present disclosure will become more apparent. Note that the same reference numerals basically refer to the same components in the embodiments of the present disclosure.
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Best Mode for Carrying Out the Invention
[0013] With reference to the accompanying drawings that illustrate embodiments of the present disclosure, some embodiments will be described in more detail. However, the present disclosure can be implemented in various manners and should not be construed as being limited to the embodiments disclosed herein.
[0014] In one embodiment, one or more components of the system can solve highly technical problems using hardware and / or software or both (e.g., vectorizing textual units of a source language and a target language, determining the similarity between a source vector and a target vector, calculating instance weights for an example in the source language according to the vector similarity, scaling the loss function of a model for one example according to the calculated instance weights, and upscaling the relevant model gradients of a neural network during back-propagation for correcting the model prediction value for the example). These solutions are not abstract and, due to the processing power required, for example, to facilitate zero-shot cross-lingual model training, cannot be performed as a series of mental acts by a human. Further, some of the processes executed may be performed by a dedicated computer for performing defined tasks related to the training of a predictive language model. For example, a dedicated computer can be used to perform tasks related to zero-shot cross-lingual model training and the like.
[0015] Cross-lingual text-classification learning refers to performing a text classification task on data from a target language using a natural language processing (NLP) text classification model trained with labeled data in a source language. In this specification, zero-shot learning refers to training a cross-lingual transfer model using data in a labeled source language without using labeled data for the target language. Zero-shot learning represents an alternative training path for training NLP models for many languages for which there is no sufficient labeled dataset for training the NLP model. The systems and methods of the present disclosure enable zero-shot transfer learning for developing NLP text classification and sequence labeling models for languages for which there is no labeled dataset, by using instance weighting in the training of the model. The methods of the present disclosure enable the development of models by transfer learning and the enhancement of existing NLP models by backpropagation based on instance weighting.
[0016] As an example, an NLP model configured to classify book reviews in English as positive or negative can be modified using the systems and methods of the present disclosure to accurately classify book reviews in other languages (target languages) without requiring labeled target language data.
[0017] In one embodiment, labeled text units (sentences, paragraphs, documents) from the source language and unlabeled text units from the target language are each processed by a pre-trained language model to generate high-dimensional vector representations of each text unit (instance) of each language. In this embodiment, the pre-trained model (such as the BERT (Bi-directional Encoder Representation from Transformers) model or the RoBERTa (Robustly optimized BERT approach)) vectorizes the input text unit and provides a vector output for each text unit. In this embodiment, source language instances (Xs1, Xs2, Xs3, … Xsn) have associated instance labels (Ys1, Ys2, Ys3, … Ysn) that are provided as ground-truth labels for each of the source language instances during training of the model. Target language instances (Xt1, Xt2, Xt3, … Xtn) do not have associated labels. For either the source language or the target language, each instance includes a sequence of tokens (t1, t2, t3, … tn). In this embodiment, the method trains a model M using labeled source language instances from a language rich in labels (such as English) so as to be able to make predictions for inputs of target language instances.
[0018] In one embodiment, for sentiment or text classification, the model includes a fully-connect neural network layer added to the [CLS] token of the BERT or RoBERTa output. In BERT, the [CLS] token indicates the start of the input sequence. The [CLS] token is inserted at the beginning of the token sequence of each input text unit. In one embodiment, for opinion extraction, the method treats the problem as one of sequence labeling, and the structure includes a fully connected layer shared for a set of tokens. For either structure, the BERT model utilized includes 12 layers. Other more complex machine learning structures may be used for the analysis.
[0019] In one embodiment, the method calculates an instance weight for each input text unit instance of the source language. The input instance weights are determined such that source language instances having a higher similarity to the target language have larger instance weights in the final model. In one embodiment, the instance weights are determined by comparing the similarity between source language text unit instances and each target language instance. In one embodiment, the method determines the similarity between a source language text unit instance and a limited batch of target language text unit instances. As an example, the method determines the similarity between a single source language text unit instance and each set consisting of 4 target language text unit instances. In one embodiment, the method randomly selects source language text units and target language text units from among the available sets of source language text units and target language text units. The available set of source language text units may include text units of each of a plurality of source languages. The method randomly selects a set of target language text units and compares them with the selected source language units.
[0020] This method determines the similarity between the source language example vector hsi selected from the source language input data batch Ds and each target language example vector htj of the corresponding batch Dt of target language examples. Using the determined vector similarity, the example weight for each hsi is determined. The determined example weights affect the change in the network gradient through the gradient descent function applied to update the node weights of both the pre-trained language neural network and the fully connected neural network.
[0021] In one embodiment, this method determines a similarity score si for each example xi in the set Xs of source language input examples. The set of source language examples can include examples from a single language or multiple languages. The input examples within the set are labeled data. For each input data example, this method determines a similarity score between the vector representation hsi of the input example and at least a batch subset of the target language input examples htj. This method sums the similarity scores s = score(i, j) between hsi and each htj of the target language input batch. In this embodiment, this method normalizes the set of summed scores for all source language examples. This method applies the normalized value wi, which is the example weight of the current example, in the gradient descent back-propagation calculation. These calculations update the node weights of both the pre-trained language neural network and the fully connected task neural network. The fully connected neural network predicts the output for the target language examples. This method utilizes a scoring function (such as a cosine similarity function, Euclidean distance function, vector correlation alignment function, or other vector similarity functions) for each pair of vectors his, htj.
[0022] Table 1 shows an example of scoring two English source language input cases (one positive and one negative) against a French target language case, where the French case is finally judged to be positive. The score for the positive English case is 0.5056, which is higher than the score of 0.3647 for the negative English case. Therefore, in this example, as a method to improve the cross-lingual accuracy of the final model, the case weight of the positive English case is set higher than that of the negative English case, so that the positive English case has a greater impact on the training of the neural network.
[0023] [Table 1]
[0024] Table 2 shows an example of scoring two English source language input cases (one positive and one negative) against a French target language case, where the French case is finally judged to be negative. The score for the negative English case is 0.4828, which is higher than the score of 0.3436 for the positive English case. Therefore, in this example, as a method to improve the cross-lingual accuracy of the final model, the case weight of the negative English case is set higher than that of the positive English case, so that the negative English case has a greater impact on the training of the neural network.
[0025] [Table 2]
[0026] Figure 1 is a diagram showing a data flow when training a prediction model of the method of the present disclosure according to an embodiment of the present invention. As shown in the figure, input text units from each of the source language 110 and the target language 120 are passed to a pre-trained language model 130 (such as BERT, RoBERTa, or other known pre-trained language models). The pre-trained language model 130 outputs a vector representation hs of the source language input text unit and a vector representation ht of the target language text unit. At block 140, the method receives the vector representations hs and ht and determines the vector similarity between them. In one embodiment, the method determines the similarity by a cosine similarity function and sums the similarity between the vector representation hs of the source language input text unit 110 and the vector representation ht of each batch of the target language input text unit 120. The method normalizes the summed similarity score. The method applies the normalized similarity score value of each source language case to the instance as an instance weight 142 for scaling the network loss function value of the case. The method uses the instance weight during the calculation of gradient descent used in backpropagation to train the network and minimize the prediction loss function value 155 of the case. The prediction loss function value 155 relates to the difference between the predicted value of the model for the source language input case and the correct value provided for the labeled source language input case. The method uses the instance weight 142 when training both the pre-trained language neural network 130 and the downstream predictive neural network. The gradients of each network are adjusted to minimize the loss function of the downstream predictive neural network that receives input from the language neural network. The training of the network by backpropagation and gradient descent is repeated until the loss function value of the instance no longer improves. After the iteration of gradient descent training including the instance weight is completed, the method provides a combination of the pre-trained language neural network 130 and the downstream predictive neural network 150 as a prediction neural network model for use in an inter-language NLP task using the target language.
[0027] FIG. 2 is a schematic diagram showing an example of network resources related to the implementation of the invention of the present disclosure. The present invention can be implemented in any processor of the elements of the present disclosure that processes an instruction stream. As shown, the network-connected client device 210 wirelessly connects to the server subsystem 202. The client device 204 wirelessly connects to the server subsystem 202 via the network 214. The client devices 204 and 210 are provided with an interlanguage model training program (not shown) and computing resources (processor, memory, network communication hardware) sufficient to execute the program. As shown in FIG. 2, the server subsystem 202 includes a server computer 250. FIG. 2 is a block diagram of the components of the server computer 250 within the network-connected computer system 2000 according to an embodiment of the present invention. Note that FIG. 2 merely illustrates one embodiment and does not imply any limitation regarding the environment in which different embodiments can be implemented. Many changes are possible to the illustrated environment.
[0028] The server computer 250 can include one or more processors 254, a memory 258, a persistent storage 270, a communication unit 252, one or more input / output (I / O) interfaces 256, and a communication fabric 240. The communication fabric 240 enables communication between the cache 262, the memory 258, the persistent storage 270, the communication unit 252, and the input / output (I / O) interface 256. The communication fabric 240 can be implemented in any architecture designed to pass data or control information or both between processors (such as microprocessors, communication and network processors), system memory, peripheral devices, and other hardware components within the system. For example, the communication fabric 240 can be implemented with one or more buses.
[0029] Memory 258 and persistent storage 270 are computer-readable storage media. In this embodiment, memory 258 includes RAM 260. Generally, memory 258 can include any suitable volatile or non-volatile computer-readable storage media. Cache 262 is fast memory and improves the performance of processor 254 by holding recently accessed data and data close to recently accessed data from memory 258.
[0030] The program instructions and data (e.g., inter-language model training program 275) used to implement embodiments of the present invention are stored in persistent storage 270 for each of one or more processors 254 of server computer 250 to execute or access or both via cache 262. In this embodiment, persistent storage 270 includes a magnetic hard disk drive. Instead of or in addition to a magnetic hard disk, persistent storage 270 can include a solid-state hard drive, a semiconductor memory device, a ROM, an erasable programmable ROM (EPROM), a flash memory, or any other computer-readable storage media capable of storing program instructions or digital information.
[0031] The media used by persistent storage 270 may be removable. For example, a removable hard drive may be used for persistent storage 270. Other examples include optical disks, magnetic disks, thumb drives, and smart cards, which are inserted into a drive for transfer to another computer-readable storage media that is also part of persistent storage 270.
[0032] In these examples, the communication unit 252 enables communication with other data processing systems or devices (including the resources of the client computing devices 204, 210). In these examples, the communication unit 252 includes one or more network interface cards. The communication unit 252 may enable communication using either or both physical and wireless communication links. Software distribution programs, as well as other programs and data used to implement the present invention, may be downloaded via the communication unit 252 to the persistent storage 270 of the server computer 250.
[0033] The I / O interface 256 enables input and output of data with other devices connectable to the server computer 250. For example, the I / O interface 256 can enable connection with one or more external devices 290 such as a keyboard, keypad, touch screen, microphone, digital camera, or other suitable input device or combination thereof. The external device 290 can also include, for example, portable computer-readable storage media such as a thumb drive, portable optical disk, portable magnetic disk, and memory card. Software and data (e.g., the interlanguage model program 275 on the server computer 250) used to implement embodiments of the present invention can be stored on such portable computer-readable storage media and loaded into the persistent storage 270 via the I / O interface 256. The I / O interface 256 is also connected to the display 280.
[0034] The display 280 realizes a mechanism for displaying data to the user and can be, for example, a computer monitor. The display 280 can also function as a touch screen (such as the display of a tablet computer).
[0035] Figure 3 is a flowchart 300 showing an example of operations related to the implementation of the present disclosure. After program start, at block 310, the interlanguage model training program 275 of FIG. 2 inputs a set of text units from each of the source language and the target language into a pre-trained language neural network. In one embodiment, text units from multiple source languages are input into the pre-trained language neural network along with text units of the target language. The text units of the source language are labeled data, and the text units of the target language are unlabeled data. In this embodiment, by using source language text units from multiple source languages, a more accurate prediction model for the target language can be obtained. The method randomly selects a single source language text unit from one of the multiple source languages and determines the instance weight of this randomly selected source language text unit using a set of randomly selected unlabeled target language text units as described above.
[0036] The method generates, at block 320, vector representations of the input text units (vector representation hs of the source language and vector representation ht of the target language). The method passes, at block 330, the generated hs and ht to an instance weighting calculator. The instance weight is calculated by first determining the similarity between each hs and a batch of ht. The set of similarities of hs is summed and normalized. The method passes the final value of hs along with the scalar instance weight used for the hs in the backpropagation gradient descent learning of the pre-trained language neural network and the downstream prediction neural network.
[0037] This method evaluates the loss function of the downstream prediction neural network for case hs using the case weight value for case hs at block 340, scales the error, or adjusts the loss function value. This method iteratively adjusts the node weights of each of the pre-trained neural network and the downstream prediction neural network using backpropagation gradient descent.
[0038] As an example, for a binary class, it is as follows. That is, when the predicted label is 0 and the correct label is 0, (in this case, since the prediction is correct) the misclassification error (loss) is 0. When the predicted label is 1 and the correct label is 1, (in this case, since the prediction is correct) the misclassification loss is 0. When the predicted label is 0 and the correct label is 1, (in this case, since the prediction is incorrect) the misclassification loss is 1. When the predicted label is 1 and the correct label is 0, (in this case, since the prediction is incorrect) the misclassification loss is 1.
[0039] Generating a scaled error by case weight, that is, scaling the loss function value, means multiplying the loss (whatever loss is determined from the above rules) by the case weight. For example, in the case of case weight 0.5, predicted label 0, and correct label 1, the scaled error is 0.5. In each iteration, the training / learning algorithm focuses on obtaining a correct prediction for the misclassified cases with the largest case weight values.
[0040] The method returns to block 310 for further input examples and repeats the evaluation of the loss function. This iteration continues until the improvement of the loss function value stops or becomes very small, indicating that the model is sufficiently trained. After training the pre-trained neural network and the downstream prediction neural network, the method provides, at block 350, the combination of the pre-trained neural network and the downstream prediction neural network to the user as a prediction model for predicting the text classification of text sequences in the target language. The method enables zero-shot learning of the prediction model in the target language. The training of the model is completed without using labeled text units in the target language.
[0041] The present disclosure includes a detailed description regarding cloud computing, but it should be understood that the implementation forms of the teachings described herein are not limited to a cloud computing environment. Rather, the embodiments of the present invention can be implemented in combination with any other type of computing environment known now or developed later.
[0042] Steps of the inter-language model training program 275, vectorization of input language examples across the source language and the target language, calculation of example weights for source language inputs, training of the pre-trained neural network and the downstream prediction neural network using backpropagation gradient descent corrected by the calculated example weights, etc. require access to large-scale computing resources. The local computing environment may lack sufficient resources for the program, and the use of edge cloud or cloud resources is required to efficiently execute the necessary programming steps.
[0043] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services), which can be rapidly provisioned and released with minimal management effort or service provider interaction. This cloud model can include at least five characteristics, at least three service models, and at least four deployment models.
[0044] The characteristics are as follows. On-demand self-service: Cloud consumers can unilaterally provision computing capabilities such as server time and network storage automatically as needed, without the need for human interaction with the service provider. Broad network access: Computing capabilities are available over the network and can be accessed via standard mechanisms, which promotes use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, PDAs). Resource pooling: Provider computing resources are pooled and offered to multiple consumers using a multi-tenant model. Various physical and virtual resources are dynamically assigned and reassigned according to demand. In general, consumers have a sense of location independence since they do not manage or have knowledge of the exact location of the provided resources, although they may be able to specify location at a higher level of abstraction (e.g., country, state, data center). Rapid Elasticity: Computing capabilities can be quickly and flexibly provisioned, so that in some cases they can automatically scale out immediately and can be quickly released and scale in immediately. To consumers, the computing capabilities available for provisioning often seem unlimited and can be purchased in any quantity at any time. Measured Services: Cloud systems utilize measurement functions at a certain level of abstraction suitable for service types (e.g., storage, processing, bandwidth, active user accounts) to automatically control and optimize resource usage. It is possible to monitor, control, and report resource usage amounts to provide transparency to both the providers and consumers of the services being utilized.
[0045] The service model is as follows. Software as a Service (SaaS): The functionality provided to consumers is that they can utilize the provider's applications running on the cloud infrastructure. The applications can be accessed from various client devices via a client interface such as a web browser (e.g., webmail). Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, and even individual application functions. However, this does not apply to limited settings of user-specific application configurations. Platform as a Service (PaaS): The functionality provided to consumers is that they can deploy the applications created or obtained by the consumers using the programming languages and tools supported by the provider onto the cloud infrastructure. Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, and storage, but can control the deployed applications and, in some cases, also control the configuration of the hosting environment. Infrastructure as a Service (IaaS): The functions provided to consumers are to prepare processors, storage, networks, and other basic computing resources that enable consumers to deploy and run any software, including operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but can control the operating system, storage, and deployed applications, and in some cases, can partially control some network components (e.g., host firewalls).
[0046] The deployment models are as follows. Private Cloud: This cloud infrastructure is operated exclusively for a specific organization. This cloud infrastructure can be managed by the organization or a third party and can exist on-premises or off-premises. Community Cloud: This cloud infrastructure is shared by multiple organizations and supports a specific community with common concerns (e.g., mission, security requirements, policies, and compliance). This cloud infrastructure can be managed by the organization or a third party and can exist on-premises or off-premises. Public Cloud: This cloud infrastructure is provided to an unspecified number of people or large industry groups and is owned by an organization that sells cloud services. Hybrid Cloud: This cloud infrastructure is a combination of two or more cloud models (private, community, or public). Each model-specific entity is retained, but is bound by standard or individual technologies to achieve data and application portability (e.g., cloud bursting for load balancing between clouds).
[0047] A cloud computing environment is a service-oriented environment that emphasizes statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is an infrastructure that includes a network of interconnected nodes.
[0048] Here, FIG. 4 illustrates an exemplary cloud computing environment 50. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10. In contrast, local computer devices used by cloud consumers (e.g., a PDA or mobile phone 54A, a desktop computer 54B, a laptop computer 54C, or an automotive computer system 54N or a combination thereof, etc.) can communicate. The nodes 10 can communicate with each other. The nodes 10 can be physically or virtually grouped (not shown) in one or more networks, such as, for example, the private, community, public, or hybrid clouds described above or a combination thereof. Thereby, the cloud computing environment 50 can provide infrastructure, platform, software, or a combination thereof as a service, and cloud consumers do not need to maintain resources on local computer devices. It should be understood that the types of computer devices 54A to N shown in FIG. 4 are merely examples, and the computing nodes 10 and the cloud computing environment 50 can communicate with any type of electronic device via any type of network or network-addressable connection (e.g., using a web browser) or both.
[0049] Here, a set of functional abstraction layers provided by the cloud computing environment 50 (FIG. 4) is shown in FIG. 5. It should be understood in advance that the components, layers, and functions shown in FIG. 5 are merely illustrative, and the embodiments of the present invention are not limited thereto. As shown in the figure, the following layers and corresponding functions are provided.
[0050] The hardware and software layer 60 includes hardware components and software components. Examples of hardware components include mainframes 61, servers 62 based on reduced instruction set computer (RISC) architectures, servers 63, blade servers 64, storage devices 65, and networks and network components 66. In some embodiments, the software components include network application server software 67 and database software 68.
[0051] The virtualization layer 70 provides an abstraction layer. From this layer, virtual entities such as virtual servers 71, virtual storage 72, virtual networks 73 including virtual private networks, virtual applications and operating systems 74, and virtual clients 75 can be provided.
[0052] As an example, the management layer 80 can provide the following functions. Resource preparation 81 enables the dynamic procurement of computing resources and other resources used to execute tasks within a cloud computing environment. Metering and pricing 82 enables cost tracking when resources are utilized within a cloud computing environment and billing or invoicing for the consumption of these resources. As an example, these resources may include licenses for application software. Security enables not only the protection of data and other resources but also the identification and verification of cloud consumers and tasks. The user portal 83 provides access to the cloud computing environment for consumers and system administrators. Service level management 84 enables the allocation and management of cloud computing resources so that the required service level is met. Planning and fulfillment of service quality assurance (SLA) 85 enables the advance arrangement and procurement of cloud computing resources that are expected to be needed in the future according to the SLA.
[0053] The workload layer 90 provides examples of functions available for use in a cloud computing environment. Examples of workloads and functions that can be provided from this layer include mapping and navigation 91, software development and life cycle management 92, delivery of virtual classroom education 93, data analysis processing 94, transaction processing 95, and interlanguage model training program 275.
[0054] The present invention can be a system, method, computer program product, or a combination thereof integrated at any possible technical detail level. The present invention can be beneficially implemented in any single or parallel system that processes instruction streams. The computer program product may include a computer-readable storage medium storing computer-readable program instructions for causing a processor to execute aspects of the present invention.
[0055] A computer-readable storage medium can be a tangible device that holds and stores instructions for use by an instruction execution device. A computer-readable storage medium can be, by way of example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or a suitable combination thereof. As a more specific example of a computer-readable storage medium, there can be a portable computer diskette, a hard disk, a RAM, a ROM, an EPROM (or flash memory), an SRAM, a CD-ROM, a DVD, a memory stick, a floppy disk, a punched card, a mechanically encoded device having instructions recorded thereon such as a raised structure in a groove, and suitable combinations thereof. A computer-readable storage medium or computer-readable storage device as used herein should not be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.
[0056] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computer devices / processing devices. Alternatively, they can be downloaded to an external computer or external storage device via a network (e.g., the Internet, a LAN, a WAN, or a wireless network, or a combination thereof). The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface within each computer device / processing device receives the computer-readable program instructions from the network and transfers them for storage in a computer-readable storage medium within each respective computer device / processing device.
[0057] The computer-readable program instructions for carrying out the operations of the present invention may be source code or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or object-oriented programming languages such as Smalltalk and C++, and procedural programming languages such as the "C" programming language and similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer as a stand-alone software package, or partially on the user's computer. Alternatively, it may be executed partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a LAN or WAN, or to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), can execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to customize the electronic circuit for the purpose of implementing aspects of the present invention.
[0058] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. Each block in the flowchart illustrations and / or block diagrams, and combinations of multiple blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0059] These computer-readable program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of such computer or other programmable data processing apparatus create means for implementing the functions / operations specified in one or more blocks in a flowchart and / or block diagram. These computer-readable program instructions can further be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, or other device to function in a particular manner, such that the computer-readable storage medium storing the instructions constitutes a manufacture including instructions for implementing the mode of operation of the functions / operations specified in one or more blocks in a flowchart and / or block diagram.
[0060] Alternatively, the computer-readable program instructions may be loaded onto a computer, other programmable apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-executed process, such that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / operations specified in one or more blocks in a flowchart and / or block diagram.
[0061] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram can represent a module, segment, or portion of instructions that includes one or more executable instructions for performing a particular logical function. In some other implementations, the functions shown within a block may be executed in an order different from that shown in each figure. For example, depending on the relevant functions, two consecutively shown blocks may actually be executed substantially simultaneously, or the blocks may be executed in the reverse order in some cases. It should be noted that each block in a block diagram, flowchart, or both, and combinations of multiple blocks in a block diagram, flowchart, or both, can be executed by a dedicated hardware-based system that performs a particular function or operation, or executes a combination of dedicated hardware and computer instructions.
[0062] As used herein, when referring to "one embodiment", "an embodiment", "an example embodiment", etc., it indicates that although the described embodiment may include certain features, structures, or characteristics, not all embodiments necessarily include these specific features, structures, or characteristics. Further, such phrases do not necessarily refer to the same embodiment. Additionally, when a particular feature, structure, or characteristic is described in relation to one embodiment, it is considered within the knowledge of those skilled in the art to relate such feature, structure, or characteristic to other embodiments, whether or not explicitly described.
[0063] The terms used in this specification are for the purpose of describing particular embodiments only and are not intended to limit the present invention. In this specification, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. Further, in this specification, the terms "comprise", "comprising", or both are used to define the presence of the described features, integers, steps, operations, elements, or components, or combinations thereof, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, or groups thereof, or combinations thereof.
[0064] Although various embodiments of the present invention have been described by way of example, they are not intended to be exhaustive or to limit the invention to these embodiments. As will be apparent to those skilled in the art, many modifications and variations are possible without departing from the scope of the present invention. The terms used in this specification are chosen in order to best explain the principles of the embodiments, the practical application, or the technical improvement to the technology as observed in the market, or to enable those skilled in the art to understand each embodiment disclosed in this specification.
Claims
1. A method for providing a prediction model for a target language, comprising: determining, by one or more computer processors, an example weight of a labeled source language text unit according to a set of unlabeled target language text units; scaling, by the one or more computer processors, an error between a predicted label for the one source language text unit and a correct label for the one source language text unit according to the example weight; updating, by the one or more computer processors, network parameters of the prediction model for the target language according to the error; providing, by the one or more computer processors, the prediction model for the target language to a user; and a method including the above.
2. Updating the network parameters of the prediction model for the target language includes: updating, by the one or more computer processors, network parameters of a downstream neural network model; updating, by the one or more computer processors, network parameters of a language neural network model; and the method according to Claim 1 including the above.
3. The method according to Claim 1, further including using, by the one or more computer processors, the prediction model for the target language as a document classification model for the target language. The method according to Claim 1.
4. The method according to Claim 1, further including using, by the one or more computer processors, the prediction model for the target language as a sequence labeling model for the target language. The method according to Claim 1.
5. further comprising determining, by the one or more computer processors, an example weight of a labeled source language text unit randomly selected from one of a plurality of source languages according to a set of randomly selected unlabeled target language text units The method according to claim 1.
6. A method of providing a prediction model for a target language, comprising: inputting, by the one or more computer processors, source language text units and target language text units into a pre-trained language neural network; generating, by the one or more computer processors, a source language vector from the source language text units; generating, by the one or more computer processors, a target language vector from the target language text units; measuring, by the one or more computer processors, a similarity between one source language vector of the source language vectors and a set of target language vectors of the target language vectors; determining, by the one or more computer processors, a scalar example weight for each source language vector according to the similarity; generating, by the one or more computer processors, a scaled error between a predicted label for a source language text unit and a correct label for the source language text unit using the scalar example weight for the source language vector; calculating, by the one or more computer processors, an update to network parameters of a downstream prediction neural network according to the scaled error; calculating, by the one or more computer processors, an update to network parameters of the language neural network according to the scaled error; providing, by the one or more computer processors, a prediction model for the target language, including the downstream prediction neural network and the language neural network; A method comprising. **Claim 7** The method according to claim 6, wherein measuring the similarity includes summing the similarities between the source language vector and each target language vector in the set of target language vectors. **Claim 8** The method according to claim 6, further comprising inputting text units from a plurality of source languages. **Claim 9** The method according to claim 6, wherein the source language text unit includes labeled data. **Claim 10** The method according to claim 6, wherein the target language text unit includes unlabeled data. **Claim 11** The method according to claim 6, further comprising using the prediction model for the target language as a document classification model for the target language. **Claim 12** A method for providing a prediction model for a target language, comprising: inputting, by the one or more computer processors, a labeled source language text unit and an unlabeled target language text unit into a language neural network; determining, by the one or more computer processors, a scalar instance weight of a source language text unit using the language neural network; scaling, by the one or more computer processors, an error between a label for the source language text unit predicted by a downstream neural network and the correct label for the source language text unit; Calculating, by the one or more computer processors, updates to the network parameters of the downstream neural network and to the network parameters of the language neural network; Providing, by the one or more computer processors, a prediction model for the target language that includes the language neural network and the downstream neural network; A method comprising.
13. Determining the scalar case weight of the one source language text unit comprises Generating, by the one or more computer processors, a source language vector from the source language text unit; Generating, by the one or more computer processors, a target language vector from the target language text unit; Measuring, by the one or more computer processors, a similarity between one source language vector and a set of target language vectors; Determining, by the one or more computer processors, the scalar case weight of the one source language vector according to the similarity; The method according to claim 12, comprising.
14. The method according to claim 12, wherein calculating, by the one or more computer processors, updates to the network parameters of the downstream neural network and to the network parameters of the language neural network includes calculating the updates according to the scaled error.
15. The method according to claim 12, wherein the labeled source language text unit includes labeled text units from a plurality of source languages.
16. A computer program product for providing a prediction model for a target language, comprising one or more computer-readable storage devices and program instructions collectively stored in the one or more computer-readable storage devices, the stored program instructions being Program instructions for determining an example weight of one labeled source language text unit according to a set of unlabeled target language text units; Program instructions for scaling an error between a predicted label for the one source language text unit and a correct label for the one source language text unit according to the example weight; Program instructions for updating network parameters of the prediction model for the target language according to the error; Program instructions for providing the prediction model for the target language to a user; A computer program product comprising the same.
17. The program instructions for updating network parameters of the prediction model for the target language are Program instructions for updating network parameters of a downstream neural network model; Program instructions for updating network parameters of a language neural network model; The computer program product according to claim 16, comprising the same.
18. The stored program instructions further comprise program instructions for using the prediction model for the target language as a document classification model for the target language, the computer program product according to claim 16.
19. The stored program instructions further comprise program instructions for using the prediction model for the target language as a sequence labeling model for the target language, the computer program product according to claim 16.
20. The stored program instructions further include program instructions for determining an instance weight of a labeled source language text unit randomly selected from one of a plurality of source languages according to a set of unlabeled target language text units, the computer program product according to claim 16.
21. A computer system for providing a prediction model for a target language, One or more computer processors, One or more computer-readable storage devices, Program instructions stored in the one or more computer-readable storage devices for execution by the one or more computer processors, the stored program instructions including: Program instructions for determining an instance weight of a labeled source language text unit according to a set of unlabeled target language text units; Program instructions for scaling an error between a predicted label for the one source language text unit and a correct label for the one source language text unit according to the instance weight; Program instructions for updating network parameters of the prediction model for the target language according to the error; Program instructions for providing the prediction model for the target language to a user; A computer system including the above.
22. The program instructions for updating network parameters of the prediction model for the target language include: Program instructions for updating network parameters of a downstream neural network model; Program instructions for updating network parameters of a language neural network model; The computer system according to claim 21, comprising **Claim 23** The computer system according to claim 21, wherein the stored program instructions further include program instructions for using the prediction model for the target language as a document classification model for the target language. **Claim 24** The computer system according to claim 21, wherein the stored program instructions further include program instructions for using the prediction model for the target language as a sequence labeling model for the target language. **Claim 25** The computer system according to claim 21, wherein the stored program instructions further include program instructions for determining an example weight of a labeled source language text unit randomly selected from one of a plurality of source languages according to a set of randomly selected unlabeled target language text units.
Citation Information
Patent Citations
Rotary floor cleaning apparatus
JP1985077727A
Neural machine translation model training method and device, and computer program therefor
JP2019149018A
Techniques for evaluation, building and / or retraining of a classification model
US20130254153A1