Word list alignment method and device, equipment, medium and product

Through dynamic programming and probability distribution assignment methods, the accuracy and accuracy of word list alignment in heterogeneous model collaborative training is improved, the problem of low alignment accuracy and accuracy in the existing technology is solved, and the transfer effect and prediction performance of model knowledge are improved.

CN120068865APending Publication Date: 2025-05-30WEBANK (CHINA)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510129587.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing vocabulary alignment methods have problems with low alignment accuracy and accuracy in heterogeneous model collaborative training, which affects the transfer effect of model knowledge.

Method used

By obtaining the first word sequence and the second word sequence, and the second predicted probability, dynamic programming is used to trace the word sequence matching result at the minimum matching cost from back to front, and based on this result, the word segmentation position in the first word sequence is assigned probability distribution to obtain the word list alignment result.

Benefits of technology

Improve the accuracy and accuracy of word list alignment, thereby improving the transfer effect of model knowledge during collaborative learning, and optimizing the prediction performance of the model in cross-domain or cross-language tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068865A_ABST
    Figure CN120068865A_ABST
Patent Text Reader

Abstract

The invention discloses a word list alignment method and device, equipment, a medium and a product, relates to the technical field of natural language processing, and is applied to first equipment, the first equipment is in communication connection with second equipment, a first language model is deployed in the first equipment, and a second language model is deployed in the second equipment. A first word sequence, a second word sequence and a second prediction probability are obtained, the first word sequence is obtained by performing word segmentation on a training text through a first word list by a first language model, and the second word sequence is obtained by performing word segmentation on the training text through a second word list by a second language model; dynamically planning the first word sequence and the second word sequence, and determining a word segmentation sequence matching result; on the basis of the word segmentation sequence matching result, probability distribution assignment is carried out on all word segmentation positions in the first word sequence according to the second prediction probability, a word list alignment result is obtained, the precision and accuracy of the word list alignment result can be improved, and therefore the knowledge migration effect during collaborative learning is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of natural language processing, and in particular, to a method, apparatus, device, medium, and product for vocabulary alignment. Background Art

[0002] With the rapid development of large language model (LLM) technology, their powerful performance has brought significant productivity improvements in various industries. However, due to the lack of a unified standard for current large model technologies in the industry, there are differences in training methods, data, etc. among different companies, so the models of each company are different. Therefore, collaborative learning of multiple models is a new requirement. Especially in fields with higher requirements for data circulation security, in scenarios where private data cannot leave the threshold and computing power is limited, the collaborative learning technology of federated large and small models is also an urgent need. However, when co-training heterogeneous models, because the vocabularies of the two models are inconsistent, in an algorithm that requires predicting the probability distribution of the vocabulary, challenges in aligning the vocabulary after sentence tokenization will be encountered, and existing vocabulary alignment methods have problems of low alignment accuracy and accuracy, which can easily affect the transfer effect of model knowledge during collaborative learning.

[0003] Therefore, it is necessary to propose a solution to improve the alignment accuracy and accuracy of the vocabulary alignment method to improve the transfer effect of model knowledge during collaborative learning.

[0004] The above content is only used to assist in understanding the technical solution of this application, and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main purpose of this application is to provide a method, apparatus, device, medium, and product for vocabulary alignment, aiming to improve the alignment accuracy and accuracy of the vocabulary alignment method to improve the transfer effect of model knowledge during collaborative learning.

[0006] To achieve the above object, this application provides a vocabulary alignment method, which is applied to a first device. The first device is communicatively connected to a second device. A first language model is deployed in the first device, and a second language model is deployed in the second device. The first language model and the second language model are heterogeneous. The method includes:

[0007] Obtain a first word sequence, a second word sequence, and a second prediction probability, where the first word sequence is obtained by the first language model tokenizing a training text through a first vocabulary, the second word sequence is obtained by the second language model tokenizing the training text through a second vocabulary, and the second prediction probability is a probability distribution vector corresponding to each token position in the second word sequence and each text component unit in the second vocabulary;

[0008] Perform dynamic programming on the first word sequence and the second word sequence, and backtrack from the end to the beginning to obtain the word segmentation sequence matching result under the minimum matching cost, where the word segmentation sequence matching result includes at least one of a unique matching relationship, a continuous word matching a single word relationship, a single word matching multiple words relationship, and a subsequence matching relationship;

[0009] Based on the word segmentation sequence matching result, assign a probability distribution to each word segmentation position in the first word sequence according to the second prediction probability to obtain a word table alignment result.

[0010] In one embodiment, the step of assigning a probability distribution to each word segmentation position in the first word sequence according to the second prediction probability based on the word segmentation sequence matching result to obtain a word table alignment result includes:

[0011] According to the word segmentation sequence matching result, determine the second subsequence in the first word sequence that has the subsequence matching relationship with the first subsequence in the second word sequence, and execute a sequence assignment process for each word segmentation position in the second subsequence; and / or,

[0012] Traverse each word segmentation position in the second word sequence, identify the matching relationship between the text composition units of each word segmentation position in the second word sequence and the text composition units of each word segmentation position in the first word sequence, and assign a probability distribution to each word segmentation position in the first word sequence according to the matching relationship to obtain the word table alignment result.

[0013] In one embodiment, the step of assigning a probability distribution to each word segmentation position in the first word sequence according to the matching relationship includes:

[0014] Determine whether there is a word segmentation sequence matching result between the text composition units of each word segmentation position in the second word sequence and the text composition units of each word segmentation position in the first word sequence according to the matching relationship;

[0015] If there is a word segmentation sequence matching result between the text composition units of each word segmentation position in the second word sequence and the text composition units of each word segmentation position in the first word sequence, then assign a probability distribution to the corresponding word segmentation positions in the first word sequence according to the word segmentation sequence matching result.

[0016] In one embodiment, the step of, if there is a word segmentation sequence matching result between the text composition units of each word segmentation position in the second word sequence and the text composition units of each word segmentation position in the first word sequence, then assign a probability distribution to the corresponding word segmentation positions in the first word sequence according to the word segmentation sequence matching result includes:

[0017] If the word segmentation sequence matching result is a unique matching relationship, set the word segmentation positions of the text composition units in the first word sequence corresponding to the text composition units at each word segmentation position in the second word sequence as the probability distribution vectors corresponding to the text composition units at the corresponding word segmentation positions in the second word sequence;

[0018] If the word segmentation sequence matching result is a continuous word matching a single word relationship, perform a sequence assignment process on the word segmentation positions of the text composition units in the first word sequence corresponding to the text composition units at the word segmentation positions in the second word sequence;

[0019] If the word segmentation sequence matching result is a single word matching multiple words relationship, set the probability distribution vector of the word segmentation position of the text composition unit in the first word sequence corresponding to the text composition unit at the word segmentation position in the second word sequence as a one-hot encoding vector.

[0020] In one embodiment, the steps of the sequence assignment process include:

[0021] Set the probability distribution vector of the position in the first word sequence corresponding to the first position of the second subsequence as the probability distribution vector of the corresponding word segmentation position in the second word sequence, and set the probability distribution vectors of the word segmentation positions in the second subsequence or the first word sequence except the first position of the second subsequence as one-hot encoding vectors; and / or,

[0022] Set the probability distribution vector of the position corresponding to the matching result of the text composition unit at each word segmentation position in the second subsequence or the first word sequence in the second word sequence as the probability distribution vector of the corresponding word segmentation position in the second word sequence.

[0023] In one embodiment, the steps of setting the probability distribution vector of the corresponding word segmentation position in the second word sequence include:

[0024] For the probability distribution vectors of the text composition units in the second word list at the positions being matched in the second word sequence, traverse each text composition unit to determine whether each text composition unit appears in the first word list;

[0025] For the text composition units in the second word list that appear in the first word list, perform corresponding probability assignment at the corresponding positions in the first word list;

[0026] For text components in the second vocabulary that do not appear in the first vocabulary, perform word segmentation using the first vocabulary to obtain a word segmentation result, calculate the mean word vector of the text components based on the word segmentation result, and select the text component in the first vocabulary with the smallest distance from the target text component; use the probability corresponding to the text component as the probability of the text component with the smallest distance.

[0027] In addition, to achieve the above object, the present application also proposes a vocabulary alignment device, which is applied to a first device. The first device is communicatively connected to a second device. A first language model is deployed in the first device, and a second language model is deployed in the second device. The first language model and the second language model are heterogeneous. The vocabulary alignment device includes:

[0028] An acquisition module, configured to acquire a first word sequence, a second word sequence, and a second prediction probability. The first word sequence is obtained by the first language model performing word segmentation on a training text through a first vocabulary. The second word sequence is obtained by the second language model performing word segmentation on the training text through a second vocabulary. The second prediction probability is a probability distribution vector corresponding to each word segmentation position in the second word sequence and each text component in the second vocabulary.

[0029] A matching module, configured to perform dynamic programming on the first word sequence and the second word sequence, and backtrack from the back to obtain a word segmentation sequence matching result under the minimum matching cost. The word segmentation sequence matching result includes at least one of a unique matching relationship, a continuous word matching a single word relationship, a single word matching multiple words relationship, and a subsequence matching relationship.

[0030] An assignment module, configured to, based on the word segmentation sequence matching result, assign a probability distribution to each word segmentation position in the first word sequence according to the second prediction probability to obtain a vocabulary alignment result.

[0031] In addition, to achieve the above object, the present application also proposes a vocabulary alignment device. The device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. The computer program is configured to implement the steps of the vocabulary alignment method as described above.

[0032] In addition, to achieve the above object, the present application also proposes a storage medium. The storage medium is a computer-readable storage medium. A computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the vocabulary alignment method as described above.

[0033] In addition, to achieve the above object, the present application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the vocabulary alignment method described above.

[0034] One or more technical solutions proposed by the present application have at least the following technical effects:

[0035] Apply the vocabulary alignment method to a first device, which is communicatively connected to a second device. A first language model is deployed in the first device, and a second language model is deployed in the second device. The first language model and the second language model are heterogeneous. By obtaining a first word sequence, a second word sequence, and a second prediction probability, where the first word sequence is obtained by the first language model segmenting a training text through a first vocabulary, the second word sequence is obtained by the second language model segmenting the training text through a second vocabulary, and the second prediction probability is a probability distribution vector corresponding to each segmentation position in the second word sequence and each text composition unit in the second vocabulary; perform dynamic programming on the first word sequence and the second word sequence, and backtrack from back to front to obtain a segmentation sequence matching result with the minimum matching cost, where the segmentation sequence matching result includes at least one of a unique matching relationship, a continuous word matching a single word relationship, a single word matching multiple words relationship, and a subsequence matching relationship; based on the segmentation sequence matching result, assign a probability distribution to each segmentation position in the first word sequence according to the second prediction probability to obtain a vocabulary alignment result. Through vocabulary alignment, the heterogeneous first language model and the second language model can effectively share knowledge. By performing dynamic programming on the first word sequence and the second word sequence and backtracking from back to front to obtain a segmentation sequence matching result with the minimum matching cost, and then based on the segmentation sequence matching result, assign a probability distribution to each segmentation position in the first word sequence according to the second prediction probability to obtain a vocabulary alignment result, the accuracy and correctness of the vocabulary alignment result can be improved, which helps to improve the knowledge transfer effect during collaborative learning; through reasonable assignment of the probability distribution, the prediction performance of the model can be optimized when dealing with cross-domain or cross-language tasks. Description of the Drawings

[0036] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0037] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0038] Figure 1 This is a schematic flowchart provided for the first embodiment of the method for aligning word lists in this application;

[0039] Figure 2 This is a schematic flowchart provided for the second embodiment of the method for aligning word lists in this application;

[0040] Figure 3 This is a schematic diagram of the module structure of the word list alignment device in the embodiment of this application;

[0041] Figure 4 This is a schematic diagram of the device structure of the hardware operating environment involved in the method for aligning word lists in the embodiment of this application.

[0042] The realization of the purpose, functional features and advantages of this application will be further described in conjunction with the embodiments with reference to the accompanying drawings. Detailed implementation manners

[0043] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.

[0044] In order to better understand the technical solutions of this application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0045] In the prior art, assume that the first language model A needs to learn the prediction probability distribution of the second language model B, and the word segmentation results of the two models are as follows:

[0046] Model A: [tokenA(1)…tokenB(n)]

[0047] Model B: [tokenA(1)…tokenB(m)]

[0048] At the same time, for each position after word segmentation of model B (each position is a token, that is, a text composition unit), there is a probability distribution corresponding to the size of the word list. That is, each position is a probability distribution vector [p1, p2…p(vocab_size)], where vocab_size is the size of the word list of model B. First, use the following dynamic programming method to calculate the minimum matching cost of the results after word segmentation of the two models:

[0049] f(i,j) = min(f(i - 1,j), f(i,j - 1), f(i - 1,j - 1)) + cost(tokenA(i),tokenB(j))

[0050] cost(i,j) = editdistance(tokenA(i),tokenB(j)

[0051] After the dynamic programming process is completed, backtrack from (n, m) to obtain the matching path. For the i-th position of the sentence in Model A, if during the minimum matching cost process, tokenA(i) is matched with a unique tokenB(j), then the probability distribution of tokenB(j) can be assigned to the i-th position. Otherwise, there will be a one-to-many situation during the transfer of the dynamic programming equation. For this situation, directly set the predicted probability distribution of the i-th position to a 0 / 1 vector, [0, …1…, 0], that is, only the position of tokenA(i) in the vocabulary is 1, and all other positions are 0. The length of the vector is the size of the vocabulary.

[0052] During the probability distribution assignment process, when assigning the vocabulary prediction probability of the j-th position in Model B to the i-th position in Model A, the steps include:

[0053] (1) Traverse each position in the vocabulary of Model B. If the text composition vector corresponding to this position appears in Model A, then the probability corresponding to this text composition vector is assigned to the position where this text composition vector appears in Model A;

[0054] (2) If during the traversal of the vocabulary of Model B, a certain text composition vector does not exist in the vocabulary of Model A, at this time, select the text composition unit with the smallest edit distance from the vocabulary of Model A to this vector, and then assign the corresponding probability to the text composition unit with the smallest edit distance found in Model A.

[0055] It can be seen from this that the problems existing in using the edit distance for probability distribution assignment in the prior art are: a. It cannot solve the Chinese matching problem. Because when segmenting Chinese words, if segmented into single characters, the edit distance between two characters is 2, which easily causes the entire dynamic programming process to be misaligned; b. Although the edit distance of some text composition unit pairs is close, the semantic space is far, and forced matching has no value, such as <we, fe>, although only one character is different, the expressed semantics are completely different.

[0056] In an embodiment of the present application, a solution is provided. The vocabulary alignment method is applied to a first device, which is communicatively connected to a second device. A first language model is deployed in the first device, and a second language model is deployed in the second device. The first language model and the second language model are heterogeneous. By obtaining a first word sequence, a second word sequence, and a second prediction probability, wherein the first word sequence is obtained by the first language model segmenting a training text through a first vocabulary, the second word sequence is obtained by the second language model segmenting the training text through a second vocabulary, and the second prediction probability is a probability distribution vector corresponding to each word segmentation position in the second word sequence and each text composition unit in the second vocabulary; performing dynamic programming on the first word sequence and the second word sequence, and backtracking from back to front to obtain a word sequence matching result with the minimum matching cost, wherein the word sequence matching result includes at least one of a unique matching relationship, a continuous word matching a single word relationship, a single word matching multiple words relationship, and a subsequence matching relationship; based on the word sequence matching result, assigning a probability distribution to each word segmentation position in the first word sequence according to the second prediction probability to obtain a vocabulary alignment result. Through vocabulary alignment, the heterogeneous first language model and the second language model can effectively share knowledge. By performing dynamic programming on the first word sequence and the second word sequence, and backtracking from back to front to obtain a word sequence matching result with the minimum matching cost, and then based on the word sequence matching result, assigning a probability distribution to each word segmentation position in the first word sequence according to the second prediction probability to obtain a vocabulary alignment result, the accuracy and correctness of the vocabulary alignment result can be improved, which helps to improve the knowledge transfer effect during collaborative learning; through reasonable assignment of the probability distribution, the prediction performance of the model can be optimized when dealing with cross-domain or cross-language tasks.

[0057] The following presents the first embodiment of the vocabulary alignment method of the present application. Refer to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the vocabulary alignment method of the present application.

[0058] In this embodiment, the vocabulary alignment method includes steps S10 to S30:

[0059] Step S10, obtaining a first word sequence, a second word sequence, and a second prediction probability, wherein the first word sequence is obtained by the first language model segmenting a training text through a first vocabulary, the second word sequence is obtained by the second language model segmenting the training text through a second vocabulary, and the second prediction probability is a probability distribution vector corresponding to each word segmentation position in the second word sequence and each text composition unit in the second vocabulary;

[0060] The vocabulary alignment method in the embodiments of the present application is applied to a first device, which is communicatively connected to a second device. Both the first device and the second device can be a computing service device with data processing, network communication, and program running functions, such as a server, a tablet computer, a personal computer, a mobile phone, etc. The first device and the second device can be devices deployed in different enterprises, that is, each enterprise uses its own device to jointly train a language model to perform natural language processing tasks through the trained language model. In the specific implementation, specific natural language processing tasks can be set according to the specific requirements of each enterprise. For example, in a dialogue scenario, a task of generating a reply to the above text can be set, and in a summary generation scenario, a task of generating a text summary for a text can be set. In the following embodiments, the two devices are distinguished as "first" and "second". In some application scenarios, the first device can also be referred to as a server or a coordination end, etc., and the second device can be referred to as a client or a participating end, etc.

[0061] The first language model is deployed in the first device, and the second language model is deployed in the second device. A language model is a model that can be used to perform natural language processing tasks, such as LLaMa2 13B, LLaMa2 7B, OPT 7B, etc. The first language model and the second language model in this embodiment both belong to language models. The language models that each enterprise can deploy or support are mostly heterogeneous. Based on this, in this embodiment, the first language model deployed in the first device and the second language model deployed in the second device are heterogeneous. Heterogeneous means that the model architectures of the two language models are different, and the scales of the model parameters are also different. For example, in a feasible implementation, LLaMa2 13B can be deployed in the first device, and OPT 7B can be deployed in the second device. The first language model and the second language model may have different internal architectures. For example, one may be a recurrent neural network (RNN) structure, and the other may be a Transformer-based model. The two models may be trained on different datasets, resulting in differences in the language features and patterns they learn. The first language model may be a large language model with a large number of parameters, while the second language model may be a small language model, which is smaller and more lightweight. The training process and the computing resources required during the operation process after the large language model is trained are also more than those of the small language model. Therefore, the computing resources of the first device and the second device used can be different. The computing resources of the first device can be better than those of the second device, that is, the computing power is stronger than that of the second device. The advantages and disadvantages of the computing resources are specifically reflected by some parameters. For example, the number of CPU cores, clock frequency, memory bandwidth and capacity, etc. This embodiment does not limit this. In the specific implementation process, an enterprise with powerful computing resources can provide a device as the first device, and an enterprise with text data in each subdivision field but limited computing resources can provide a device as the second device. Each enterprise jointly trains the language model and uses the advantages of both parties to jointly train a language model with higher natural language processing capabilities. For example, enterprises such as banks and insurance companies need to train language models to perform natural language processing tasks in their business scenarios to improve the business service effect. Although enterprises such as banks and insurance companies have text data in subdivision fields such as the banking business field and the insurance business field, their computing resources are limited. It is difficult for these enterprises to train large language models that require powerful computing power with their own devices, making it difficult for them to successfully deploy language models that can improve the business service effect in the business scenario. Enterprises or institutions with powerful computing power lack the text data in these subdivision fields and are difficult to train language models with stronger natural language processing capabilities. Therefore, the devices of these enterprises can be jointly used to train the language model collaboratively, and the advantages of both parties can be used to jointly train a language model with higher natural language processing capabilities.

[0062] The first language model and the second language model may exhibit different performance characteristics when processing different types of language tasks. For example, the first language model performs better in text classification, while the second language model is more excellent in machine translation. Another example is that the first language model may have deeper knowledge in specific fields (such as medical and legal), while the second language model may perform better in a wider general field.

[0063] When co-training heterogeneous models, due to the inconsistency of the vocabulary lists between the first language model and the second language model, in an algorithm that requires using the probability distribution of the prediction vocabulary list, challenges in aligning the vocabulary lists after sentence tokenization will be encountered. Through the vocabulary list alignment method in the embodiments of the present application, the alignment accuracy of the vocabulary lists between the first language model and the second language model can be improved, and then the aligned vocabulary lists can be used to co-train the first language model and the second language model, further improving the co-training effect.

[0064] In the embodiments of the present application, taking the example of assigning probability distributions to the first vocabulary list of the first language model according to the second vocabulary list of the second language model for illustration. Similarly, in other embodiments, probability distributions can also be assigned to the second vocabulary list of the second language model according to the first vocabulary list of the first language model, which will not be elaborated in this embodiment.

[0065] Step S20: Perform dynamic programming on the first word sequence and the second word sequence, and backtrack from the back to the front to obtain the token sequence matching result with the minimum matching cost, where the token sequence matching result includes at least one of a unique matching relationship, a continuous word matching a single word relationship, a single word matching multiple words relationship, and a subsequence matching relationship.

[0066] Furthermore, after obtaining the first word sequence of the first language model and the second word sequence of the second language model, dynamic programming can be performed on the first word sequence and the second word sequence.

[0067] Optionally, a text component unit (token) is the basic unit that composes a text and is artificially divided. Each text component unit divided based on a set of division rules constitutes a word sequence. Since the first language model and the second language model in this embodiment are heterogeneous, the division rules used by the first language model and the second language model may be different, that is, the vocabulary lists used by the first language model and the second language model may be different, that is, the number of text component units included in the two divided word sequences may be different, and the text component units included in the two vocabulary lists may be different themselves. For example, for Chinese texts, one word sequence is divided with each character as a text component unit, and the other word sequence is divided with each word (a word may include multiple characters) as a text component unit.

[0068] Due to the differences between the first vocabulary and the second vocabulary, there are differences in the corresponding prediction probabilities. Therefore, in this embodiment, the first device can align the text composition units of the first word sequence and the second word sequence, correspond the same text composition units, and correspond the different text composition units according to a certain mapping rule, so that the text composition units involved in the first vocabulary and the second vocabulary can be corresponding. That is, according to the prediction probabilities corresponding to the word segmentation positions in the second word sequence, probability distribution assignment is performed on the word segmentation positions in the first word sequence to obtain the vocabulary alignment result, so that the heterogeneous first language model and the second language model can also be co-trained.

[0069] Optionally, dynamic programming is the core in the vocabulary alignment process, which involves using the dynamic programming algorithm to identify the best matching relationship between two word sequences. This process starts from the ends of the two word sequences and gradually traces back forward to minimize the overall matching cost. The matching results may include various relationships, such as unique matching, continuous words matching a single word, a single word matching multiple words, and continuous word sequence matching, etc. Each matching relationship corresponds to different processing strategies to ensure that the final vocabulary alignment result is both accurate and effective, thereby maximizing the collaborative effect between the two heterogeneous models while retaining their respective language characteristics and advantages.

[0070] Step S30, based on the word segmentation sequence matching result, perform probability distribution assignment on each word segmentation position in the first word sequence according to the second prediction probability to obtain the vocabulary alignment result.

[0071] Furthermore, after performing dynamic programming on the first word sequence and the second word sequence and tracing back from the back to the front to obtain the word segmentation sequence matching result under the minimum matching cost, probability distribution assignment can be performed on each word segmentation position in the first word sequence according to the second prediction probability based on the word segmentation sequence matching result to obtain the vocabulary alignment result.

[0072] Optionally, after completing the word segmentation sequence matching, analyze the obtained result to identify different types of matching relationships, including unique matching, continuous words matching a single word, a single word matching multiple words, and continuous subsequence matching. If it is a one-to-one matching, traverse each word table position of the second language model. If the text composition unit TokenB at this word table position appears in the first language model, set the probability of this position in the first word table of the first language model to TokenB; otherwise, find the position of the word with the closest semantics in the first language model for probability assignment: segment TokenB using the tokenizer of the first language model, then set the word vector of TokenB to the mean of the word vectors after segmentation, and find the position of the text composition unit corresponding to the closest word vector in the word vector space corresponding to the first word table for probability distribution assignment to obtain the vocabulary alignment result.

[0073] In this embodiment, through the above solution, specifically, the vocabulary alignment method is applied to the first device, the first device is communicatively connected to the second device, the first language model is deployed in the first device, the second language model is deployed in the second device, the first language model and the second language model are heterogeneous. By obtaining a first word sequence, a second word sequence, and a second prediction probability, wherein the first word sequence is obtained by the first language model segmenting a training text through a first vocabulary, the second word sequence is obtained by the second language model segmenting the training text through a second vocabulary, and the second prediction probability is a probability distribution vector corresponding to each token position in the second word sequence and each text component unit in the second vocabulary; performing dynamic programming on the first word sequence and the second word sequence, and backtracking from back to front to obtain a token sequence matching result with the minimum matching cost, wherein the token sequence matching result includes at least one of a unique matching relationship, a continuous word matching a single word relationship, a single word matching multiple words relationship, and a subsequence matching relationship; based on the token sequence matching result, assigning a probability distribution to each token position in the first word sequence according to the second prediction probability to obtain a vocabulary alignment result. Through vocabulary alignment, the heterogeneous first language model and the second language model can effectively share knowledge. By performing dynamic programming on the first word sequence and the second word sequence, and backtracking from back to front to obtain a token sequence matching result with the minimum matching cost, and then based on the token sequence matching result, assigning a probability distribution to each token position in the first word sequence according to the second prediction probability to obtain a vocabulary alignment result, the accuracy and correctness of the vocabulary alignment result can be improved, which helps to improve the knowledge transfer effect during collaborative learning; through reasonable assignment of the probability distribution, the prediction performance of the model can be optimized when dealing with cross-domain or cross-language tasks.

[0074] Based on the first embodiment above, the second embodiment of the vocabulary alignment method of the present application is proposed. In this embodiment, the content that is the same as or similar to the above first embodiment can be referred to the above introduction and will not be repeated hereinafter. In this embodiment, the matching process of the word sequences of the segmented first language model and the second language model is optimized to find the optimal vocabulary matching between two heterogeneous language models.

[0075] The sequence matching method in the embodiments of the present application is as follows:

[0076] f(i,j) = min(min(f(i - 1,j), f(i,j - 1), f(i - 1,j - 1)) + cost(tokenA(i),tokenB(j)), f(i′,j - 1) - oo | [concate(tokenA(i′ + 1)…tokenA(i)) = tokenB(j)])

[0077] Optionally, trace back from the end to the beginning to find the matching relationship under the minimum matching cost. At this time, for the matching relationship of the tokenized sequence A, there are the following three cases:

[0078] (1) The text component unit tokenA(i) in sequence A uniquely matches the text component unit tokenB(j) in sequence B. If tokenA(i) = tokenB(j), then assign the probability distribution at position j to the i-th position of the first language model A;

[0079] (2) A continuous segment of text component units in sequence A matches a text component unit in sequence B: Then execute the sequence assignment process;

[0080] (3) A text component unit in sequence A matches multiple text component units in sequence B. At this time, the matching has no value, and directly set the probability distribution at this position to a 0 / 1 vector, where the position of tokenA(i) in the vocabulary is 1, and all other positions are 0.

[0081] Optionally, cost calculates the distance between the word vectors of tokenA(i) and TokenB(j). For tokenB(j), if the word tokenB(j) exists in the first language model A, then the word vector is the corresponding word vector of the first language model A. Otherwise, the first language model A tokenizes tokenB(j). Assume the tokenization result is [t1, t2... tk], and sum and average the word vectors corresponding to t1, t2..tk in the first language model A, that is, the word vector of tokenB(j) is (Embidding(t1) +... Embedding(tk)) / k. The distance formula can use cosine and / or Euclidean distance, etc.

[0082] [concate(tokenA(i′ + 1)... tokenA(i)) = tokenB(j)]

[0083] This indicates that the concatenation of words in the first language model A is equal to the word in the second language model B. -oo represents an infinitesimal cost, that is, if there is a concatenation of subsequences equal to a single word matching relationship, in this embodiment, the subsequence matching relationship is preferentially considered.

[0084] Optionally, in this embodiment, a cost matrix can be created, and its size is the product of the lengths of the two word sequences. Each element in the matrix represents the matching cost between the corresponding words or subsequences. Starting from the lower right corner of the matrix (i.e., the last elements of the two word sequences), the cost of each cell is calculated one by one. Using the strategy of dynamic programming, the cost of each cell is calculated based on the costs of its adjacent cells. Generally, this involves considering the costs of three operations: matching, insertion, and deletion. During the calculation of the matching cost, for each text unit in the tokenization position, calculate the similarity or distance between them (such as edit distance, cosine similarity, etc.) as the matching cost. If two text units are the same or similar enough, the matching cost is low; if they are completely different, the matching cost is high. During the backward trace from the end to the beginning, once the cost matrix is filled, start the backward trace from the lower right corner to find the path with the minimum cost, which represents the optimal matching sequence. This path will define the optimal alignment between the two word sequences. During the backward trace, different types of matching relationships can be identified:

[0085] Unique matching relationship: The corresponding text units in the two word sequences directly match;

[0086] Continuous words match a single word relationship: Multiple consecutive text units in the first word sequence match multiple consecutive text units in the second word sequence;

[0087] Single word matches multiple words relationship: One text unit in the first word sequence matches multiple text units in the second word sequence;

[0088] Subsequence matching relationship: There are consecutive subsequence matches in the two word sequences.

[0089] Optionally, record the optimal matching path during the backward trace for use in the subsequent probability distribution assignment step. For text units without direct matches, heuristic methods may be needed to find the best matching items, such as based on semantic similarity or edit distance. Finally, output the optimal matching path obtained from the backward trace as the token sequence matching result, which will be used as the basis for probability distribution assignment.

[0090] Through the above solution, this embodiment, specifically by adding support for continuous field splicing matching, has the following advantages: a. For exactly the same words, forced matching will be obtained at this time; b. For some Chinese words, some model vocabularies may not contain this word, but split it into multiple characters. This matching method can recall these words, especially some idioms; c. For some models, such as llama2, the vocabulary is small, and some words will be split into multiple words, such as utilize will be split into util and ize, while for some vocabularies with richer content, utilize will be regarded as a single word directly. At this time, this type of matching relationship can be recalled.

[0091] Based on the above first and / or second embodiments, a third embodiment of the language model training method of the present application is proposed. In this embodiment, for the same or similar content as in the above first and second embodiments, reference can be made to the above introduction and will not be repeated hereinafter. Referring to Figure 2 , Figure 2 is a schematic flowchart of the second embodiment of the word table alignment method of the present application. In this embodiment, step S30 may include step S31 and / or step S32:

[0092] Step S31: According to the word segmentation sequence matching result, determine a second subsequence in the first word sequence that has a subsequence matching relationship with a first subsequence in the second word sequence, and perform a sequence assignment process on each word segmentation position in the second subsequence;

[0093] Optionally, for example, sum and average the probability distribution of subsequence B, that is, the predicted probability distribution vector of the positions [j'+1,..., j]. After directly summing and taking the average, the matching relationship is converted into the case of continuous word matching a single word at this time, and then the sequence assignment process can be performed.

[0094] Optionally, in this embodiment, the case of two consecutive subsequence matches is preferentially considered. Using the word segmentation sequence matching result obtained by the dynamic programming algorithm, identify the continuous text units in the first word sequence that match the subsequence in the second word sequence. Collect the predicted probabilities of each text composition unit related to the matching subsequence from the second word sequence. Sum the predicted probabilities of each text unit within the matching subsequence in the second word sequence, and then calculate the average value to obtain the comprehensive predicted probability of the entire subsequence. According to the probability distribution obtained by summing and averaging, perform probability assignment on the corresponding continuous text units in the first word sequence. This may involve distributing the average probability distribution to each text unit, or adjusting according to specific weights or rules. Adjust the assigned probability distribution to ensure that the sum of the probabilities of all text units is 1, meeting the requirements of the probability distribution. Thereby, the continuous text units in the first word sequence can obtain an updated probability distribution, which will improve the prediction accuracy of the model for the probability of these text units appearing in a specific context.

[0095] Step S32: Traverse each word segmentation position in the second word sequence, identify the matching relationship between the text composition units at each word segmentation position in the second word sequence and the text composition units at each word segmentation position in the first word sequence, and perform probability distribution assignment on each word segmentation position in the first word sequence according to the matching relationship to obtain the word table alignment result.

[0096] Optionally, for the application of the word segmentation sequence matching result, starting from the first text unit of the second word sequence, check its matching relationship with the text units in the first word sequence one by one. If a matching relationship is found, assign probabilities to the corresponding text units in the first word sequence according to the word segmentation sequence matching result: If there is no matching relationship, segment the text units in the second word sequence, and then assign probabilities to the text units in the first word sequence based on the segmentation result.

[0097] Optionally, the step of assigning a probability distribution to each word segmentation position in the first word sequence according to the matching relationship includes:

[0098] Determine whether there is a word segmentation sequence matching result between the text composition units at each word segmentation position in the second word sequence and the text composition units at each word segmentation position in the first word sequence according to the matching relationship;

[0099] If there is a word segmentation sequence matching result between the text composition units at each word segmentation position in the second word sequence and the text composition units at each word segmentation position in the first word sequence, assign a probability distribution to the corresponding word segmentation positions in the first word sequence according to the word segmentation sequence matching result.

[0100] Optionally, if there is a word segmentation sequence matching result between the text composition units at each word segmentation position in the second word sequence and the text composition units at each word segmentation position in the first word sequence, the step of assigning a probability distribution to the corresponding word segmentation positions in the first word sequence according to the word segmentation sequence matching result includes:

[0101] If the word segmentation sequence matching result is a unique matching relationship, set the word segmentation position of the text composition unit in the first word sequence corresponding to the text composition unit at each word segmentation position in the second word sequence as the probability distribution vector corresponding to the text composition unit at the corresponding word segmentation position in the second word sequence;

[0102] If the word segmentation sequence matching result is a continuous word matching a single word relationship, execute a sequence assignment process on the word segmentation position of the text composition unit in the first word sequence corresponding to the text composition unit at the word segmentation position in the second word sequence;

[0103] If the word segmentation sequence matching result is a single word matching multiple words relationship, set the probability distribution vector of the word segmentation position of the text composition unit in the first word sequence corresponding to the text composition unit at the word segmentation position in the second word sequence as a one-hot encoding vector.

[0104] Optionally, the steps of the sequence assignment process include:

[0105] Set the probability distribution vector of the position corresponding to the first position of the second subsequence in the first word sequence as the probability distribution vector of the corresponding word segmentation position in the second word sequence, and set the probability distribution vectors of the word segmentation positions in the second subsequence or the first word sequence except the first position of the second subsequence as one-hot encoding vectors; and / or,

[0106] Set the probability distribution vector of the corresponding position of the matching result of the text composition unit at each word segmentation position in the second subsequence or the first word sequence in the second word sequence as the probability distribution vector of the corresponding word segmentation position in the second word sequence.

[0107] Optionally, the step of setting the probability distribution vector of the corresponding word segmentation position in the second word sequence includes:

[0108] For the probability distribution vector of the text composition unit of the second word list at the matched position in the second word sequence, traverse each text composition unit to determine whether each text composition unit appears in the first word list;

[0109] For the text composition unit of the second word list that appears in the first word list, perform corresponding probability assignment at the corresponding position in the first word list;

[0110] For the text composition unit of the second word list that does not appear in the first word list, use the first word list for word segmentation to obtain a word segmentation result, and calculate the word vector mean of the text composition unit according to the word segmentation result, and select the text composition unit in the first word list with the smallest distance from the target text composition unit; take the probability corresponding to the text composition unit as the probability of the text composition unit with the smallest distance.

[0111] Optionally, for text units in the second vocabulary that do not appear in the first vocabulary, the text units can be segmented, and then probability values can be assigned to the text units in the first word sequence based on the segmentation results. Use the tokenizer corresponding to the first vocabulary to segment the text units in the second vocabulary that do not appear in the first vocabulary. Calculate the mean of the word vectors of each text unit in the segmentation result to obtain the word vector representation of the target text unit. Calculate the minimum distance between the text unit in the first vocabulary and the target text unit. Take the text unit in the first vocabulary with the minimum distance from the target text unit as the best matching item. Assign the prediction probability of the target text unit to the corresponding best matching text unit in the first vocabulary. Probability assignment is performed by finding the position of the word with the closest semantics in the first language model A, specifically including: using the tokenizer of the first language model A to segment TokenB, and then setting the word vector of TokenB to the mean of the word vectors after segmentation. In the word vector space corresponding to the vocabulary of the first language model A, find the word vector with the closest distance. The distance calculation here can use common distance formulas such as cosine and / or Euclidean distance.

[0112] Optionally, traverse each vocabulary position of the second language model B. If the word TokenB at this vocabulary position appears in the first vocabulary of the first language model A, set the probability of this position in the first vocabulary to TokenB, that is, if it is a unique matching relationship, the prediction probability of the matching text unit in the second word sequence can be directly assigned to the corresponding text unit in the first word sequence.

[0113] Optionally, for the relationship where a continuous word matches a single word, probability assignment can be performed by finding the position of the word with the closest semantics in the first language model A, specifically including: using the tokenizer of the first language model A to segment TokenB, and then setting the word vector of TokenB to the mean of the word vectors after segmentation. In the word vector space corresponding to the word sequence of A, find the closest one. The distance calculation here can use common distance formulas such as cosine and / or Euclidean distance.

[0114] Optionally, for a single word matching multiple word relationships, the matching at this time has no value, and the probability distribution at this position can be directly set to a one-hot encoding vector, that is, a 0 / 1 vector. The position of tokenA(i) in the vocabulary is 1, and the other positions are all 0. One-hot encoding is a numerical representation method commonly used in computer science and information theory, which is used to convert categorical variables into a format that can be better processed by machine learning algorithms. The Chinese name of the one-hot encoding vector is usually called "one-hot encoding vector" or "one-hot vector", and sometimes it is simply called "one-hot vector" or "one-hot encoding". The one-hot encoding vector is a binary (0 and 1) vector, most of whose elements are 0, and only one element is 1. In natural language processing, each word can be converted into a one-hot encoding vector, and the length of the vector is usually the same as the size of the vocabulary. For example, if there are 10,000 words in the vocabulary, then each word will be represented as a vector of length 10,000, where only one position is 1, indicating the index of the word in the vocabulary, and the rest of the positions are 0. One-hot encoding enables the original discrete data to be used for model training. In the embodiments of the present application, when a word in sequence A matches multiple words in sequence B, setting the probability distribution at this position to a one-hot encoding vector is to simplify the processing and retain the independent representation of this word in model A, ignoring all other possible matching words.

[0115] For example, for the word sequence: we utilize the dynamic programming approach to align tokens, after tokenizing this sentence, the resulting sequence is assumed to have matching results such as [we,we], [(util,ize),utilize]... At this time, it is necessary to determine how to copy the probability of the second word sequence in each specific matching result. For we or utilize, there are probability vectors of the size of the vocabulary. At this time, the first model needs to learn to predict the probability distribution of the text composition units at each position in this sentence.

[0116] For the vocabulary probability vector, for each position in the word sequence, the large model has a complete probability table, which represents the output probability of each text composition unit in the vocabulary for the current position, and then the one with the highest probability is taken as the true output of this position.

[0117] Optionally, the process of assigning the probability distribution vector needs to involve N probabilities, where N is the size of the vector (also the size of the second vocabulary). For each text component unit of the second vocabulary (vector), if this text component unit is in the first vocabulary, find the position in the first vocabulary and then assign a value. If it does not appear, segment this text component unit using the first vocabulary, and then the word vector is the mean of these segmentation results. After obtaining the mean word vector X, for each text component unit in the first vocabulary, check which text component unit has the closest word vector to X and then assign a value.

[0118] Optionally, during the process of performing the sequence assignment process, for the assignment of positions [i'+1…i] in A, there are the following two possible situations:

[0119] (1) Only assign a value at the position of i'+1. The one-to-one matching assignment process can refer to the foregoing implementation method. For the positions [i'+2,…i], use the One-hot vector mode, that is, for the word vector at each position, the position of the word that appears in the vocabulary is 1, and the others are 0;

[0120] (2) Execute the one-to-one matching process in the foregoing implementation method for each position in this sequence segment.

[0121] Through the above solution in this embodiment, specifically, by determining, according to the word segmentation sequence matching result, the second subsequence in the first word sequence that has the subsequence matching relationship with the first subsequence in the second word sequence, and performing a sequence assignment process on each word segmentation position in the second subsequence; and / or, traversing each word segmentation position in the second word sequence, identifying the matching relationship between the text component units at each word segmentation position in the second word sequence and the text component units at each word segmentation position in the first word sequence, and performing a probability distribution assignment on each word segmentation position in the first word sequence according to the matching relationship to obtain the vocabulary alignment result. By proposing the optimal sequence matching function, the matching problem of Chinese model sentences can be solved. At the same time, the recall for English sentences will also be enhanced, and the mismatch rate will be reduced. In addition, the distance formula uses the distance between word vectors, considering the semantic similarity. Modifying the vocabulary probability distribution assignment to the method of the closest word vector also makes the probability distribution learning consider the semantic similarity in the semantic space as much as possible, avoiding the irrationality problem that using the edit distance hard coding may cause semantic misalignment and may bring negative effects in probability distribution learning, and is also more friendly to Chinese and the like.

[0122] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the vocabulary alignment method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.

[0123] The present application also provides a vocabulary alignment device. Please refer to Figure 3 , which is applied to a first device. The first device is communicatively connected to a second device. A first language model is deployed in the first device, and a second language model is deployed in the second device. The first language model and the second language model are heterogeneous. The device includes:

[0124] An acquisition module, configured to acquire a first word sequence, a second word sequence, and a second prediction probability. Wherein, the first word sequence is obtained by the first language model segmenting a training text through a first vocabulary, the second word sequence is obtained by the second language model segmenting the training text through a second vocabulary, and the second prediction probability is a probability distribution vector corresponding to each segmentation position in the second word sequence and each text composition unit in the second vocabulary;

[0125] A matching module, configured to perform dynamic programming on the first word sequence and the second word sequence, and backtrack from back to front to obtain a segmentation sequence matching result with the minimum matching cost. Wherein, the segmentation sequence matching result includes at least one of a unique matching relationship, a continuous word matching a single word relationship, a single word matching multiple words relationship, and a subsequence matching relationship;

[0126] An assignment module, configured to perform probability distribution assignment on each segmentation position in the first word sequence according to the second prediction probability based on the segmentation sequence matching result, to obtain a vocabulary alignment result.

[0127] The vocabulary alignment device provided by the present application adopts the vocabulary alignment method in the above embodiment, and can solve the technical problem of vocabulary alignment. Compared with the prior art, the beneficial effects of the vocabulary alignment device provided by the present application are the same as those of the vocabulary alignment method provided by the above embodiment, and other technical features in the vocabulary alignment device are the same as the features disclosed in the method of the above embodiment, and will not be elaborated here.

[0128] The present application provides a vocabulary alignment device. The vocabulary alignment device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the vocabulary alignment method in the first embodiment above.

[0129] Next, refer to Figure 4, which shows a schematic structural diagram of a vocabulary alignment device suitable for implementing the embodiments of the present application. The vocabulary alignment device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 4 The shown vocabulary alignment device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0130] As Figure 4 shown, the vocabulary alignment device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM: Read Only Memory) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the vocabulary alignment device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the vocabulary alignment device to communicate with other devices wirelessly or wireline to exchange data. Although the figure shows a vocabulary alignment device with various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be implemented or had alternatively.

[0131] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by a processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.

[0132] The vocabulary alignment device provided by the present application adopts the vocabulary alignment method in the above embodiment and can solve the technical problem of vocabulary alignment. Compared with the prior art, the beneficial effects of the vocabulary alignment device provided by the present application are the same as those of the vocabulary alignment method provided by the above embodiment, and other technical features in the vocabulary alignment device are the same as those disclosed in the method of the previous embodiment and will not be elaborated here.

[0133] It should be understood that the various parts disclosed in the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0134] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0135] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the vocabulary alignment method in the above embodiment.

[0136] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0137] The above computer-readable storage medium can be included in the vocabulary alignment device; or it can exist separately without being assembled into the vocabulary alignment device.

[0138] The above computer-readable storage medium carries one or more programs. When the one or more programs are executed by the vocabulary alignment device, the vocabulary alignment device is caused to: obtain the first vocabulary of the first language model and the second vocabulary of the second language model; perform dynamic programming on each text component unit in the first vocabulary and the second vocabulary, and backtrack from the end to the beginning to obtain the word segmentation sequence matching result with the minimum matching cost; based on the word segmentation sequence matching result, assign a probability distribution to each text component unit in the first vocabulary according to the prediction probability corresponding to each text component unit in the second vocabulary, to obtain the vocabulary alignment result.

[0139] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0140] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0141] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.

[0142] The readable storage medium provided by this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned vocabulary alignment method, and can solve the technical problem of vocabulary alignment. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the vocabulary alignment method provided by the above embodiments, and will not be elaborated here.

[0143] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the vocabulary alignment method as described above.

[0144] The computer program product provided by the present application can solve the technical problem of vocabulary alignment. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the vocabulary alignment method provided by the above embodiments, and will not be elaborated here.

[0145] The above are only partial embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A vocabulary alignment method, characterized in that: Applied to a first device, the first device is in communication connection with a second device, a first language model is deployed in the first device, a second language model is deployed in the second device, the first language model and the second language model are heterogeneous, and the method includes: Obtaining a first word sequence, a second word sequence, and a second prediction probability, wherein the first word sequence is obtained by segmenting a training text by the first language model through a first vocabulary, the second word sequence is obtained by segmenting the training text by the second language model through a second vocabulary, and the second prediction probability is a probability distribution vector corresponding to each word segmentation position in the second word sequence and each text component unit in the second vocabulary; Dynamically programming the first word sequence and the second word sequence, and backtracking from back to front to find a word segmentation sequence matching result with a minimum matching cost, wherein the word segmentation sequence matching result includes at least one of a unique matching relationship, a continuous word matching a single word relationship, a single word matching multiple words relationship, and a subsequence matching relationship; Based on the word segmentation sequence matching result, probability distribution is assigned to each word segmentation position in the first word sequence according to the second prediction probability to obtain a word list alignment result.

2. The method according to claim 1, characterized in that The step of assigning probability distribution to each word segmentation position in the first word sequence based on the word segmentation sequence matching result according to the second predicted probability to obtain the word list alignment result comprises: According to the word segmentation sequence matching result, determining a second subsequence in the first word sequence that has the subsequence matching relationship with the first subsequence in the second word sequence, and performing a sequence assignment process on each word segmentation position in the second subsequence; and / or, Traverse each word segmentation position in the second word sequence, identify the matching relationship between the text component units of each word segmentation position in the second word sequence and the text component units of each word segmentation position in the first word sequence, and assign probability distribution to each word segmentation position in the first word sequence according to the matching relationship to obtain the word list alignment result.

3. The method according to claim 2, characterized in that The step of assigning probability distribution to each word segmentation position in the first word sequence according to the matching relationship comprises: Determine whether there is a segmentation sequence matching result between the text component units at each segmentation position in the second word sequence and the text component units at each segmentation position in the first word sequence according to the matching relationship; If there is a word segmentation sequence matching result between the text component units of each word segmentation position in the second word sequence and the text component units of each word segmentation position in the first word sequence, a probability distribution is assigned to each corresponding word segmentation position in the first word sequence according to the word segmentation sequence matching result.

4. The method according to claim 3, characterized in that If there is a word segmentation sequence matching result between the text component units of each word segmentation position in the second word sequence and the text component units of each word segmentation position in the first word sequence, the step of assigning probability distribution values ​​to the corresponding word segmentation positions in the first word sequence according to the word segmentation sequence matching result includes: If the word segmentation sequence matching result is a unique matching relationship, the word segmentation position of the text component unit in the first word sequence corresponding to the text component unit of each word segmentation position in the second word sequence is set to the probability distribution vector corresponding to the text component unit of the corresponding word segmentation position in the second word sequence; If the segmentation sequence matching result is a continuous word matching single word relationship, a sequence assignment process is performed on the segmentation position of the text component unit in the first word sequence corresponding to the text component unit of the segmentation position in the second word sequence; If the word segmentation sequence matching result is a single word matching multiple word relationship, the probability distribution vector of the word segmentation position of the text component unit in the first word sequence corresponding to the text component unit at the word segmentation position of the second word sequence is set to a one-bit valid encoding vector.

5. The method according to any one of claims 2 or 4, characterized in that The steps of the sequence assignment process include: The probability distribution vector of the position corresponding to the first position of the second subsequence in the first word sequence is set as the probability distribution vector of the corresponding word segmentation position in the second word sequence, and the probability distribution vector of the word segmentation position other than the first position of the second subsequence in the second subsequence or the first word sequence is set as a one-bit effective coding vector; and / or, The probability distribution vector of the position corresponding to the matching result of the text component unit at each word segmentation position in the second subsequence or the first word sequence in the second word sequence is set as the probability distribution vector of the corresponding word segmentation position in the second word sequence.

6. The method according to claim 5, characterized in that The step of setting the probability distribution vector of the corresponding word segmentation position in the second word sequence comprises: For the probability distribution vector of the text component unit of the second vocabulary at the matched position in the second word sequence, traverse each text component unit to determine whether each text component unit appears in the first vocabulary; For the text component units in the second vocabulary that appear in the first vocabulary, assign corresponding probabilities at corresponding positions in the first vocabulary; For text constituent units that do not appear in the second vocabulary in the first vocabulary, the first vocabulary is used to perform word segmentation to obtain a segmentation result, and the word vector mean of the text constituent unit is calculated based on the segmentation result, and the text constituent unit in the first vocabulary with the smallest distance to the target text constituent unit is selected; the probability corresponding to the text constituent unit is used as the probability of the text constituent unit with the smallest distance.

7. A vocabulary alignment device, characterized in that: Applied to a first device, the first device is communicatively connected with a second device, a first language model is deployed in the first device, a second language model is deployed in the second device, the first language model and the second language model are heterogeneous, and the apparatus includes: an acquisition module, configured to acquire a first word sequence, a second word sequence, and a second prediction probability, wherein the first word sequence is obtained by segmenting a training text by the first language model through a first vocabulary, the second word sequence is obtained by segmenting the training text by the second language model through a second vocabulary, and the second prediction probability is a probability distribution vector corresponding to each word segmentation position in the second word sequence and each text component unit in the second vocabulary; A matching module, configured to dynamically program the first word sequence and the second word sequence, and backtrack from back to front to obtain a word segmentation sequence matching result with a minimum matching cost, wherein the word segmentation sequence matching result includes at least one of a unique matching relationship, a continuous word matching a single word relationship, a single word matching multiple words relationship, and a subsequence matching relationship; An assignment module is used to perform probability distribution assignment on each word segmentation position in the first word sequence based on the word segmentation sequence matching result and the second predicted probability to obtain a word list alignment result.

8. A vocabulary alignment device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the vocabulary alignment method according to any one of claims 1 to 6.

9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the vocabulary alignment method according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the vocabulary alignment method according to any one of claims 1 to 6 are implemented.