Text processing method, model training method, device, equipment and product
By introducing a combination model into a large language model for data fusion, the problem of difficulty in adapting large language models in professional fields is solved, and rapid adaptation and efficient text processing are achieved under limited resources, improving the text processing effect in professional fields.
Patent Information
- Application Number
- CN202510191200.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-07-01
AI Technical Summary
When the prior art adapts large language models to professional fields, there are problems such as difficult updates and poor text processing, especially when it is difficult to achieve fast adaptation and efficient text processing under limited resources.
Using a combination model, data fusion is carried out by combining the first language model suitable for the target professional field and the second language model suitable for the general field. The feature representation between the two is fused through the combination model, reducing parameter adjustments and improving text processing capabilities.
With limited resources, the rapid adaptation of large language models in the professional field has been achieved, which improves text processing effect, reduces training difficulty and cost, and enhances the performance of text processing models in the professional field.
Smart Images

Figure CN120235147A_ABST
Abstract
Description
Technical Field
[0001] This application is applied to the field of natural language technology, and particularly relates to a text processing method, a model training method, a device, a device and a product. Background Art
[0002] With the rise of chatbots, large language models (LLMs) have gradually become a research hotspot in the field of artificial intelligence. How to quickly implement large language models in various professional fields (such as the medical field) is a question that relevant practitioners have been thinking about.
[0003] In related technologies, training data in a professional field is collected, and through the training data, further pre-training is performed on a general large language model, or the parameters of the general large language model are fine-tuned, so that the large language model can be adapted to this professional field.
[0004] However, with the enhancement of the performance of large language models, it is difficult to update the large language models in the professional field obtained by the above methods, resulting in poor text processing effects in the professional field. Summary of the Invention
[0005] To solve the above problems, this application proposes a text processing method, a model training method, a device, a device and a product, which can improve the text processing effect of large language models in professional fields.
[0006] The first aspect of this application provides a text processing method, including obtaining a target text to be processed in a target professional field; processing the target text through a text processing model to obtain a processing result; wherein, the text processing model includes a first large language model, a second large language model and a combination model, the first large language model is applicable to the target professional field, the second large language model is applicable to the general field, and the combination model is used for data fusion between the first large language model and the second large language model.
[0007] In some embodiments, the combination model is used for feature fusion between the feature representations output by N selected network layers in the first large language model and the feature representations output by M selected network layers in the second large language model, N is greater than or equal to 1, and M is greater than or equal to 1.
[0008] In some embodiments, the process of fusing the feature representation output by the \(i\)-th selected network layer among the \(N\) selected network layers with the feature representation output by the \(j\)-th selected network layer among the \(M\) selected network layers through the combined model includes: in the combined model, transforming the feature representation output by the \(i\)-th selected network layer and the feature representation output by the \(j\)-th selected network layer into feature representations with the same dimension; after the dimension transformation, fusing the feature representation output by the \(i\)-th selected network layer and the feature representation output by the \(j\)-th selected network layer to obtain a fused feature; where the value range of \(i\) is from 1 to \(N\), and the value range of \(j\) is from 1 to \(M\).
[0009] In some embodiments, the combined model includes a projection layer; the transforming the feature representation output by the \(i\)-th selected network layer and the feature representation output by the \(j\)-th selected network layer into feature representations with the same dimension includes: inputting the feature representation output by the \(i\)-th selected network layer into the projection layer; in the projection layer, performing a linear transformation on the feature representation output by the \(i\)-th selected network layer to obtain a projection representation, and the dimension of the projection representation is the same as the dimension of the feature representation output by the \(j\)-th selected network layer.
[0010] In some embodiments, the combined model further includes a cross-attention network; the fusing the feature representation output by the \(i\)-th selected network layer and the feature representation output by the \(j\)-th selected network layer to obtain a fused feature includes: inputting the projection representation and the feature representation output by the \(j\)-th selected network layer into the cross-attention network; in the cross-attention network, fusing the projection representation and the feature representation output by the \(j\)-th selected network layer to obtain the fused feature.
[0011] In some embodiments, the cross-attention network includes \(H\) cross-attention heads, where \(H\) is greater than or equal to 1; the fusing the projection representation and the feature representation output by the \(j\)-th selected network layer in the cross-attention network to obtain the fused feature includes: determining the key vector and the value vector of the \(k\)-th cross-attention head in the cross-attention layer according to the projection representation, where the value range of \(k\) is from 1 to \(H\); determining the query vector of the \(k\)-th cross-attention head according to the feature representation output by the \(j\)-th selected network layer; performing a cross-attention operation on the key vector, the value vector and the query vector through a cross-attention mechanism to obtain the output data of the \(k\)-th cross-attention head; determining the fused feature according to the output data.
[0012] In some embodiments, the cross-attention mechanism includes a cross-attention mask; the cross-attention mask is determined through the following process: According to the tokenizer in the first large language model, determine the character position information corresponding to each of the multiple tokens in the first token sequence, where the first token sequence is obtained by tokenizing the target text through the tokenizer in the first large language model; According to the tokenizer in the second large language model, determine the character position information corresponding to each of the multiple tokens in the second token sequence, where the second token sequence is obtained by tokenizing the target text through the tokenizer in the second large language model; Generate the cross-attention mask according to the character position information corresponding to each of the multiple tokens in the first token sequence and the character position information corresponding to each of the multiple tokens in the second token sequence.
[0013] In some embodiments, the generating the cross-attention mask according to the character position information corresponding to each of the multiple tokens in the first token sequence and the character position information corresponding to each of the multiple tokens in the second token sequence includes: Determine a first valid token according to the attention mask of the first large language model, where the first valid token is a valid token in the first token sequence; Determine a second valid token according to the attention mask of the second large language model, where the second valid token is a valid token in the second token sequence; In the character position information corresponding to each of the multiple tokens in the first token sequence, find the character position information of the first valid token; In the character position information corresponding to each of the multiple tokens in the second token sequence, find the character position information of the second valid token; Compare the character position information of the first valid token with the character position information of the second valid token to obtain a comparison result; Generate the cross-attention mask according to the comparison result.
[0014] In some embodiments, the determining the key vector and the value vector of the k-th cross-attention head in the cross-attention layer according to the projected representation includes: Multiply the first weight matrix corresponding to the k-th cross-attention head by the projected representation to obtain the key vector; Multiply the second weight matrix corresponding to the k-th cross-attention head by the projected representation to obtain the value vector; The determining the query vector of the k-th cross-attention head according to the feature representation output by the j-th selected network layer includes: Multiply the third weight matrix corresponding to the k-th cross-attention head by the feature representation output by the j-th selected network layer to obtain the query vector.
[0015] The second aspect of the present application provides a model training method, including: obtaining training texts in a target professional field; processing the training texts through a text processing model to obtain a processing result, where the text processing model includes a first large language model, a second large language model, and a combination model, the first large language model is applicable to the target professional field, the second large language model is applicable to the general field, and the combination model is used for data fusion between the first large language model and the second large language model; adjusting the parameters of the combination model according to the processing result to obtain the text processing model after one round of training.
[0016] In some embodiments, the combination model includes a cross-attention network, and the parameters of the cross-attention network include the weight matrix corresponding to the cross-attention heads. The adjusting the parameters of the combination model according to the processing result to obtain the text processing model after one round of training includes: adjusting the weight matrix corresponding to the cross-attention heads according to the processing result to obtain the text processing model after one round of training.
[0017] The third aspect of the present application provides a text processing device, including: an obtaining unit for obtaining a target text to be processed in a target professional field; a processing unit for processing the target text through a text processing model to obtain a processing result; where the text processing model includes a first large language model, a second large language model, and a combination model, the first large language model is applicable to the target professional field, the second large language model is applicable to the general field, and the combination model is used for data fusion between the first large language model and the second large language model.
[0018] The fourth aspect of the present application provides a model training device, including: an obtaining unit for obtaining training texts in a target professional field; a processing unit for processing the training texts through a text processing model to obtain a processing result, where the text processing model includes a first large language model, a second large language model, and a combination model, the first large language model is applicable to the target professional field, the second large language model is applicable to the general field, and the combination model is used for data fusion between the first large language model and the second large language model; an adjusting unit for adjusting the parameters of the combination model according to the processing result to obtain the text processing model after one round of training.
[0019] The fifth aspect of the present application provides an electronic device, including a memory and a processor; the memory is connected to the processor and is used for storing programs; the processor is used for implementing the text processing method as described in the first aspect or any embodiment of the first aspect, or implementing the model training method as described in the second aspect or any embodiment of the second aspect by running the programs in the memory.
[0020] The sixth aspect of the present application provides a chip, including a processor and a data interface. The processor reads and runs a program stored on a memory through the data interface to execute the text processing method described in the first aspect or any embodiment of the first aspect, or to execute the model training method described in the second aspect or any embodiment of the second aspect.
[0021] The seventh aspect of the present application provides a computer program product, including a computer program which, when executed by a processor, implements the text processing method described in the first aspect or any embodiment of the first aspect, or implements the model training method described in the second aspect or any embodiment of the second aspect.
[0022] The eighth aspect of the present application provides a storage medium, on which a computer program is stored. When the computer program is run by a processor, it implements the text processing method described in the first aspect or any embodiment of the first aspect, or implements the model training method described in the second aspect or any embodiment of the second aspect.
[0023] According to a text processing method, model training method, device, equipment and product provided by the present application, in a text processing model, through a combined model, data fusion is performed on a large language model in the general domain and a large language model in the professional domain, so that the text processing model has the text processing capabilities of the large language model in the general domain and the large language model in the professional domain, improving the performance of the text processing model in text processing tasks in the professional domain and improving the text processing effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0025] Figure 1 It is a schematic diagram of the implementation environment involved in the embodiments of the present application;
[0026] Figure 2 It is a schematic flowchart of the text processing method provided by the embodiments of the present application;
[0027] Figure 3 It is a schematic flowchart of the fusion processing of the feature representation output by the i-th selected network layer in the N selected network layers and the feature representation output by the j-th selected network layer in the M selected network layers in the text processing method provided by the embodiments of the present application;
[0028] Figure 4 It is a structural example diagram of a text processing model provided according to an embodiment of the present application;
[0029] Figure 5 It is a schematic flowchart of a model training method provided according to an embodiment of the present application;
[0030] Figure 6 It is a structural schematic diagram of a text processing device provided according to an embodiment of the present application;
[0031] Figure 7 It is a structural schematic diagram of a model training device provided according to an embodiment of the present application;
[0032] Figure 8 It is a structural schematic diagram of an electronic device provided according to an embodiment of the present application. Detailed implementation manners
[0033] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0034] In the related art, for the adaptation of the base large language model (a large language model applicable to the general field) to a specific field, it is achieved by further pre-training or fine-tuning the large language model using the training data of the specific field. However, on the one hand, the structure of the base large language model is becoming more and more huge and complex, and on the other hand, the corpus data of the specific field is relatively scarce. It is difficult, inefficient and costly to pre-train or adjust the parameters on the complex base large language model to obtain a large language model adapted to the specific field.
[0035] There are multiple ways to achieve the adaptation of large language models in specific domains with limited computing resources. Method 1: Parameter-efficient fine-tuning (PEF). For example, the low-rank adaptation of large language models (LoRA) method keeps the structure of the large language model unchanged and realizes the efficient fine-tuning of the large language model in a specific domain by introducing a small number of trainable parameters, aiming to reduce the amount of parameter adjustment during model training, lower the computing cost, reduce memory occupancy, and maintain good performance of the large language model on specific tasks. Method 2: Model Merging. For example, the task vector method combines the capabilities of different models by weighted averaging of different models.
[0036] However, in Method 1, PEF is mainly applied to similar domains within the domain of the large language model (or similar tasks processed by the large language model). If it is a completely different professional domain or a brand-new task, the capabilities of the large language model obtained through the PEF method are poor, and the model performance is inferior to that of the large language model obtained through full-parameter fine-tuning or other more flexible methods. In Method 2, Model Merging is applicable to the case where the models participating in the merging are well-aligned (such as the same model structure and the same number of parameters), and it is not applicable to diverse large language models.
[0037] Therefore, to achieve the rapid adaptation of large language models in specific professional domains with limited resources and improve the performance of large language models in specific professional domains, the embodiments of this application propose a text processing method, a model training method, a device, a device, and a product. In the embodiments of this application, the text processing model includes a first large language model applicable to the target professional domain, a second large language model applicable to the general domain, and a combination model, enabling the text processing model to have the text processing capabilities of the general domain and the target professional domain. During the training process of the text processing model, the parameters of the first large language model and the second large language model do not need to change, and the parameters of the combination model are adjusted, reducing the number of parameters in model training. Thus, the text processing capabilities of the large language model in specific domains are improved, and the rapid adaptation of large language models in specific professional domains is achieved with limited resources.
[0038] Exemplary implementation environment
[0039] Please refer to Figure 1 , Figure 1It is a schematic diagram of the implementation environment involved in the embodiments of the present application. In the implementation environment involved in the present application, it includes a text processing device 100 and a model training device 101. A text processing model can be trained on the model training device 101. The text processing model includes a first large language model, a second large language model, and a combined model; the trained text processing model is deployed on the text processing device 100. In the actual application process, on the text processing device 100, the text in the professional field can be processed through the text processing model.
[0040] Among them, the text processing device 100 and the model training device 101 can be a terminal or a server. Figure 1 Taking the text processing device 100 and the model training device 101 as servers as an example. When the text processing device 100 is a server, the implementation environment may further include a user terminal 103. The text processing device 100 can receive the text to be processed input by the user through the user terminal 103, and after processing the text to be processed through the text processing model, return the processing result to the user terminal 103.
[0041] As an example, the above implementation environment is an intelligent question - answering environment, an intelligent translation environment, etc. in the professional field. In the intelligent question - answering environment, the text processing model performs a question - answering task. The user inputs a question text, and the text processing model outputs an answer text; in the intelligent translation environment, the text processing model performs a translation task. The user inputs the text to be translated, and the text processing model outputs the translation result text.
[0042] Exemplary method
[0043] Please refer to Figure 2 , in an exemplary embodiment, a text processing method is provided. The text processing method includes:
[0044] S201, obtain the target text to be processed in the target professional field.
[0045] In this embodiment, the target text input by the user can be received, or the target text sent by other devices (such as the terminal where the user is located) can be received. Or, the task information of the next task can be obtained from the task list, and the target text can be obtained from the task information.
[0046] In one example, the target professional field is the medical field.
[0047] When the target field is the medical field, the target text is medical - related text, and the tasks performed by the text processing model can be medical question - answering tasks, medical term translation / interpretation tasks.
[0048] In the medical field, medical documents such as medical examination reports and clinical diagnosis and treatment records contain complex medical terms. After receiving these medical documents, patients may not be able to understand the specific meanings. The target text can be the text contained in the medical documents. The target text can be translated from the professional description statements in the medical field into simple and easy-to-understand statements through a text processing model.
[0049] S202, through the text processing model, process the target text to obtain a processing result. The text processing model includes a first large language model, a second large language model, and a combination model. The first large language model is applicable to the target professional field, the second large language model is applicable to the general field, and the combination model is used for data fusion between the first large language model and the second large language model.
[0050] Among them, the first large language model and the second large language model can be models with different structures. The model complexity of the second large language model is greater than that of the first large language model. Compared with the first large language model, the second large language model has stronger text processing capabilities in the general field.
[0051] If the second large language model is further adjusted to be a large language model adapted to the target professional field by means of parameter fine-tuning, the number of parameters of the second large language model is large, and the training process will consume high time costs and computing resources. Moreover, there is a need for sufficient corpus in the target professional field, and the data collection difficulty is large. In this embodiment, the first large language model and the second large language model are combined through the combination model. By means of the text processing capabilities of the first large language model in the target professional field, the disadvantage that the second large language model has weak text processing capabilities in the professional field is made up for; the first large language model is a model trained well in the target professional field, and the second large language model is a model trained well in the general field. During the training process of the text processing model, the parameters of the combination model are adjusted, and the parameters of the first large language model and the second large language model do not need to be adjusted, reducing the training difficulty of the text processing model and reducing the time costs and computing resources consumed by the training of the text processing model.
[0052] Among them, the combination model can be used to fuse the feature representations output by the network layer in the first large language model and the feature representations output by the network layer in the second large language model to obtain fused features. The fused features can be continuously input into the next network layer in the second large language model. Thus, by fusing the features extracted by the first large language model into the features of the second large language model, the second large language model is made to have the text processing capabilities of the first large language model in the target professional field, so that the text processing model has the capabilities of the first large language model and the second large language model and has strong text processing capabilities in the target professional field.
[0053] In this embodiment, the target text is respectively input into the first large language model and the second large language model included in the text processing model. The network layer in the first large language model is used to extract features from the target text to obtain the feature representation output by this network layer. The network layer in the second large language model is used to extract features from the target text to obtain the feature representation output by this network layer. The feature representation output by the network layer in the first large language model and the feature representation output by the network layer in the second large language model can be input into the combined model. In the combined model, the feature representation output by the network layer in the first large language model and the feature representation output by the network layer in the second large language model are fused to obtain the fused features output by the combined model. The input data for the next network layer of the second large language model can be determined according to the fused features output by the combined model, and this input data is input into this next network layer to continue feature processing in this next network layer. In this way, the above process can be executed once or multiple times, and finally the processing result corresponding to the target text is obtained.
[0054] In one example, the target text is a consultation question to be answered in the medical field, and the processing result is the reply text corresponding to this consultation question; or, the target text is a professional statement to be explained in the medical field, and the processing result is the explanation text corresponding to this professional statement.
[0055] For example, in the scenario of interpreting diagnostic results, the target text input into the text processing model is "The patient is diagnosed with acute myocardial infarction and requires urgent coronary intervention", and the explanation text output by the text processing model is "Your condition is a heart attack, and it is recommended to immediately perform cardiovascular treatment to restore blood flow"; another example is, in the scenario of interpreting laboratory indicators, the target text input into the text processing model is "The glucose level is 220 mg / dL, and the patient shows obvious diabetes characteristics", and the explanation text output by the text processing model is "Your blood sugar is high, which indicates a possible risk of diabetes and requires further diet control and treatment".
[0056] In the embodiments of the present application, in a text processing model, through a combined model, the feature representations output by the network layer of a first large language model and the feature representations output by a second large language model are fused. Since the first large language model has the text processing ability in the target professional field, and the second large language model has strong text processing ability in the general field, this feature fusion operation enables the second large language model to obtain the features learned by the first large language model, that is, to obtain the text processing ability of the first large language model in the target professional field. Thus, by combining a language processing model with a relatively low model complexity in the target professional field and a language processing model with a relatively high complexity in the general field, without the need for large-scale parameter training, the text processing model becomes a large language model applicable to a specific professional field and with strong capabilities, achieving the rapid adaptation of the large language model in the specific professional field with limited resources.
[0057] Next, embodiments are provided for the fusion processing of the feature representations output by the network layer in the first large language model and the feature representations output by the network layer in the second large language model by the combined model.
[0058] In some embodiments, the combined model is used for feature fusion between the feature representations output by N selected network layers in the first large language model and the feature representations output by M selected network layers in the second large language model, where N is greater than or equal to 1 and M is greater than or equal to 1.
[0059] In this embodiment, among the multiple feature extraction network layers included in the first large language model, N feature extraction network layers (for the sake of distinction, called N selected network layers, which can be a part of the feature extraction network layers included in the first large language model or all of the feature extraction network layers included in the first large language model) can be selected to participate in feature fusion, and among the multiple feature extraction network layers included in the second large language model, M feature extraction network layers (for the sake of distinction, called M selected network layers, which can be a part of the feature extraction network layers included in the second large language model or all of the feature extraction network layers included in the second large language model) can be selected to participate in feature fusion.
[0060] For example, there are N A feature network layers in the first large language model and N B feature network layers in the second large language model. N selected network layers are selected from the N A feature network layers, and M selected network layers are selected from the N B feature network layers. The N selected network layers are denoted as L A , that is, |L A | = N, and the M selected network layers are denoted as L B , that is, |L B | = M.
[0061] In one example, N is equal to M. The selected network layers in the first large language model correspond one-to-one with the selected network layers in the second large language model. The combined model is used for feature fusion between the feature representations output by the first selected network layer in the first large language model and the feature representations output by the second selected network layer in the second large language model. The first selected network layer is any one of the N selected network layers, and the second selected network layer is the selected network layer corresponding to the first selected network layer among the M selected network layers.
[0062] In this example, the feature representations output by the first selected network layer and the feature representations output by the second selected network layer can be input into the combined model. In the combined model, the feature representations output by the first selected network layer and the feature representations output by the second selected network layer are fused to obtain fused features; the fused features are input into the next network layer of the second selected network layer in the second large language model to continue feature processing in this next network layer until the feature representations output by the next selected network layer of the second selected network are obtained; then, the feature representations output by the next selected network layer of the first selected network layer and the feature representations output by the next selected network layer of the second selected network layer are input into the combined model. In the combined model, the feature representations output by the next selected network layer of the first selected network layer and the feature representations output by the next selected network layer of the second selected network layer are fused to obtain fused features, and the fused features are input into the next network layer of the next selected network layer of the second selected network layer. In this way, until the processing result output by the output layer of the second selected network is obtained. In this way, through the one-to-one correspondence of the selected network layers between the first large language model and the second large language model, the features of different granularities between the first large language model and the second large language model are fused one-to-one, improving the feature fusion effect, and further improving the writing processing ability of the text processing model in the target professional field.
[0063] In one example, the N selected network layers are equally spaced in the first large language model, the M selected network layers are equally spaced in the second large language model, and the number of network layers between adjacent selected network layers in the N selected network layers is equal to the number of network layers between adjacent selected network layers in the M selected network layers, further enabling the feature fusion between the first large language model and the second large language model to be accurately fused in an orderly manner according to the feature extraction granularity, improving the feature fusion effect.
[0064] For example, among the N selected network layers, they are successively represented as l a1 、l a2 、l a3 、……、l aN ,and among the M selected network layers, they are successively represented as l b1 、l b2 、……、lbM , (l a2 -l a1 ) = (l a3 -l a2 ) = …… = (l a(N--1) -l aN ) = (l b2 -l b1 ) = (l b3 -l b2 ) = …… = (l b(M-1) -l bM ).
[0065] For example, in the first large language model, the N selected network layers are respectively the first feature extraction network layer, the second feature extraction network layer, the third feature extraction network layer, ……, the Nth feature extraction network layer; in the second large language model, the M selected network layers are respectively the first feature extraction network layer, the second feature extraction network layer, the third feature extraction network layer, ……, the Mth feature extraction network layer, and N is equal to M; at this time, the one-to-one correspondence between the N selected network layers and the M selected network layers is: the first feature extraction network layer in the first large language model corresponds to the first feature extraction network layer in the second large language model, the second feature extraction network layer in the first large language model corresponds to the second feature extraction network layer in the second large language model, ……, and so on.
[0066] Please refer to Figure 3 , in another exemplary embodiment, in the process of fusing the feature representations output by the ith selected network layer among the N selected network layers and the feature representations output by the jth selected network layer among the M selected network layers through the combined model, it may include:
[0067] S301, in the combined model, transform the feature representations output by the ith selected network layer among the N selected network layers of the first large language model and the feature representations output by the jth selected network layer among the M selected network layers of the second large language model into feature representations with the same dimension.
[0068] Wherein, the value range of i is from 1 to N, and the value range of j is from 1 to M. The jth selected network layer among the M selected network layers corresponds to the ith selected network layer among the N selected network layers.
[0069] Wherein, in the first large language model and the second large language model, the feature representation output by the network layer refers to the feature vector output by the network layer, and the dimension of the feature representation is the vector dimension of the feature vector.
[0070] In this embodiment, the first large language model and the second large language model are different language models. Therefore, the dimensions of the feature representations output by the network layers in the first large language model are different from those of the feature representations output by the network layers in the second large language model. The feature representation output by the $i$-th selected network layer among the $N$ selected network layers of the first large language model and the feature representation output by the $j$-th selected network layer among the $M$ selected network layers of the second large language model can be transformed into feature representations with the same dimension, so as to unify the dimensions of the feature representations participating in the fusion process and improve the feature fusion processing effect of the combined model.
[0071] In one example, the dimension of the feature representation output by the $i$-th selected network layer among the $N$ selected network layers is transformed into the dimension of the feature representation output by the $j$-th selected network layer among the $M$ selected network layers. In this way, the dimension of the feature representation obtained after fusion is the dimension of the feature representation output by the $j$-th selected network layer, which is convenient for directly inputting it into the next network layer of the $j$-th selected network layer for processing.
[0072] S302. After the dimension transformation, the feature representation output by the $i$-th selected network layer and the feature representation output by the $j$-th selected network layer are fused to obtain a fused feature.
[0073] In this embodiment, after unifying the feature representation output by the $i$-th selected network layer in the first large language model and the feature representation output by the $j$-th selected network layer in the second large language model into the same dimension through dimension transformation, that is, after aligning the feature representation output by the $i$-th selected network layer with the feature representation output by the $i$-th selected network layer, the feature representation output by the $i$-th selected network layer and the feature representation output by the $j$-th selected network layer are fused, such as performing corresponding weighted operations, to obtain the fused feature corresponding to the $i$-th selected network layer and the $j$-th selected network layer.
[0074] In the embodiment of the present application, for the fusion between the feature representation output by the $i$-th selected network layer among the $N$ selected network layers and the feature representation output by the $j$-th selected network layer among the $M$ selected network layers, first, the feature representation output by the $i$-th selected network layer and the feature representation output by the $j$-th selected network layer are dimensionally aligned, and then the two dimensionally aligned feature representations are fused, which improves the fusion effect of the feature representations from two different large language models.
[0075] Next, an exemplary implementation manner is provided for the steps of the combined model and Figure 3 the embodiments shown.
[0076] In some embodiments, the combined model includes a projection layer. As Figure 3As shown, S301 may include: S3011, inputting the feature representation output by the i-th selected network layer into the projection layer; S3012, in the projection layer, performing a linear transformation on the feature representation output by the i-th selected network layer to obtain a projection representation, and the dimension of the projection representation is the same as the dimension of the feature representation output by the j-th selected network layer. Thus, by means of performing a linear transformation in the projection layer, the dimension alignment of the feature representation output by the i-th selected network layer in the first large language model and the feature representation output by the j-th selected network layer in the second large language model is achieved while keeping the parameters of the first large language model and the second large language model unchanged.
[0077] Among them, the projection layer contains a projection function, and the projection function is a linear transformation function. During the training process of the combined model, the parameters of the projection function can be adjusted to improve the accuracy of the linear transformation of the feature representation by the projection function.
[0078] Among them, the combined model may contain multiple projection layers, and one projection layer may correspond to one selected network layer in the first large language model and one selected network layer in the second large language model.
[0079] In this embodiment, the feature representation output by the i-th selected network layer in the first large language model can be input into the projection layer corresponding to the i-th selected network layer (which is also the projection layer corresponding to the j-th selected network layer in the second large language model). In this projection layer, a linear transformation is performed on the feature representation output by the i-th selected network layer through the projection function to obtain a projection representation.
[0080] In one example, the projection function can be expressed as Among them, It means that the dimension of the feature representation output by the feature extraction network layer of the first large language model is D A ; It means that the dimension of the feature representation output by the feature extraction network layer of the second large language model is D B .
[0081] In one example, the linear transformation of the feature representations output by N selected network layers can be expressed as:
[0082] f proj (H A ) = {f proj (H A1 ), f proj (H A2 ),..., f proj (H AN )}
[0083] Among them, H ADenote the feature representations output by N selected network layers. Among the feature representations output by the N selected network layers, H A1 Denote the feature representation output by the first selected network layer, H A2 Denote the feature representation output by the second selected network layer, H AN Denote the feature representation output by the Nth selected network layer.
[0084] In some embodiments, the combined model further includes a cross-attention network. As Figure 3 shown, S302 includes: S3021, input the projection representation and the feature representation output by the jth selected network layer into the cross-attention network; S3022, in the cross-attention network, perform a fusion process on the projection representation and the feature representation output by the jth selected network layer to obtain a fused feature. Thus, through the cross-attention network, more important information in the projection representation and the feature representation output by the jth selected network layer can be focused on, improving the fusion effect of the projection representation and the feature representation output by the jth selected network layer.
[0085] Among them, the projection representation can refer to the description of the foregoing examples and will not be elaborated here.
[0086] In this embodiment, after inputting the projection representation and the feature representation output by the jth selected network layer into the cross-attention network, the projection representation and the feature representation output by the jth selected network layer can be used to determine the attention vectors of the cross-attention network. The attention vectors can include a key vector, a value vector, and a query vector; through the cross-attention mechanism, a weighted attention operation can be performed on the attention vectors to obtain a fused feature. Thus, based on the projection representation and the feature representation output by the jth selected network layer, the rationality and accuracy of the attention vectors are improved, and more important information in the projection representation and the feature representation output by the jth selected network layer can be focused on during the fusion process, improving the feature fusion effect.
[0087] In one example, the cross-attention network includes H cross-attention heads, where H is greater than or equal to 1; S3022 can include: determining the key vector and the value vector of the kth cross-attention head in the cross-attention layer according to the projection representation, where the value range of k is from 1 to H; determining the query vector of the kth cross-attention head according to the feature representation output by the jth selected network layer; performing a cross-attention operation on the key vector of the kth cross-attention head, the value vector of the kth cross-attention head, and the query vector of the kth cross-attention head through the cross-attention mechanism to obtain the output data of the kth cross-attention head; and determining the fused feature according to the output data.
[0088] In this example, in the cross-attention network, the query vector represents the information that the cross-attention network needs to focus on, the key vector represents the information provided for the cross-attention network to search for, and the value vector represents the importance of the key vector. According to the feature representation output by the j-th selected network layer, the query vector of the k-th cross-attention head is determined so that the k-th cross-attention head focuses on the feature representation output by the j-th selected network layer; according to the projection representation, the key vector of the k-th cross-attention head and the value vector of the k-th cross-attention head are determined so that the k-th cross-attention head can find the information to be focused on in the key vector according to the value vector, that is, find the information to be focused on in the projection representation. Then, through the cross-attention mechanism, cross-attention operations are performed on the key vector of the k-th cross-attention head, the value vector of the k-th cross-attention head, and the query vector of the k-th cross-attention head to obtain the output data of the k-th cross-attention head; the output data of the H cross-attention heads can be concatenated to obtain a concatenation result, and then the concatenation result is input into a linear layer for linear processing to obtain the fused feature output by the cross-attention network, that is, the fused feature of the feature representation output by the i-th selected network layer in the first large language model and the feature representation output by the j-th selected network layer in the second large language model.
[0089] Thus, in the cross-attention network, through H cross-attention heads, mainly focusing on the feature representation output by the j-th selected network layer in the second large language model, the relevant information in the feature representation output by the i-th selected network layer in the first large language model is fused with the feature representation output by the j-th selected network layer in the second large language model, so that the second large language model can obtain the important features extracted by the first large language model and learn the text processing ability of the first large language model in the target professional field.
[0090] Optionally, multiply the first weight matrix corresponding to the k-th cross-attention head by the projection representation to obtain the key vector of the k-th cross-attention head; multiply the second weight matrix corresponding to the k-th cross-attention head by the projection representation to obtain the value vector of the k-th cross-attention head. Correspondingly, multiply the third weight matrix corresponding to the k-th cross-attention head by the feature representation output by the j-th selected network layer to obtain the query vector of the k-th cross-attention head.
[0091] Among them, during the training process of the text processing model store, the first weight matrix, the second weight matrix, and the third weight matrix can be adjusted to improve the accuracy of the first weight matrix, the second weight matrix, and the third weight matrix, and further improve the accuracy of the key vector, the value vector, and the query vector.
[0092] In this optional method, the key vector of the k-th cross-attention head can be expressed as:
[0093]
[0094] The value vector of the k-th cross-attention head can be expressed as:
[0095]
[0096] The query vector of the k-th cross-attention head can be expressed as:
[0097]
[0098] where, K Ak represents the key vector of the k-th cross-attention head, H Ai represents the feature representation output by the i-th selected network layer in the first large language model, f proj (H Ai ) represents the projection representation corresponding to the feature representation output by the i-th selected network layer in the first large language model, represents the first weight matrix corresponding to the k-th cross-attention head, i.e., has a dimension of D B × D B / H; V Ak represents the value vector of the k-th cross-attention head, represents the second weight matrix corresponding to the k-th cross-attention head, Q Bk represents the query vector of the k-th cross-attention head, H Bj represents the feature representation output by the j-th selected network layer in the second large language model, represents the third weight matrix corresponding to the k-th cross-attention head,
[0099] The attention formula of the k-th cross-attention head can be expressed as:
[0100] head k = Atten(Q Bk , K Ak , V Ak )
[0101] where, head k represents the fused feature output by the k-th cross-attention head, and Atten() represents the attention formula corresponding to the cross-attention mechanism in the k-th cross-attention head.
[0102] Optionally, the output data of H cross-attention heads are concatenated to obtain a concatenation result, and then the concatenation result is input into a linear layer for linear processing to obtain the fused feature output by the cross-attention network, which can be expressed by the following formula:
[0103] f cross (f proj (H Ai ),H Bj ) = Concat .k (head k )W o
[0104] where f cross () represents a cross-attention network, and the value range of k is from 1 to H. Therefore, Concat .k (head k ) represents concatenating the output data of H attention heads, and W o represents the weight parameter of the linear layer, which is a learnable weight matrix in the training process of the text processing model. That is, this weight parameter can be adjusted during the training process of the text processing model.
[0105] In one example, the cross-attention mechanism includes a cross-attention mask. That is, the cross-attention mask is included in the above attention formula. Thus, during the fusion process, the cross-attention network can focus on the features with non-zero corresponding mask values and ignore the features with zero corresponding mask values, and fuse the features that need to be focused on to improve the feature fusion effect.
[0106] In one possible implementation, the cross-attention mask can be determined through the following process:
[0107] According to the tokenizer in the first large language model, determine the character position information corresponding to each token in the first token sequence. The first token sequence is obtained by tokenizing the target text through the tokenizer in the first large language model; according to the tokenizer in the second large language model, determine the character position information corresponding to each token in the second token sequence. The second token sequence is obtained by tokenizing the target text through the tokenizer in the second large language model; generate a cross-attention mask according to the character position information corresponding to each token in the first token sequence and the character position information corresponding to each token in the second token sequence.
[0108] where the tokenizer is used to divide one or more texts into multiple tokens (a token can contain one or more characters, and the characters can be words, punctuation marks, etc., or special tokens, such as the special token bos_token used to represent the beginning of a sentence n and the special token eos_token used to represent the end of a sentence), to obtain a token sequence composed of multiple tokens.
[0109] Among them, the character position information corresponding to a token may include the starting character position of the token and the ending character position of the token. The starting character position may be the sorting position of the first character in the token in the target text or the sorting position of the character preceding the first character in the target text; the ending character position may be the sorting position of the last character in the token in the target text or the sorting position of the character following the last character in the target text.
[0110] In this implementation manner, during the process of inputting the target text into the text processing model, the target text can be input into the tokenizer in the first large language model, and the target text is tokenized in the tokenizer to obtain the first token sequence; the target text is input into the tokenizer in the second large language model, and the target text is tokenized in the tokenizer to obtain the second token sequence. According to the positions of the characters in the target text of the multiple tokens in the first token sequence, the character position information corresponding to each of the multiple tokens is constructed; according to the positions of the characters in the target text of the multiple tokens in the second token sequence, the character position information corresponding to each of the multiple tokens is constructed. According to the character position information corresponding to the multiple tokens in the first token sequence and the character position information corresponding to the multiple tokens in the second token sequence, the mask values in the cross-attention mask are determined.
[0111] Optionally, the character position information corresponding to the multiple tokens in the first token sequence can be represented as a position mapping table corresponding to the first token sequence, and the position mapping table contains the character position information corresponding to the multiple tokens. Among them, the data form of the position mapping table is a two-dimensional array, and the character position information corresponding to the multiple tokens is the elements in the two-dimensional array.
[0112] For example, for the target text "The patient was diagnosed with acute myocardial infarction and requires emergency coronary intervention", tokenization is performed to obtain the following tokens in the token sequence: "patient", "diagnosed", "with", "acute", "myocardial infarction", "(", "acute", "myocardial", "infarction", ")", ",", "requires", "emergency", "coronary artery", "intervention", "treatment". Then, starting from the character position 0, a mapping from each token to the character position is created: patient-(0, 2), diagnosed-(3, 5), with-(6, 7), acute-(8, 10),... Finally, the mapping table {(0, 2), (3, 5), (6, 7), (8, 10),...} can be constructed.
[0113] Optionally, after obtaining the first token sequence, the sequence length of the first token sequence can be adjusted according to the context length corresponding to the first large language model, so that the sequence length of the adjusted first token sequence conforms to the context length corresponding to the first large language model. After obtaining the second token sequence, the sequence length of the second token sequence can be adjusted according to the context length corresponding to the second large language model, so that the sequence length of the adjusted second token sequence conforms to the context length corresponding to the second large language model.
[0114] In this optional method, after the same text is processed by different tokenizers, token sequences of different lengths may be obtained. Different large language models can consider different maximum numbers of tokens when processing the input token sequences. The context length corresponding to the first large language model reflects the maximum number of tokens that the first large language model can consider when processing the input token sequence. The sequence length of the first token sequence can be padded or truncated according to the context length corresponding to the first large language model, so that the sequence length of the adjusted first token sequence conforms to the context length corresponding to the first large language model; the context length corresponding to the second large language model reflects the maximum number of tokens that the second large language model can consider when processing the input token sequence. The sequence length of the second token sequence can be padded or truncated according to the context length corresponding to the second large language model, so that the sequence length of the adjusted second token sequence conforms to the context length corresponding to the second large language model.
[0115] Among them, the length of the token sequence can be padded by adding meaningless tokens to the token sequence.
[0116] Optionally, according to the character position information corresponding to multiple tokens in the first token sequence and the character position information corresponding to multiple tokens in the second token sequence, a cross-attention mask is generated, including: determining a first valid token according to the attention mask of the first large language model, where the first valid token is a valid token in the first token sequence; determining a second valid token according to the attention mask of the second large language model, where the second valid token is a valid token in the second token sequence; searching for the character position information of the first valid token in the character position information corresponding to multiple tokens in the first token sequence; searching for the character position information of the second valid token in the character position information corresponding to multiple tokens in the second token sequence; comparing the character position information of the first valid token with the character position information of the second valid token to obtain a comparison result; generating a cross-attention mask according to the comparison result. Thus, by screening out the comparison of the character position information between valid tokens, the accuracy of the cross-attention mask is improved, so that the valid tokens and the context information of the valid tokens are focused on in the cross-attention network, and other information is ignored.
[0117] Among them, the attention mask of the first large language model is determined according to the tokenizer of the first large language model, and the attention mask of the second large language model is determined according to the tokenizer of the second large language model. The tokenizer of the large language model reflects the tokenization process of the target text and can reflect which of the obtained token sequences are valid tokens and which are invalid tokens.
[0118] Among them, valid tokens refer to tokens that have meaning in the large language model, rather than padding or meaningless tokens. For example, a space can be determined as an invalid token, and a token of the word type is a valid token; another example is that a meaningless token added during the process of padding the sequence length is an invalid token.
[0119] In this optional method, both the first token sequence and the second token sequence may contain invalid tokens. In order to focus on valid tokens during the feature fusion process, the valid tokens in the first token sequence (i.e., the first valid tokens) and the valid tokens in the second token sequence (i.e., the second valid tokens) can be identified first. This identification process needs to rely on the attention mask of the first large language model and the attention mask of the second large language model respectively. Specifically, a token with a corresponding mask value of 1 or TRUE in the attention mask of the first large language model can be determined as the first valid token; a token with a corresponding mask value of 1 or TRUE in the attention mask of the second large language model can be determined as the second valid token. Compare the character position information of the first valid token with the character position information of the second valid token to obtain a comparison result, which reflects the character position relationship between the first valid token and the second valid token; according to the comparison result, that is, according to the character position relationship between the first valid token and the second valid token, generate a cross-attention mask so that valid tokens can focus on their own context information in the cross-attention.
[0120] Further, the cross-attention mask includes attention masks in two directions. One is the attention mask pointing from the first large language model to the second large language model. For the sake of distinction, this attention mask is called the first attention mask. The other is the attention mask pointing from the second large language model to the first large language model. For the sake of distinction, this attention mask is called the second attention mask. According to the comparison result of the character position information of the first valid token and the character position information of the second valid token, a cross-attention mask is generated, including: if the character end position in the second valid token in the comparison result is before the character end position in the first valid token, then determine that the mask value corresponding to the second valid token in the first attention mask is 1 or true to focus on the second valid token. Otherwise, it can be determined that the mask value corresponding to the second valid token in the first attention mask is 0 or false (False), indicating not to focus on the second valid token; if the character end position in the first valid token in the comparison result is before the character end position in the second valid token, then determine that the mask value corresponding to the first valid token in the second attention mask is 1 or true to focus on the first valid token. Otherwise, it can be determined that the mask value corresponding to the first valid token in the second attention mask is 0 or false, indicating not to focus on the first valid token. Finally, the constructed mask matrix is applied to the self-attention calculation. In this way, when calculating the attention weight of each token in the cross-attention mechanism, it will be restricted by the cross-attention mask and can only focus on other tokens before each token, rather than other tokens after it.
[0121] In the combined model, after fusing the feature representation output by the i-th selected network layer and the feature representation output by the j-th selected network layer, the fused feature of the feature representation output by the i-th selected network layer and the feature representation output by the j-th selected network layer output by the combined model is obtained. After that, according to this fused feature, the input data of the next network layer of the j-th selected network layer can be determined.
[0122] Next, an exemplary embodiment is provided for determining the input data of the next network layer of the j-th selected network layer according to the fused feature.
[0123] In some embodiments, the fused feature is added to the feature representation output by the j-th selected network layer in a residual manner to obtain the input data of the next network layer of the j-th selected network layer. Thus, the fused feature is supplemented to the second large language model in a residual manner, improving the text processing ability of the second large language model in the target professional field.
[0124] In one example, the calculation formula for the input data of the next network layer of the j-th selected network layer is as follows:
[0125]
[0126] Among them, represents the input data of the next network layer of the j-th selected network layer, represents the output data of the j-th selected network layer, f cross (f proj (H Ai ), H Bj ) represents the fused feature obtained by fusing the feature representation output by the i-th selected network layer and the feature representation output by the j-th selected network layer through the cross-attention network.
[0127] As an example, please refer to Figure 4 , which provides a structural example diagram of the text processing model. As Figure 4 shown, m A represents the first large language model, m B represents the second large language model, l Ai represents the i-th selected network layer in the first large language model, l Bj represents the j-th selected network layer in the second large language model. There is a combined model between m A and m B , and the combined model includes a projection layer and a cross-attention network. In the combined model, the feature representation output by l Ai is input into the projection layer, and through the linear transformation of the projection layer, the corresponding projection representation is obtained; based on the projection representation, the key vector W K and the value vector W V of the attention head in the cross-attention network are determined. Based on the feature representation output by l Bj , the query vector W Q of the attention head in the cross-attention network is determined. According to the key vector, the value vector, and the query vector, the data output by the attention head is obtained, and finally the output data of the cross-attention network is obtained; the output data of the cross-attention network is added to the feature representation output by l Bj in a residual manner to obtain the output data of the l Bj+1 layer (i.e., the (j + 1)-th network layer in the second large language model, and this network layer is not necessarily a selected network layer).
[0128] In some embodiments, the text processing model can perform text processing in an autoregressive decoding manner. Specifically, the target text is input into the text processing model, and the target text is processed in the text processing model to obtain the output data of the first time step; the output data of the first time step is added to the target text, and the target text is input into the text processing model, and the target text is processed in the text processing model to obtain the output data of the second time step. In this way, the output data of multiple time steps is obtained. According to the output data of multiple time steps, the processing result is obtained.
[0129] For example, output data y is obtained at t time steps t , and y t is added to the input data x at the t-th time step t to obtain the input data x at the (t + 1)-th time step t+1 ,
[0130] Please refer to Figure 5 , in another exemplary embodiment, a model training method is provided, and the model training method may include:
[0131] S501, obtaining training texts in a target professional field.
[0132] In this embodiment, the training texts can be obtained from a training database corresponding to the target professional field. Among them, the training database contains pre-collected training data.
[0133] In one example, the target professional field is the medical field, and the training texts are relevant texts for training in the medical field.
[0134] S502, processing the training texts through a text processing model to obtain a processing result. The text processing model includes a first large language model, a second large language model, and a combination model. The first large language model is applicable to the target professional field, the second large language model is applicable to the general field, and the combination model is used for data fusion between the first large language model and the second large language model.
[0135] Among them, the text processing model can refer to the description of the foregoing embodiment and will not be elaborated here.
[0136] In this embodiment, the training texts are respectively input into the first large language model and the second large language model included in the text processing model. Feature extraction is performed on the training texts through the network layer in the first large language model to obtain the feature representation output by this network layer. Feature extraction is performed on the training texts through the network layer in the second large language model to obtain the feature representation output by this network layer. The feature representation output by the network layer in the first large language model and the feature representation output by the network layer in the second large language model can be input into the combination model. In the combination model, fusion processing is performed on the feature representation output by the network layer in the first large language model and the feature representation output by the network layer in the second large language model to obtain the fusion feature output by the combination model. The input data for the next network layer of the second large language model can be determined according to the fusion feature output by the combination model, and this input data is input into this next network layer to continue feature processing in this next network layer. In this way, the above process can be executed once or multiple times, and finally the processing result corresponding to the training texts is obtained.
[0137] Among them, the data fusion process of the combined model can refer to the description of the foregoing embodiments and will not be elaborated herein.
[0138] S503. According to the processing result, adjust the parameters of the combined model to obtain the text processing model after the first training.
[0139] In this embodiment, the training data may include the label data corresponding to the training text. The training error can be determined by comparing the label data with the processing result. According to the training error, the parameters of the combined model are adjusted to obtain the text processing model after the first training. In this way, the text processing model can be trained multiple times through multiple training texts. In each training, the parameters of the combined model are adjusted to obtain the text processing model after multiple trainings.
[0140] In the embodiments of the present application, a text processing model including a first large language model, a second large language model, and a combined model is designed. By introducing a small number of trainable parameters into the intermediate layer representations of the two large language models, the text processing model obtained with a small training cost can not only retain the model capabilities of the first large language model and the second large language model participating in the combination, but also perform better when facing text processing tasks in the target professional field.
[0141] In some embodiments, the combined model includes a projection layer. Adjusting the parameters of the combined model according to the processing result to obtain the text processing model after the first training includes: adjusting the parameters of the projection function in the projection layer according to the processing result to obtain the text processing model after the first training. Thus, the parameters of the projection function are learned during the training process to improve the accuracy of dimension transformation of the feature representation output by the network layer in the first large language model through the projection function.
[0142] In some embodiments, the combined model includes a cross-attention network. The parameters of the cross-attention network include the weight matrices corresponding to the cross-attention heads (including the first weight matrix, the second weight matrix, and the third weight matrix in the foregoing embodiments). Adjusting the parameters of the combined model according to the processing result to obtain the text processing model after the first training includes: adjusting the weight matrices corresponding to the cross-attention heads according to the processing result to obtain the text processing model after the first training. Thus, the weight matrices corresponding to the cross-attention heads are learned during the training process to improve the fusion effect of the cross-attention network on the feature representations from different large language models.
[0143] Optionally, the weight parameters of the linear layer (for linearly processing the concatenation result of the output data of H cross-attention heads) are included in the cross-attention network. Adjusting the parameters of the combined model further includes: adjusting the weight parameters of the linear layer in the cross-attention network.
[0144] Exemplary device
[0145] Correspondingly, an embodiment of the present application further provides a text processing device.
[0146] Please refer to Figure 6 , in an exemplary embodiment, a text processing device 600 is provided. The text processing device 600 includes: an acquisition unit 601 for acquiring a target text to be processed in a target professional field; a processing unit 602 for processing the target text through a text processing model to obtain a processing result; the text processing model includes a first large language model, a second large language model, and a combination model. The first large language model is applicable to the target professional field, the second large language model is applicable to the general field, and the combination model is used for data fusion between the first large language model and the second large language model.
[0147] In some embodiments, the combination model is used for feature fusion between the feature representations output by N selected network layers in the first large language model and the feature representations output by M selected network layers in the second large language model, where N is greater than or equal to 1 and M is greater than or equal to 1.
[0148] In some embodiments, the process of performing fusion processing on the feature representation output by the i-th selected network layer among the N selected network layers and the feature representation output by the j-th selected network layer among the M selected network layers through the combination model includes: in the combination model, transforming the feature representation output by the i-th selected network layer and the feature representation output by the j-th selected network layer into feature representations with the same dimension; after the dimension transformation, performing fusion processing on the feature representation output by the i-th selected network layer and the feature representation output by the j-th selected network layer to obtain a fused feature; where the value range of i is from 1 to N, and the value range of j is from 1 to M.
[0149] In some embodiments, the combination model includes a projection layer; transforming the feature representation output by the i-th selected network layer and the feature representation output by the j-th selected network layer into feature representations with the same dimension includes: inputting the feature representation output by the i-th selected network layer into the projection layer; in the projection layer, performing a linear transformation on the feature representation output by the i-th selected network layer to obtain a projection representation, and the dimension of the projection representation is the same as the dimension of the feature representation output by the j-th selected network layer.
[0150] In some embodiments, the combination model further includes a cross-attention network; performing fusion processing on the feature representation output by the i-th selected network layer and the feature representation output by the j-th selected network layer to obtain a fused feature includes: inputting the projection representation and the feature representation output by the j-th selected network layer into the cross-attention network; in the cross-attention network, performing fusion processing on the projection representation and the feature representation output by the j-th selected network layer to obtain a fused feature.
[0151] In some embodiments, the cross-attention network includes H cross-attention heads, where H is greater than or equal to 1; in the cross-attention network, the projection representation and the feature representation output by the j-th selected network layer are fused to obtain a fused feature, including: determining the key vector and the value vector of the k-th cross-attention head in the cross-attention layer according to the projection representation, where the value range of k is from 1 to H; determining the query vector of the k-th cross-attention head according to the feature representation output by the j-th selected network layer; performing cross-attention operations on the key vector, the value vector, and the query vector through the cross-attention mechanism to obtain the output data of the k-th cross-attention head; and determining the fused feature according to the output data of the k-th cross-attention head.
[0152] In some embodiments, the cross-attention mechanism includes a cross-attention mask; the cross-attention mask is determined through the following process: determining the character position information corresponding to each of the multiple tokens in the first token sequence according to the tokenizer in the first large language model, where the first token sequence is obtained by tokenizing the target text through the tokenizer in the first large language model; determining the character position information corresponding to each of the multiple tokens in the second token sequence according to the tokenizer in the second large language model, where the second token sequence is obtained by tokenizing the target text through the tokenizer in the second large language model; and generating the cross-attention mask according to the character position information corresponding to each of the multiple tokens in the first token sequence and the character position information corresponding to each of the multiple tokens in the second token sequence.
[0153] In some embodiments, generating the cross-attention mask according to the character position information corresponding to each of the multiple tokens in the first token sequence and the character position information corresponding to each of the multiple tokens in the second token sequence includes: determining a first valid token according to the attention mask of the first large language model, where the first valid token is a valid token in the first token sequence; determining a second valid token according to the attention mask of the second large language model, where the second valid token is a valid token in the second token sequence; searching for the character position information of the first valid token among the character position information corresponding to each of the multiple tokens in the first token sequence; searching for the character position information of the second valid token among the character position information corresponding to each of the multiple tokens in the second token sequence; comparing the character position information of the first valid token with the character position information of the second valid token to obtain a comparison result; and generating the cross-attention mask according to the comparison result.
[0154] In some embodiments, determining the key vector and the value vector of the k-th cross-attention head in the cross-attention layer according to the projection representation includes: multiplying the first weight matrix corresponding to the k-th cross-attention head by the projection representation to obtain the key vector; multiplying the second weight matrix corresponding to the k-th cross-attention head by the projection representation to obtain the value vector; determining the query vector of the k-th cross-attention head according to the feature representation output by the j-th selected network layer, including: multiplying the third weight matrix corresponding to the k-th cross-attention head by the feature representation output by the j-th selected network layer to obtain the query vector.
[0155] The text processing apparatus 600 provided in this embodiment belongs to the same inventive concept as the text processing method provided in the foregoing embodiments of the present application, can execute the text processing method provided in any of the foregoing embodiments of the present application, and has the corresponding functional modules and beneficial effects for executing the text processing method. For technical details not described in detail in this embodiment, reference may be made to the specific processing content of the text processing method provided in the foregoing embodiments of the present application, which will not be elaborated herein.
[0156] Please refer to Figure 7 , in an exemplary embodiment, a model training apparatus 700 is provided. The model training apparatus 700 includes: an acquisition unit 701, configured to acquire training texts in a target professional field; a processing unit 702, configured to process the training texts through a text processing model to obtain a processing result, where the text processing model includes a first large language model, a second large language model, and a combination model, the first large language model is applicable to the target professional field, the second large language model is applicable to the general field, and the combination model is used for data fusion between the first large language model and the second large language model; an adjustment unit 703, configured to adjust the parameters of the combination model according to the processing result to obtain a text processing model after one training.
[0157] In some embodiments, the combination model includes an attention network, and the parameters of the attention network include the weight matrices corresponding to the cross-attention heads. The adjustment unit 703 is specifically configured to: adjust the weight matrices corresponding to the cross-attention heads according to the processing result to obtain a text processing model after one training.
[0158] The model training apparatus 700 provided in this embodiment belongs to the same inventive concept as the model training method provided in the foregoing embodiments of the present application, can execute the model training method provided in any of the foregoing embodiments of the present application, and has the corresponding functional modules and beneficial effects for executing the model training method. For technical details not described in detail in this embodiment, reference may be made to the specific processing content of the model training method provided in the foregoing embodiments of the present application, which will not be elaborated herein.
[0159] The functions implemented by each unit in the above devices (text processing device, model training device) can be implemented by the same or different processors, and the embodiments of the present application do not make any limitations.
[0160] It should be understood that the units in the above devices can be implemented in the form of a processor calling software. For example, the device includes a processor, the processor is connected to a memory, and instructions are stored in the memory. The processor calls the instructions stored in the memory to implement any of the above methods or the functions of each unit of the device. The processor can be a general-purpose processor, such as a CPU or a microprocessor, etc., and the memory can be a memory inside the device or a memory outside the device. Alternatively, the units in the device can be implemented in the form of a hardware circuit, and the functions of some or all of the units can be implemented through the design of the hardware circuit. The hardware circuit can be understood as one or more processors. For example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above units are implemented through the design of the logical relationship of the components in the circuit. Another example is that in another implementation, the hardware circuit can be implemented by a PLD. Taking an FPGA as an example, it can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured through a configuration file, so as to implement the functions of some or all of the above units. All units of the above devices can be all implemented in the form of a processor calling software, or all implemented in the form of a hardware circuit, or some implemented in the form of a processor calling software, and the remaining part implemented in the form of a hardware circuit.
[0161] In the embodiments of the present application, a processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and running capabilities, such as a CPU, a microprocessor, a GPU, or a DSP, etc. In another implementation, the processor can implement certain functions through the logical relationship of a hardware circuit, and the logical relationship of the hardware circuit is fixed or can be reconstructed. For example, the processor is a hardware circuit implemented by an ASIC or a PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the configuration of the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as a type of ASIC, such as an NPU, a TPU, a DPU, etc.
[0162] It can be seen that each unit in the above devices can be one or more processors (or processing circuits) configured to implement the above methods. For example: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.
[0163] In addition, all or part of the units in the above device can be integrated together or can be implemented independently. In one implementation, these units are integrated together and implemented in the form of an SOC. The SOC may include at least one processor for implementing any of the above methods or implementing the functions of the units of the device. The types of the at least one processor can be different. For example, it includes a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.
[0164] Exemplary electronic device
[0165] Another embodiment of the present application also proposes an electronic device. Refer to Figure 8 As shown, the electronic device may include: a memory 800 and a processor 810; wherein, the memory 800 is connected to the processor 810 for storing programs; the processor 810 is configured to implement the text processing method or the model training method disclosed in any of the above embodiments by running the programs stored in the memory 800.
[0166] Specifically, the above electronic device may further include: a bus, a communication interface 820, an input device 830, and an output device 840.
[0167] The processor 810, the memory 800, the communication interface 820, the input device 830, and the output device 840 are interconnected through the bus. Among them:
[0168] The bus may include a path for transmitting information between various components of the computer system.
[0169] The processor 810 may be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the solution of the present application. It may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0170] The processor 810 may include a main processor and may also include a baseband chip, a modem, etc.
[0171] The program for implementing the technical solution of this application is stored in the memory 800, and the operating system and other key services may also be stored. Specifically, the program may include program code, and the program code includes computer operation instructions. More specifically, the memory 800 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk memory, a flash memory, and so on.
[0172] The input device 830 may include devices for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor, etc.
[0173] The output device 840 may include devices for allowing information to be output to a user, such as a display screen, a printer, a speaker, etc.
[0174] The communication interface 820 may include devices of any transceiver type for communicating with other devices or communication networks, such as an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.
[0175] The processor 810 executes the program stored in the memory 800 and calls other devices, and can be used to implement each step of any one of the text processing methods or any one of the model training methods provided in the above embodiments of this application.
[0176] An embodiment of this application also proposes a chip, which includes a processor and a data interface. The processor reads and runs the program stored on the memory through the data interface to execute any one of the text processing methods or any one of the model training methods provided in the above embodiments. The specific processing process and its beneficial effects can be referred to the embodiment introduction of the above text processing method or model training method.
[0177] Exemplary computer program product and storage medium
[0178] In addition to the above methods and devices, an embodiment of this application may also be a computer program product, which includes computer program instructions. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the text processing method or the model training method according to various embodiments of this application described in any of the above embodiments of this specification.
[0179] The computer program product can be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The programming code can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0180] In addition, an embodiment of the present application can also be a storage medium on which a computer program is stored, and the computer program is executed by a processor to perform the steps in the text processing method or the model training method according to various embodiments of the present application described in any of the above embodiments of this specification.
[0181] For the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0182] It should be noted that the embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0183] The steps in the methods of the embodiments of the present application can be adjusted, combined, and deleted according to actual needs, and the technical features recorded in each embodiment can be replaced or combined.
[0184] The modules and sub-modules in the devices and terminals in the embodiments of the present application can be combined, divided, and deleted according to actual needs.
[0185] In several embodiments provided by the present application, it should be understood that the disclosed terminals, devices, and methods can be implemented in other ways. For example, the terminal embodiments described above are merely illustrative. For example, the division of modules or sub-modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple sub-modules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be an indirect coupling or communication connection through some interfaces, devices, or modules, and can be in electrical, mechanical, or other forms.
[0186] The modules or sub-modules described as separate components may or may not be physically separated. The components as modules or sub-modules may or may not be physical modules or sub-modules, that is, they can be located in one place, or can be distributed to multiple network modules or sub-modules. Some or all of the modules or sub-modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0187] In addition, in each embodiment of the present application, the functional modules or sub-modules can be integrated in a processing module, or each module or sub-module can exist physically alone, or two or more modules or sub-modules can be integrated in one module. The above-mentioned integrated modules or sub-modules can be implemented in the form of hardware, or in the form of software functional modules or sub-modules.
[0188] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0189] The steps of the methods or algorithms described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software units executed by a processor, or a combination of the two. The software units can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0190] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.
[0191] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A text processing method, characterized in that: include: Obtain the target text to be processed in the target professional field; Processing the target text through a text processing model to obtain a processing result; The text processing model includes a first language model, a second language model and a combination model. The first language model is applicable to the target professional field, the second language model is applicable to the general field, and the combination model is used for data fusion between the first language model and the second language model.
2. The text processing method according to claim 1, characterized in that: The combined model is used for feature fusion between feature representations of N selected network layer outputs in the first large language model and feature representations of M selected network layer outputs in the second large language model, where N is greater than or equal to 1 and M is greater than or equal to 1.
3. The text processing method according to claim 2, characterized in that: The process of fusing the feature representation output by the i-th selected network layer in the N selected network layers with the feature representation output by the j-th selected network layer in the M selected network layers through the combined model includes: In the combined model, the feature representation output by the i-th selected network layer and the feature representation output by the j-th selected network layer are transformed into feature representations with the same dimension; After the dimension transformation, the feature representation output by the i-th selected network layer and the feature representation output by the j-th selected network layer are fused to obtain a fused feature; Among them, the value range of i is 1 to N, and the value range of j is 1 to M.
4. The text processing method according to claim 3, characterized in that: The combined model includes a projection layer; The step of transforming the feature representation output by the i-th selected network layer and the feature representation output by the j-th selected network layer into feature representations with the same dimension includes: Inputting the feature representation output by the i-th selected network layer into the projection layer; In the projection layer, a linear transformation is performed on the feature representation output by the i-th selected network layer to obtain a projection representation, and the dimension of the projection representation is the same as the dimension of the feature representation output by the j-th selected network layer.
5. The text processing method according to claim 4, characterized in that: The combined model also includes a cross-attention network; The fusing the feature representation output by the i-th selected network layer and the feature representation output by the j-th selected network layer to obtain the fused feature includes: Input the projection representation and the feature representation output by the j-th selected network layer into the cross attention network; In the cross-attention network, the projection representation and the feature representation output by the j-th selected network layer are fused to obtain the fused feature.
6. The text processing method according to claim 5, characterized in that: The cross attention network includes H cross attention heads, where H is greater than or equal to 1; In the cross attention network, the projection representation and the feature representation output by the j-th selected network layer are fused to obtain the fused feature, including: Determine, according to the projection representation, a key vector of a k-th cross-attention head and a value vector of the k-th cross-attention head in the cross-attention layer, where k ranges from 1 to H; Determining a query vector for the kth cross attention head according to the feature representation output by the jth selected network layer; Performing a cross attention operation on the key vector, the value vector, and the query vector through a cross attention mechanism to obtain output data of the kth cross attention head; The fusion feature is determined according to the output data.
7. The text processing method according to claim 6, characterized in that: The criss-cross attention mechanism includes a criss-cross attention mask; The criss-cross attention mask is determined by the following process: Determining, according to the word segmenter in the first large language model, character position information corresponding to a plurality of word units in a first word unit sequence, wherein the first word unit sequence is obtained by performing word segmentation processing on the target text by the word segmenter in the first large language model; Determining, according to the word segmenter in the second largest language model, character position information corresponding to each of a plurality of word units in a second word unit sequence, wherein the second word unit sequence is obtained by performing word segmentation processing on the target text by the word segmenter in the second largest language model; The cross-attention mask is generated according to the character position information respectively corresponding to the multiple words in the first word-gram sequence and the character position information respectively corresponding to the multiple words in the second word-gram sequence.
8. The text processing method according to claim 7, characterized in that: The step of generating the cross attention mask according to the character position information respectively corresponding to the multiple word-grams in the first word-gram sequence and the character position information respectively corresponding to the multiple word-grams in the second word-gram sequence comprises: Determine a first valid word-gram according to the attention mask of the first large language model, where the first valid word-gram is a valid word-gram in the first word-gram sequence; Determine a second valid word-gram according to the attention mask of the second largest language model, where the second valid word-gram is a valid word-gram in the second word-gram sequence; Searching for the character position information of the first valid word-unit in the character position information respectively corresponding to the plurality of word-units in the first word-unit sequence; Searching for the character position information of the second valid word-unit in the character position information respectively corresponding to the plurality of word-units in the second word-unit sequence; Comparing the character position information of the first valid word unit with the character position information of the second valid word unit to obtain a comparison result; Based on the comparison result, the cross attention mask is generated.
9. The text processing method according to any one of claims 6 to 8, characterized in that: Determining, according to the projection representation, a key vector of a k-th cross-attention head in the cross-attention layer and a value vector of the k-th cross-attention head, comprises: Multiplying the first weight matrix corresponding to the k-th cross attention head by the projection representation to obtain the key vector; Multiplying the second weight matrix corresponding to the k-th cross attention head by the projection representation to obtain the value vector; The step of determining the query vector of the kth cross attention head according to the feature representation output by the jth selected network layer comprises: The third weight matrix corresponding to the k-th cross attention head is multiplied by the feature representation output by the j-th selected network layer to obtain the query vector.
10. A model training method, characterized in that: The invention is characterized by comprising: Obtain training texts in the target professional field; Processing the training text by a text processing model to obtain a processing result, wherein the text processing model includes a first language model, a second language model, and a combination model, wherein the first language model is applicable to the target professional field, the second language model is applicable to a general field, and the combination model is used for data fusion between the first language model and the second language model; According to the processing result, the parameters of the combined model are adjusted to obtain the text processing model after one training.
11. The model training method according to claim 10, characterized in that: The combined model includes a cross-attention network, and the parameters of the cross-attention network include a weight matrix corresponding to the cross-attention head. The combined model is adjusted according to the processing result to obtain the text processing model after one training, including: According to the processing results, the weight matrix corresponding to the cross attention head is adjusted to obtain the text processing model after one training.
12. A text processing device, characterized in that: include: An acquisition unit, used for acquiring a target text to be processed in a target professional field; A processing unit, used to process the target text through a text processing model to obtain a processing result; The text processing model includes a first language model, a second language model and a combination model. The first language model is applicable to the target professional field, the second language model is applicable to the general field, and the combination model is used for data fusion between the first language model and the second language model.
13. A model training device, characterized in that: include: An acquisition unit, used to acquire training texts in a target professional field; a processing unit, configured to process the training text through a text processing model to obtain a processing result, wherein the text processing model comprises a first language model, a second language model and a combination model, wherein the first language model is applicable to the target professional field, the second language model is applicable to a general field, and the combination model is used for data fusion between the first language model and the second language model; An adjustment unit is used to adjust parameters of the combined model according to the processing result to obtain the text processing model after one training.
14. An electronic device, characterized in that: including memory and processor; The memory is connected to the processor and is used to store programs; The processor is used to implement the text processing method described in any one of claims 1 to 9 or the model training method described in any one of claims 10 to 11 by running the program in the memory.
15. A computer program product, characterized in that It includes a computer program, which, when executed by a processor, implements the text processing method as described in any one of claims 1 to 9 or the model training method as described in any one of claims 10 to 11.