Domain knowledge injection method, device, equipment and medium for language embedding model
By copying and inserting the target knowledge layer and using preset domain knowledge data for training and comparison learning, the knowledge forgetting problem of the language embedding model in the professional field corpus is solved, and efficient knowledge transfer and retrieval ability improvement is achieved.
Patent Information
- Application Number
- CN202510005352.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-03
AI Technical Summary
In the prior art, language embedding models are prone to catastrophic forgetting of knowledge when dealing with professional field corpus, and low-order adaptation (LoRA) methods have problems with excessive freedom of parameter selection and parameter adjustment.
By obtaining the target language model with the same architecture as the embedding model, selecting and copying the target knowledge layer, inserting it into the original transformer layer, pre-training and small sample comparison learning using preset domain knowledge data, and generating a new embedding model.
Knowledge migration is realized, catastrophic forgetting is avoided, and only a small number of domain knowledge samples are needed for parameter training, with very few training parameters, low hardware memory requirements, improved retrieval capabilities on professional data, and strong operability.
Smart Images

Figure CN119398109B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of embedding model technology, and in particular to a method, device, equipment and medium for injecting domain knowledge based on a language embedding model. Background Art
[0002] Generative large language models (LLMs) have been widely used and developed in various industries. Many large models have been released at home and abroad, and different technical implementations and attempts have continued to emerge. With the implementation of LLM applications, whether in the "Agent" mode or the Retrieval Enhanced Generation (RAG) mode, the language embedding model has become an indispensable and important aspect. In the construction of LLM-Agent scenarios, the embedding model serves as a bridge connecting LLM and external data, and can complete the extraction of key information, memory recall and other functions; in the RAG mode, the embedding model plays an important role in retrieving massive external data and optimizing question and answer information (such as corpus inversion). In addition, as an important tool for extending LLM to handle long texts and semantic dependencies, the embedding model is indispensable in existing implementation applications.
[0003] Due to the need to process knowledge in different fields when building applications, the performance of directly using the published embedding model in vertical fields is often unsatisfactory. This is mainly because the professional semantic expressions, professional vocabulary, and the association of sentences in specific professional contexts contained in different professional fields are different from general corpora. Therefore, when processing professional domain corpora, the embedding model often needs to be injected with domain knowledge first.
[0004] In the related technology, a full process from pre-training to contrastive learning training is proposed, and all the weights of the model are trained and updated, or new layers are added to the output side of the model for training. However, it must be pointed out that if only professional corpus is used to train and update all weights, it will often directly destroy the knowledge learned in the pre-training stage. This phenomenon is called "catastrophic forgetting". If a new network layer is added, there are often certain requirements for the amount of data in the professional field. If the amount of data is small, not only will it be difficult for the model to learn professional knowledge, but also knowledge forgetting will occur due to insufficient training of the new layer. In addition, the fine-tuning method using low-order adaptation (LoRA) can update fewer parameters and requires relatively less data, but it is still prone to catastrophic forgetting difficulties, and the use of LoRA has the problem of excessive freedom in parameter selection and parameter adjustment. Summary of the invention
[0005] The embodiments of the present application provide a method, apparatus, device and medium for injecting domain knowledge based on a language embedding model, aiming to solve technical problems such as catastrophic forgetting of knowledge in language embedding models existing in related technologies.
[0006] In a first aspect, an embodiment of the present application provides a method for injecting domain knowledge based on a language embedding model, comprising:
[0007] Obtain a target language model having the same architecture as the embedding model, wherein the target language model includes a plurality of first original transformer layers;
[0008] Selecting and copying a first target knowledge layer from the plurality of first original transformer layers, and inserting the copied first target knowledge layer into the plurality of first original transformer layers of the target language model to generate a new language model;
[0009] Pre-training the first target knowledge layer of the new language model using preset domain knowledge data to determine the first target domain knowledge layer;
[0010] Inserting the first target domain knowledge layer into a corresponding position of the embedding model to obtain a new first embedding model;
[0011] The preset domain knowledge data is used to perform small sample comparative learning on the new first embedding model to obtain a target embedding model.
[0012] In one embodiment, optionally, before obtaining a target language model having the same architecture as the embedding model, the method further includes:
[0013] determining whether there is a target language model having the same architecture as the embedding model;
[0014] When it is determined that there is a target language model having the same architecture as the embedding model, acquiring the target language model;
[0015] When it is determined that there is no target language model having the same architecture as the embedding model, directly selecting and copying a second target knowledge layer from the multiple second original transformer layers of the embedding model, and inserting the copied second target knowledge layer into the multiple second original transformer layers of the embedding model to generate a new second embedding model;
[0016] Pre-training the second target knowledge layer of the new second embedding model using the preset domain knowledge data to determine the second target domain knowledge layer;
[0017] The preset domain knowledge data is used to perform small sample comparative learning on the second target domain knowledge layer of the new second embedding model to obtain a target embedding model.
[0018] In one embodiment, optionally, the method further comprises:
[0019] receiving a search instruction containing preset domain knowledge;
[0020] According to the search instruction, the target embedding model is used to search for corresponding preset domain knowledge to obtain corresponding search results.
[0021] In one embodiment, optionally, selecting and copying a first target knowledge layer from the plurality of first original transformer layers includes:
[0022] According to a preset knowledge layer selection rule, a first target knowledge layer is selected and copied from the plurality of first original transformer layers, wherein the preset knowledge layer selection rule includes a selection position and a selection quantity of the target knowledge layer.
[0023] In one embodiment, optionally, inserting the copied first target knowledge layer into a plurality of first original transformer layers of the target language model to generate a new language model includes:
[0024] Insert the copied first target knowledge layer into a preset position of multiple first original transformer layers of the target language model to generate a new language model, wherein the preset position includes a position adjacent to and subsequent to the selected position of the first target knowledge layer.
[0025] In one embodiment, optionally, pre-training the first target knowledge layer of the new language model using preset domain knowledge data to determine the first target domain knowledge layer includes:
[0026] Freeze the parameter values of the multiple first original transformer layers, and pre-train the first target knowledge layer of the new language model using preset domain knowledge data to determine a first target parameter value corresponding to the first target knowledge layer;
[0027] The using the preset domain knowledge data to perform small sample comparative learning on the new first embedding model to obtain a target embedding model includes:
[0028] Freeze the parameter values of other layers of the new first embedding model, use the preset domain knowledge data to perform small sample comparative learning on the inserted first target domain knowledge layer to determine the second target parameter value corresponding to the first target knowledge layer, and obtain the target embedding model.
[0029] In one embodiment, optionally, the method further comprises:
[0030] Acquire different numbers of continuous candidate knowledge layers from different positions of multiple first original transformer layers of the target language model respectively;
[0031] Respectively copying the candidate knowledge layers, and inserting the copied candidate knowledge layers into preset positions in a plurality of first original transformer layers of the target language model to generate corresponding candidate language models;
[0032] Freeze the parameter values of the multiple first original transformer layers, and pre-train the candidate knowledge layer of each candidate language model using preset domain knowledge data to determine the candidate parameter values corresponding to the candidate knowledge layer, and obtain a trained candidate language model;
[0033] Inserting the candidate knowledge layer into the corresponding position of the embedding model, and performing small sample comparative learning using the preset domain knowledge data to obtain a candidate embedding model;
[0034] When using the candidate embedding model to retrieve the preset domain knowledge, calculate the retrieval accuracy corresponding to each candidate embedding model;
[0035] Based on the candidate embedding model with the highest retrieval accuracy, the preset knowledge layer selection rules are determined.
[0036] In one embodiment, optionally, the method further comprises:
[0037] Receive input knowledge layer selection rule setting command;
[0038] The preset knowledge layer selection rule is set according to the knowledge layer selection rule setting command.
[0039] In one embodiment, optionally, the embedding model includes a Decoder-Only architecture embedding model and an Encoder architecture embedding model.
[0040] In a second aspect, an embodiment of the present application provides a domain knowledge injection device based on a language embedding model, comprising:
[0041] An acquisition module, configured to acquire a target language model having the same architecture as the embedding model, wherein the target language model includes a plurality of first original transformer layers;
[0042] A generating module, configured to select and copy a first target knowledge layer from the plurality of first original transformer layers, and insert the copied first target knowledge layer into the plurality of first original transformer layers of the target language model to generate a new language model;
[0043] A training module, used for pre-training the first target knowledge layer of the new language model using preset domain knowledge data to determine the first target domain knowledge layer;
[0044] An insertion module, configured to insert the first target domain knowledge layer into a corresponding position of the embedding model to obtain a new first embedding model;
[0045] A contrastive learning module is used to perform small sample contrastive learning on the new first embedding model using the preset domain knowledge data to obtain a target embedding model.
[0046] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned domain knowledge injection method based on the language embedding model when executing the computer program.
[0047] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned domain knowledge injection method based on the language embedding model are implemented.
[0048] In the above scheme implemented by the domain knowledge injection method, device, equipment and medium based on the language embedding model, a target language model with the same architecture as the embedding model is obtained, wherein the target language model includes multiple first original transformer layers; a first target knowledge layer is selected and copied from the multiple first original transformer layers, and the copied first target knowledge layer is inserted into the multiple first original transformer layers of the target language model to generate a new language model; the first target knowledge layer of the new language model is pre-trained using preset domain knowledge data to determine the first target domain knowledge layer; the first target domain knowledge layer is inserted into the corresponding position of the embedding model to obtain a new first embedding model; the new first embedding model is subjected to small sample comparative learning using the preset domain knowledge data to obtain a target embedding model. Through the above technical scheme of the present invention, the knowledge layer of the language model or the embedding model itself can be copied and trained to realize knowledge transfer, and full parameter training is not required, and only a small number of domain knowledge samples are used for parameter training of the copied knowledge layer, the training parameters are very few, the hardware memory requirement is low, the retrieval ability on professional data is well improved, the operability is strong, and catastrophic forgetting is avoided. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0050] Figure 1 A schematic flow chart of a method for injecting domain knowledge based on a language embedding model according to an embodiment of the present application is shown.
[0051] Figure 2 A schematic flow chart of a method for injecting domain knowledge based on a language embedding model according to another embodiment of the present application is shown.
[0052] Figure 3 A schematic flow chart of a method for injecting domain knowledge based on a language embedding model according to yet another embodiment of the present application is shown.
[0053] Figure 4 A schematic flow chart of a method for determining a preset knowledge layer selection rule according to an embodiment of the present application is shown.
[0054] Figure 5 A schematic flowchart of a method for injecting domain knowledge into a Decoder-Only architecture embedding model according to an embodiment of the present application is shown.
[0055] Figure 6 A schematic diagram of C-Eval scores of network layer weights of different depths under a Decoder-Only architecture embedding model according to an embodiment of the present application is shown.
[0056] Figure 7 A schematic flowchart of a method for injecting domain knowledge into an Encoder embedding model according to an embodiment of the present application is shown.
[0057] Figure 8 A schematic diagram of C-MTEB-T2Retrieval scores of network layer weights at different depths under the Encoder embedding model according to an embodiment of the present application is shown.
[0058] Fig. 9 A block diagram of a domain knowledge injection device based on a language embedding model according to an embodiment of the present application is shown.
[0059] Fig.10 A block diagram of a computer device according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0060] In order to better understand the technical solution of the present application, the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0061] It should be clear that the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.
[0062] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms "a", "said" and "the" used in the embodiments of the present application and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.
[0063] In conjunction with the accompanying drawings, some embodiments of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.
[0064] Terminology explanation:
[0065] "Catastrophic forgetting": For language models, multimodal models, or embedding models, due to some kind of training on newly provided corpus data or inappropriate parameter updates, they largely forget the knowledge they have learned in the pre-training stage during the reasoning process. This phenomenon is called "catastrophic forgetting".
[0066] Injection of domain knowledge: Generally speaking, after pre-training, instruction learning, and contrastive learning training on a wider range of corpora, the embedding model can distinguish the similarity of general corpus content, but it performs poorly on vertical field content with a certain degree of professionalism. At this time, it is necessary to consider pre-training or fine-tuning the model on professional field corpora to improve the performance of the model on the corpus in this field. This process is called "injection of domain knowledge" into the model.
[0067] See also Figure 1 , Figure 1 A schematic flow chart of a method for injecting domain knowledge based on a language embedding model according to an embodiment of the present application is shown.
[0068] like Figure 1 As shown, the process of the domain knowledge injection method based on the language embedding model according to an embodiment of the present application includes:
[0069] Step S101, obtaining a target language model having the same architecture as the embedding model, wherein the target language model includes a plurality of first original transformer layers;
[0070] In one embodiment, optionally, the embedding model includes a Decoder-Only architecture model and an Encoder architecture model.
[0071] Among them, the AI model that completes the conversion of corpus content into vector operations is called an embedding model.
[0072] When the embedding model is a Decoder-Only architecture model, its corresponding target language model is a Decoder-Only architecture large language model.
[0073] Step S102, selecting and copying a first target knowledge layer from the plurality of first original transformer layers, and inserting the copied first target knowledge layer into the plurality of first original transformer layers of the target language model to generate a new language model;
[0074] In one embodiment, optionally, selecting and copying a first target knowledge layer from the plurality of first original transformer layers includes:
[0075] According to a preset knowledge layer selection rule, a first target knowledge layer is selected and copied from the plurality of first original transformer layers, wherein the preset knowledge layer selection rule includes a selection position and a selection quantity of the target knowledge layer.
[0076] In this embodiment, for multiple first original transformer layers, there will be original transformer layer numbers, such as numbers 1-10, the selected position can be a specific number that can be selected, and the selected quantity can be the number of selected consecutive transformer layers.
[0077] In one embodiment, optionally, inserting the copied first target knowledge layer into a plurality of first original transformer layers of the target language model to generate a new language model includes:
[0078] Insert the copied first target knowledge layer into a preset position of multiple first original transformer layers of the target language model to generate a new language model, wherein the preset position includes a position adjacent to and subsequent to the selected position of the first target knowledge layer.
[0079] After copying several knowledge layers in the language model, they are inserted into a preset position, which can be inserted after the target knowledge layer in the original transformer layer, to form a new language model.
[0080] Step S103, training the first target knowledge layer of the new language model using preset domain knowledge data to determine the first target domain knowledge layer;
[0081] In one embodiment, optionally, step S103 includes:
[0082] Freeze the parameter values of the plurality of first original transformer layers, and train the first target knowledge layer of the new language model using preset domain knowledge data to determine a first target parameter value corresponding to the first target knowledge layer;
[0083] In this step, the domain knowledge data is directly used to pre-train the new language model. This is different from the conventional large language model pre-training in that only the weights of the copied and inserted knowledge layer are updated, while the parameters of other layers of the original language model are frozen. In this way, there is no need to train all parameters, and only a small number of domain knowledge samples are used to train the parameters of the copied knowledge layer. At the same time, the knowledge of the embedded model is transferred through the trained language model, which not only retains the accuracy of the language model in parsing semantic expressions, but also reduces training parameters, simplifies training steps, and improves the processing capacity of the model.
[0084] Step S104, inserting the first target domain knowledge layer into the corresponding position of the embedding model to obtain a new first embedding model;
[0085] Since the embedding model and the language model have the same architecture, the first target domain knowledge layer of the new language model is copied together with its trained parameters and then inserted into the corresponding position of the embedding model. For example, if it is located after the second transformer layer in the new language model, when it is inserted into the embedding model, it is also inserted after the second transformer layer of the embedding model.
[0086] Step S105: Use the preset domain knowledge data to perform small sample comparative learning on the new first embedding model to obtain a target embedding model.
[0087] Wherein, step S105 includes:
[0088] Freeze the parameter values of other layers of the new first embedding model, use the preset domain knowledge data to perform small sample comparative learning on the inserted first target domain knowledge layer to determine the second target parameter value corresponding to the first target knowledge layer, and obtain the target embedding model.
[0089] In this step, the parameter values of other layers of the new first embedding model can also be frozen, and the preset domain knowledge data can be used to perform small sample comparative learning on the inserted first target domain knowledge layer, so that only the parameter values of the first target domain knowledge layer are updated to obtain the target embedding model. In this way, not only can the accuracy of the obtained model in processing domain knowledge be guaranteed, but also full parameter training is not required, the retrieval ability on professional data is greatly improved, the operability is strong, and catastrophic forgetting is avoided.
[0090] like Figure 2 As shown, in one embodiment, optionally, before step S101, the method further includes:
[0091] Step S201, determining whether there is a target language model having the same architecture as the embedding model;
[0092] Step S202, when it is determined that there is a target language model having the same architecture as the embedding model, executing step S101;
[0093] Step S203, when it is determined that there is no target language model having the same architecture as the embedding model, directly select and copy the second target knowledge layer from the multiple second original transformer layers of the embedding model, and insert the copied second target knowledge layer into the multiple second original transformer layers of the embedding model to generate a new second embedding model;
[0094] Step S204, using the preset domain knowledge data to train the second target knowledge layer of the new second embedding model to determine the second target domain knowledge layer;
[0095] Step S205: Use the preset domain knowledge data to perform small sample comparative learning on the second target domain knowledge layer of the new second embedding model to obtain a target embedding model.
[0096] In this embodiment, if the embedding model does not have a target language model with the same architecture, the second target knowledge layer is directly selected and copied from the multiple second original transformer layers of the embedding model, and the copied second target knowledge layer is inserted into the multiple second original transformer layers of the embedding model to generate a new target embedding model. When inserting the second target knowledge layer, it can also be inserted into a preset position of the multiple second original transformer layers, wherein the preset position includes a position adjacent to and after the selected position of the second target knowledge layer. When selecting the second target knowledge layer, it can also be selected according to the preset knowledge layer selection rules, such as selecting according to the preset selection position and selection quantity.
[0097] like Figure 3 As shown, in one embodiment, optionally, the method further includes:
[0098] Step S301, receiving a search instruction containing preset domain knowledge;
[0099] Step S302: According to the search instruction, the target embedding model is used to search for corresponding preset domain knowledge to obtain corresponding search results.
[0100] In this embodiment, the preset domain knowledge is retrieved through the trained target embedding model, which can improve the retrieval ability of the target embedding model on the preset domain data, has strong operability, and avoids catastrophic forgetting.
[0101] like Figure 4 As shown, in one embodiment, optionally, the method further includes:
[0102] Step S401, obtaining different numbers of continuous candidate knowledge layers from different positions of multiple first original transformer layers of the target language model respectively;
[0103] Step S402, respectively copying the candidate knowledge layers, and inserting the copied candidate knowledge layers into preset positions in a plurality of first original transformer layers of the target language model to generate corresponding candidate language models;
[0104] Step S403, freezing the parameter values of the multiple first original transformer layers, and training the candidate knowledge layer of each of the candidate language models using preset domain knowledge data to determine the candidate parameter values corresponding to the candidate knowledge layer, and obtaining a trained candidate language model;
[0105] Step S404, inserting the candidate knowledge layer into the corresponding position of the embedding model, and performing small sample comparative learning using the preset domain knowledge data to obtain a candidate embedding model;
[0106] Step S405, when using the candidate embedding model to perform preset domain knowledge retrieval, calculate the retrieval accuracy rate corresponding to each candidate embedding model;
[0107] Step S406, determining a preset knowledge layer selection rule based on the candidate embedding model with the highest retrieval accuracy.
[0108] In this embodiment, different numbers of continuous candidate knowledge layers can be obtained from different positions of multiple original transformer layers of the target language model, and then copied and inserted into corresponding positions to generate corresponding candidate language models. After training with preset domain knowledge data, a trained candidate language model is obtained. After that, the trained candidate knowledge layer is inserted into the corresponding position of the embedding model, and a candidate embedding model is obtained after small sample comparison learning. When using the candidate embedding model to retrieve preset domain knowledge, the retrieval accuracy corresponding to each candidate embedding model is calculated, and the candidate embedding model with the highest retrieval accuracy is selected. Then, the selected candidate knowledge layer and the number of selections corresponding to the candidate embedding model can be used as the preset knowledge layer selection rules of the embedding model. In this way, the accuracy of the retrieval results of the obtained embedding model can be guaranteed.
[0109] In one embodiment, optionally, the method further comprises:
[0110] Receive input knowledge layer selection rule setting command;
[0111] The preset knowledge layer selection rule is set according to the knowledge layer selection rule setting command.
[0112] In this embodiment, the user or manufacturer can also set the transformer layer selection rules as needed. Different transformer layer selection rules can be set for different language embedding models to meet the different requirements of different language embedding models.
[0113] The following takes a language embedding model that is a Decoder-Only architecture embedding model as an example to describe the above technical solution of the present invention in detail.
[0114] like Figure 5 As shown in the figure, the process of the domain knowledge injection method of the Decoder-Only architecture embedding model includes:
[0115] Step S501, copy several transformer layers in the large language model of the Decoder-Only architecture, and insert the copied transformer layers according to the selected positions of the transformer layers to form a new large language model.
[0116] Among them, regarding the selection of the transformer layer, in the Decoder-Only architecture, the transformer layer closest to the input end in the Decoder-Only large language model can be selected. C-Eval is a benchmark evaluation dataset for evaluating large language models in various fields of Chinese, and can be used to evaluate the "knowledge mastery" of the large language model itself. Figure 6As shown, an experiment was conducted on a current open source large language model. After copying and inserting different depth network layers, C-Eval evaluation was performed. The weights of the deeper layers were copied, and the scores were close to the original model scores.
[0117] Step S502, using preset domain knowledge data to train the new large language model, wherein only the weights of the inserted transformer layer are updated, while the parameters of other layers are frozen, thereby obtaining a domain knowledge layer.
[0118] Step S503: insert the trained domain knowledge layer into the corresponding position in the embedding model to obtain a new embedding model.
[0119] Step S504: perform small sample comparative learning on the new embedding model to complete knowledge transfer.
[0120] The following takes the language embedding model as an Encoder architecture model as an example to explain the above technical solution of the present invention in detail. Unlike Decoder-only, the Encoder architecture directly outputs the embedding vector through the pooler layer, and there is no large language model of the corresponding architecture.
[0121] like Figure 7 As shown in the figure, the process of the domain knowledge injection method of the Encoder architecture embedding model includes:
[0122] Step S701, copy some transformer layers from its own BertEncoder and insert them sequentially after the copied layers.
[0123] Step S702, directly use the preset domain corpus to pre-train the embedding model. During the training process, the original weights of each layer are kept unchanged, and only the weights of the copied knowledge layer are updated.
[0124] Step S703, performing small sample comparative learning based on preset domain knowledge data on the embedding model that completes step S702, and only updating the knowledge layer weights.
[0125] like Figure 8 As shown in the figure, the C-MTEB dataset is used to evaluate the retrieval ability of the embedding model on a general dataset. For the Encoder architecture embedding model, the open source embedding model bge-large-zh1.5 is used. The evaluation results on the C-MTEB dataset are as follows Figure 8 As shown in the figure, it can be seen that after copying the shallow weights for retrieval ability evaluation, its various indicators are close to the original model, while after copying the deep weights, the indicators drop significantly.
[0126] Through the above-mentioned technical scheme of the present invention, the language model or the knowledge layer of the embedded model itself can be copied and trained to realize knowledge transfer, and there is no need to perform full parameter training. Only a small number of domain knowledge samples are needed for parameter training of the copied knowledge layer. The training parameters are very few, the hardware memory requirement is low, the retrieval ability on professional data is greatly improved, the operability is strong, and catastrophic forgetting is avoided.
[0127] Fig. 9 A block diagram of a domain knowledge injection device based on a language embedding model according to an embodiment of the present application is shown.
[0128] like Fig. 9 As shown, in the second aspect, the embodiment of the present application provides a domain knowledge injection device 90 based on a language embedding model, comprising:
[0129] An acquisition module 91 is used to acquire a target language model having the same architecture as the embedding model, wherein the target language model includes a plurality of first original transformer layers;
[0130] A generating module 92, configured to select and copy a first target knowledge layer from the plurality of first original transformer layers, and insert the copied first target knowledge layer into the plurality of first original transformer layers of the target language model to generate a new language model;
[0131] A training module 93, configured to train the first target knowledge layer of the new language model using preset domain knowledge data to determine the first target domain knowledge layer;
[0132] An inserting module 94, configured to insert the first target domain knowledge layer into a corresponding position of the embedding model to obtain a new first embedding model;
[0133] The contrastive learning module 95 is used to perform small sample contrastive learning on the new first embedding model using the preset domain knowledge data to obtain a target embedding model.
[0134] In one embodiment, optionally, the device further comprises:
[0135] A determination module, used for determining whether there is a target language model having the same architecture as the embedding model before obtaining the target language model having the same architecture as the embedding model;
[0136] The acquisition module is used to: when it is determined that there is a target language model having the same architecture as the embedding model, acquire the target language model;
[0137] The generating module is further configured to: when it is determined that there is no target language model having the same architecture as the embedding model, directly select and copy a second target knowledge layer from the plurality of second original transformer layers of the embedding model, and insert the copied second target knowledge layer into the plurality of second original transformer layers of the embedding model to generate a new second embedding model;
[0138] The training module is also used to: train the second target knowledge layer of the new second embedding model using the preset domain knowledge data to determine the second target domain knowledge layer;
[0139] The contrastive learning module is also used to: use the preset domain knowledge data to perform small sample contrastive learning on the second target domain knowledge layer of the new second embedding model to obtain a target embedding model.
[0140] In one embodiment, optionally, the device further comprises:
[0141] A first receiving module, used for receiving a search instruction containing preset domain knowledge;
[0142] The processing module is used to retrieve the corresponding preset domain knowledge using the target embedding model according to the retrieval instruction to obtain the corresponding retrieval result.
[0143] In one embodiment, optionally, the generating module is used to:
[0144] According to a preset knowledge layer selection rule, a first target knowledge layer is selected and copied from the plurality of first original transformer layers, wherein the preset knowledge layer selection rule includes a selection position and a selection quantity of the target knowledge layer.
[0145] In one embodiment, optionally, the generating module is further configured to:
[0146] Insert the copied first target knowledge layer into a preset position of multiple first original transformer layers of the target language model to generate a new language model, wherein the preset position includes a position adjacent to and subsequent to the selected position of the first target knowledge layer.
[0147] In one embodiment, optionally, the training module is used to:
[0148] Freeze the parameter values of the plurality of first original transformer layers, and train the first target knowledge layer of the new language model using preset domain knowledge data to determine a first target parameter value corresponding to the first target knowledge layer;
[0149] The contrastive learning module is used to:
[0150] Freeze the parameter values of other layers of the new first embedding model, use the preset domain knowledge data to perform small sample comparative learning on the inserted first target domain knowledge layer to determine the second target parameter value corresponding to the first target knowledge layer, and obtain the target embedding model.
[0151] In one embodiment, optionally,
[0152] The acquisition module is also used to: respectively acquire different numbers of continuous candidate knowledge layers from different positions of multiple first original transformer layers of the target language model;
[0153] The generating module is further used to: respectively copy the candidate knowledge layers, and insert the copied candidate knowledge layers into preset positions in the plurality of first original transformer layers of the target language model to generate corresponding candidate language models;
[0154] The training module is also used to: freeze the parameter values of the multiple first original transformer layers, and use the preset domain knowledge data to train the candidate knowledge layer of each candidate language model to determine the candidate parameter value corresponding to the candidate knowledge layer, and obtain the trained candidate language model;
[0155] The contrastive learning module is also used to insert the candidate knowledge layer into the corresponding position of the embedding model, and use the preset domain knowledge data to perform small sample contrastive learning to obtain a candidate embedding model;
[0156] A calculation module, used to calculate the retrieval accuracy corresponding to each candidate embedding model when using the candidate embedding model to perform preset domain knowledge retrieval;
[0157] The rule determination module is used to determine the preset knowledge layer selection rules based on the candidate embedding model with the highest retrieval accuracy.
[0158] In one embodiment, optionally, the device further comprises:
[0159] A second receiving module is used to receive an input knowledge layer selection rule setting command;
[0160] The setting module is used to set the preset knowledge layer selection rule according to the knowledge layer selection rule setting command.
[0161] In one embodiment, optionally, the embedding model includes a Decoder-Only architecture model and an Encoder architecture model.
[0162] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned domain knowledge injection method based on the language embedding model when executing the computer program.
[0163] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned domain knowledge injection method based on the language embedding model are implemented.
[0164] It should be noted that technical personnel in the relevant field can clearly understand that, for the convenience and conciseness of description, the specific working process of the domain knowledge injection device based on the language embedding model and each module described above can refer to the corresponding process in the aforementioned embodiment of the domain knowledge injection method based on the language embedding model, and will not be repeated here.
[0165] It should be noted that technical personnel in the relevant field can clearly understand that, for the convenience and conciseness of description, the specific working process of the model training device and each module described above can refer to the corresponding process in the aforementioned embodiment of the domain knowledge injection method based on the language embedding model, and will not be repeated here.
[0166] The above-mentioned domain knowledge injection device based on language embedding model can be implemented in the form of a computer program. Figure 5 Runs on the computer device shown.
[0167] Fig.10 A block diagram of a computer device according to an embodiment of the present application is shown.
[0168] See also Fig.10 The computer device includes a processor, a memory and a network interface connected through a system bus, wherein the memory may include a storage medium and an internal memory.
[0169] The storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any of the domain knowledge injection methods based on language embedding models for multi-source data provided in the embodiments of the present application.
[0170] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.
[0171] The internal memory provides an environment for the operation of the computer program in the storage medium. When the computer program is executed by the processor, the processor can execute any infectious disease transmission path analysis method or prediction neural network training method. The storage medium can be non-volatile or volatile.
[0172] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that Fig.10 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0173] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0174] In addition, an embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute the method steps described in the embodiment of the first aspect.
[0175] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or electronic device can refer to the relevant description in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0176] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0177] It should be understood that, although the terms first, second, etc. may be used to describe the setting unit in the embodiments of the present application, these setting units should not be limited to these terms. These terms are only used to distinguish the setting units from each other. For example, without departing from the scope of the embodiments of the present application, the first setting unit may also be referred to as the second setting unit, and similarly, the second setting unit may also be referred to as the first setting unit.
[0178] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.
[0179] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0180] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0181] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0182] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A domain knowledge injection method based on a language embedding model, characterized in that: The method comprises: Obtain a target language model having the same architecture as the embedding model, wherein the target language model includes a plurality of first original transformer layers; According to a preset knowledge layer selection rule, a first target knowledge layer is selected and copied from the plurality of first original transformer layers, and the copied first target knowledge layer is inserted into the plurality of first original transformer layers of the target language model to generate a new language model, wherein the preset knowledge layer selection rule includes a selection position and a selection quantity of the target knowledge layer; Freezing parameter values of the plurality of first original transformer layers, and pre-training a first target knowledge layer of the new language model using preset domain knowledge data to determine a first target domain knowledge layer; Inserting the first target domain knowledge layer into a corresponding position of the embedding model to obtain a new first embedding model; Freeze parameter values of other layers of the new first embedding model, and use the preset domain knowledge data to perform small sample comparative learning on the new first embedding model to obtain a target embedding model; The method further comprises: Acquire different numbers of continuous candidate knowledge layers from different positions of multiple first original transformer layers of the target language model respectively; Respectively copying the candidate knowledge layers, and inserting the copied candidate knowledge layers into preset positions in a plurality of first original transformer layers of the target language model to generate corresponding candidate language models; Pre-training the candidate knowledge layer of each candidate language model using preset domain knowledge data to determine the candidate parameter values corresponding to the candidate knowledge layer, and obtaining a trained candidate language model; Inserting the candidate knowledge layer into the corresponding position of the embedding model to obtain a candidate embedding model; When using the candidate embedding model to retrieve the preset domain knowledge, calculate the retrieval accuracy corresponding to each candidate embedding model; Determining the preset knowledge layer selection rule according to the candidate embedding model with the highest retrieval accuracy; The method further comprises: receiving a search instruction containing preset domain knowledge; According to the search instruction, the target embedding model is used to search for corresponding preset domain knowledge to obtain corresponding search results.
2. The method according to claim 1, characterized in that The method of pre-training the candidate knowledge layer of each candidate language model using the preset domain knowledge data to determine the candidate parameter value corresponding to the candidate knowledge layer to obtain the trained candidate language model includes: Freeze the parameter values of the multiple first original transformer layers, and pre-train the candidate knowledge layer of each candidate language model using preset domain knowledge data to determine the candidate parameter values corresponding to the candidate knowledge layer, and obtain a trained candidate language model; Inserting the candidate knowledge layer into the corresponding position of the embedding model to obtain the candidate embedding model includes: The candidate knowledge layer is inserted into the corresponding position of the embedding model, and small sample comparative learning is performed using the preset domain knowledge data to obtain a candidate embedding model.
3. The method according to claim 1, characterized in that Before obtaining a target language model having the same architecture as the embedding model, the method further includes: determining whether there is a target language model having the same architecture as the embedding model; When it is determined that there is a target language model having the same architecture as the embedding model, acquiring the target language model; When it is determined that there is no target language model having the same architecture as the embedding model, directly selecting and copying a second target knowledge layer from the multiple second original transformer layers of the embedding model, and inserting the copied second target knowledge layer into the multiple second original transformer layers of the embedding model to generate a new second embedding model; Pre-training the second target knowledge layer of the new second embedding model using the preset domain knowledge data to determine the second target domain knowledge layer; The preset domain knowledge data is used to perform small sample comparative learning on the second target domain knowledge layer of the new second embedding model to obtain a target embedding model.
4. The method according to claim 1, characterized in that The step of inserting the copied first target knowledge layer into a plurality of first original transformer layers of the target language model to generate a new language model comprises: Insert the copied first target knowledge layer into a preset position of multiple first original transformer layers of the target language model to generate a new language model, wherein the preset position includes a position adjacent to and subsequent to the selected position of the first target knowledge layer.
5. The method according to claim 1, characterized in that Pre-training the first target knowledge layer of the new language model using preset domain knowledge data to determine the first target domain knowledge layer includes: Pre-training a first target knowledge layer of the new language model using preset domain knowledge data to determine a first target parameter value corresponding to the first target knowledge layer; The using the preset domain knowledge data to perform small sample comparative learning on the new first embedding model to obtain a target embedding model includes: The preset domain knowledge data is used to perform small sample comparative learning on the inserted first target domain knowledge layer to determine the second target parameter value corresponding to the first target knowledge layer and obtain a target embedding model.
6. A domain knowledge injection device based on a language embedding model, characterized in that: include: An acquisition module, configured to acquire a target language model having the same architecture as the embedding model, wherein the target language model includes a plurality of first original transformer layers; A generation module, configured to select and copy a first target knowledge layer from the plurality of first original transformer layers according to a preset knowledge layer selection rule, and insert the copied first target knowledge layer into the plurality of first original transformer layers of the target language model to generate a new language model; wherein the preset knowledge layer selection rule includes a selection position and a selection quantity of the target knowledge layer; A training module, used to freeze the parameter values of the plurality of first original transformer layers, and pre-train the first target knowledge layer of the new language model using preset domain knowledge data to determine the first target domain knowledge layer; An insertion module, configured to insert the first target domain knowledge layer into a corresponding position of the embedding model to obtain a new first embedding model; A contrastive learning module, used to freeze parameter values of other layers of the new first embedding model, and use the preset domain knowledge data to perform small sample contrastive learning on the new first embedding model to obtain a target embedding model; The acquisition module is also used to: respectively acquire different numbers of continuous candidate knowledge layers from different positions of multiple first original transformer layers of the target language model; The generating module is further used to: respectively copy the candidate knowledge layers, and insert the copied candidate knowledge layers into preset positions in the plurality of first original transformer layers of the target language model to generate corresponding candidate language models; The training module is also used to: train the candidate knowledge layer of each candidate language model using preset domain knowledge data to determine the candidate parameter value corresponding to the candidate knowledge layer, and obtain the trained candidate language model; The contrastive learning module is further used to insert the candidate knowledge layer into the corresponding position of the embedding model to obtain a candidate embedding model; A calculation module, used to calculate the retrieval accuracy corresponding to each candidate embedding model when using the candidate embedding model to perform preset domain knowledge retrieval; A rule determination module, used to determine the preset knowledge layer selection rule according to the candidate embedding model with the highest retrieval accuracy; A first receiving module, used for receiving a search instruction containing preset domain knowledge; The processing module is used to retrieve the corresponding preset domain knowledge using the target embedding model according to the retrieval instruction to obtain the corresponding retrieval result.
7. A computer device, characterized in that: include: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, and the instructions are configured to execute the method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: Computer executable instructions are stored, and the computer executable instructions are used to execute the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Dialogue model training method in low data scene and computer equipment
CN115329062A
Text information identification method and device, storage medium and electronic equipment
CN117744760A