Processing method and device of dense encoder realized based on large language model

By transforming the decoder of the large language model into a bidirectional encoder and fine-tuning it, the decoder lacks bidirectional encoding capability in text-intensive vector encoding is solved, and the retrieval accuracy and recall rate are improved.

CN120353916AActive Publication Date: 2025-07-22BEIJING DP TECH CO LTD

Patent Information

Application Number
CN202510467306.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-22
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The existing decoders of large language models lack bidirectional encoding capabilities when encoding text dense vectors, and cannot effectively capture the overall content and structure of the document, resulting in insufficient retrieval accuracy and recall.

Method used

By transforming the causal mask matrix of the decoder into a full 1 matrix, a bidirectional encoder is built, and the dense encoder is fine-tuned through the masked word prediction task and the unsupervised comparison learning mechanism to form a dense encoder with bidirectional encoding capability.

Benefits of technology

Improve the accuracy and recall of text retrieval tasks, and enhance the understanding of document content and structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353916A_ABST
    Figure CN120353916A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a processing method and device for a dense encoder realized based on a large language model, and the method comprises the steps: selecting a large language model which has completed pre-training and NLP task fine tuning and is realized based on a pure decoder architecture as a target model, the bidirectional encoders are obtained through a transformation mode of solidifying a causal mask matrix used by a target model decoder in a reasoning process into an all-one matrix, and the embedded encoding module of the target model and the plurality of bidirectional encoders are connected in sequence to form a dense encoder; performing first-stage fine adjustment on the dense encoder through a shielding word prediction task, and performing second-stage fine adjustment on the dense encoder through an unsupervised contrast learning mechanism; and after the fine tuning is finished, constructing a document vector library for a target document library appointed by the user by utilizing the dense encoder, and providing retrieval service for the target document library based on the document vector library and the dense encoder. The dense encoder is used for processing a text retrieval task, so that the retrieval accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to a processing method and device for a dense encoder implemented based on a large language model. Background Art

[0002] Large Language Models (LLMs) implemented based on the decoder (denoted as Decoder-A) framework of the Transformer model, also known as Decoder-only LLMs, are the current mainstream autoregressive generation models, such as the GPT series models, LLaMA series models, OPT series models, GLM series models, DeepSeek series models, and so on.

[0003] The decoder module of Decoder-only LLMs is denoted as Decoder-B. The difference between Decoder-A / B is that Decoder-B does not contain the multi-head attention layer (Multi-Head Attention, MHA) for performing cross-attention operations with the encoder output in Decoder-A, as well as its corresponding residual connection and layer normalization unit (Add&Norm). That is to say, Decoder-A is sequentially connected by a masked multi-head attention layer (Masked Multi-Head Attention, MMHA)+Add&Norm, MHA+Add&Norm, and a feedforward neural network (Feedforward Neural Network, FNN)+Add&Norm, while Decoder-B is only sequentially connected by MMHA+Add&Norm and FNN+Add&Norm. In addition, the decoding logic of the decoder network (formed by connecting multiple Decoder-Bs) of Decoder-only LLMs is consistent with the decoding logic of the Transformer model.

[0004] The decoder vectors output by the decoder network of Decoder-only LLMs after pre-training and fine-tuning for natural language processing (NLP) tasks can be well used for a series of autoregressive prediction tasks, such as the next word prediction task, sequence-to-sequence prediction tasks, etc. However, if the decoded vector is used as a high-dimensional semantic expression vector (referred to as a dense vector) of the input text, it is still insufficient, because from the decoding logic of the Transformer model, it can be seen that the masked multi-head attention layer in Decoder-B will use a set mask vector (also called causal mask vector) to ensure that the model will only calculate in a set single-line direction (from left to right or from right to left) when performing scaled dot-product attention (Scaled Dot-ProductAttention). This will make the sub-vectors corresponding to each word on the final output decoding vector only related to the previous text (also called the previous text) of the current word position, and not related to the subsequent text (also called the following text) of the current word position. That is to say, although the decoded vectors output by the decoder network of Decoder-only LLMs after pre-training and NLP task fine-tuning can be well used for a series of autoregressive prediction tasks, they lack bidirectional encoding capabilities if used as dense vectors of input text.

[0005] The technical problem to be solved by the present invention is how to transform and fine-tune the decoder of a Decoder-only LLMs that has completed pre-training and NLP task fine-tuning so that it can be used as a dense vector encoder with bidirectional encoding capabilities (denoted as encoder-B). Because encoder-B comes from a trained autoregressive model, it can not only complete bidirectional encoding work during the encoding process, but also inherit the long-range dependencies learned by the original model during the training process, which is something that the traditional encoder (denoted as encoder-A) based on the Transformer model encoder structure does not have. In principle, in the retrieval scenario, compared with encoder-A, the long-range dependency capture ability and bidirectional encoding ability of encoder-B should be able to better understand the overall content and structure of the document and generate more expressive text dense vectors. Retrieval based on the text dense vector output by encoder-B can further improve the accuracy and recall rate of retrieval. Summary of the invention

[0006] The object of the present invention is to provide a processing method, device, electronic device and computer-readable storage medium for a dense encoder implemented based on a large language model, aiming at the defects of the existing technology. The present invention first selects a large language model that has completed pre-training and NLP task fine-tuning and is implemented based on a pure decoder architecture as the corresponding target model, and uses the decoder of the target model as the corresponding target decoder, and obtains a bidirectional encoder by transforming the causal mask matrix used by the target decoder during the inference process into a matrix of all 1s, and a corresponding bidirectional encoding network is composed of multiple bidirectional encoders connected in sequence, and a dense encoder is composed of the embedding encoding module of the target model and the bidirectional encoding network connected; then, a corresponding model training framework is constructed for the dense encoder, and the dense encoder is fine-tuned in the first stage through the masked word prediction task of the model training framework, and after the first stage of fine-tuning, the dense encoder is fine-tuned in the second stage through an unsupervised contrast learning mechanism; after the second stage of fine-tuning, the user-specified document library is used as the target document library, and the dense encoder is used to construct a document vector library for the target document library, and a retrieval service is provided for the target document library based on the document vector library and the dense encoder. The present invention obtains a dense encoder with bidirectional encoding ability by transforming the decoder of the target model, ensures that the dense encoder can successfully inherit the ability to capture long-range dependencies through the first stage of fine-tuning, and further improves the semantic representation ability of the dense encoder through the second stage of fine-tuning. Processing text retrieval tasks based on the dense encoder provided by the present invention can effectively improve the accuracy and recall rate of retrieval.

[0007] To achieve the above object, the first aspect of the embodiments of the present invention provides a processing method for a dense encoder implemented based on a large language model, the method comprising:

[0008] Select a large language model that has completed pre-training and NLP task fine-tuning and is implemented based on a pure decoder architecture as the corresponding target model; and use the decoder of the target model as the corresponding target decoder; and obtain a bidirectional encoder by transforming the causal mask matrix used by the target decoder during the inference process into a matrix of all 1s; and a corresponding bidirectional encoding network is composed of multiple bidirectional encoders connected in sequence, and a dense encoder is composed of the embedding encoding module of the target model and the bidirectional encoding network connected; the NLP tasks at least include text generation tasks, information extraction tasks, and question answering tasks;

[0009] Construct a corresponding model training framework for the dense encoder; and fine-tune the dense encoder in the first stage through the masked word prediction task of the model training framework; after the first stage of fine-tuning, fine-tune the dense encoder in the second stage through an unsupervised contrast learning mechanism;

[0010] After the two-stage fine-tuning is completed, the document library specified by the user is used as the target document library; and a document vector library is constructed for the target document library by using the dense encoder; and a retrieval service is provided for the target document library based on the document vector library and the dense encoder; the target document library includes a plurality of first documents; the document vector library includes a plurality of first dense vectors, and the first dense vectors correspond to the first documents one by one.

[0011] Preferably, the dense encoder is used to perform dense vector encoding processing on the input text of the encoder and output the corresponding text dense vector Y;

[0012] The dense encoder is formed by connecting the embedding encoding module and the bidirectional encoding network;

[0013] The embedding encoding module is used to identify whether the input text contains a preset masking word "[MASK]"; if it does not contain the preset masking word, the input text is segmented according to the word segmentation rule of the target model and according to the preset word list of the target model to obtain the corresponding current segmentation sequence; if it contains the preset masking word, the input text is segmented into multiple text segments without the preset masking word with each preset masking word as the cut-off position, and each text segment is segmented according to the word segmentation rule of the target model and according to the preset word list to obtain the corresponding segment segmentation sequence, and each preset masking word is used as an independent masking segment, and all the segment segmentations of all the segment segmentation sequences and all the masking segments are sorted according to the arrangement order in the input text to obtain the corresponding current segmentation sequence; and a preset start segment "[SOS]" is added at the start position of the sequence of the current segmentation sequence, and the current segmentation sequence after adding the start segment is used as the corresponding text segmentation sequence W; and the text segmentation sequence W is subjected to embedding encoding processing according to the embedding encoding rule of the target model to obtain the corresponding text embedding vector X and sent to the bidirectional encoding network; the text segmentation sequence W is composed of multiple segments w i sorted, 1 ≤ index i ≤ N W , N W is the total number of segments of the text segmentation sequence; the text embedding vector X is composed of N W segment embedding vectors x i s, and the segment embedding vector x i corresponds to the segment w i one by one; when the input text contains the preset masking word, the total number of the preset masking words is denoted as N M , and the segment embedding vector x i corresponding to each preset masking word is denoted as the corresponding masking word embedding vector 1 ≤ index j ≤ NM ;

[0014] The bidirectional encoding network is composed of N X bidirectional encoders connected in sequence, and the total number of encoders N X is a preset positive integer greater than 1; the bidirectional encoding network is used to continuously bidirectionally encode the text embedding vector X through the N X bidirectional encoders inside to obtain the corresponding text dense vector Y and output it; the text dense vector Y is composed of N W token dense vectors y i ; the token dense vector y i corresponds one-to-one with the token embedding vector x i ; when the total number N M is not zero, the token dense vector y corresponding to each masked word embedding vector i is denoted as the corresponding masked word dense vector

[0015] The model structures of the bidirectional encoder and the target decoder are the same, and both are sequentially connected by a masked multi-head attention layer, a residual connection and layer normalization unit, a feed-forward neural network, and another residual connection and layer normalization unit;

[0016] The inference process of the masked multi-head attention layer of each bidirectional encoder is as follows:

[0017]

[0018] 1 ≤ index n ≤ N X ; X n-1 , X n are the input and output vectors of the nth bidirectional encoder respectively; d n-1 is the feature dimension of the input vector X n-1 ; Q n , K n , V n are the query vector, key vector, and value vector corresponding to the nth bidirectional encoder respectively; are the corresponding weights of the query vector Q n , the key vector K n , and the value vector V n respectively; is the inverted vector of the key vector K n ; M all-ones is a matrix of all 1s; the parameter set G n of each bidirectional encoder is composed of the corresponding weights ;

[0019] The only difference between the inference process of the bidirectional encoder and that of the target decoder is that the masked multi-head attention layer of the bidirectional encoder uses the all-ones matrix M during inference. all-ones while the masked multi-head attention layer of the target decoder uses the causal mask matrix M during inference. Causal The causal mask matrix M corresponding to the decoding order of the target model from left to right word by word Causal is a mask matrix with all ones in the lower left triangle and all zeros in the upper right triangle. The causal mask matrix M corresponding to the decoding order of the target model from right to left word by word Causal is a mask matrix with all ones in the lower right triangle and all zeros in the upper left triangle.

[0020] Preferably, the model training framework is used to predict the original word text masked by the preset masking word in the masked text input to the framework and output the corresponding original word prediction sequence;

[0021] The masked text is a string without punctuation marks and special characters, which contains at least one preset masking word and is not entirely composed of the preset masking words;

[0022] The original word prediction sequence is composed of one or more predicted original word texts; the predicted original word texts in the original word prediction sequence correspond one by one to the preset masking words in the masked text; the order of all the predicted original word texts in the original word prediction sequence matches the order of all the preset masking words in the masked text;

[0023] The model training framework is sequentially connected by the dense encoder, the linear layer, the Softmax function layer and the output module;

[0024] The dense encoder is used to take the masked text as the corresponding input text, perform dense vector encoding processing on the input text and output the corresponding text dense vector Y to send to the linear layer;

[0025] The linear layer is used to perform word table feature vector conversion on each token dense vector y i in the text dense vector Y to obtain the corresponding token feature vector h i ; and all the obtained token feature vectors h i are sequentially sorted to form a vector sequence H and sent to the Softmax function layer; the token feature vector h i corresponds one by one to the token dense vector y i ; the token feature vector h corresponding to each masked word dense vector i is the previous token feature vector h i-1Denoted as the corresponding previous word feature vector

[0026] The Softmax function layer is used to calculate the probability distribution of the word list according to each of the previous word feature vectors and use the calculation result as the corresponding masked word probability vector and form a vector sequence P by sorting all the obtained masked word probability vectors in order and send it to the output module; each of the masked word probability vectors consists of multiple word segmentation probabilities, and the word segmentation probabilities correspond one by one to the word segmentations in the preset word list;

[0027] The output module is used to use the word segmentation in the table corresponding to the maximum word segmentation probability of each of the masked word probability vectors as the corresponding predicted original word text; and form the corresponding original word prediction sequence by sorting all the obtained predicted original word texts in order and output it.

[0028] Preferably, the dense encoder is fine-tuned in one stage through the masked word prediction task of the model training framework, which specifically includes:

[0029] Step 41, randomly extract multiple sentences from the historical text processed by the target model, and use each extracted sentence as a corresponding first extracted text; and delete the punctuation marks and special characters in each of the first extracted texts;

[0030] Step 42, implant a low-rank matrix into the parameter set G n of each of the bidirectional encoders of the model training framework to obtain a corresponding parameter set

[0031] wherein, the parameter set consists of weights ;

[0032] The masked multi-head attention layer of the bidirectional encoder is based on the parameter set and the inference process is:

[0033]

[0034] is a pair of low-rank matrices corresponding to the weight ; is a pair of low-rank matrices corresponding to the weight ; is a pair of low-rank matrices corresponding to the weight ; the low-rank matrix corresponding to the parameter set is Form a corresponding set of low-rank matrix parameters

[0035] Step 43, initialize the current masking rate to a preset first masking rate;

[0036] Wherein, 0 < first masking rate < 1;

[0037] Step 44, perform pre-segmentation on each of the first extraction texts based on the word segmentation rule of the target model and according to the preset word list to obtain corresponding pre-segmented sequences; and randomly replace the word segments of each pre-segmented sequence according to the current masking rate using the preset masking words to obtain corresponding replaced word segment sequences; and perform string splicing on all the word segments of each replaced word segment sequence to obtain corresponding first training texts; and count the total number of the preset masking words in all the first training texts to obtain a corresponding total number N K ;

[0038] Wherein, the sequence word segments of the replaced word segment sequence contain the preset masking words, and the total number of the preset masking words is the floor value of the product of the total number of sequence word segments and the current masking rate;

[0039] Step 45, input each of the first training texts as the corresponding masked text into the model training framework for prediction processing, and record the probability vectors of each of the masked words calculated by the Softmax function layer in this processing as a corresponding prediction vector p k , 1 ≤ index k ≤ N K ; and based on each prediction vector p k set a corresponding label vector for the corresponding original word text and form a corresponding first prediction-label pair from each prediction vector p k and the corresponding label vector Wherein, the label vector

[0040] consists of multiple label word segment probabilities, and the label word segment probabilities correspond one by one to the word segments in the preset word list; only one of the label word segment probabilities in the label vector is 1, and the rest of the label word segment probabilities are 0; the word segment in the label vector corresponding to the label word segment probability of 1 matches the corresponding original word text;

[0041] Step 46, the N K obtainedfirst prediction-label pairs Calculate using the preset first model loss function L1 to obtain the corresponding first loss value;

[0042] Among them, the first model loss function L1 is:

[0043]

[0044] Step 47, identify whether the first loss value meets the preset first loss value range; if it meets, go to step 48; if it does not meet, identify whether the current masking rate is the first masking rate. If so, based on the preset first model optimizer, modulate all the low-rank matrix parameter sets of the bidirectional encoding network in the direction of minimizing the first model loss function L1 and the linear layer parameters of the linear layer for one round of modulation. Otherwise, based on the preset second model optimizer, modulate all the low-rank matrix parameter sets for one round of modulation, and return to step 45 to continue training after this round of modulation ends;

[0045] Among them, both the first and second model optimizers at least include the Adam optimizer and the SGD optimizer;

[0046] Step 48, identify whether the current masking rate is the first masking rate; if so, reset the current masking rate to the preset second masking rate and return to step 44 to continue training after resetting; if not, stop training and confirm that the first-stage fine-tuning ends;

[0047] Among them, 0 < first masking rate < second masking rate < 1.

[0048] Preferably, the two-stage fine-tuning of the dense encoder through the unsupervised contrastive learning mechanism specifically includes:

[0049] Step 51, randomly extract multiple sentences from the historical texts processed by the target model, and use each extracted sentence as a corresponding second extracted text; delete the punctuation marks and special characters in each second extracted text; and count the total number of the second extracted texts to obtain the corresponding total number N U ;

[0050] Step 52, use each second extracted text as the corresponding input text to input into the dense encoder for dense vector encoding processing, and record the text dense vector Y output by this processing as the corresponding first sample;

[0051] Step 53: Use each of the second extracted texts as the corresponding input text and input it into the dense encoder again for dense vector encoding. During this processing, randomly select multiple residual connection and layer normalization units from the bidirectional encoding network as corresponding random scrambling units. At the end of the inference process of each random scrambling unit, randomly scramble some sub-vectors of the current unit output vector, and pass the scrambled unit output vector downstream; and denote the text dense vector Y finally output by this dense vector encoding processing as the corresponding second sample.

[0052] Among them, the scrambling method of the random scrambling is: set the selected sub-vectors to all-zero vectors or add Gaussian noise to the selected sub-vectors.

[0053] Step 54: Denote any one of the first samples as the corresponding sample s u , 1 ≤ index u ≤ N U ; and denote the other first samples different from each sample s u as the negative samples of the current sample s u 1 ≤ index v ≤ N -1; and form a corresponding negative sample set from the N U -1 negative samples corresponding to each sample s u ; and denote the second sample corresponding to each sample s U as the positive sample of the current sample s ; and form a corresponding sample data group from each sample s and its corresponding positive sample u and the negative sample set u ; and form a corresponding sample data group ; and form a corresponding sample data group from each sample s u and its corresponding positive sample and the negative sample set ; and form a corresponding sample data group

[0054] Step 55: Input the N U obtained sample data groups into a preset second model loss function L2 for calculation to obtain the corresponding second loss value.

[0055] Among them, the second model loss function L2 is:

[0056]

[0057] λ is a preset adjustment parameter; sim() is the cosine similarity function of vectors, is the cosine similarity between the sample s u and the corresponding positive sample ; For the sample s u and the corresponding negative sample cosine similarity;

[0058] Step 56, identify whether the second loss value satisfies a preset second loss value range; if not, based on a preset third model optimizer, modulate all low-rank matrix parameter sets of the dense encoder in the direction of minimizing the second model loss function L2 for one round of modulation, and after the end of this round of modulation, return to Step 52 to continue training; if satisfied, stop training and confirm the end of two-stage fine-tuning;

[0059] The third model optimizer includes at least an Adam optimizer and an SGD optimizer.

[0060] Preferably, using the dense encoder to construct a document vector library for the target document library specifically includes:

[0061] Take each of the first documents in the target document library as the corresponding input text and input it into the dense encoder for dense vector encoding processing, and use the text dense vector output by the current processing as a corresponding first dense vector; and form the corresponding document vector library from all the obtained first dense vectors.

[0062] Preferably, providing a retrieval service for the target document library based on the document vector library and the dense encoder specifically includes:

[0063] Take the query text input by the user each time as the corresponding current query text; and take the current query text as the corresponding input text and input it into the dense encoder for dense vector encoding processing, and use the text dense vector output by the current processing as the corresponding current dense vector; and calculate the cosine similarity between the current dense vector and each of the first dense vectors in the document vector library to obtain the corresponding first similarity; and mark the first dense vectors whose first similarity exceeds a preset similarity threshold as the corresponding retrieval vectors; and when the total number of the retrieval vectors is not zero, mark the first documents in the target document library corresponding to each of the retrieval vectors as the corresponding retrieval documents; and sort all the obtained retrieval documents in descending order of the first similarity to form the corresponding retrieval document sequence and feedback it to the current user.

[0064] A second aspect of the embodiments of the present invention provides an apparatus for implementing the processing method of the dense encoder implemented based on the large language model described in the first aspect above. The apparatus includes: a model transformation module, a model fine-tuning module, and a retrieval application module;

[0065] The model transformation module is used to select a large language model that has completed pre-training and NLP task fine-tuning and is implemented based on a pure decoder architecture as the corresponding target model; and use the decoder of the target model as the corresponding target decoder; and obtain a bidirectional encoder by transforming the causal mask matrix used by the target decoder during the inference process into a matrix of all 1s; and a corresponding bidirectional encoding network is composed of multiple such bidirectional encoders connected in sequence, and a dense encoder is composed of the embedding encoding module of the target model and the bidirectional encoding network; the NLP tasks at least include text generation tasks, information extraction tasks, and question answering tasks.

[0066] The model fine-tuning module is used to construct a corresponding model training framework for the dense encoder; and perform a first-stage fine-tuning on the dense encoder through the masked word prediction task of the model training framework; after the first-stage fine-tuning, perform a second-stage fine-tuning on the dense encoder through an unsupervised contrast learning mechanism.

[0067] The retrieval application module is used to, after the second-stage fine-tuning, use the user-specified document library as the target document library; and use the dense encoder to construct a document vector library for the target document library; and provide a retrieval service for the target document library based on the document vector library and the dense encoder; the target document library includes multiple first documents; the document vector library includes multiple first dense vectors, and the first dense vectors correspond to the first documents one by one.

[0068] A third aspect of the embodiments of the present invention provides an electronic device, including: a memory, a processor, and a transceiver;

[0069] The processor is used to be coupled with the memory, read and execute the instructions in the memory to implement the method steps described in the first aspect above;

[0070] The transceiver is coupled with the processor, and the processor controls the transceiver to perform message sending and receiving.

[0071] A fourth aspect of the embodiments of the present invention provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed by a computer, the computer is caused to execute the instructions of the method described in the first aspect above.

[0072] An embodiment of the present invention provides a processing method, apparatus, electronic device, and computer-readable storage medium for a dense encoder implemented based on a large language model. As can be seen from the above content, in the embodiment of the present invention, a large language model that has completed pre-training and NLP task fine-tuning and is implemented based on a pure decoder architecture is first selected as the corresponding target model, and the decoder of the target model is used as the corresponding target decoder. By transforming the causal mask matrix used by the target decoder during the inference process into an all-1 matrix, a bidirectional encoder is obtained. A corresponding bidirectional encoding network is composed of multiple bidirectional encoders connected in sequence, and a dense encoder is formed by connecting the embedding encoding module of the target model and the bidirectional encoding network. Then, a corresponding model training framework is constructed for the dense encoder, and the dense encoder is fine-tuned in the first stage through the masked word prediction task of the model training framework. After the first-stage fine-tuning is completed, the dense encoder is fine-tuned in the second stage through an unsupervised contrast learning mechanism. After the second-stage fine-tuning is completed, the user-specified document library is used as the target document library, and the dense encoder is used to construct a document vector library for the target document library, and a retrieval service is provided for the target document library based on the document vector library and the dense encoder. In the embodiment of the present invention, a dense encoder with bidirectional encoding ability is obtained by transforming the decoder of the target model, and the first-stage fine-tuning ensures that the dense encoder can successfully inherit the ability to capture long-range dependencies, and the second-stage fine-tuning is used to further improve the semantic representation ability of the dense encoder. The dense encoder in the embodiment of the present invention effectively improves the retrieval accuracy and recall rate of the document retrieval task. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] Figure 1 FIG. is a schematic diagram of a processing method for a dense encoder implemented based on a large language model provided in Embodiment 1 of the present invention;

[0074] Figure 2 FIG. is a module structure diagram of Decoder-A / B provided in Embodiment 1 of the present invention;

[0075] Figure 3 FIG. is a module structure diagram of the model training framework provided in Embodiment 1 of the present invention;

[0076] Figure 4 FIG. is a module structure diagram of a processing apparatus for a dense encoder implemented based on a large language model provided in Embodiment 2 of the present invention;

[0077] Figure 5 FIG. is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0078] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0079] Embodiment 1 of the present invention provides a processing method for a dense encoder implemented based on a large language model. As Figure 1 shown in the schematic diagram of the processing method for a dense encoder implemented based on a large language model provided in Embodiment 1 of the present invention, the method mainly includes the following steps:

[0080] Step 1, select a large language model that has completed pre-training and NLP task fine-tuning and is implemented based on a pure decoder architecture as the corresponding target model; and use the decoder of the target model as the corresponding target decoder; and obtain a bidirectional encoder by transforming the causal mask matrix used by the target decoder during the inference process into a matrix of all 1s; and form a corresponding bidirectional encoding network by sequentially connecting multiple bidirectional encoders, and form a dense encoder by connecting the embedding encoding module of the target model and the bidirectional encoding network.

[0081] Here, the target model of the embodiment of the present invention is any large language model that has completed pre-training and NLP task fine-tuning and is implemented based on a pure decoder architecture, such as GPT series models, LLaMA series models, OPT series models, GLM series models, DeepSeek series models, etc. The NLP tasks of the target model include at least text generation tasks, information extraction tasks, question answering tasks, text classification tasks, text translation tasks, etc.

[0082] As can be seen from the foregoing background description, the target decoder in the embodiment of the present invention is Decoder-B mentioned in the foregoing background description. As Figure 2 shown in the module structure diagram of Decoder-A / B provided in Embodiment 1 of the present invention, the structural difference between Decoder-B and the decoder (Decoder-A) of the Transformer model lies in that Decoder-B does not include the multi-head attention layer for cross-attention operation with the output of the Transformer encoder and the residual connection and layer normalization unit corresponding to the attention layer.

[0083] The dense encoder in the embodiment of the present invention is used to perform dense vector encoding processing on the input text of the encoder and output the corresponding text dense vector Y. The dense encoder is formed by connecting the embedding encoding module and the bidirectional encoding network. As Figure 3As shown in the module structure diagram of the model training framework provided in the first embodiment of the present invention.

[0084] The embedding encoding module of the dense encoder in the embodiment of the present invention is used to identify whether the input text contains a preset masked word "[MASK]"; if it does not contain the preset masked word, the input text is tokenized according to the tokenization rules of the target model and the preset vocabulary of the target model to obtain the corresponding current token sequence; if it contains the preset masked word, the input text is split into multiple text segments without the preset masked word with each preset masked word as the cut-off position, and each text segment is tokenized according to the tokenization rules of the target model and the preset vocabulary to obtain the corresponding segment token sequence, and each preset masked word is used as an independent masked token, and all the segment tokens of all segment token sequences and all masked tokens are sorted in the order in the input text to obtain the corresponding current token sequence; and a preset start token "[SOS]" is added at the start position of the obtained current token sequence, and the current token sequence after adding the start token is used as the corresponding text token sequence W; and the text token sequence W is subjected to embedding encoding processing according to the embedding encoding rules of the target model to obtain the corresponding text embedding vector X and sent to the bidirectional encoding network.

[0085] Here, the preset vocabulary in the embodiment of the present invention is the large language model vocabulary accumulated by the target model through pre-training and NLP task fine-tuning; the text token sequence W in the embodiment of the present invention is composed of multiple tokens w i sorted, 1 ≤ index i ≤ N W , N W being the total number of tokens in the text token sequence; the text embedding vector X is composed of N W token embedding vectors x i s, and the token embedding vector x i corresponds to the token w i one by one; when the input text contains a preset masked word, the total number of preset masked words is denoted as N M , and the token embedding vector x i corresponding to each preset masked word is denoted as the corresponding masked word embedding vector 1 ≤ index j ≤ N M .

[0086] The bidirectional encoding network of the dense encoder in the embodiment of the present invention is composed of N X bidirectional encoders connected in sequence, and the total number of encoders N X is a preset positive integer greater than 1; this bidirectional encoding network is used to continuously bidirectionally encode the text embedding vector X through the internal N X bidirectional encoders to obtain the corresponding text dense vector Y and output it.

[0087] Here, the text dense vector Y of the embodiment of the present invention consists of N W token dense vectors y i and the token dense vectors y i correspond one-to-one with the token embedding vectors x i ; when the total number N M is not zero, the token dense vectors y corresponding to the respective masked word embedding vectors i are denoted as the corresponding masked word dense vectors

[0088] It should be noted that the bidirectional encoder of the embodiment of the present invention is obtained by transforming the target decoder. The transformation scheme of the embodiment of the present invention does not involve model structure adjustment. It only replaces the causal mask matrix M Causal used by the target decoder in the masked multi-head attention layer with a matrix of all 1s M all-ones .

[0089] The role of the causal mask matrix M Causal is to guide the target decoder to perform unidirectional attention inference through the setting of a semi-triangular matrix. Specifically: when the decoding order of the target model is word-by-word decoding from left to right, the corresponding causal mask matrix M Causal is a mask matrix with all 1s in the lower left triangle and all 0s in the upper right triangle. At this time, the target decoder only performs unidirectional attention inference based on the left context of the current token; when the decoding order is word-by-word decoding from right to left, the corresponding causal mask matrix M Causal is a mask matrix with all 1s in the lower right triangle and all 0s in the upper left triangle. At this time, the target decoder only performs unidirectional attention inference based on the right context of the current token.

[0090] The embodiment of the present invention enables the bidirectional encoder to perform attention inference based on the matrix of all 1s M all-ones because the matrix of all 1s M all-ones does not perform 0-value mask shielding on any previous context (above context) or subsequent context (below context). Naturally, this makes the bidirectional encoder not perform unidirectional shielding on the previous context (above context) or subsequent context (below context) of each token, and thus has the bidirectional encoding ability.

[0091] As can be seen from the previous text, the model structures of the bidirectional encoder and the target decoder in the embodiment of the present invention are the same. Both are sequentially connected by a masked multi-head attention layer, a layer of residual connection and layer normalization unit, a layer of feed-forward neural network, and another layer of residual connection and layer normalization unit, as Figure 2 , 3 shown.

[0092] The inference process of the masked multi-head attention layer of each bidirectional encoder in the embodiment of the present invention is as follows:

[0093]

[0094] where 1 ≤ index n ≤ N X ; X n-1 and X n are the input and output vectors of the n-th bidirectional encoder respectively; d n-1 is the feature dimension of the input vector X n-1 ; Q n and K n and V n are the query vector, key vector and value vector corresponding to the n-th bidirectional encoder respectively; are the corresponding weights of the query vector Q n , key vector K n and value vector V n respectively; is the inverted vector of the key vector K n ; M all-ones is a matrix of all 1s; the parameter set G of each bidirectional encoder n is composed of the corresponding weights .

[0095] The only difference between the inference process of the bidirectional encoder in the embodiment of the present invention and the target decoder is that the masked multi-head attention layer of the bidirectional encoder uses the all-1 matrix M all-ones during inference, while the masked multi-head attention layer of the target decoder uses the causal mask matrix M Causal during inference.

[0096] Step 2: Construct a corresponding model training framework for the dense encoder; and perform a first-stage fine-tuning on the dense encoder through the masked word prediction task of the model training framework; after the first-stage fine-tuning, perform a second-stage fine-tuning on the dense encoder through an unsupervised contrastive learning mechanism;

[0097] Specifically, it includes: Step 21: Construct a corresponding model training framework for the dense encoder;

[0098] Here, the model training framework of the embodiment of the present invention is used to predict the original word text masked by the preset masked word in the masked text input to the framework and output the corresponding original word prediction sequence, as Figure 3 shown;

[0099] where the masked text is a string without punctuation marks and special characters, which contains at least one preset masked word and is not entirely composed of preset masked words; the original word prediction sequence is composed of one or more predicted original word texts; the predicted original word texts in the original word prediction sequence correspond one-to-one to the preset masked words in the masked text; the order of all the predicted original word texts in the original word prediction sequence matches the order of all the preset masked words in the masked text;

[0100] As Figure 3 shown, the model training framework of the embodiment of the present invention is sequentially connected by a dense encoder, a linear layer, a Softmax function layer, and an output module; where:

[0101] 1) The dense encoder is used to take the masked text as the corresponding input text, perform dense vector encoding processing on the input text, and output the corresponding text dense vector Y to send to the linear layer;

[0102] 2) The linear layer is used to perform vocabulary feature vector conversion on each token dense vector y i in the text dense vector Y to obtain the corresponding token feature vector h i ; and all the obtained token feature vectors h i are sorted in order to form a vector sequence H and sent to the Softmax function layer;

[0103] Among them, the token feature vector h i corresponds one-to-one with the token dense vector y i ; the previous token feature vector h corresponding to each masked word dense vector i is denoted as the corresponding previous word feature vector i-1 ;

[0104] 3) The Softmax function layer is used to calculate the vocabulary probability distribution according to each previous word feature vector and use the calculation result as the corresponding masked word probability vector ; and all the obtained masked word probability vectors are sorted in order to form a vector sequence P and sent to the output module;

[0105] Among them, each masked word probability vector is composed of multiple token probabilities, and the token probabilities correspond one-to-one with the tokens in the preset vocabulary;

[0106] 4) The output module is used to take the token in the preset vocabulary corresponding to the maximum token probability of each masked word probability vector as the corresponding predicted original word text; and all the obtained predicted original word texts are sorted in order to form the corresponding original word prediction sequence and output;

[0107] Step 22, and perform one-stage fine-tuning on the dense encoder through the masked word prediction task of the model training framework;

[0108] Specifically, it includes: Step 221, randomly extracting multiple sentences from the historical text processed by the target model, and taking each extracted sentence as a corresponding first extracted text; and deleting punctuation marks and special characters in each first extracted text;

[0109] Step 222, implanting a low-rank matrix into the parameter set G of each bidirectional encoder in the model training framework to obtain a corresponding parameter set n Here, the parameter set after implanting the low-rank matrix

[0110] is composed of weights ;

[0111] The masked multi-head attention layer of the bidirectional encoder is based on the parameter set The inference process is as follows:

[0112]

[0113] is a pair of low-rank matrices corresponding to the weights ; is a pair of low-rank matrices corresponding to the weights ; is a pair of low-rank matrices corresponding to the weights ; The low-rank matrices corresponding to the parameter set form a corresponding low-rank matrix parameter set The low-rank matrix parameter sets of each bidirectional encoder are the implanted additional parameter sets; as can be seen from the following, the fine-tuning rule of each bidirectional encoder in the embodiment of the present invention is: keeping the original parameter set G n unchanged and only modulating the implanted additional parameter set, that is, the low-rank matrix parameter set ;

[0114] Step 223, initializing the current masking rate to a preset first masking rate;

[0115] where 0 < first masking rate < 1;

[0116] Here, the first masking rate in the embodiment of the present invention is a preset ratio parameter, such as 15%, 20%, etc.;

[0117] ​​Step 224: Based on the word segmentation rules of the target model and according to the preset word list, perform pre-word segmentation on each first extraction text to obtain the corresponding pre-word segmentation sequence; use the preset masking words to randomly replace the word segments in each pre-word segmentation sequence according to the current masking rate to obtain the corresponding replaced word segmentation sequence; concatenate all the word segments in each replaced word segmentation sequence to obtain the corresponding first training text; and count the total number of preset masking words in all the first training texts to obtain the corresponding total number N K ;

[0118] Among them, the sequence word segments of the replaced word segmentation sequence contain preset masking words, and the total number of preset masking words is the integer value obtained by rounding down the product of the total number of sequence word segments and the current masking rate;

[0119] For example, given the first masking rate = 20%, the preset masking word is "[MASK]", the first extraction text is "The cat loves to eat sardines and pumpkins", and the pre-word segmentation sequence is {"The cat", "loves", "to eat", "sardines", "and", "pumpkins"}; then, the total number of sequence word segments = 6, the total number of preset masking words = int(6 × 20%) = 1, int() is the rounding-down function, take the third word segment "sardines" in the pre-word segmentation sequence as the random replacement word segment this time, and the obtained replaced word segmentation sequence is {"The cat", "loves", "to eat", "[MASK]", "and", "pumpkins"}, and the obtained first training text is "The cat loves to eat [MASK] and pumpkins";

[0120] Step 225: Use each first training text as the corresponding masked text to input into the model training framework for prediction processing, and record the probability vectors of each masking word calculated by the Softmax function layer in this processing as a corresponding prediction vector p k , 1 ≤ index k ≤ N K ; and based on the original word text corresponding to each prediction vector p k set a corresponding label vector and form a corresponding first prediction-label pair from each prediction vector p k and the corresponding label vector Among them, the label vector

[0121] is composed of multiple label word segment probabilities, and the label word segment probabilities correspond one by one to the word segments in the preset word list; there is only one label word segment probability equal to 1 in the label vector and the rest of the label word segment probabilities are all 0; the word segment in the preset word list corresponding to the label word segment probability equal to 1 in the label vector matches the corresponding original word text;

[0122] ​Step 226: Bring the obtained N K first prediction-label pairs into the preset first model loss function L1 for calculation to obtain the corresponding first loss value;

[0123] Here, the first model loss function L1 in the embodiment of the present invention is:

[0124]

[0125] Step 227: Identify whether the first loss value meets the preset first loss value range; if it meets, go to Step 228; if it does not meet, identify whether the current masking rate is the first masking rate. If it is, modulate all the low-rank matrix parameter sets of the bi-directional encoding network and the linear layer parameters of the linear layer in the direction of minimizing the first model loss function L1 based on the preset first model optimizer. If not, modulate all the low-rank matrix parameter sets in the direction of minimizing the first model loss function L1 based on the preset second model optimizer, and return to Step 225 to continue training after this round of modulation;

[0126] Here, the first loss value range in the embodiment of the present invention is a preset numerical range; both the first and second model optimizers in the embodiment of the present invention at least include the Adam optimizer and the SGD optimizer;

[0127] It should be noted that in the first-stage fine-tuning, the embodiment of the present invention performs two rounds of fine-tuning by setting two masking rates (the first masking rate, the second masking rate); in the first round of fine-tuning, the bi-directional encoding network and the linear layer of the model training framework are synchronously trained, and the modulation parameter objects include all the low-rank matrix parameter sets of the bi-directional encoding network and the linear layer parameters of the linear layer; in the first round of fine-tuning, the linear layer parameters of the linear layer are no longer modulated, and only all the low-rank matrix parameter sets of the bi-directional encoding network are modulated; the optimizer used in the first round of fine-tuning is the first model optimizer, and the optimizer used in the second round of fine-tuning is the second model optimizer;

[0128] Step 228: Identify whether the current masking rate is the first masking rate; if it is, reset the current masking rate to the preset second masking rate and return to Step 224 to continue training after the reset; if not, stop training and confirm the end of the first-stage fine-tuning;

[0129] where, 0 < the first masking rate < the second masking rate < 1;

[0130] Here, the second masking rate of the embodiment of the present invention is a preset ratio parameter, and the second masking rate should be greater than the first masking rate, such as 80%, 90%, etc.;

[0131] It should be noted that, as can be seen from the foregoing, in the first-stage fine-tuning of the embodiment of the present invention, two masking rates are set to perform two rounds of fine-tuning. In practical applications, more masking rates can also be set to perform more rounds of fine-tuning. For example: four masking rates are provided (the first masking rate = 20%, the second masking rate = 40%, the third masking rate = 60%, the fourth masking rate = 80%) to perform four rounds of fine-tuning; it must be noted that whether two or more than two masking rates are set to perform two or more rounds of fine-tuning, only the linear layer parameters of the linear layer are modulated in the first round of fine-tuning, and in the subsequent second round or more rounds of fine-tuning, only the bidirectional encoding network is modulated and the linear layer is no longer modulated;

[0132] Step 23, after the first-stage fine-tuning is completed, the dense encoder is fine-tuned in the second stage through an unsupervised contrastive learning mechanism;

[0133] Specifically, it includes: Step 231, randomly extract multiple sentences from the historical texts processed by the target model, and use each extracted sentence as a corresponding second extracted text; and delete the punctuation marks and special characters in each second extracted text; and count the total number of second extracted texts to obtain the corresponding total number N U ;

[0134] Step 232, use each second extracted text as the corresponding input text to input the dense encoder for dense vector encoding processing, and record the text dense vector Y output by this processing as the corresponding first sample;

[0135] Step 233, use each second extracted text as the corresponding input text to input the dense encoder for dense vector encoding processing again, and randomly select multiple residual connection and layer normalization units from the bidirectional encoding network as the corresponding random scrambling units during this processing, and randomly scramble some sub-vectors of the current unit output vector at the end of the inference process of each random scrambling unit, and pass the scrambled unit output vector downstream; and record the text dense vector Y finally output by this dense vector encoding processing as the corresponding second sample;

[0136] Here, the scrambling method of random scrambling in the embodiment of the present invention is: setting the selected sub-vector to a zero vector or adding Gaussian noise to the selected sub-vector;

[0137] Step 234, record any one of the first samples as the corresponding sample s u , 1 ≤ index u ≤ N U ; and for each sample su The different other first samples are denoted as the current sample s u negative samples 1 ≤ index v ≤ N U -1; and by each sample s u corresponding N U -1 negative samples form a corresponding negative sample set And the second samples corresponding to each sample s u are denoted as the current sample s u positive samples And by each sample s u and its corresponding positive samples and the negative sample set form a corresponding sample data group

[0138] Step 235, bring the obtained N U sample data groups into the preset second model loss function L2 for calculation to obtain the corresponding second loss value;

[0139] Here, the second model loss function L2 of the embodiment of the present invention is:

[0140]

[0141] where λ is a preset adjustment parameter;

[0142] sim() is the cosine similarity function of vectors, is the cosine similarity between the sample s u and the corresponding positive sample ; is the cosine similarity between the sample s u and the corresponding negative sample ;

[0143] Step 236, identify whether the second loss value satisfies the preset second loss value range; if not, based on the preset third model optimizer, modulate all the low-rank matrix parameter sets of the dense encoder in the direction of minimizing the second model loss function L2 for one round, and after the end of this round of modulation, return to Step 232 to continue training; if satisfied, stop training and confirm the end of the second-stage fine-tuning;

[0144] Here, the second loss value range of the embodiment of the present invention is a preset numerical range; the third model optimizer of the embodiment of the present invention at least includes the Adam optimizer and the SGD optimizer.

[0145] Step 3, after the second-stage fine-tuning is completed, use the document library specified by the user as the target document library; use the dense encoder to construct a document vector library for the target document library; and provide a retrieval service for the target document library based on the document vector library and the dense encoder;

[0146] Specifically, it includes: Step 31, after the second-stage fine-tuning is completed, use the document library specified by the user as the target document library;

[0147] Here, the target document library in the embodiment of the present invention includes multiple first documents;

[0148] Step 32, use the dense encoder to construct a document vector library for the target document library;

[0149] Here, the document vector library in the embodiment of the present invention includes multiple first dense vectors, and the first dense vectors correspond to the first documents one by one;

[0150] Specifically, it includes: taking each first document in the target document library as the corresponding input text and inputting it into the dense encoder for dense vector encoding processing, and taking the text dense vector output by the current processing as a corresponding first dense vector; and forming the corresponding document vector library from all the obtained first dense vectors;

[0151] Step 33, provide a retrieval service for the target document library based on the document vector library and the dense encoder;

[0152] Specifically, it includes: taking the query text input by the user each time as the corresponding current query text; taking the current query text as the corresponding input text and inputting it into the dense encoder for dense vector encoding processing, and taking the text dense vector output by the current processing as the corresponding current dense vector; calculating the cosine similarity between the current dense vector and each first dense vector in the document vector library to obtain the corresponding first similarity; taking the first dense vector with the first similarity exceeding the preset similarity threshold as the corresponding retrieval vector; when the total number of retrieval vectors is not zero, taking the first document in the target document library corresponding to each retrieval vector as the corresponding retrieval document; and sorting all the obtained retrieval documents in descending order of the first similarity to form the corresponding retrieval document sequence and feedback it to the current user. Here, the similarity threshold in the embodiment of the present invention is a preset threshold parameter.

[0153] Figure 4 This is the module structure diagram of a processing device of a dense encoder implemented based on a large language model provided in the second embodiment of the present invention. This device is a terminal device or a server for implementing the foregoing method embodiment, or can also be a device that enables the foregoing terminal device or server to implement the foregoing method embodiment. For example, this device can be a device or a chip system of the foregoing terminal device or server. As Figure 4As shown in the figure, the device includes: a model transformation module 201, a model fine-tuning module 202, and a retrieval application module 203.

[0154] The model transformation module 201 is used to select a large language model that has completed pre-training and NLP task fine-tuning and is implemented based on a pure decoder architecture as the corresponding target model; and use the decoder of the target model as the corresponding target decoder; and obtain a bidirectional encoder by transforming the causal mask matrix used by the target decoder during the inference process into a full 1 matrix; and a corresponding bidirectional encoding network is composed of multiple bidirectional encoders connected in sequence, and a dense encoder is composed of the embedding encoding module of the target model and the bidirectional encoding network; the NLP tasks at least include text generation tasks, information extraction tasks, and question answering tasks.

[0155] The model fine-tuning module 202 is used to construct a corresponding model training framework for the dense encoder; and perform a first-stage fine-tuning on the dense encoder through the masked word prediction task of the model training framework; after the first-stage fine-tuning is completed, perform a second-stage fine-tuning on the dense encoder through an unsupervised contrast learning mechanism.

[0156] The retrieval application module 203 is used to, after the second-stage fine-tuning is completed, use the document library specified by the user as the target document library; and use the dense encoder to construct a document vector library for the target document library; and provide a retrieval service for the target document library based on the document vector library and the dense encoder; the target document library includes multiple first documents; the document vector library includes multiple first dense vectors, and the first dense vectors correspond to the first documents one by one.

[0157] The processing device for a dense encoder implemented based on a large language model provided by an embodiment of the present invention can execute the method steps in the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here.

[0158] It should be noted that it should be understood that the division of each module of the above device is only a division of logical functions. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by processing elements; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. For example, the model transformation module can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and the function of the above determined module can be called and executed by a certain processing element of the above device. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together or can be independently implemented. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the processor element or the instruction in the form of software.

[0159] For example, the above modules can be one or more integrated circuits configured to implement the above method, such as: one or more Application Specific Integrated Circuits (ASICs), or one or more Digital Signal Processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a processing element scheduling program code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call program code. Again, these modules can be integrated together and implemented in the form of a System-on-a-chip (SOC).

[0160] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the foregoing method embodiments are generated in whole or in part. The above computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The above computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the above computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wireless (such as infrared, wireless, Bluetooth, microwave, etc.). The above computer-readable storage medium may be any available medium that a computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The above available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0161] Figure 5 FIG. 4 is a schematic structural diagram of an electronic device provided in Embodiment 3 of the present invention. The electronic device may be a terminal device or a server for implementing the method of the foregoing embodiments, or may be a terminal device or a server for implementing the method of the foregoing embodiments connected to the foregoing terminal device or server. As Figure 5 shown, the electronic device may include: a processor 301 (such as a CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiver operations of the transceiver 303. Various instructions may be stored in the memory 302 to complete various processing functions and implement the processing steps described in the foregoing method embodiments. Preferably, the electronic device according to the embodiments of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to implement communication connections between components. The above communication port 306 is used for the electronic device to connect and communicate with other peripherals.

[0162] In Figure 5The system bus 305 mentioned above can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The system bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used to implement communication between the database access device and other devices (such as clients, read-write libraries, and read-only libraries). The memory may include Random Access Memory (RAM), and may also include non-volatile memory, such as at least one disk memory.

[0163] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a Graphics Processing Unit (GPU), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0164] It should be noted that the embodiments of the present invention also provide a computer-readable storage medium, in which instructions are stored. When it runs on a computer, it causes the computer to execute the methods and processing procedures provided in the above embodiments.

[0165] An embodiment of the present invention provides a processing method, apparatus, electronic device, and computer-readable storage medium for a dense encoder implemented based on a large language model. As can be seen from the above, in the embodiment of the present invention, a large language model that has completed pre-training and NLP task fine-tuning and is implemented based on a pure decoder architecture is first selected as the corresponding target model, and the decoder of the target model is used as the corresponding target decoder. A bidirectional encoder is obtained by transforming the causal mask matrix used by the target decoder during the inference process into a matrix of all 1s. A corresponding bidirectional encoding network is formed by sequentially connecting multiple bidirectional encoders, and a dense encoder is formed by connecting the embedding encoding module of the target model and the bidirectional encoding network. Then, a corresponding model training framework is constructed for the dense encoder, and the dense encoder is fine-tuned in the first stage through the masked word prediction task of the model training framework. After the first-stage fine-tuning is completed, the dense encoder is fine-tuned in the second stage through an unsupervised contrast learning mechanism. After the second-stage fine-tuning is completed, the user-specified document library is used as the target document library, and the dense encoder is used to construct a document vector library for the target document library, and a retrieval service is provided for the target document library based on the document vector library and the dense encoder. In the embodiment of the present invention, a dense encoder with bidirectional encoding ability is obtained by transforming the decoder of the target model, and the first-stage fine-tuning ensures that the dense encoder can successfully inherit the ability to capture long-range dependencies, and the second-stage fine-tuning is used to further improve the semantic representation ability of the dense encoder. The dense encoder in the embodiment of the present invention effectively improves the retrieval accuracy and recall rate of the document retrieval task.

[0166] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented in hardware, software modules executed by a processor, or a combination thereof. The software modules may be located in random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0167] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A processing method for a dense encoder implemented based on a large language model, characterized in that, The method includes: Selecting a large language model that has completed pre-training and NLP task fine-tuning and is implemented based on a pure decoder architecture as the corresponding target model; and using the decoder of the target model as the corresponding target decoder; and obtaining a bidirectional encoder by transforming the causal mask matrix used by the target decoder during the inference process into a full 1 matrix; and sequentially connecting multiple bidirectional encoders to form a corresponding bidirectional encoding network, and connecting the embedding encoding module of the target model and the bidirectional encoding network to form a dense encoder; the NLP tasks at least include text generation tasks, information extraction tasks, and question answering tasks; Constructing a corresponding model training framework for the dense encoder; and performing a first-stage fine-tuning on the dense encoder through the masked word prediction task of the model training framework; after the first-stage fine-tuning is completed, performing a second-stage fine-tuning on the dense encoder through an unsupervised contrast learning mechanism; After the second-stage fine-tuning is completed, using the document library specified by the user as the target document library; and using the dense encoder to construct a document vector library for the target document library; and providing a retrieval service for the target document library based on the document vector library and the dense encoder; the target document library includes multiple first documents; the document vector library includes multiple first dense vectors, and the first dense vectors correspond to the first documents one by one.

2. The processing method of the dense encoder implemented based on the large language model according to claim 1, wherein The dense encoder is used to perform dense vector encoding processing on the input text of the encoder and output the corresponding text dense vector Y; The dense encoder is connected by the embedding encoding module and the bidirectional encoding network; The embedding encoding module is used to identify whether the input text contains a preset masked word "[MASK]"; if it does not contain the preset masked word, perform word segmentation processing on the input text according to the word segmentation rules of the target model and according to the preset word list of the target model to obtain the corresponding current word segmentation sequence; if it contains the preset masked word, use each of the preset masked words as a cut-off position to cut the input text into multiple text segments without the preset masked word, and perform word segmentation processing on each text segment according to the word segmentation rules of the target model and according to the preset word list to obtain the corresponding segment word segmentation sequence, and use each of the preset masked words as an independent masked word segmentation, and sort all the segment word segmentations of all the segment word segmentation sequences and all the masked word segmentations in the order of arrangement in the input text to obtain the corresponding current word segmentation sequence; and add a preset start word segmentation "[SOS]" at the start position of the sequence of the current word segmentation sequence, and use the current word segmentation sequence after adding the start word segmentation as the corresponding text word segmentation sequence W; and perform embedding encoding processing on the text word segmentation sequence W according to the embedding encoding rules of the target model to obtain the corresponding text embedding vector X and send it to the bidirectional encoding network; The text token sequence W is composed of multiple tokens w i sorted, where 1 ≤ index i ≤ N W , and N W is the total number of tokens in the text token sequence; the text embedding vector X is composed of N W token embedding vectors x i , and the token embedding vector x i corresponds one-to-one with the token w i ; when the input text has the preset masking words, the total number of the preset masking words is denoted as N M , and the token embedding vector x i corresponding to each preset masking word is denoted as the corresponding masking word embedding vector where 1 ≤ index j ≤ N M ; The bidirectional encoding network consists of N X bidirectional encoders connected in sequence, and the total number of encoders N X is a preset positive integer greater than 1; the bidirectional encoding network is used to continuously bidirectionally encode the text embedding vector X through the N X bidirectional encoders inside to obtain the corresponding text dense vector Y and output it; the text dense vector Y consists of N W token dense vectors y i ; the token dense vector y i corresponds one-to-one with the token embedding vector x i ; when the total number N M is not zero, the token dense vector y corresponding to each masked word embedding vector i is denoted as the corresponding masked word dense vector The model structures of the bidirectional encoder and the target decoder are the same, and both are sequentially connected by a masked multi-head attention layer, a residual connection and layer normalization unit, a feed-forward neural network, and another residual connection and layer normalization unit; The inference process of the masked multi-head attention layer of each bidirectional encoder is as follows: 1 ≤ index n ≤ N X ; X n-1 、X n are the input and output vectors of the n-th bidirectional encoder respectively; d n-1 is the feature dimension of the input vector X n-1 ; Q n 、K n 、V n are the query vector, key vector and value vector corresponding to the n-th bidirectional encoder respectively; are the corresponding weights of the query vector Q n 、the key vector K n and the value vector V n respectively; is the inverted vector of the key vector K n ; M all-ones is a matrix of all 1s; the parameter set G of each bidirectional encoder n is composed of the corresponding weights ; The only difference between the inference process of the bidirectional encoder and that of the target decoder is that the masked multi-head attention layer of the bidirectional encoder uses a matrix M of all 1s during inference. all-ones While the masked multi-head attention layer of the target decoder uses a causal mask matrix M during inference. Causal The causal mask matrix M corresponding to the decoding order of the target model when decoding word by word from left to right Causal is a mask matrix with all 1s in the lower left triangle and all 0s in the upper right triangle. The causal mask matrix M corresponding to the decoding order of the target model when decoding word by word from right to left Causal is a mask matrix with all 1s in the lower right triangle and all 0s in the upper left triangle.

3. The processing method of the dense encoder implemented based on the large language model according to claim 2, wherein The model training framework is used to predict the original word text masked by the preset masked word in the masked text input to the framework and output the corresponding original word prediction sequence; The masked text is a string without punctuation marks and special characters, which contains at least one of the preset masked words and is not entirely composed of the preset masked words; The original word prediction sequence is composed of one or more predicted original word texts; the predicted original word texts in the original word prediction sequence correspond one by one to the preset masked words in the masked text; the order of all the predicted original word texts in the original word prediction sequence matches the order of all the preset masked words in the masked text; The model training framework is sequentially connected by the dense encoder, a linear layer, a Softmax function layer, and an output module; The dense encoder is used to take the masked text as the corresponding input text, perform dense vector encoding processing on the input text, and send the corresponding text dense vector Y to the linear layer; The linear layer is used to perform a vocabulary feature vector conversion on each of the token dense vectors y in the text dense vector Y i to obtain the corresponding token feature vector h i ; and all the obtained token feature vectors h i are sorted in sequence to form a vector sequence H and sent to the Softmax function layer; the token feature vector h i corresponds one-to-one with the token dense vector y i ; and the previous token feature vector h corresponding to each masked word dense vector i is denoted as the corresponding previous word feature vector i-1 ​ The Softmax function layer is used to calculate the probability distribution of the vocabulary according to each of the previous word feature vectors and use the calculation result as the corresponding masked word probability vector All the obtained masked word probability vectors are sorted in order to form a vector sequence P and sent to the output module; each masked word probability vector consists of multiple token probabilities, and the token probabilities correspond one by one to the tokens in the preset vocabulary; The output module is used to take the word segmentation in the table corresponding to the maximum word segmentation probability of each of the shielding word probability vectors as the corresponding predicted original word text; and the obtained predicted original word texts are sorted in order to form the corresponding original word prediction sequence and output it.

4. The processing method of the dense encoder implemented based on the large language model according to claim 3, wherein The one-stage fine-tuning of the dense encoder through the masked word prediction task of the model training framework specifically includes: Step 41, randomly extract multiple sentences from the historical texts processed by the target model, and use each extracted sentence as a corresponding first extracted text; and delete the punctuation marks and special characters in each first extracted text; Step 42, implant a low-rank matrix into the parameter set G of each of the bidirectional encoders of the model training framework n to obtain a corresponding parameter set Among them, the parameter set consists of weights ; The masked multi-head attention layer of the bidirectional encoder is based on the parameter set The inference process is as follows: for the weight a corresponding pair of low-rank matrices for the weight a corresponding pair of low-rank matrices for the weight a corresponding pair of low-rank matrices; the parameter set the corresponding low-rank matrix forms a corresponding low-rank matrix parameter set Step 43, initialize the current masking rate to a preset first masking rate; Wherein, 0 < first masking rate < 1; Step 44, perform pre-segmentation on each of the first extraction texts based on the word segmentation rules of the target model and according to the preset word list to obtain corresponding pre-segmented sequences; randomly replace the word segments of each of the pre-segmented sequences according to the current masking rate using the preset masking words to obtain corresponding replaced word segment sequences; perform string concatenation on all the word segments of each of the replaced word segment sequences to obtain corresponding first training texts; and count the total number of the preset masking words in all the first training texts to obtain the corresponding total number N K ; Wherein, the sequence segmentation of the replaced segmented sequence contains the preset masked word, and the total number of the preset masked words is the downward rounded value of the product of the total number of sequence segmentations and the current masking rate; Step 45, use each of the first training texts as the corresponding masked text to input into the model training framework for prediction processing, and use each masked word probability vector calculated by the Softmax function layer in this processing as a corresponding prediction vector p k , where 1 ≤ index k ≤ N K ; and based on each prediction vector p k set a corresponding label vector for the corresponding original word text and form a corresponding first prediction-label pair from each prediction vector p k and the corresponding label vector ​ Among them, the label vector is composed of multiple label word segmentation probabilities, and the label word segmentation probabilities correspond one-to-one to the in-table word segmentations of the preset word list; the label vector has only one of the label word segmentation probabilities being 1, and the remaining label word segmentation probabilities are all 0; the label vector the in-table word segmentation corresponding to the label word segmentation probability that is 1 in it matches the corresponding original word text; Step 46: Take the obtained N K first prediction-label pairs and calculate the corresponding first loss value by bringing them into a preset first model loss function L1; Wherein, the first model loss function L1 is: Step 47, identify whether the first loss value meets a preset first loss value range; if it meets, go to Step 48; if it does not meet, identify whether the current occlusion rate is the first occlusion rate, and if so, based on a preset first model optimizer, modulate all the low-rank matrix parameter sets of the bidirectional encoding network in the direction of minimizing the first model loss function L1 and the linear layer parameters of the linear layer for one round of modulation, otherwise, based on a preset second model optimizer, modulate all the low-rank matrix parameter sets for one round of modulation, and after the end of this round of modulation, return to Step 45 to continue training; Wherein, both the first and second model optimizers at least include an Adam optimizer and an SGD optimizer; Step 48, identify whether the current masking rate is the first masking rate; if so, reset the current masking rate to a preset second masking rate and return to step 44 to continue training after resetting; if not, stop training and confirm the end of one-stage fine-tuning; Wherein, 0 < first masking rate < second masking rate < 1.

5. The processing method of the dense encoder implemented based on the large language model according to claim 2, wherein, The two-stage fine-tuning of the dense encoder through an unsupervised contrast learning mechanism specifically includes: Step 51, randomly extract multiple sentences from the historical text processed by the target model, and use each extracted sentence as a corresponding second extracted text; delete punctuation marks and special characters in each of the second extracted texts; and count the total number of the second extracted texts to obtain a corresponding total number N U ; Step 52, use each second extracted text as the corresponding input text to input into the dense encoder for dense vector encoding processing, and record the text dense vector Y output by this processing as the corresponding first sample; Step 53: Take each of the second extracted texts as the corresponding input text and input it into the dense encoder again for dense vector encoding processing. During this processing, randomly select multiple residual connection and layer normalization units from the bidirectional encoding network as corresponding random scrambling units. At the end of the inference process of each random scrambling unit, randomly scramble some sub-vectors of the current unit output vector, and pass the scrambled unit output vector downstream; and denote the text dense vector Y finally output by this dense vector encoding processing as the corresponding second sample; Among them, the scrambling method of the random scrambling is: set the selected sub-vector to a zero vector or add Gaussian noise to the selected sub-vector; Step 54, denote any one of the first samples as the corresponding sample s u , where 1 ≤ index u ≤ N U ; and denote the other first samples different from each sample s u as the negative samples of the current sample s u where 1 ≤ index v ≤ N - 1; and form a corresponding negative sample set from the N U - 1 negative samples corresponding to each sample s u U And denote the second samples corresponding to each sample s u as the positive samples of the current sample s u And form a corresponding sample data group from each sample s u and its corresponding positive samples and the negative sample set Step 55, bringing the obtained N U sample data groups into the preset second model loss function L2 for calculation to obtain the corresponding second loss value; Among them, the second model loss function L2 is: λ is a preset adjustment parameter; sim() is the cosine similarity function of vectors, for the sample s u and the corresponding positive sample cosine similarity, for the sample s u and the corresponding negative sample cosine similarity; Step 56, identify whether the second loss value meets a preset second loss value range; if not, based on a preset third model optimizer, perform one round of modulation on all low-rank matrix parameter sets of the dense encoder in the direction of minimizing the second model loss function L2, and return to step 52 to continue training after the end of this round of modulation; if it meets, stop training and confirm the end of the second-stage fine-tuning; and if it meets, stop training and confirm the end of the second-stage fine-tuning; The third model optimizer includes at least Adam optimizer and SGD optimizer.

6. The processing method of the dense encoder implemented based on the large language model according to claim 1, wherein, The use of the dense encoder to construct a document vector library for the target document library specifically includes: Take each of the first documents in the target document library as the corresponding input text and input it into the dense encoder for dense vector encoding processing, and take the text dense vector output by the current processing as a corresponding first dense vector; and form the corresponding document vector library from all the obtained first dense vectors.

7. The processing method of the dense encoder implemented based on the large language model according to claim 1, characterized in that, The providing of retrieval services for the target document library based on the document vector library and the dense encoder specifically includes: Take the query text input by the user each time as the corresponding current query text; and take the current query text as the corresponding input text and input it into the dense encoder for dense vector encoding processing, and take the text dense vector output by the current processing as the corresponding current dense vector; and calculate the cosine similarity between the current dense vector and each of the first dense vectors in the document vector library to obtain the corresponding first similarity; and denote the first dense vectors whose first similarity exceeds the preset similarity threshold as the corresponding retrieval vectors; and when the total number of the retrieval vectors is not zero, denote the first documents in the target document library corresponding to each of the retrieval vectors as the corresponding retrieval documents; and sort all the obtained retrieval documents in descending order of the first similarity to form the corresponding retrieval document sequence and feedback it to the current user.

8. An apparatus for performing the processing method of the dense encoder implemented based on the large language model according to any one of claims 1-7, characterized in that, The device includes: a model transformation module, a model fine-tuning module, and a retrieval application module; The model transformation module is used to select a large language model that has completed pre-training and NLP task fine-tuning and is implemented based on a pure decoder architecture as the corresponding target model; and take the decoder of the target model as the corresponding target decoder; and obtain a bidirectional encoder through the transformation method of solidifying the causal mask matrix used by the target decoder during the inference process into a full 1 matrix; and sequentially connect multiple bidirectional encoders to form a corresponding bidirectional encoding network, and connect the embedding encoding module of the target model and the bidirectional encoding network to form a dense encoder; the NLP tasks include at least text generation tasks, information extraction tasks, and question answering tasks; The model fine-tuning module is used to construct a corresponding model training framework for the dense encoder; and perform a first-stage fine-tuning on the dense encoder through the masked word prediction task of the model training framework; after the first-stage fine-tuning is completed, perform a second-stage fine-tuning on the dense encoder through an unsupervised contrastive learning mechanism; The retrieval application module is used to, after the second-stage fine-tuning is completed, use the document library specified by the user as the target document library; and use the dense encoder to construct a document vector library for the target document library; and provide a retrieval service for the target document library based on the document vector library and the dense encoder; the target document library includes multiple first documents; the document vector library includes multiple first dense vectors, and the first dense vectors correspond to the first documents one by one.

9. An electronic device, characterized in that, Including: A memory, a processor, and a transceiver; The processor is used to be coupled with the memory, read and execute the instructions in the memory to implement the method according to any one of claims 1-7; The transceiver is coupled with the processor, and the processor controls the transceiver to perform message sending and receiving.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computer, the computer is caused to execute the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Document retrieval method based on dense pseudo query vector representation

    CN112732864A

  • Information retrieval method and device based on text extension, electronic equipment and medium

    CN117992573A

  • End-to-end image description generation method based on attention enhancement mechanism

    CN118736575A

  • Large language model processing method and device introducing dense vector retriever

    CN119398193A

  • Method and device for processing large language model with introduction of conditional constraints

    CN119476208A

Cited By

  • Large language model reasoning method, device, equipment, medium and program product

    CN121615762A