Multi-text data processing method for fine tuning, fine tuning method and electronic equipment

By concatenating prompts with multiple text data and setting a causal attention mechanism, the problem of wasted computational resources in fine-tuning text search models is solved, achieving resource conservation and improved model efficiency. This method is applicable to text retrieval and code retrieval in multiple scenarios.

CN122019752APending Publication Date: 2026-05-12BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2026-01-23
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing text search models consume significant computational resources during fine-tuning, and known optimization techniques are not very effective.

Method used

A multi-text data processing method is adopted, which concatenates the prompt word with multiple text data, adds preset symbols to form an overall concatenation structure, and sets a causal attention mechanism in the model so that each text only focuses on itself and the prompt word, reducing the waste of computing resources.

Benefits of technology

By extracting features from multiple prompt words and text pairs through a single model inference, computational resources are saved and the efficiency of model training and inference is improved. It is especially suitable for multi-scenario text retrieval and code retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122019752A_ABST
    Figure CN122019752A_ABST
Patent Text Reader

Abstract

The invention provides a multi-text data processing method for fine tuning, a fine tuning method and electronic equipment, and relates to the field of artificial intelligence model training, in particular to the field of text search and computing power resource optimization. According to the specific implementation scheme, cue words and multiple pieces of text data are obtained, wherein the multiple pieces of text data are ranked according to a preset mode or ranked randomly; splicing the cue word and the multiple pieces of text data, adding a preset symbol for the cue word during splicing, adding a preset symbol at the tail of the multiple pieces of text data, and adding a preset symbol between at least two adjacent pieces of text data, the data obtained after splicing processing is used for conducting fine adjustment on the large model to generate a text search model. According to the invention, computing resources can be saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence model training technology, as well as the field of text search and computing resource optimization technology, specifically to multi-text data processing methods, large model fine-tuning methods, devices, electronic devices, and program products for fine-tuning. Background Technology

[0002] By fine-tuning a pre-trained large model using labeled prompts and corresponding text data, a text search model (or text retrieval model) can be obtained. However, such fine-tuning currently suffers from high computational resource consumption, and known optimization techniques are not very effective. Summary of the Invention

[0003] This disclosure provides a method for fine-tuning multi-text data processing, a fine-tuning method, an electronic device, a storage medium, and a program product.

[0004] According to one aspect of this disclosure, a multi-text data processing method for fine-tuning is provided, comprising: acquiring prompt words and multiple text data, wherein the multiple text data are sorted or randomly sorted according to a predetermined method; concatenating the prompt words and the multiple text data, wherein during concatenation, a preset symbol is added to the prompt words, a preset symbol is added to the end of the multiple text data, and a preset symbol is added between at least two adjacent text data, and the data obtained after concatenation is used to fine-tune a large model to generate a text search model.

[0005] According to another aspect of this disclosure, a method for fine-tuning a large model is provided, comprising: inputting preprocessed cue words and multiple text data into a large model, wherein the preprocessing is performed according to the multi-text data processing method for fine-tuning as described above; wherein the attention mechanism of the large model is configured to conform to a causal attention mechanism and, when performing attention calculation, a single text only focuses on itself and the cue words.

[0006] According to another aspect of this disclosure, a multi-text data processing apparatus for fine-tuning is provided, comprising: an acquisition module for acquiring prompt words and multiple text data, wherein the multiple text data are sorted or randomly ordered according to a predetermined method; and a splicing module for splicing the prompt words and the multiple text data, wherein during splicing, the splicing module adds a preset symbol to the prompt words, adds a preset symbol to the end of the multiple text data, and adds a preset symbol between at least two adjacent text data, and the data obtained after splicing is used to fine-tune a large model to generate a text search model.

[0007] According to another aspect of this disclosure, a large model fine-tuning apparatus is provided, comprising: an input module for inputting preprocessed prompt words and multiple text data into a large model, wherein the preprocessing is performed according to the multi-text data processing method for fine-tuning as described above; wherein the attention mechanism of the large model is configured to conform to a causal attention mechanism and, when performing attention calculation, a single text only focuses on itself and the prompt words.

[0008] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a multi-text data processing method or fine-tuning method for fine-tuning as described above.

[0009] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the multi-text data processing method or fine-tuning method for fine-tuning as described above.

[0010] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the multi-text data processing method or fine-tuning method for fine-tuning as described above.

[0011] This disclosure proposes a novel data concatenation method that allows for multi-text fusion processing. Using this method, a novel concatenation structure of prompt words and multi-text data can be constructed, which can be used to fine-tune large models to generate text retrieval models. Based on this concatenation structure, features of multiple prompt word-text pairs can be extracted through a single model inference, eliminating the need to repeatedly initiate feature extraction for the same prompt word, thus saving computational resources. This disclosure is particularly suitable for multi-scenario text retrieval and code retrieval.

[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0013] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 This is a schematic diagram of a multi-text data processing method for fine-tuning according to a first embodiment of the present disclosure; Figure 2 This is a schematic diagram illustrating the concatenation of prompt words and multiple text data according to the first embodiment of this disclosure; Figure 3 This is a schematic diagram of a fine-tuning model of the text retrieval model according to the first embodiment of this disclosure; Figure 4 This is a schematic diagram of a large model fine-tuning method according to the second embodiment of this disclosure; Figure 5 This is a schematic diagram of a fine-tuning model of the text retrieval model according to the second embodiment of this disclosure; Figure 6 yes Figure 5 A partial schematic diagram within the dashed box; Figure 7 This is a schematic diagram of a multi-text data processing apparatus for fine-tuning according to an embodiment of the present disclosure; Figure 8 This is a schematic diagram of a large model fine-tuning device according to an embodiment of the present disclosure; Figure 9 This is a block diagram of an electronic device used to implement the multi-text data processing method or fine-tuning method for fine-tuning embodiments of the present disclosure. Detailed Implementation

[0014] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0015] The terms used in the text, such as first, second, third, etc., are only used to distinguish one entity (or operation) from another entity (or operation), and are not intended to require or imply a specific order or relationship between these entities (or operations).

[0016] Before describing the contents of this disclosure, a brief explanation of the general meanings or common operations of the technical terms that may be involved in this disclosure will be given first.

[0017] Large Language Model (LLM): The hidden layer feature extraction capability of LLM can be used to obtain the hidden layer features of the data.

[0018] Multilayer Perception Module (MLP): MLP can use multiple network layers to compress features and use the sigmoid activation function to obtain the similarity probability (or similarity value, relevance, relevance score) of text.

[0019] Loss module: The contrastive learning loss function is used to calculate the loss value of the similarity probability of positive and negative samples, and to train the model. There are various contrastive learning loss functions, and you can choose to use them according to the needs of the model or data.

[0020] The prompts for the text search model can be short words, phrases, code snippets, or long sentences, paragraphs, entire books, code sets, etc. When fine-tuning a large model using labeled data (at least one prompt word query and multiple text documents), to extract the hidden layer features of the query and each document to calculate similarity, the fine-tuning structure uses query + document followed by a token to form a "query-document pair." The number of "query-document pairs" equals the number of input text data. Then, a pre-trained large model, such as a large language model (LLM), extracts hidden layer features, where the LLM executes a causal attention mechanism. Afterward, a multilayer perceptron (MLP) layer calculates the similarity value between the query and document, which is used for loss function parameter tuning. This process is repeated for each "query-document pair" until the desired fine-tuning effect is achieved.

[0021] In other words, the fine-tuning process of existing text search models requires repeatedly starting the encoding and calculation process of LLM feature extraction and similarity calculation between the same query and different documents.

[0022] This research has revealed the following problem: the same query undergoes multiple encoding calculations as described above, and each calculation is concatenated with different documents without interference. Therefore, for the same query, each calculation requires LLM to extract hidden layer features, resulting in a waste of computational resources.

[0023] This disclosure provides a method for fine-tuning multi-text data processing, referencing... Figure 1 This includes the following processing: S101, Obtain prompt words and multiple text data, wherein the multiple text data are sorted in a predetermined manner or randomly; S102, the prompt word and the plurality of text data are concatenated. During concatenation, a preset symbol is added to the prompt word, a preset symbol is added to the end of the plurality of text data, and a preset symbol is added between at least two adjacent text data. The data obtained after concatenation is used to fine-tune the large model to generate a text search model.

[0024] This disclosure provides an implementation method that allows for the fusion processing of features from multiple text data. Unlike the previous method of independently concatenating a prompt word with a piece of text data, this disclosure concatenates a prompt word with multiple pieces of text data together. In addition to adding a predetermined symbol at the end, predetermined symbols are also added between the prompt word and the text data, as well as between two adjacent pieces of text data, forming a concatenation structure of a prompt word and multiple pieces of text data. This concatenation method enables the features of multiple text data to be fused and processed in the model.

[0025] The embodiments of this disclosure can be used to construct a novel concatenation structure of prompt words and multiple text data, which can be used to fine-tune large models to generate text retrieval models. Based on the concatenation structure of this disclosure, the features of prompt words and multiple text data can be extracted in a single model inference. For the same prompt word, it is no longer necessary to repeatedly start model inference to extract features, thus saving computing resources and improving the overall model training and inference efficiency.

[0026] The predetermined symbols in this embodiment can take various forms, such as separators, custom terms, custom numbers, tokens from a known token set, tokens outside the token set, etc., and can be handled by large models.

[0027] Optionally, the number of symbols added during splicing can be determined as needed. It is preferable to add symbols between every two adjacent text data to enable the large model to accurately distinguish between prompt words and different text data, resulting in more accurate inference results.

[0028] As an example, Figure 2 The illustration schematically shows the concatenation structure of a prompt word and multiple text data in an embodiment of this disclosure, wherein one prompt word "query" is concatenated with n text data "document_1", "document_2", ..., "document_n". During concatenation, a custom symbol is added after "query". Figure 2 In addition to adding a special token (referred to as a special token) to the end of all text data, a special token is also added between every pair of adjacent documents, forming a concatenated structure of 1 query and n documents, that is: query+specialtoken+document_1+specialtoken+document_2+specialtoken+...+specialtoken+document_n+specialtoken. The above is an illustrative representation of... Figure 2There are no special restrictions on the form of the token, for example, the token can be number 10 or 100.

[0029] There are no special requirements for the initial order of multiple text data. They can be sorted according to the labeled attributes or randomly sorted. After the model reasoning of this embodiment, a specific sorting will be formed.

[0030] During subsequent processing, it can be Figure 2 The spliced ​​structure formed in the example is input into a pre-trained large model such as an LLM for feature extraction. Figure 3 The illustration shows the basis Figure 2 A schematic diagram of the fine-tuning model structure of the text retrieval model in the embodiment, according to Figure 3 By constructing a fine-tuned model and performing feature fusion, similarity calculation, and loss calculation, fine-tuning training can be completed.

[0031] Furthermore, the attention mechanism of the large model can be configured according to the actual situation, so that the model only focuses on the specified data and ignores the data that it does not want to focus on, which can further reduce the amount of computation.

[0032] refer to Figure 4 This disclosure also provides a method for fine-tuning large models, including: S201, the preprocessed prompt words and multiple text data are input into the large model, preferably using the aforementioned multi-text data processing method for preprocessing; wherein, the attention mechanism of the large model is set to conform to the causal attention mechanism and when performing attention calculation, each text only focuses on itself and the prompt words.

[0033] For example, during fine-tuning, this disclosure will be followed. Figure 1The concatenated structure of prompt words and multi-text data generated in the embodiment is input into a large model, and special customized settings are added to the attention mechanism. Specifically, it first conforms to the causal attention mechanism, and further, it ensures that each text only focuses on its own information and the prompt word. That is, when calculating attention, each text can "see" itself and the prompt word, while it cannot "see" any other text input at the same time. The purpose of this setting is not only to reduce the amount of computation, but more importantly, to shield as much as possible the interference that may be introduced into the model when multiple texts are concatenated and input at the same time. Therefore, although the fine-tuning method of this disclosure uses the newly proposed concatenated structure of prompt words and multi-text data as training data, it will not cause mutual interference between multi-text data during model inference calculation, and will not cause significant deviations in the model's vector feature extraction and similarity score calculation. The inference process and results are controllable. Therefore, the large model fine-tuning method of this embodiment can save computing resources due to the use of the new concatenated structure of prompt words and multi-text data, and further reduces the amount of model computation due to the setting of a suitable attention mechanism, which can further save computing resources and will not lead to a deterioration in model training and inference performance.

[0034] Optionally, in some embodiments, the attention mechanism of the large model can be further configured so that the cue word is only visible to itself and the current text when performing attention calculations, that is: ① When calculating attention, the current text can see the cue words, and the cue words can also see themselves; ② It should also be noted that when splicing, the prompt word comes first and the text data comes later. Because it conforms to causality, the prompt word (which comes first) cannot see the text data that comes later.

[0035] The reason for this setting is that this disclosure aims to extract features of cue words and multiple text data in a single LLM inference. The features of the same cue word are then fused with the features of different text data. During the encoding calculation process of feature extraction, according to point ①, the current text can see the cue word, and the cue word can see itself. At the same time, as mentioned in S201, the current text can see itself, which is beneficial for extracting the correlation features between the cue word and the current text for subsequent feature fusion. Meanwhile, according to point ②, the cue word cannot see the current text data. The reason for this setting is that this is the original setting in the current pre-trained large model self-attention mechanism. To avoid other unknown problems, the embodiments of this disclosure can preferably use this setting to improve the reliability of the optimization in this stage.

[0036] The above provides a detailed explanation of the implementation process of the text data processing method and the large model fine-tuning method proposed in this disclosure. Using this disclosure, it can be applied to suitable models and scenarios according to needs, such as various text retrieval models, especially code retrieval models. By adopting implementation methods adapted to this disclosure, the goal of saving computing resources can be achieved.

[0037] Several alternative implementation methods are provided below.

[0038] Optionally, the concatenated prompt words and multiple text data are input into a large model. Taking LLM as an example, LLM extracts hidden state features, obtaining hidden state features for each data item, including: hidden state features of the prompt word's token, hidden state features of the tokens for each text data item, and hidden state features of each special token. That is, the hidden state features of all these items can be extracted in a single inference. LLM does not need to repeatedly calculate for multiple "prompt word-text pairs" (query-document pairs) or repeatedly extract hidden state features for the same query.

[0039] Optionally, based on the foregoing limitations, embodiments of this disclosure can further optimize the attention mechanism of the large model. Therefore, an attention mask is constructed according to an optimizable attention mechanism. During fine-tuning, the concatenated cue words and multiple text data are fed into the LLM along with the custom attention mask. As an example, the custom attention mask used in embodiments of this disclosure can be schematically represented as follows:

[0040] Where Q represents the cue word matrix, D represents the text data matrix, and the subscripts are... i Indicates the location of the token being used to calculate attention, the indices. j The indicator shows the location of the viewed token; the subscripts m and k represent the number of the text data. The first row contains the causal mask, meaning that the query token cannot see the subsequent (or future) documents (causality) when calculating attention. The second line indicates that the query token can see itself and can be seen by the document when calculating attention (query visible). The third and fourth lines represent the document being visible to itself (visible to the same text) and not visible to other texts (isolated across texts), respectively. In other words, when calculating attention, the document's token can see itself but not other documents. In summary, the above attention masks M i,j This mechanism simultaneously achieves the following when calculating attention: "a single document only focuses on itself and the query" and "the query is only visible to itself and the current document." Under this attention mechanism, on the one hand, the query's token can see itself when calculating attention, allowing the query to compare itself with a document when it sees it. On the other hand, each document's token can see itself and the query's token when calculating attention, without seeing other documents. This avoids interference between different document contents and prevents adverse effects on feature extraction caused by multi-text concatenation and fusion processing. It saves computational power without compromising the overall training and inference performance of the model.

[0041] Optionally, in one embodiment of this disclosure, the large model is coupled to a preset attention pooling module. The hidden layer features of the token of the prompt word, the hidden layer features of the special token following the prompt word, the hidden layer features of the token of each text data, and the hidden layer features of the special token following the first text data are extracted by the attention pooling module LLM and weighted fusion calculation is performed to obtain the fusion features of the prompt word and the query-document pair of each text data respectively.

[0042] Optionally, when calculating the fusion features of the prompt word and different text data, the hidden layer features of the prompt word's token are reused, and / or, the hidden layer features of the special token following the prompt word are reused. Here, "following" indicates what follows, and the special token following the prompt word is the same as the special token that follows the prompt word.

[0043] During computation, the attention pooling module (i.e., attention pooling layer) in this embodiment of the disclosure performs weighted fusion of the same prompt word with each text data from multiple text datasets. For n text data, n fused features can be obtained. Therefore, the hidden layer features of the prompt word's token can be reused without repeated extraction, and the hidden layer features of the special token following the prompt word can also be reused without repeated extraction. This saves computational resources, conserves token resources, improves model training and inference efficiency, and does not affect training and inference performance.

[0044] Thus, this disclosure can obtain weighted fusion features of n query-document pairs in a single fine-tuning inference. The weighted fusion features are matrices (or vectors). After the aforementioned calculation and processing, the matrix (or vector) corresponding to each document carries the deep meaning information between the query and the current document, which is used to calculate the relevance score.

[0045] Optionally, the relevance score between the prompt words and each piece of text data can be calculated based on the weighted fusion features of the prompt words and each piece of text data; the text data can be sorted according to the size of the corresponding relevance score; the ranking loss function can be used to tune the parameters, and when the ranking output by the model meets the preset requirements with the ranking of the annotation, the fine-tuning training can be stopped to obtain the text search model.

[0046] The various embodiments of this disclosure and their advantages have been described in detail above. The specific processing steps of the embodiments of this disclosure are described below through specific examples.

[0047] As an example, Figure 5 The text retrieval model fine-tuning structure of this embodiment is illustrated schematically. During fine-tuning, the training data is adjusted according to... Figure 2 The example performs concatenation processing to obtain a concatenated structure of query, n documents, and n+1 special tokens. This structure, along with a custom attention mask, is fed into an LLM for hidden layer feature extraction, yielding hidden states of all tokens (including query tokens, tokens of each document, and each special token). The hidden states of the query and each document are collected (i.e., the hidden states of query tokens and the hidden states of tokens of each document during computation). Weighted fusion calculation is performed through an attention pool layer to obtain the fusion weights of each hidden state, generating a weighted fusion feature. In this example, it is a fusion feature of n query + document pairs.

[0048] Figure 6 schematically shown Figure 5The attention pool within the dashed box processes the query and the first text data, document_1, from multiple text datasets. Optionally, the attention pool can use a fully connected layer to calculate weights for the hidden state of the query, the hidden state of the special token following the query, the hidden state of document_1, and the hidden state of the special token following document_1. After normalization, these weights are summed and fused to obtain the fused features of the query and document_1. Figure 5 and Figure 6 The query+document_1fusion feature shown.

[0049] For query and document_2, execute as follows Figure 6 A similar process is used to perform a weighted fusion of the hidden state of the query, the hidden state of document_2, and the hidden state of the special tokens following the query and document_2, generating a fusion feature of the query and document_2 (query+document_2 fusionfeature).

[0050] Through the above processing, it can be clearly seen that for different documents, the hidden state of the query token can be reused during weighted fusion calculation, as can the hidden state of the special token after the query. When fine-tuning training, there are many texts. Following the above processing can greatly reduce the waste of token resources and save computing resources.

[0051] For the weighted fusion features of n query-document pairs, refer to Figure 5As an example, the fusion features of n query-document pairs can be input into an MLP layer for computation, yielding n relevance values ​​(representing the relevance or similarity between the query and the current document). These n relevance values ​​are then sorted and compared to the labeled text ranking. A ranking loss function, such as ListMLE, is used to adjust and optimize the model parameters until a training stopping condition is met (e.g., the text ranking obtained by the model inference is similar to, partially consistent with, or completely consistent with the labeled text ranking). After fine-tuning, a text search model is obtained, which can recall multiple text data with similar meanings (or relevance) based on a given input query and rank them according to relevance.

[0052] This disclosure provides a multi-text data processing method and a fine-tuning method for fine-tuning. Using various embodiments of this disclosure, a concatenated structure of query and multi-text data can be constructed as fine-tuning data for a text search model, which can reduce repetitive feature extraction operations and save computing resources. Furthermore, based on the above concatenated structure, this disclosure also optimizes the overall structure of the fine-tuning model. By refining the calculation of the attention mechanism and attention pooling layer, a complete set of multi-text feature fusion fine-tuning models for text search is constructed, which can achieve the purpose of saving computing power and reducing token resource waste, without significantly affecting the model inference process and training effect.

[0053] The following simplified illustrative example illustrates the advantages of this disclosure's embodiments in terms of model inference computing power and token resource consumption: ① Data to be processed: 1 query and 10 documents; ② Required computing power and token resources: Assuming that the query itself requires 30 tokens of computation, and each of the 10 documents requires an average of 70 tokens of computation; ③ Previous approach: 1 query is paired with 1 document, forming a total of 10 independent query-document pairs. The same query needs to be repeatedly started 10 times for feature extraction, weighted fusion and relevance calculation, which requires at least (30+70)×10=1000 tokens of computation. ④ The processing method of at least one embodiment of the present disclosure is adopted: 1 query and 10 documents are concatenated together. The features of the 1 query and 10 documents can be extracted in a single LLM inference. The similarity of the 10 query-document pairs is calculated during weighted fusion. The features of the query token and the following special token can be reused during fusion. The total computation of about 700 tokens is consumed. ⑤ Analysis conclusion: Compared with previous solutions, the present invention can save at least 30% of token computing resources, reduce the number of LLM inference startups from 10 to 1, and reduce the complexity of model inference.

[0054] The embodiments of this disclosure propose an optimization idea for a multi-text feature fusion structure for text retrieval based on a large language model. It incorporates an isolated attention mask mechanism and a multi-text isolated feature fusion mechanism, which enables the extraction of features from multiple document-query pairs in a single LLM inference, thereby reducing the number of inference initiations and the amount of token wasted. This results in better training and inference performance, improving the efficiency of training and inference, and is particularly suitable for code retrieval applications.

[0055] According to embodiments of this disclosure, this disclosure also provides a multi-text data processing apparatus 100 for fine-tuning, see reference to Figure 7 It includes: The acquisition module 110 is used to acquire prompt words and multiple text data, wherein the multiple text data are sorted in a predetermined manner or randomly. The splicing module 120 is used to splice the prompt words and the multiple text data. During splicing, the splicing module adds a preset symbol to the prompt words, adds a preset symbol to the end of the multiple text data, and adds a preset symbol between at least two adjacent text data. The data obtained after splicing is used to fine-tune the large model to generate a text search model.

[0056] According to embodiments of this disclosure, this disclosure also provides a large model fine-tuning device 200, with reference to... Figure 8 It includes: Input module 210 is used to input preprocessed prompt words and multiple text data into a large model, wherein the preprocessing is performed according to the multi-text data processing method for fine-tuning as described above; wherein the attention mechanism of the large model is set to conform to the causal attention mechanism and when performing attention calculation, a single text only focuses on itself and the prompt words.

[0057] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0058] Figure 9 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0059] like Figure 9 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.

[0060] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0061] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as a multi-text data processing method or fine-tuning method for fine-tuning. For example, in some embodiments, the multi-text data processing method or fine-tuning method for fine-tuning may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the multi-text data processing method or fine-tuning method for fine-tuning described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured by any other suitable means (e.g., by means of firmware) to perform a multi-text data processing method or a fine-tuning method for fine-tuning.

[0062] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0063] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0064] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0065] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0066] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0067] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0068] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0069] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for fine-tuning multi-text data processing, comprising: Obtain prompt words and multiple text data, wherein the multiple text data are sorted in a predetermined manner or randomly; The prompt word and the multiple text data are concatenated. During concatenation, a preset symbol is added to the prompt word, a preset symbol is added to the end of the multiple text data, and a preset symbol is added between at least two adjacent text data. The data obtained after concatenation is used to fine-tune the large model to generate a text search model.

2. The method according to claim 1, wherein, Adding a preset symbol between at least two adjacent text data includes adding a preset symbol between every two adjacent text data.

3. A method for fine-tuning a large model, comprising: The preprocessed prompt words and multiple text data are input into the large model, wherein the preprocessing is performed according to the method of claim 1 or 2; The attention mechanism of the large model is configured to conform to the causal attention mechanism, and when performing attention calculation, a single text only focuses on itself and the prompt word.

4. The method according to claim 3, wherein, The attention mechanism of the large model is further configured so that when performing attention calculations, the cue words are only visible to themselves and the current text.

5. The method according to claim 3 or 4, wherein, The preset symbol uses a custom special token; The method further includes: extracting hidden layer features from the input prompt words and multiple text data using the large model to obtain the hidden layer features of the prompt word token, the hidden layer features of the tokens of each text data, and the hidden layer features of each special token.

6. The method according to claim 5, wherein, The large model is coupled to an attention pooling module; The method further includes: performing weighted fusion calculation on the hidden layer features of the prompt word's token, the hidden layer features of each text data's token, and the hidden layer features of each special token through the attention pooling module, so as to obtain the fusion features of the prompt word and each text data respectively.

7. The method according to claim 6, wherein, The plurality of text data includes the first text data; The weighted fusion calculation includes: the attention pooling module performing weight calculations on the hidden layer features of the prompt word's token, the hidden layer features of the special token following the prompt word, the hidden layer features of the first text data's token, and the hidden layer features of the special token following the first text data; normalizing the calculated weights and then performing weighted summation calculations to obtain the fusion features of the prompt word and the first text data.

8. The method according to claim 6 or 7, wherein, When calculating the fusion features of the prompt word with different text data, the hidden layer features of the prompt word's token are reused, and / or, the hidden layer features of the special token that follows the prompt word are reused.

9. The method according to any one of claims 6-8, further comprising: Based on the weighted fusion features of the prompt words and each piece of text data, the relevance scores between the prompt words and each piece of text data are calculated. Sort each text data according to its corresponding relevance score; By using the ranking loss function for parameter tuning, the fine-tuning training stops when the ranking output by the model meets the preset requirements with the ranking of the labeled data, thus obtaining the text search model.

10. The method according to any one of claims 3-9, wherein, The formalized expression of the attention mask for the large model is as follows: in, M i,j This represents the attention mask, Q represents the cue word matrix, D represents the text matrix, and the subscripts are... i Indicates the location of the token being used to calculate attention, the indices. j The indicator shows the location of the token being viewed; the subscripts m and k represent the number of the text data.

11. A multi-text data processing apparatus for fine-tuning, comprising: The acquisition module is used to acquire prompt words and multiple text data, wherein the multiple text data are sorted in a predetermined manner or randomly. The concatenation module is used to concatenate the prompt words and the multiple text data. During concatenation, the concatenation module adds a preset symbol to the prompt words, adds a preset symbol to the end of the multiple text data, and adds a preset symbol between at least two adjacent text data. The data obtained after concatenation is used to fine-tune the large model to generate a text search model.

12. A large model fine-tuning device, comprising: The input module is used to input preprocessed prompt words and multiple text data into the large model, wherein the preprocessing is performed according to the method of claim 1 or 2; The attention mechanism of the large model is configured to conform to the causal attention mechanism, and when performing attention calculation, a single text only focuses on itself and the prompt word.

13. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method of any one of claims 1-10.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-10.

15. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-10.