Text retrieval method and apparatus, model training method and apparatus, and vector compression method and apparatus
Patent Information
- Application Number
- PCT/CN2026/084263
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-27
- Filing Date
- 2026-03-18
- Publication Date
- 2026-10-01
Smart Images

Figure CN2026084263_01102026_PF_FP_ABST
Abstract
Description
A method and apparatus for text retrieval, model training, and vector compression.
[0001] This application claims priority to Chinese Patent Application No. 202510386474.2, filed on March 27, 2025, with the China National Intellectual Property Administration, entitled “A Method and Apparatus for Text Retrieval, Model Training and Vector Compression”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of artificial intelligence, and more specifically, to a method and apparatus for text retrieval, model training, and vector compression. Background Technology
[0003] With the rapid development of artificial intelligence, language models, using text as a medium, have performed excellently in various text understanding and text generation tasks. However, both large-scale and small-scale language models have limitations in their text representation capabilities, which leads to unsatisfactory performance in retrieval tasks.
[0004] Current methods for enhancing the representational capabilities of language models suffer from high costs or difficulties in handling complex retrieval tasks after enhancement. Summary of the Invention
[0005] This application provides a method and apparatus for text retrieval, model training, and vector compression, which can enhance the representational capabilities of language models at a lower cost.
[0006] Firstly, a text retrieval method is provided, comprising: determining a text retrieval request to be processed; inputting the text retrieval request into a first small language model to obtain retrieval results, wherein the first small language model is obtained by training a second small language model based on training data and a loss value, the training data including input text, output text, and task vector, the loss value being determined by a large language model based on the input vector, the task vector, and the output vector, the input vector being a vector generated by the second small language model to represent the input text, and the output vector being a vector generated by the large language model based on a standard answer to represent the output text.
[0007] For example, the first small language model used for text retrieval is trained on the second small language model; the standard answer is the token identifier of the output text; the second small language model trains itself by using a loss value determined based on the input vector it generates and the output vector generated by the large language model, which enables the representation ability of the second small language model to be improved by leveraging the powerful generative capabilities of the large language model.
[0008] For example, the input text, output text, and task vector belong to a set of training data; the first small language model is obtained by training a second small language model based on multiple sets of training data and loss values, and the set of training data belongs to the multiple sets of training data.
[0009] Determining the loss value involves: the large language model determining the inference result based on the input vector and the task vector; and the large language model determining the loss value based on the inference result and the output vector. The goal of training the second small language model using this loss value is to input the input vector generated by the trained first small language model into the large language model, resulting in a loss value less than or equal to the loss threshold (or, in other words, a result close to the output vector). Training the second small language model using this loss value allows the representational ability of the second small language model to be improved through the generative capabilities of the large language model.
[0010] It is understood that the loss threshold can be set based on the requirements for model performance or the nature of the task to be processed by the trained model. This application embodiment does not limit the specific setting of the loss threshold value.
[0011] Large and small language models hold great promise in text retrieval, and the performance of language models in text retrieval tasks depends on their representational capabilities. Large language models possess powerful generative capabilities, but their representational capabilities still have significant room for improvement. While fine-tuning large language models can enhance their representational capabilities, it is computationally expensive and requires substantial training data. Small language models have limited representational capabilities. Training small language models with massive amounts of unsupervised data often results in models that struggle to capture deep semantic information when handling complex tasks, limiting their representational capabilities. Furthermore, limited information compression rates make these models inefficient when processing large-scale data. Alternatively, training small language models with massive amounts of weakly supervised contrastive learning data can improve their representational capabilities, but data acquisition and the construction of weakly supervised contrastive learning data are costly.
[0012] Based on the solution provided in this application, a first small language model trained based on loss values is used to process text retrieval needs. On the one hand, compared with processing text retrieval needs through fine-tuning of a large language model, the cost of training a second language model is lower. On the other hand, compared with processing text retrieval needs through a small language model trained based on massive amounts of data, the second small language model improves its representation ability through loss values and the generation ability of the large language model, resulting in lower cost / quantity of training data acquisition. Furthermore, the trained first small language model can handle complex retrieval tasks well, and the trained model performs well in processing text retrieval needs.
[0013] In conjunction with the first or second aspect, in some possible implementations of the first or second aspect, the loss value is determined by the large language model based on the input vector, the task vector, and the output vector, including: the loss value is determined by the large language model based on a first compression result, which is obtained by compressing the sequence length of the input vector.
[0014] For example, determining the loss value may include: the large language model determining the inference result based on the first compression result and the task vector; and the large language model determining the loss value based on the inference result and the output vector. The computational cost of determining the inference result based on the first compression result is less than the computational cost of determining the inference result based on the input vector.
[0015] In combination with the first, second, or fourth aspects, in some possible implementations of the first, second, or fourth aspects, the first compression result includes the product of a first length and a compression dimension, wherein the first length is the result of compressing the sequence length of the input vector; and the compression dimension is determined based on the decoder embedding dimension of the large language model.
[0016] For example, the input vector is the product of the sequence length of the input vector and the encoder dimension (the encoder dimension of the input vector is equal to the encoder dimension of the small language model); adjusting the encoder dimension of the input vector to a compressed dimension can improve the accuracy of the large language model in understanding the input vector (or, based on the loss value determined by the first compression result, it can improve the training effect of the small language model); compressing the sequence length of the input vector to a first length that is smaller than the sequence length of the input vector (or, based on the loss value determined by the first compression result), can reduce the computational cost of model training.
[0017] In conjunction with the first or second aspect, in some possible implementations of the first or second aspect, the loss value is determined by the large language model based on the input vector, the task vector, and the output vector, including: the loss value is determined by the large language model based on a second compression result, which is obtained by compressing the sequence length of the task vector.
[0018] For example, determining the loss value may include: the large language model determining the inference result based on the second compression result and the input vector; and the large language model determining the loss value based on the inference result and the output vector.
[0019] In combination with the first, second, or fourth aspects, in some possible implementations of the first, second, or fourth aspects, the second compression result includes the product of a second length and the compression dimension, the second length being the result of compressing the sequence length of the task vector based on a reference vector generated by the large language model; the compression dimension being determined based on the decoder embedding dimension of the large language model.
[0020] For example, the task vector is the product of the sequence length of the task vector and the decoder embedding dimension (the decoder embedding dimension is the decoder embedding dimension of the large language model). Since the input text, task vector, and output text belong to the same set of training data (or, in other words, the input text, task vector, and output text have a corresponding relationship), if the compression result of the task vector is related to the encoder of the second small language model when determining the loss value, it may affect the input vector generated by the second small language model, thereby affecting the training effect of the second small language model.
[0021] Based on the solution provided in the embodiments of this application, the second small language model is trained by using the loss value determined based on the second compression result. Since the second length in the second compression result is related to the reference vector in the encoder of the large language model, the influence of the task vector on the input vector generated by the second small language model during the training process can be avoided, thereby optimizing the representation ability of the first small language model trained based on the loss value.
[0022] In conjunction with the first aspect, in some possible implementations of the first aspect, the text retrieval request is input into a first small language model to obtain retrieval results, including: the first small language model generating a first input vector based on the text retrieval request; a first compression model outputting a second input vector based on the first input vector, the second input vector being a compression result of the first input vector; the first small language model processing the text retrieval request based on the second input vector, wherein the first compression model is obtained by training a second compression model based on the training data and the loss value, and the second compression model is used to determine the first compression result and / or the second compression result.
[0023] For example, the second compression model and the second small language model are trained together using the training data and the loss value.
[0024] Based on the solution provided in the embodiments of this application, by compressing the first input vector in the text retrieval request processing module, on the one hand, the input / output load and computing resource consumption of the system in processing text retrieval requests can be reduced; on the other hand, since the second compression model is trained together with the first small language model, the two are used together to process text retrieval requests, which can improve the accuracy of the retrieval results.
[0025] In combination with the first, second, third, or fourth aspects, in some possible implementations of the first, second, or fourth aspects, the loss value includes the loss corresponding to the output vector.
[0026] The loss calculated by the large language model can include the loss at the corresponding position of the output vector, as well as the loss at the corresponding positions of the input vector and the task vector; during the training of the second small language model, training only the loss corresponding to the output vector is sufficient to achieve the purpose of training the second small language model.
[0027] Based on the solution provided in the embodiments of this application, by including the loss corresponding to the output vector in the loss value, compared to including the loss corresponding to the output vector and the loss corresponding to the positions of the input vector and the task vector, computational resources can be saved.
[0028] In combination with the first, second, third, or fourth aspects, in some possible implementations of the first, second, third, or fourth aspects, the set of training data belongs to multiple sets of training data used to train the small language model, and the multiple sets of training data include question-answering data and / or summary data.
[0029] For example, training data includes question-answering data and / or summary data.
[0030] Based on the solution provided in the embodiments of this application, by enriching the training data, on the one hand, the representation ability of the trained small language model can be optimized compared with single training data; on the other hand, compared with question-answering data, the input or output of summary data is longer, and there will be a higher data compression ratio in the process of training the second small language model based on summary data. Training data including summary data can obtain better model training results.
[0031] Secondly, a method for training a model is provided, which is applied to a small language model. The method includes: receiving input text, the input text, output text, and task vector belonging to a set of training data; generating an input vector to represent the input text; and training the small language model based on a loss value determined by a large language model based on the input vector, the task vector, and the output vector, wherein the output vector is a vector generated by the large language model based on a standard answer to represent the output text.
[0032] Based on the solution provided in the embodiments of this application, by training a small language model based on loss values, on the one hand, the training cost is lower compared to fine-tuning a large language model; on the other hand, compared to training a small language model based on massive amounts of data, the cost / quantity of acquiring training data is lower because the trained small language model improves its representation ability through loss values and the generative ability of the large language model, and the trained small language model can handle complex retrieval tasks well, and the trained model performs well in handling text retrieval needs.
[0033] In conjunction with the second or third aspect, in some possible implementations of the second or third aspect, the loss value is also used to train a compression model; the compression model is used to determine the product of the compression dimension and the first length, the first length being the result of compressing the sequence length of the input vector; and / or, the compression model is used to determine the product of the compression dimension and the second length, the second length being the result of compressing the sequence length of the task vector based on a reference vector generated by the large language model; the compression dimension is determined based on the decoder embedding dimension of the large language model.
[0034] For example, the product of the compression dimension and the first length can be referred to as the first compression result; and / or, the product of the compression dimension and the second length can be referred to as the second compression result.
[0035] The possible implementation methods and their technical effects of the second aspect above can be referred to the description of the first aspect and any of its implementation methods, and will not be repeated here.
[0036] Thirdly, a method for compressing vectors is provided, which includes: compressing or increasing the encoder dimension of an input vector to a compressed dimension, the compressed dimension being determined based on the decoder embedding dimension of a large language model, the input vector being generated by a small language model; compressing the sequence length of the input vector to a first length, the product of the compressed dimension and the first length being used to train the small language model.
[0037] Based on the solution provided in the embodiments of this application, on the one hand, by adjusting the encoder dimension of the input vector to the compression dimension, the accuracy of the large language model in understanding the input vector can be improved; on the other hand, by compressing the sequence length of the input vector to a first length smaller than the sequence length of the input vector, the computational cost of training the small language model can be reduced.
[0038] In conjunction with the third aspect, some possible implementations of the third aspect also include: compressing the sequence length of the task vector to a second length based on a reference vector generated by the large language model, and using the product of the compressed dimension and the second length to train the small language model.
[0039] In conjunction with the third aspect, in some possible implementations of the third aspect, the input vector is generated based on the input text; the input text, the output text, and the task vector belong to a set of training data, and the output vector is a vector generated by the large language model based on the standard answer to represent the output text; the product of the compression dimension and the first length, the product of the compression dimension and the second length, and the output vector are used by the large language model to determine the loss value, which is used to train the small language model.
[0040] The possible implementation methods and their technical effects of the third aspect above can be referred to the description of the first aspect and any of its implementation methods, and will not be repeated here.
[0041] Fourthly, a method for model training is provided, which includes: a small language model receiving input text, the input text, output text, and task vector belonging to a set of training data; the small language model generating an input vector to represent the input text; a large language model generating an output vector to represent the output text based on a standard answer; the large language model determining a loss value based on the input vector, the task vector, and the output vector; and the small language model training itself based on the loss value.
[0042] Based on the solution provided in the embodiments of this application, a small language model is trained by using the loss value determined by the large language model. On the one hand, compared with fine-tuning the large language model, the training cost is lower. On the other hand, compared with training the small language model based on massive amounts of data, the cost / quantity of training data is lower because the trained small language model improves its representation ability through the loss value and the generation ability of the large language model. Moreover, the trained small language model can handle complex retrieval tasks well, and the trained model performs well in handling text retrieval needs.
[0043] In conjunction with the fourth aspect, in some possible implementations of the fourth aspect, the large language model determines the loss value based on the input vector, the task vector, and the output vector, including: the large language model determines the loss value based on a first compression result and / or a second compression result, wherein the first compression result is obtained by compressing the sequence length of the input vector, and the second compression result is obtained by compressing the sequence length of the task vector.
[0044] In conjunction with the fourth aspect, in some possible implementations of the fourth aspect, the loss value is also used to train a compression model; the compression model is used to determine the first compression result and / or the second compression result.
[0045] The possible implementation methods and their technical effects of the fourth aspect above can be referred to the description of the first aspect and any of its implementation methods, and will not be repeated here.
[0046] Fifthly, a text retrieval apparatus is provided, comprising: a processing unit configured to determine a text retrieval request to be processed; the processing unit is further configured to input the text retrieval request into a first small language model to obtain retrieval results, wherein the first small language model is obtained by training a second small language model based on training data and a loss value, the training data including input text, output text, and a task vector, the loss value being determined by a large language model based on the input vector, the task vector, and the output vector, the input vector being a vector generated by the second small language model to represent the input text, and the output vector being a vector generated by the large language model based on a standard answer to represent the output text.
[0047] In conjunction with the fifth aspect, in some possible implementations of the fifth aspect, the apparatus further includes a receiving unit for receiving data, and a processing unit for determining the text retrieval request to be processed from the data received by the receiving unit.
[0048] In conjunction with the fifth aspect, in some possible implementations of the fifth aspect, the receiving unit is a receiving module or a receiver.
[0049] In conjunction with the fifth aspect, in some possible implementations of the fifth aspect, the processing unit is a processing module or a processor.
[0050] The possible implementation methods and their technical effects of the fifth aspect above can be referred to the description of the first aspect and any of its implementation methods, and will not be repeated here.
[0051] In a sixth aspect, an apparatus for model training is provided, comprising: a receiving unit for receiving input text, wherein the input text, output text, and task vector belong to a set of training data; a processing unit for generating an input vector for representing the input text; the processing unit is further configured to train a small language model based on a loss value determined by a large language model based on the input vector, the task vector, and the output vector, wherein the output vector is a vector generated by the large language model based on a standard answer for representing the output text.
[0052] In conjunction with the sixth aspect, in some possible implementations of the sixth aspect, the receiving unit is a receiving module or a receiver.
[0053] In conjunction with the sixth aspect, in some possible implementations of the sixth aspect, the processing unit is a processing module or a processor.
[0054] The possible implementation methods and their technical effects of the sixth aspect above can be referred to the description of the second aspect and any of its implementation methods, and will not be repeated here.
[0055] In a seventh aspect, an apparatus for compressing a vector is provided, comprising: a compression unit for compressing or increasing the encoder dimension of an input vector to a compression dimension, the compression dimension being determined based on the decoder embedding dimension of a large language model, the input vector being generated by a small language model; the compression unit is further configured to compress the sequence length of the input vector to a first length, the product of the compression dimension and the first length being used to train the small language model.
[0056] In conjunction with the seventh aspect, in some possible implementations of the seventh aspect, the compression unit is also used to compress the sequence length of the task vector to a second length based on a reference vector generated by the large language model, and the product of the compression dimension and the second length is used to train the small language model.
[0057] In conjunction with the seventh aspect, in some possible implementations of the seventh aspect, the compression unit is a compression module or a processor.
[0058] The possible implementation methods and their technical effects of the seventh aspect above can be referred to the description of the third aspect and any of its implementation methods, and will not be repeated here.
[0059] Eighthly, a system for training a model is provided, comprising: a first apparatus for deploying a small language model for receiving input text, the input text, output text, and a task vector belonging to a set of training data; the small language model is also used to generate an input vector representing the input text; a second apparatus for deploying a large language model for generating an output vector representing the output text based on a standard answer; the large language model is also used to determine a loss value based on the input vector, the task vector, and the output vector; and the small language model is also used to train itself based on the loss value.
[0060] In conjunction with the eighth aspect, in some possible implementations of the eighth aspect, the large language model is also used to determine the loss value based on a first compression result and / or a second compression result, the first compression result being obtained by compressing the sequence length of the input vector and the second compression result being obtained by compressing the sequence length of the task vector.
[0061] In conjunction with the eighth aspect, in some possible implementations of the eighth aspect, the loss value is also used to train the compression model; the system also includes a third device for deploying the compression model for determining the first compression result and / or the second compression result.
[0062] The possible implementation methods and their technical effects of the eighth aspect above can be referred to the description of the fourth aspect and any of its implementation methods, and will not be repeated here.
[0063] A ninth aspect provides a computing device, including a processor and a memory, and optionally, an input / output interface. The processor controls the input / output interface to send and receive information, the memory stores a computer program, and the processor retrieves and runs the computer program from the memory, causing the computing device to execute the method of the first aspect or any possible implementation thereof, or to execute the method of the second aspect or any possible implementation thereof, or to execute the method of the third aspect or any possible implementation thereof, or to execute the method of the fourth aspect or any possible implementation thereof.
[0064] Optionally, the processor can be a general-purpose processor, which can be implemented in hardware or software. When implemented in hardware, the processor can be a logic circuit, integrated circuit, etc.; when implemented in software, the processor can be a general-purpose processor that reads software code stored in memory. This memory can be integrated into the processor or located outside the processor and exist independently.
[0065] A tenth aspect provides a computing device cluster, including at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the computing device cluster to perform a method in the first aspect or any possible implementation of the first aspect, or to cause the computing device cluster to perform a method in the second aspect or any possible implementation of the second aspect, or to cause the computing device cluster to perform a method in the third aspect or any possible implementation of the third aspect, or to cause the computing device cluster to perform a method in the fourth aspect or any possible implementation of the fourth aspect.
[0066] Eleventhly, a chip is provided that acquires and executes instructions to implement the methods in the first aspect and any implementation thereof, or acquires and executes instructions to implement the methods in the second aspect and any implementation thereof, or acquires and executes instructions to implement the methods in the third aspect and any implementation thereof, or acquires and executes instructions to implement the methods in the fourth aspect and any implementation thereof.
[0067] Optionally, as one implementation, the chip includes a processor and a data / communication interface. The processor reads instructions stored in the memory through the data / communication interface to cause the chip to execute the methods in the first aspect and any implementation thereof, or to execute the methods in the second aspect and any implementation thereof, or to execute the methods in the third aspect and any implementation thereof, or to execute the methods in the fourth aspect and any implementation thereof.
[0068] Optionally, as one implementation, the chip may further include a memory storing instructions, and the processor is configured to execute the instructions stored in the memory. When the instructions are executed, the chip may perform the methods of the first aspect and any of the implementations thereof, or perform the methods of the second aspect and any of the implementations thereof, or perform the methods of the third aspect and any of the implementations thereof, or perform the methods of the fourth aspect and any of the implementations thereof.
[0069] In a twelfth aspect, a computer program product comprising instructions is provided, which, when executed by a computing device / processor, cause the computing device / processor to perform a method as described in the first aspect and any implementation thereof, or cause the computing device / processor to perform a method as described in the second aspect and any implementation thereof, or cause the computing device / processor to perform a method as described in the third aspect and any implementation thereof, or cause the computing device / processor to perform a method as described in the fourth aspect and any implementation thereof.
[0070] In a thirteenth aspect, a computer program product containing instructions is provided, which, when executed by a cluster of computing devices, cause the cluster of computing devices to perform a method as described in the first aspect and any implementation thereof, or cause the cluster of computing devices to perform a method as described in the second aspect and any implementation thereof, or cause the cluster of computing devices to perform a method as described in the third aspect and any implementation thereof, or cause the cluster of computing devices to perform a method as described in the fourth aspect and any implementation thereof.
[0071] In a fourteenth aspect, a computer-readable storage medium is provided, including computer program instructions that, when executed by a computing device / processor, perform a method as described in the first aspect and any implementation thereof, or perform a method as described in the second aspect and any implementation thereof, or perform a method as described in the third aspect and any implementation thereof, or perform a method as described in the fourth aspect and any implementation thereof.
[0072] As examples, these computer-readable storage devices include, but are not limited to, one or more of the following: read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), flash memory, electrically EPROM (EEPROM), and hard drive.
[0073] Alternatively, as one implementation method, the aforementioned storage medium can specifically be a non-volatile storage medium.
[0074] In a fifteenth aspect, a computer-readable storage medium is provided, comprising computer program instructions that, when executed by a cluster of computing devices, perform a method as described in the first aspect and any implementation thereof, or perform a method as described in the second aspect and any implementation thereof, or perform a method as described in the third aspect and any implementation thereof, or perform a method as described in the fourth aspect and any implementation thereof.
[0075] As examples, these computer-readable storages include, but are not limited to, one or more of the following: ROM, PROM, EPROM, Flash memory, EEPROM, and hard drive.
[0076] Alternatively, as one implementation method, the aforementioned storage medium can specifically be a non-volatile storage medium. Attached Figure Description
[0077] Figure 1 is a schematic diagram of the architecture of a data processing system applicable to an embodiment of this application.
[0078] Figure 2 is a schematic block diagram of a cloud scenario applicable to an embodiment of this application.
[0079] Figure 3 is a schematic diagram of the interaction between a tenant and an artificial intelligence basic development platform applicable to an embodiment of this application.
[0080] Figure 4 is a schematic diagram of a model training method provided in an embodiment of this application.
[0081] Figure 5 is a schematic diagram of a method for compressing vectors provided in an embodiment of this application.
[0082] Figure 6 is a flowchart illustrating a training model provided in an embodiment of this application.
[0083] Figure 7 is a schematic diagram of a text retrieval method provided in an embodiment of this application.
[0084] Figure 8 is a schematic block diagram of an apparatus provided in an embodiment of this application.
[0085] Figure 9 is a schematic diagram of the architecture of a computing device provided in an embodiment of this application.
[0086] Figure 10 is a schematic diagram of the architecture of a computing device cluster provided in an embodiment of this application. Detailed Implementation
[0087] The technical solutions of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort should fall within the scope of protection of this application.
[0088] To facilitate understanding of the embodiments of this application, the terms involved in this application will be briefly explained first.
[0089] It should be understood that the related conceptual explanations may be limited by the specific circumstances of the embodiments of this application, but it does not mean that this application can only be limited to the specific circumstances. The specific circumstances of different embodiments may also differ, which are not limited here.
[0090] (1) Artificial intelligence (AI)
[0091] AI (Artificial Intelligence) is the theory, methods, technology, and application systems that utilize digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. Artificial intelligence studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities. Research in the field of artificial intelligence includes robotics, natural language processing, computer vision, decision-making and reasoning, human-computer interaction, recommendation and search, and fundamental AI theories.
[0092] The basic principle of AI is to combine massive amounts of data with powerful computing capabilities and intelligent algorithms to build an AI model that solves specific problems. This AI model can automatically summarize and learn potential patterns or features from the data, thereby achieving a way of thinking that is close to that of humans.
[0093] AI models, also known as AI algorithms (or AI operators), are a collective term for mathematical algorithms built upon the principles of artificial intelligence. They form the foundation for using AI to solve specific problems. Depending on the specific methods and / or technologies used to implement artificial intelligence, AI models can also be called machine learning models, deep learning models, or reinforcement learning models.
[0094] (2) Machine Learning
[0095] Machine learning is a method for achieving artificial intelligence. The goal of this method is to design and analyze algorithms (i.e., models) that allow computers to "learn" automatically. The designed algorithms are called machine learning models.
[0096] Machine learning models are algorithms that automatically analyze data to identify patterns and use those patterns to predict unknown data. Machine learning models are diverse and can be categorized based on whether their training relies on the labels of the training data: 1. Supervised learning models; 2. Unsupervised learning models.
[0097] 1. Supervised Learning Models: These are models obtained by determining the parameters of an initial AI model based on data from a given training dataset and the labels corresponding to each data point. The process of determining the parameters of the initial AI model using the data and their labels in the training dataset is also called supervised learning (or supervised training). The labels on the data in the training dataset are usually manually labeled to indicate the correct answer for a specific task. Typical supervised learning models include: Support Vector Machines, Neural Network Models, Logistic Regression Models, Decision Trees, Naive Bayes Models, and Gaussian Discriminant Models. Supervised learning models are commonly used for classification or regression.
[0098] 2. Unsupervised Learning Models: These are models obtained by determining the parameters of an initial AI model using unlabeled data from a given training dataset. The process of determining the parameters of the initial AI model using unlabeled training data is also called unsupervised learning (or unsupervised training). Through unsupervised learning, the model can discover meaningful information and correlations in the data, thereby making predictions. There are many types of unsupervised learning models, some of the more commonly used ones being: clustering models, principal component analysis (PCA), anomaly detection models, autoencoders, and generative adversarial networks (GANs).
[0099] (3) Deep Learning
[0100] Deep learning is a new technological field that emerged during machine learning research. Specifically, deep learning is a method in machine learning based on deep representation learning of data. Deep learning interprets data by building neural networks that simulate the human brain's analytical learning process.
[0101] In the field of AI, deep learning is a learning technique based on deep neural network algorithms. A deep learning model consists of an input layer, hidden layers, and an output layer, and it uses multiple nonlinear transformations to process data.
[0102] In machine learning methods, almost all features need to be determined by industry experts and then encoded. However, deep learning algorithms attempt to learn features from data themselves; algorithms designed based on the principles of deep learning are called deep learning models.
[0103] The typical structure of current deep learning models is a deep neural network. A neural network is a mathematical or computational model that mimics the structure and function of biological neural networks (the central nervous system of animals, especially the brain). Neural networks consist of a large number of interconnected neurons performing computations. A neural network can include multiple layers with different functions, each layer containing parameters and computational rules. Depending on the computational formula or function, different layers in a neural network have different names; for example, the layer performing convolution calculations is called a convolutional layer, which is often used for feature extraction from input signals (e.g., text). A neural network can also be composed of multiple sub-neural networks. Different neural network structures can be applied to different scenarios (e.g., classification, recognition, representation, retrieval) or provide different results when used in the same scenario. The specific differences in neural network structures include one or more of the following: different numbers of network layers, different order of network layers, and different weights, parameters, or computational formulas in each network layer. Various high-accuracy neural networks exist in the industry for applications such as recognition, classification, retrieval, or recommendation. Some neural networks can be trained on specific datasets and used alone to complete a task or combined with other neural networks (or other functional modules) to complete a task.
[0104] In other words, deep learning models are actually machine learning models with complex neural network structures. Based on whether deep learning models need to rely on the labels of the training data during training, they can also be divided into supervised learning models and unsupervised learning models, which will not be elaborated upon here. Classic deep learning models include convolutional neural networks (CNNs), recurrent neural networks (RNNs), and recursive neural networks (RNNs).
[0105] (4) Language Model
[0106] A language model (LM) is a type of machine learning model that can be used to process and predict natural language data. Language models can include various types, such as large language models, compact language models, or small language models, for processing and predicting natural language data.
[0107] Large language models (LLMs) are neural network models trained on massive corpora with a large number of parameters. They are capable of understanding and generating natural language text. Specifically, LLMs are typically based on neural network techniques, learning the syntax, semantics, and contextual information of a language through training on large amounts of text data. During training, the model continuously optimizes its parameters to improve its text understanding and generation capabilities. Due to their powerful abilities in understanding and generating natural language, LLMs have been widely applied in many fields to solve natural language understanding and generation problems. They have broad applications in artificial intelligence, such as natural language processing, machine translation, and dialogue systems.
[0108] Small language models (SLMs) typically refer to models with a relatively small number of parameters. They are designed with efficiency and practicality in mind, are small in size, and easily adaptable to resource-constrained and computationally limited environments, such as mobile devices or embedded systems. Compared to LLMs, SLMs have a significantly reduced number of parameters, making them more economical in terms of storage and computing resource requirements. Due to the simplified parameters and model structure, SLMs often respond faster when processing requests. SLMs may be optimized for specific application scenarios or tasks to provide good performance with limited resources. Despite their smaller number of parameters, SLMs can still achieve generalization capabilities across a variety of tasks through careful design and training.
[0109] In the embodiments of this application, a large language model and a small language model can refer to two language models with a significantly different number of parameters. For example, a large language model is a model with tens of billions of parameters, while a small language model is a model with hundreds of millions of parameters; or, the parameter difference between a large language model and a small language model can be hundreds of times.
[0110] (5) Representational ability
[0111] Representational ability refers to a model's capacity to transform input data (such as text and images) into meaningful numerical representations (such as vectors). These numerical representations capture the key features and semantic information of the data. Good representational ability means that the vectors generated by the model can be effectively used for downstream tasks, such as classification, retrieval, and clustering. Currently, the representational ability of LLMs is not proportional to their powerful generative capabilities, and small-scale language models with relatively few parameters have limitations in representational ability.
[0112] The architecture applicable to the embodiments of this application is described below with reference to Figures 1 to 3.
[0113] Figure 1 is a schematic diagram of the architecture 100 of a data processing system applicable to an embodiment of this application.
[0114] As shown in Figure 1, the data acquisition device 160 is used to collect training data and store the training data (data stream) into the database 130. The training device 120 trains the target model / rule 101 based on the training data maintained in the database 130.
[0115] It should be noted that the data acquisition device 160, the execution device 110, and the training device 120 may be the same or different devices.
[0116] Model training refers to the process of using a specified initial model to calculate training data, and then adjusting the parameters of the initial model based on the calculation results, so that the model gradually learns certain patterns and acquires specific functions. After training, a model with stable functions can be used for representation, text retrieval, etc.
[0117] The following will describe in more detail how the training device 120 obtains the target model / rule 101 based on the training data. The training data may be multiple sets of training data provided in the embodiments of this application, including input text, task vectors, output text, etc.
[0118] It should be noted that, in practical applications, the training data maintained in database 130 can come from data acquisition device 160 or other devices (for example, the training data maintained in database 130 may also include data streams from client device 140 or data from I / O interface 112). Training device 120 does not necessarily train the target model / rule 101 entirely based on the training data maintained in database 130; it can also obtain training data from the cloud or other sources for model training. The above description should not be construed as limiting the embodiments of this application. The target model / rule 101 trained by training device 120 can be applied to different systems or devices, for example, to execution device 110 shown in FIG. 1.
[0119] It should be understood that the execution device 110 may be a terminal. For example, the terminal may be a digital camera, video recording device, mobile phone, personal computer (PC), laptop computer, server, tablet computer, smart TV, in-vehicle terminal, mobile internet device (MID), wearable device, virtual reality (VR) device, augmented reality (AR) device, wireless terminal in industrial control, wireless terminal in self-driving, wireless terminal in remote medical surgery, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city, wireless terminal in smart home, etc., or it may be an edge device (e.g., a box carrying a chip with processing capabilities); the execution device 110 may also be a server or cloud, etc.
[0120] The execution device 110 is configured with an I / O interface 112 for data interaction with external devices. Users can input data to and / or receive output data from the I / O interface 112 via the client device 140 (in one possible implementation, the I / O interface 112 can also interact with the database 130). Furthermore, the input data can be user-inputted data, data uploaded by the user via notepad, memo, document, etc., or data from a database; this application does not limit the scope of the data input.
[0121] In one possible embodiment, the execution device 110 and the training device 120 are different processors deployed on different physical devices (e.g., servers in a server or cluster). For example, the execution device 110 may be a graphics processing unit (GPU), a central processing unit (CPU), other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or one or more integrated circuits for model training, representation, vector compression, or text retrieval according to the methods provided in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The training device 120 may be a GPU, a neural network processing unit (NPU), a microprocessor, an application-specific integrated circuit (ASIC), etc. The training device 120 can configure the trained neural network to multiple execution devices 110. Each execution device 110 utilizes the trained neural network to implement functions such as representation, vector compression, or text retrieval.
[0122] In another possible embodiment, the execution device 110 and the training device 120 are deployed on the same physical device, or the execution device 110 and the training device 120 are on the same physical device. The computing device can configure the trained neural network to itself and use the trained neural network to achieve functions such as representation, vector compression, or text retrieval.
[0123] Training device 120 is used to train a neural network using training data until the loss function in the neural network converges and its value is less than a specific threshold, at which point the neural network training is complete, thus achieving a certain level of accuracy. Alternatively, all training data in database 130 can be used for training, completing the neural network training and enabling it to perform functions such as representation, vector compression, or text retrieval. Then, training device 120 configures the trained neural network to execution device 110. Execution device 110 is used to perform representation, vector compression, or text retrieval on the input data based on the trained neural network.
[0124] The preprocessing module 113 is used to preprocess the input data received by the I / O interface 112. In this embodiment, the preprocessing module 113 can be used to obtain input data from the user (e.g., from a memo or notepad). During the preprocessing of the input data by the execution device 110, or during the calculation module 111 of the execution device 110 performing calculations or other related processing, the execution device 110 can call data, code, etc., in the data storage system 150 for corresponding processing, or store the processed data, instructions, etc., into the data storage system 150. Finally, the I / O interface 112 returns the processing result (e.g., output data) to the client device 140, thereby providing it to the user.
[0125] It should be understood that Figure 1 is only a schematic diagram of a system architecture, and the positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, in Figure 1, the data storage system 150 is an external memory relative to the execution device 110. In other cases, the data storage system 150 can also be placed in the execution device 110.
[0126] As an example, in one possible implementation, the method provided in this application embodiment can be applied to a cloud service scenario, where the method is executed by a cloud management platform within the cloud service scenario. For instance, the cloud management platform in the cloud service scenario may perform model training, representation, vector compression, or text retrieval.
[0127] For ease of description, the cloud service scenario will be described in detail below with reference to Figure 2.
[0128] Figure 2 shows a schematic block diagram of a cloud scenario applicable to an embodiment of this application. As shown in Figure 2, the cloud scenario may include: a cloud management platform 210, the Internet 220, and a client 230.
[0129] As shown in Figure 2, the cloud management platform 210 is used to manage the infrastructure that provides multiple cloud services. The infrastructure includes multiple cloud data centers, each cloud data center includes multiple servers, and each server includes cloud service resources to provide corresponding cloud services to tenants.
[0130] The cloud management platform 210 can be located in a cloud data center and provides access interfaces (such as user interfaces or application program interfaces, APIs). Tenants can use client 230 to remotely access the access interface to register a cloud account and password on the cloud management platform 210 and log in. After successful authentication of the cloud account and password on the cloud management platform 210, the tenant can further select and purchase virtual machines of specific specifications (processor, memory, disk) on the cloud management platform 210. After successful purchase, the cloud management platform 210 provides the remote login account and password for the purchased virtual machine, and client 230 can remotely log in to the virtual machine to install and run the tenant's applications. Therefore, tenants can create, manage, log in to, and operate virtual machines in the cloud data center through the cloud management platform 210. Virtual machines can also be referred to as Elastic Compute Service (ECS) or Elastic Instances (different cloud service providers may use different names).
[0131] It should be understood that cloud service tenants can be individuals, businesses, schools, hospitals, government agencies, etc.
[0132] The cloud management platform 210 includes, but is not limited to, a user console, compute management services, network management services, storage management services, authentication services, and image management services. The user console provides an interface or API for interaction with tenants. The compute management services manage servers running virtual machines and containers, as well as bare metal servers. The network management services manage network services (such as gateways and firewalls). The storage management services manage storage services (such as data bucket services). The authentication services manage tenant account passwords. The image management services manage virtual machine images. Tenants use client 230 and can log in to the cloud management platform 210 via the internet 220 to manage their rented cloud services.
[0133] As an example, this cloud service may include, but is not limited to, AI services. AI services and products in the cloud domain embody both the on-demand and purchase-based characteristics of cloud services and the abstract, diverse, and widely applicable characteristics of AI technology. AI services in the cloud domain include: Platform-as-a-Service (PaaS) type AI infrastructure development platform services.
[0134] It should be understood that the AI Infrastructure Development Platform service is a PaaS cloud service within the cloud management platform 210. It's a software platform provided to users (also known as tenants, AI developers, etc.) based on the abundant underlying resources and software capabilities of public cloud service providers. This platform assists users in building, training, and deploying AI models, as well as developing and deploying AI applications. In other words, public cloud service providers offer an AI Infrastructure Development Platform to tenants, leveraging their ample underlying resources and upper-layer AI algorithm capabilities. The AI development framework and various AI algorithms built into this platform allow tenants to quickly build and develop AI models or applications that meet their individual needs.
[0135] As shown in Figure 3, Figure 3 illustrates an interaction diagram between a tenant and an AI basic development platform applicable to an embodiment of this application. The interaction between the tenant and the AI basic development platform mainly includes: the tenant logs into the cloud management platform 210 through a webpage on the client 230, selects and purchases cloud services of the AI basic development platform in the cloud management platform 210, and after purchase, the tenant can carry out full-process AI development based on the functions provided by the AI basic development platform.
[0136] As an example, when tenants develop and train their AI models on an AI infrastructure development platform, they rely on the basic resources (primarily computing resources such as CPUs, GPUs, and NPUs) in the cloud service provider's data center. Therefore, when purchasing and using the AI infrastructure development platform, payment is mainly for the resources used. For instance, tenants need to prepay before using the AI infrastructure development platform. During prepayment, different resource specifications support different functions of the AI infrastructure development platform. Tenants mainly select the name and specifications of the resources they need based on the functions they require. Tenants can also choose the purchase duration. The cloud management platform 210 prices the package based on the user's selected resource name, specifications, and purchase duration. After purchasing the prepaid package, tenants can utilize the capabilities provided by the AI infrastructure development platform and the basic computing resources included in the prepaid package for building, training, and deploying AI models. When resource usage exceeds the current prepaid package limit, the cloud management platform 210 charges for the excess resources on a pay-as-you-go basis. In reality, the basic resources used by tenants on the AI infrastructure development platform are mainly virtualized computing resources, such as virtual machines and containers.
[0137] It should be understood that the sale of AI basic development platforms is actually a form of selling software capabilities together with hardware virtualization basic resources. Furthermore, the basic resources supporting any process in the AI basic development platform may be distributed across different physical devices. In other words, the hardware devices that actually execute a process are usually server clusters in the same data center or server clusters distributed across different data centers.
[0138] With the rapid development of artificial intelligence, language models, using text as a medium, have performed excellently in various text understanding and text generation tasks. However, both large-scale and small-scale language models have limitations in their text representation capabilities, which leads to unsatisfactory performance in retrieval tasks.
[0139] Current methods for enhancing the representational capabilities of language models suffer from high costs or difficulties in handling complex retrieval tasks after enhancement.
[0140] In view of this, embodiments of this application provide a method and apparatus for text retrieval, model training, and vector compression, which can enhance the representational capabilities of language models at a lower cost.
[0141] The following describes in detail a model training method provided by an embodiment of this application, with reference to Figure 4. Figure 4 shows a schematic diagram of a model training method 400 provided by an embodiment of this application. The method 400 may include the following steps.
[0142] S410, the small language model receives input text, which, along with the output text and the task vector, belongs to a set of training data.
[0143] S420, the small language model generates an input vector to represent the input text.
[0144] S430, the small language model trains itself based on a loss value, which is determined by the large language model based on the input vector, the task vector, and the output vector, which is a vector generated by the large language model based on the standard answer to represent the output text.
[0145] For example, the training data used to train the small language model can be in multiple sets, where each set of training data can include corresponding input text, output text, and task vectors; or, where each set of training data can include corresponding input text, output text, and task (the representation vector / representation result of the task is a task vector). In each set of training data, the output text is obtained based on the input text and the task (or task vector).
[0146] For example, the task vector, the input vector generated by the small language model, and the output vector generated by the large language model are input into the large language model. The large language model outputs an inference result based on the mask model, the task vector, and the input vector. The large language model determines the loss value between the inference result and the standard answer and outputs the loss value. The small language model optimizes its own model parameters based on the loss value.
[0147] The large language model, based on the masking model, the task vector, and the input vector, outputs an inference result that can be understood as follows: due to the masking model, the large language model generates the (M+1)th position of the inference result based on the first M positions of the task vector, the input vector, and the output vector; the (M+1)th position of the inference result corresponds to the (M+1)th position of the output vector, where M is a non-negative integer.
[0148] The output vector, generated by the large language model based on the standard answer, can be understood as follows: the output text is converted into token identifiers (token_IDs); the encoding module of the large language model generates the output vector based on these token_IDs; and the token_IDs serve as the standard answer for comparison with the inference results to determine the loss value.
[0149] The training of a small language model based on its loss value can be understood as: the small language model updates its own model parameters based on the loss value.
[0150] Taking input text a, output text b, and the small language model's representation of a as A (a and b belong to a set of training data; a and b can also be understood as a set of real input and output) as an example: the small language model adjusts its own model parameters based on the loss value, so that as the small language model is optimized, the inference result B obtained by inputting A into the large language model gradually approaches b.
[0151] For the specific implementation of training the small language model based on the loss value, please refer to relevant technologies; for the specific implementation of outputting the inference result based on the mask model, the task vector, and the input vector, please refer to relevant technologies; for the specific implementation of determining the loss value based on the task vector, input vector, and output vector, please refer to relevant technologies or the text description in Figure 6. The embodiments of this application will not be described in detail here.
[0152] By training the small language model based on the loss value determined by the inference result and the standard answer, it is possible to leverage the powerful generative capabilities of the large language model (the generative capability is manifested in the large language model generating inference results based on the input vector generated by the small language model) to optimize the representational capabilities of the small language model in generating representation vectors of the input text (or, optimize the ability of the small language model to represent the input text, so that the small language model aligns with the representation space of the large language model).
[0153] Aligning the representation space of a smaller language model with that of a larger language model can improve the representation performance of the smaller language model. Furthermore, this approach avoids the problem of not being able to construct weakly supervised contrastive learning datasets when dealing with vertical domains.
[0154] In one possible implementation, the loss value is determined by the large language model based on the input vector, the task vector, and the output vector, including: the loss value is determined by the large language model based on a first compression result and / or a second compression result, wherein the first compression result is obtained by compressing the sequence length of the input vector, and the second compression result is obtained by compressing the sequence length of the task vector.
[0155] Compressing the sequence length of the input vector can reduce the amount of computation during model training.
[0156] In one possible implementation, the first compression result includes the product of a first length and a compression dimension, the first length being the result of compressing the sequence length of the input vector; the second compression result includes the product of a second length and the compression dimension, the second length being the result of compressing the sequence length of the task vector based on a reference vector generated by the large language model; the compression dimension is determined based on the decoder embedding dimension of the large language model.
[0157] For example, the compression result (first length) of the sequence length of the input vector can be determined according to a preset standard.
[0158] For example, the compression result of the input vector sequence length is determined according to the degree of compression; the compression result that meets the degree of compression can achieve the preset standard of saving computation.
[0159] For example, the input vector includes the product of the encoder dimension and the sequence length. Adjusting the encoder dimension to a compressed dimension helps the large language model to correctly understand the input vector after it is fed into the large language model.
[0160] The compression dimension can be the decoder embedding dimension of a large language model.
[0161] For example, compressing the sequence length of task vectors can optimize the performance of model training.
[0162] For example, the second length is obtained by assigning values to the initial representation of task #1 based on the reference vector (the initial representation of task #1 can be a zero vector; the initial representation of task #1 is the task vector). The reference vector is the vector representing task #1 in the encoding module of the large language model. Assigning values to the initial representation of task #1 based on the reference vector can also be called compressing the sequence length of the task vector.
[0163] Compressing the sequence length of the input vector can reduce the dimensionality of the text representation vector, reduce computational resource consumption, and at the same time maintain the retrieval performance of the trained model, thereby reducing the cost of model training.
[0164] The reference vector representing task #1 in the encoding module of the large language model is compressed. Compared with training the small language model based on the task vectors representing the task generated by the small language model, this can avoid the influence of corresponding tasks (or task vectors) on the representation of the input text during instruction fine-tuning, making the result of the small language model representing the input text more consistent with its own representation ability.
[0165] In this application embodiment, compressing both the task vector and the input vector can also be called dual compression or dual information compression; the model that compresses the above-mentioned task vector and / or input vector can also be called a compression model; compared with the traditional compression model, the compression model provided in this application embodiment is more lightweight.
[0166] In one possible implementation, the loss value is also used to train a compression model; the compression model is used to determine the product of a first length and a compression dimension, the first length being the result of compressing the sequence length of the input vector; and / or, the compression model is used to determine the product of a second length and the compression dimension, the second length being the result of compressing the sequence length of the task vector based on a reference vector generated by the large language model; the compression dimension is determined based on the decoder embedding dimension of the large language model.
[0167] Training a compression model based on the loss value can optimize the compression result of the sequence length of the input vector; furthermore, by optimizing the small language model based on the loss value determined by the compression result, the small language model can be better aligned with the representation space of the large language model.
[0168] In one possible implementation, the loss value includes the loss corresponding to the output vector.
[0169] For example, the task vector, the input vector generated by the small language model, and the output vector generated by the large language model are input into the large language model; the large language model can determine the loss value between the inference result and the standard answer (the loss value between the inference result and the standard answer can be understood as the loss corresponding to the output vector) and the loss corresponding to the input vector.
[0170] Large language models can reduce the computational cost of loss calculation and save computational resources by only determining the loss corresponding to the output vector (not the loss corresponding to the input vector).
[0171] In one possible implementation, the set of training data belongs to multiple sets of training data used to train the small language model, and the multiple sets of training data include question-answering data and / or summary data.
[0172] For example, the training data can be question-and-answer data, where the input text represents the question, the task represents giving the answer based on the question, and the output text represents the answer given based on the question; and / or, the training data can be summary data, where the input text represents the full text, the task represents extracting key points from the input text, and the output text represents the key points extracted based on the input text.
[0173] Alternatively, for example, the training data can be question-and-answer data, where the input text represents the answer, the task represents giving a question based on the answer, and the output text represents the question given based on the answer; and / or, the training data can be summary data, where the input text represents the summary, the task represents giving the full text based on the input text, and the output text represents the full text obtained based on the summary.
[0174] Compared to question-answering data, the input and / or output of summary data in training data are longer. Training the model using summary data will result in a higher data compression ratio, thus leading to better model training results.
[0175] The embodiments of this application, through the scheme of method 400, pre-train a small language model, which can combine the powerful generation capabilities of instruction tuning and large language models, as well as a lightweight dual information compression module, to improve the retrieval model's ability to understand and represent text, solve the shortcomings of existing retrieval models in representation capabilities, and thus enhance its retrieval efficiency and accuracy in practical applications.
[0176] The solution provided in this application is suitable for large-scale retrieval enhancement pre-training. Furthermore, compared to solutions that require negative samples and train small language models using weakly supervised data, the solution provided in this application only requires positive samples and does not require an additional negative sample set, thus reducing data acquisition costs and making it more suitable for large-scale data processing.
[0177] The method for compressing vectors according to an embodiment of this application will be described in detail below with reference to Figure 5. Figure 5 shows a schematic diagram of a method 500 for compressing vectors according to an embodiment of this application. Method 500 is a possible implementation of the compression method in method 400 described above. Method 500 may include the following steps.
[0178] S510, the compression module compresses or increases the encoder dimension of the input vector to a compressed dimension, which is determined based on the decoder embedding dimension of the large language model, and the input vector is generated by the small language model.
[0179] S520, the compression module compresses the sequence length of the input vector to a first length, and the product of the compression dimension and the first length is used to train the small language model.
[0180] In one possible implementation, the method 500 may further include: S530, whereby the compression module compresses the sequence length of the task vector to a second length based on a reference vector generated by the large language model, and the product of the compression dimension and the second length is used to train the small language model.
[0181] In one possible implementation, the input vector is generated based on the input text; the input text, the output text, and the task vector belong to a set of training data, and the output vector is a vector generated by the large language model based on the standard answer to represent the output text; the product of the compression dimension and the first length, the product of the compression dimension and the second length, and the output vector are used by the large language model to determine the loss value, which is used to train the small language model.
[0182] The specific implementation and / or effects of method 500 can be found in the textual description of method 400 above. This application's embodiments will not elaborate further.
[0183] The following describes in detail, with reference to Figure 6, a flowchart of a training model provided in an embodiment of this application. This training model's flowchart can be a possible implementation of the model training method 400 described above.
[0184] As shown in Figure 6, Figure 6 illustrates a flowchart 600 of a training model provided in an embodiment of this application.
[0185] In this embodiment, the training data used to train the model includes multiple sets of training data. Each set of training data includes: input text, task vector, and output text. The task vector is a vector representation of the task, and the input and output texts are the input and output of the task represented by the task vector. For example, the input text might be "I like green.", the task vector might represent the task "translate the input into English," and the output text might be "I like green."
[0186] Taking a set of training data as an example (for ease of description, the contents of this set of training data are denoted as input text #1, task vector #1, and output text #1), in process 600, the input text #1 is input into the small language model to be trained, and the small language model outputs V_i based on the input text #1, where V_i is a vector used to represent the input text #1; V_i is input into the compression module, which obtains V_I based on V_i; the task vector #1 (V_t) is input into the compression module, which obtains V_T based on V_t; the compression module concatenates V_I and V_T and outputs the concatenated result V_D; the output text #1 is input into the LLM, and the embedding module of the LLM outputs V_o based on the output text #1, where V_o is a vector used to represent the output text #1. The concatenation result of V_D and V_o, Concat(V_D,V_o), is input into the LLM. The LLM, based on teacher forcing and cross entropy loss, outputs the loss at the output position corresponding to V_o. The small language model to be trained optimizes its own model parameters based on the loss at the output position corresponding to V_o. The optimization goal is to reduce the difference between the output text #2 and the output text #1 of the large model based on V_i.
[0187] Where V_i is a possible implementation of the input vector; V_t is a possible implementation of the task vector; V_o is a possible implementation of the output vector; the loss at the output position corresponding to V_o is a possible implementation of the loss value; and the output text #2 is a possible implementation of the inference result.
[0188] As shown in Figure 6, the specific implementation of the loss for determining the output position corresponding to V_o using the teacher forcing method and cross-entropy loss function in LLM can include: LLM generating the content of position #1 of the representation vector of output text #2 based on V_D; LLM calculating the cross-entropy loss between the content of position #1 of the representation vector of output text #2 and the content of position #1 of V_o; LLM determining the content of position #2 of the representation vector of output text #2 based on V_D and the content of position #1 of V_o; LLM calculating the cross-entropy loss between the content of position #2 of the representation vector of output text #2 and the content of position #2 of V_o; LLM determining the content of position #2 of V_D and the content of position #1 of V_o and the content of position #2 of V_o. The content of #2 is set, and the content of position #3 of the representation vector of output text #2 is determined. LLM calculates the cross-entropy loss between the content of position #3 of the representation vector of output text #2 and the content of position #3 of V_o, and so on, until LLM calculates the cross-entropy loss between the content of the last position of the representation vector of output text #2 and the content of the last position of V_o (or, based on the content of the first T positions of V_D and V_o, LLM determines the content of position #T+1 of the representation vector of output text #2; LLM calculates the cross-entropy loss between the content of position #T+1 of the representation vector of output text #2 and the content of position #T+1 of V_o; T is a non-negative integer). All cross-entropy losses are summed or weighted averaged to obtain the final loss, which is one possible implementation of the loss at the output position corresponding to V_o.
[0189] This compression module outputs V_T based on V_t, which can be understood as using the existing vectors in the encoding module of the large language model to assign values to V_t to obtain V_T.
[0190] During the training process based on teacher forcing and cross-entropy loss function, only the cross-entropy loss of the output text is trained. Since V_D is obtained based on the real input text, the solution provided in this application does not need to calculate the loss of the output position corresponding to V_D (or, the model parameters of the small language model to be trained do not need to be optimized based on the loss of the output position corresponding to V_D).
[0191] Where V_t can be a zero vector used to represent the "task" (and has a corresponding relationship with the input and output texts); the model parameters of the LLM can change or remain fixed during the above process (the fixed ones can also be called the model is frozen); Concat(V_D,V_o) can be obtained by concatenating V_D and V_o with the small language model and other models / modules other than the LLM.
[0192] In the embodiments of this application, the specific implementation of representing the "task" to obtain the representation vector V_t can refer to related technologies; the specific implementation of the determination method of Concat(V_D,V_o) can refer to related technologies; the method of training the model based on teacher forcing and cross-entropy loss function can refer to related technologies.
[0193] It should be understood that other sets of training data can repeat a similar training process as a particular set of training data, and this application will not elaborate on this further in the embodiments.
[0194] In the above process, the model parameters of the small language model are trainable / optimizable; and / or, the model parameters of the compression module are trainable / optimizable.
[0195] In some possible implementations, the small language model can be a one-way encoded model or a two-way encoded model, and the embodiments of this application do not limit this.
[0196] In this embodiment of the application, a small language model may include a language model with a smaller parameter size.
[0197] In some possible implementations, the aforementioned sets of training data can include both question-and-answer data and summary data. Question-and-answer data could be: input text "I like green.", task vector representing the task "translate the input into English," and output text "I like green."; summary data could be: input text "the full text of 'The Back View'", task vector representing the task "extract key points from the input," and output text "The article focuses on the scenes of the author's two partings with his father, highlighting the farewell at the train station: the overweight, simply dressed father insists on crossing the railway tracks and trudging up the platform to buy oranges for his son; his clumsy yet persistent back view becomes an eternal mark of paternal love."
[0198] Compared to question-answering data, the input and / or output of summary data in training data are longer. Training the model using summary data will result in a higher data compression ratio and better model training results.
[0199] For example, when question-and-answer data accounts for about 90% and summary data accounts for about 10% in the above training data sets, the trained model performs better than the model trained with other proportions.
[0200] In some possible implementations, during the model training process based on the aforementioned multiple sets of training data, at least one set of training data is processed with a certain probability (e.g., 60% or 50%) to obtain input text ', task vector ', and output text ', and the model is trained based on the input text ', task vector ', and output text '. Here, the input text ' represents the output text before processing, the task vector ' represents the inverse task represented by the task vector before processing, and the output text ' represents the input text before processing.
[0201] For example, a set of training data before processing includes: input text "I like green.", task vector representing the task "translate the input from Chinese to English", and output text "I like green."; after processing this training data, we get: input text 'I like green.', task vector ' representing the task "translate the input from English to Chinese", and output text '"I like green."
[0202] By processing the training data, the training dataset can be expanded at a lower cost.
[0203] In some possible implementations, the aforementioned multiple sets of training data can be designed for different training scenarios, enabling the trained model to be more suitable for specific scenarios.
[0204] In some possible implementations, the aforementioned sets of training data are supervised data; furthermore, these sets of training data can be obtained through a large language model, a small language model, other models, purchased data, or user-defined data. This application does not limit the source of the aforementioned sets of training data.
[0205] Through the above scheme, based on grouped data and instruction fine-tuning, the token_IDs used to determine V_o are used as the standard answer. The loss between the token_IDs and the output obtained by LLM processing V_D (the output of LLM processing V_D is: the representation vector / inference result of the LLM predicted output text when V_D is input into LLM) is compared. The small language model can optimize its own parameters through this loss, so that inputting the optimized output of the small language model into LLM yields a predicted output as close as possible to the true output. Instruction fine-tuning, also known as "supervised fine-tuning," means that there is a standard answer for the input text, and the model is trained using the standard answer.
[0206] In some possible implementations, the compression module in the above scheme can be a convolution compression method or a channel attention-based compression method.
[0207] In some other possible implementations, the compression module in the above scheme is denoted as compression module #1, and the compression method applicable to compression module #1 is as follows:
[0208] Compression module #1 (compressor_double) includes a channel projection module. This channel projection module is used to: for a V_i with dimension seq_len * embedding_dim_encoder, firstly, increase the vector dimension of V_i from embedding_dim_encoder to the input vector dimension of LLM embedding_dim_decoder through channel projection, and then compress the seq_len of V_i to embedding_len through adaptive average pooling to obtain a compressed V_I with dimension embedding_len * embedding_dim_decoder.
[0209] Specifically, increasing the vector dimension of V_i to the input vector dimension of the LLM is to enable the input text vector to be input into the LLM; compressing seq_len to embedding_len of V_i is token compression, which can give the small language model a higher compression ratio; seq represents sequence; len represents length; and dim represents dimension.
[0210] Compression module #1 is also used to output V_T based on V_t, where the dimension of V_t is task_len*embedding_dim_decoder, and V_T represents the result of initializing V_t using the embedding weights of LLM. The dimension of the initialized V_T is embedding_len*embedding_dim_decoder.
[0211] Initializing V_t means representing the task itself through LLM embedding, which enables compression of V_t, resulting in better text retrieval performance for the trained small language model.
[0212] Compression module #1 is also used to concatenate V_I and V_T to obtain V_D.
[0213] The compression method executed in compression module #1 can be one possible implementation of the compression vector method 500 mentioned above.
[0214] By using the compression module #1 to compress the seq_len of V_i and initialize V_t, dual compression of text representation and task representation is achieved. On the one hand, since the small language model does not represent the task, the influence of the task on the text representation in token compression techniques can be avoided. On the other hand, compared with traditional compression methods (convolutional compression methods or channel attention-based compression methods), this lightweight compression method can reduce the compression computation and reduce the training cost of the model, thereby improving the representation performance / retrieval effect of the model trained with the same computational cost.
[0215] The following describes in detail a text retrieval method provided by an embodiment of this application, with reference to Figure 7. Figure 7 shows a schematic diagram of a text retrieval method 700 provided by an embodiment of this application. The method 700 may include the following steps.
[0216] S710, Determine the text retrieval requirements to be processed.
[0217] S720, the text retrieval request is input into the first small language model to obtain the retrieval results. The first small language model is obtained by training the second small language model based on the training data and the loss value. The training data includes the input text, the output text, and the task vector. The loss value is determined by the large language model based on the input vector, the task vector, and the output vector. The input vector is a vector generated by the second small language model to represent the input text, and the output vector is a vector generated by the large language model based on the standard answer to represent the output text.
[0218] It is understandable that the first small language model is a possible implementation of the small language model trained by method 400 above; the second small language model is a possible implementation of the small language model trained in method 400 above. Or, in other words, the first small language model is obtained by training the second small language model based on method 400 above.
[0219] In some possible implementations, the text retrieval request is input into a first small language model to obtain retrieval results, including: the first small language model generating a first input vector based on the text retrieval request; a first compression model outputting a second input vector based on the first input vector, the second input vector being a compressed result of the first input vector; and the first small language model processing the text retrieval request based on the second input vector, wherein the first compression model is obtained by training a second compression model based on the training data and the loss value, and the second compression model is used to determine the first compression result and / or the second compression result.
[0220] In the embodiments of this application, the specific implementation of the training compression module or compression model (e.g., the first compression model) in text retrieval tasks or other tasks can also refer to related technologies, and will not be repeated in the embodiments of this application.
[0221] The text retrieval model in method 700 can be a text retrieval model obtained by fine-tuning the small language model trained by method 400 or process 600 (or, the text retrieval model in method 700 includes: a text retrieval model and a compression module obtained by fine-tuning the small language model and compression module trained by method 400 and method 500). In other words, method 700 is a method of performing downstream tasks after fine-tuning the small language model trained by method 400 or process 600 (or, a method of performing downstream tasks after fine-tuning the small language model and compression module trained by method 400 and method 500). The specific implementation of the text retrieval model in method 700 before fine-tuning can refer to method 400, method 500, or process 600, and will not be repeated in this embodiment.
[0222] The results of text retrieval based on method 700 can be found in Tables 1, 2, and 3 below. The bolded numbers in the tables represent the maximum values of the metrics for different datasets / compression methods. The text retrieval model in method 700 is a fine-tuned version of the small language model and compression module trained using methods 400, 500, or process 600.
[0223] In this embodiment of the application, the fine-tuning of the trained small language model and compression module can be achieved through contrastive learning or other fine-tuning methods, and this embodiment of the application does not limit this.
[0224] Table 1. Zero-shot MTEB benchmark text retrieval
[0225] In Table 1 or Table 2, pre-trained LLM refers to pre-training on large-scale text data, enabling the LLM to learn general language representations; pre-training the representation model refers to pre-training the small language model before text retrieval to enhance its retrieval capabilities; the values in Table 1 represent nDCG@10 (normalized discounted cumulative gain at 10); zero-shot MTEB (massive text embedding benchmark) is a standardized framework for evaluating the performance of text embedding models, covering multi-task and multilingual tests.
[0226] nDCG@10 is a metric used to evaluate the performance of ranking models. It measures the model's ability to rank highly relevant content at the top of the top 10 results. The larger the nDCG@10 value, the closer the ranking effect of the corresponding model is to the ideal state, that is, the better the model's performance.
[0227] Table 2. Supervised Microsoft Machine Reading Comprehension (MS MARCO) Text Retrieval
[0228] DEV / DL19 is a different subset of the MS MARCO dataset; MRR@10 or R@1000 are metrics for evaluating the retrieval performance of the model, and the higher the value of MRR@10 or R@1000, the better the performance of the corresponding model.
[0229] Table 3 illustrates the reduction in model training costs using different compression methods (convolution, channel attention, and compression module #1). The AVG in Table 3 represents the average of 14 nDCG@10 values obtained based on different compression methods and 14 datasets as shown in Table 1 (Arguana, ClimateFEVER, CQADupstackRetrieval, DBPedia, FEVER, FiQA2018, HotpotQA, NFCorpus, NQ, QuoraRetrieval, SCIDOCS, SciFact, Touche2020, or TRECCOVID).
[0230] Table 3
[0231] In Table 3, a higher AVG value indicates better performance of the compressed model; a lower number of parameters indicates a lower training cost for the model.
[0232] As shown in Tables 1, 2, and 3, the scheme provided in this application enhances the retrieval representation capabilities of the small language model by combining the generative capabilities of LLM, achieving superior performance of the enhanced (and fine-tuned) small language model in the Zero-shot MTEB benchmark and Supervised MS MARCO text retrieval. By introducing a lightweight double compression compressor_double, the training cost of the small language model can be significantly reduced, while ensuring that the retrieval performance of the small language model is superior under the same computational cost. Compared with self-supervised methods and weak supervised contrastive learning pre-training methods, the scheme provided in this application has a significant advantage in the cost of acquiring training data. The scheme provided in this application does not require a large amount of human experience to synthesize data, and the cost is lower when acquiring positive sample pairs, making it more suitable for large-scale retrieval enhancement pre-training.
[0233] The solutions provided in this application can be applied to fields such as information retrieval, natural language processing, recommendation, and intelligent question-answering systems. For information retrieval, applying the text representation vectors encoded by the trained small language model to document retrieval can improve the relevance and accuracy of search results. In the field of natural language processing, the solutions provided in this application can be applied to tasks such as text classification and sentiment analysis. In recommendation systems, the solutions provided in this application can improve the personalization and accuracy of recommendations. In intelligent question-answering systems, the solutions provided in this application can improve the response quality and relevance of the question-answering system.
[0234] Alternatively, in the field of information retrieval, compression module #1 can be replaced with other more efficient compression algorithms to improve retrieval speed; in the field of natural language processing, LLM can be replaced with other types of pre-trained models to adapt to different task requirements. This application does not impose any limitations on these aspects.
[0235] To facilitate understanding of the above embodiments provided in this application, the following points are made.
[0236] In this application, words such as “exemplary,” “for example,” “as,” “as an example,” and “as an example” are used to indicate examples, illustrations, or explanations. Any embodiment or design described as an “example” in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the term “example” is intended to present concepts in a concrete manner. In embodiments of this application, “of,” “corresponding, relevant,” and “corresponding” may sometimes be used interchangeably, and it should be noted that their intended meanings are consistent unless their distinction is emphasized.
[0237] The business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of model architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0238] References such as "in some possible implementations" as used in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, phrases such as "in some possible implementations" appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including, but not limited to," unless otherwise specifically emphasized.
[0239] In the various embodiments of this application, unless otherwise specified or in case of logical conflict, the terminology and / or descriptions of different embodiments are consistent and can be referenced by each other. The technical features of different embodiments can be combined to form new embodiments according to their inherent logical relationship.
[0240] In this application, "at least one" or "at least one item" refers to one or more items, and "more than one" refers to two or more items. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0241] In this application, terms such as "first," "second," "#1," and "#2" are used merely for descriptive convenience to distinguish objects and are not intended to limit the scope of the embodiments of this application. They are not used to describe the order or sequence of features. It should be understood that such described objects can be interchanged where appropriate to describe solutions other than those in the embodiments of this application.
[0242] The methods of the embodiments of this application have been described in detail above with reference to Figures 4 to 7. In order to implement the functions of the methods provided in this application, both the transmitting device and the receiving device may include hardware structures and / or software modules, and the above functions may be implemented in the form of hardware structures, software modules, or hardware structures plus software modules. Whether a certain function is implemented in the form of hardware structures, software modules, or hardware structures plus software modules depends on the specific application and design constraints of the technical solution.
[0243] The data processing apparatus of this application embodiment is described below with reference to Figures 8 to 10.
[0244] Figure 8 is a schematic block diagram of a device 800 provided in an embodiment of this application.
[0245] In one possible implementation, the device 800 can be implemented by software, hardware, or a combination of both.
[0246] In one possible implementation, the apparatus 800 provided in this application embodiment can implement the method shown in FIG7 of this application embodiment. The apparatus 800 includes a processing module 820.
[0247] The processing module 820 is used to determine the text retrieval requirement to be processed; input the text retrieval requirement into the first small language model to obtain the retrieval result. The first small language model is obtained by training a second small language model based on training data and loss value. The training data includes input text, output text and task vector. The loss value is determined by the large language model based on the input vector, the task vector and the output vector. The input vector is a vector generated by the second small language model to represent the input text. The output vector is a vector generated by the large language model based on the standard answer to represent the output text.
[0248] In some possible implementations, the device 800 may further include a transceiver module 810. The transceiver module 810 is used to receive data, which includes the text retrieval request to be processed; the processing module 820 is used to determine the text retrieval request to be processed based on the data.
[0249] In another possible implementation, the apparatus 800 provided in this application embodiment can implement the method shown in FIG4 of this application embodiment. The apparatus 800 includes a transceiver module 810 and a processing module 820.
[0250] The transceiver module 810 is used for the small language model to receive input text, which, along with the output text and the task vector, belongs to a set of training data. The processing module 820 is used for the small language model to generate an input vector that represents the input text. The processing module 820 is also used for the small language model to train itself based on a loss value, which is determined by the large language model based on the input vector, the task vector, and the output vector. The output vector is a vector generated by the large language model based on the standard answer that represents the output text.
[0251] In another possible implementation, the apparatus 800 provided in this application embodiment can implement the method shown in FIG5 of this application embodiment. The apparatus 800 includes a processing module 820.
[0252] The processing module 820 is used to compress or increase the encoder dimension of the input vector to a compressed dimension, which is determined based on the decoder embedding dimension of the large language model, and the input vector is generated by the small language model; the processing module 820 is also used to compress the sequence length of the input vector to a first length, and the product of the compressed dimension and the first length is used to train the small language model.
[0253] In some possible implementations, the processing module 820 is also used to compress the sequence length of the task vector to a second length based on a reference vector generated by the large language model, and the product of the compression dimension and the second length is used to train the small language model.
[0254] For example, the processing module 820 may include a compression module or a module with compression function (a module with compression function may also be called a compression module).
[0255] The device 800 can be used to execute the above-described methods 400, 500 or 700; the specific implementation of the above-described methods 400, 500 or 700 based on the device 800 can be referred to the textual description of the above-described methods 400, 500 or 700, and will not be repeated in this embodiment.
[0256] The device 800 here may be embodied in the form of a functional module. The term "module" here can be implemented in software and / or hardware, without specific limitation.
[0257] For example, a processing module can be a processor or a processing unit; a transceiver module can be a transceiver or a transceiver unit.
[0258] For example, a "module" can be a software program, a hardware circuit, or a combination of both that implements the above functions. For instance, the implementation of transceiver module 810 in device 800 will be described below. Similarly, the implementation of other modules in device 800, such as processing module 820, can refer to the implementation of transceiver module 810.
[0259] As an example of a software functional unit, the transceiver module 810 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned computing instance may be one or more. For example, the transceiver module 810 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.
[0260] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.
[0261] As an example of a hardware functional unit, the transceiver module 810 may include at least one computing device, such as a server. Alternatively, the transceiver module 810 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.
[0262] The transceiver module 810 includes multiple computing devices that can be distributed in the same region or in different regions. Similarly, the transceiver module 810 can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the transceiver module 810 can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0263] Therefore, the modules of the various examples described in the embodiments of this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0264] It should be noted that the above embodiments of the device 800, when executing the above methods, are only illustrative examples of the division of functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device 800 can be divided into different functional modules to complete all or part of the functions described above. For example, the transceiver module 810 can be used to execute any step in the above methods, and the processing module 820 can be used to execute any step in the above methods. The steps implemented by the transceiver module 810 and the processing module 820 can be specified as needed, and all the functions of the device 800 can be realized by implementing different steps in the above methods through the transceiver module 810 and the processing module 820 respectively.
[0265] Furthermore, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments above, which will not be repeated here.
[0266] The method provided in this application can be executed by a computing device, which can also be referred to as a computer system. It includes a hardware layer, an operating system layer running on top of the hardware layer, and an application layer running on the operating system layer. The hardware layer includes hardware such as processing units, memory, and memory control units; the functions and structure of this hardware will be described in detail later. The operating system can be any one or more computer operating systems that implement business processing through processes, such as Linux, Unix, Android, iOS, or Windows. The application layer includes applications such as browsers, address books, word processing software, and instant messaging software. Optionally, the computer system can be a handheld device such as a smartphone, or a terminal device such as a personal computer; this application does not particularly limit this, as long as the method provided in this application can be used. The executing entity of the method provided in this application can be a computing device, or a functional module within the computing device capable of calling and executing programs.
[0267] The following describes in detail, with reference to Figure 9, a computing device provided in an embodiment of this application.
[0268] Figure 9 is a schematic diagram of the architecture of a computing device 900 provided in an embodiment of this application.
[0269] The computing device 900 can be a server, a computer, or other device with computing capabilities. The computing device 900 shown in Figure 9 includes at least one processor 910 and a memory 920.
[0270] It should be understood that this application does not limit the number of processors and memories in the computing device 900.
[0271] The processor 910 executes instructions in the memory 920, causing the computing device 900 to implement the method provided in this application. Alternatively, the processor 910 executes instructions in the memory 920, causing the computing device 900 to implement the various functional modules provided in this application, thereby implementing the method provided in this application.
[0272] Optionally, the computing device 900 also includes a communication interface 930. The communication interface 930 uses a transceiver module, such as, but not limited to, a network interface card or a transceiver, to enable communication between the computing device 900 and other devices or communication networks.
[0273] Optionally, the computing device 900 also includes a system bus 940, wherein the processor 910, memory 920, and communication interface 930 are respectively connected to the system bus 940. The processor 910 can access the memory 920 through the system bus 940; for example, the processor 910 can perform data read / write or code execution in the memory 920 through the system bus 940. The system bus 940 is a peripheral component interconnect express (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The system bus 940 is divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in Figure 9, but this does not mean that there is only one bus or one type of bus.
[0274] In one possible implementation, the processor 910 primarily functions to interpret the instructions (or code) of a computer program and process data within the computer software. The instructions of the computer program and the data within the computer software can be stored in memory 920 or cache 916.
[0275] Optionally, the processor 910 may be an integrated circuit chip with signal processing capabilities. By way of example and not limitation, the processor 910 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. Among these, a general-purpose processor is a microprocessor, etc. For example, the processor 910 is a CPU.
[0276] Optionally, each processor 910 includes at least one processing unit 912 and a memory control unit 914.
[0277] Optionally, the processing unit 912, also known as the core, is the most important component of the processor. The processing unit 912 is manufactured from single-crystal silicon using a specific production process. All calculations, command reception, command storage, and data processing are performed by the core. Each processing unit independently executes program instructions, utilizing parallel computing capabilities to accelerate program execution. Various processing units have fixed logical structures; for example, a processing unit includes logical units such as L1 cache, L2 cache, execution unit, instruction-level unit, and bus interface.
[0278] In one implementation example, the memory control unit 914 controls the data interaction between the memory 920 and the processing unit 912. Specifically, the memory control unit 914 receives memory access requests from the processing unit 912 and controls access to memory based on the memory access requests. By way of example and not limitation, the memory control unit is a device such as a memory management unit (MMU).
[0279] In one implementation example, each memory control unit 914 addresses the memory 920 via the system bus. An arbitrator (not shown in Figure 9) is configured on the system bus to handle and coordinate contention for access by multiple processing units 912.
[0280] In one implementation example, the processing unit 912 and the memory control unit 914 are connected via internal chip connection lines, such as address lines, thereby enabling communication between the processing unit 912 and the memory control unit 914.
[0281] Optionally, each processor 910 also includes a cache 916, which is a buffer for data exchange (called a cache). When the processing unit 912 needs to read data, it first looks for the required data in the cache. If the data is found, it is executed directly; otherwise, it looks for the data in memory. Since the cache operates much faster than memory, its role is to help the processing unit 912 run faster.
[0282] The memory 920 provides runtime space for processes in the computing device 900. For example, the memory 920 stores the computer program (specifically, the program code) used to generate the process. After the computer program is run by the processor to generate a process, the processor allocates corresponding storage space for the process in the memory 920. Furthermore, the aforementioned storage space further includes text segments, initialized data segments, bit initialized data segments, stack segments, heap segments, etc. The memory 920 stores data generated during the process's execution, such as intermediate data or process data, in the aforementioned process-specific storage space.
[0283] Optionally, the memory, also known as RAM, is used to temporarily store the data processed by the processor 910, as well as data exchanged with external storage devices such as hard disks. As long as the computer is running, the processor 910 will load the data that needs to be processed into RAM for processing, and after the processing is completed, the processing unit 912 will send the result out.
[0284] By way of example and not limitation, memory 920 may be a non-transitory computer-readable storage medium, which may include volatile memory or non-volatile memory, or both. Non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory is random access memory (RAM) used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory 920 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0285] The structure of the computing device 900 listed above is merely illustrative and is not limited thereto. The computing device 900 in this application includes various hardware components in computer systems of related technologies. For example, the computing device 900 also includes other memories besides the memory 920, such as disk storage. Those skilled in the art should understand that the computing device 900 may also include other devices necessary for normal operation. Furthermore, depending on specific needs, those skilled in the art should understand that the above-mentioned computing device 900 may also include hardware devices for implementing other additional functions. In addition, those skilled in the art should understand that the above-mentioned computing device 900 may only include the devices necessary for implementing the embodiments of this application, and does not necessarily include all the devices shown in FIG. 9.
[0286] Figure 10 is a schematic diagram of the architecture of a computing device cluster provided in an embodiment of this application.
[0287] The computing device cluster includes at least one computing device. This computing device may be a server. In some embodiments, the computing device may also be a terminal device such as a desktop computer, laptop computer, or smartphone.
[0288] As shown in Figure 10, the computing device cluster includes at least one computing device 900. The memory 920 of one or more computing devices 900 in the computing device cluster may store the same instructions for performing the methods described above.
[0289] In some possible implementations, the memory 920 of one or more computing devices 900 in the computing device cluster may also store partial instructions for executing the above-described methods. In other words, a combination of one or more computing devices 900 can jointly execute the instructions of the above-described methods.
[0290] It should be noted that the memories 920 in different computing devices 900 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the aforementioned apparatus. That is, the instructions stored in the memories 920 of different computing devices 900 can implement the functions of one or more modules within the aforementioned apparatus.
[0291] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc.
[0292] For example, two computing devices 900A and 900B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device.
[0293] It should be understood that the functions of computing device 900A can also be performed by multiple computing devices 900. Similarly, the functions of computing device 900B can also be performed by multiple computing devices 900.
[0294] This application also provides a model training system, which may include the above-described apparatus 800 for model training, or the first apparatus and the second apparatus.
[0295] This application also provides a computer program product containing instructions, which may be a software or program product containing instructions capable of running on a computing device or stored on any usable medium. When run on a computing device, it causes the computing device to perform the methods provided above, or causes the computing device to perform the functions of the apparatus provided above.
[0296] This application also provides a computer-readable storage medium, which can be any available medium capable of being stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that, when executed on a computing device, cause the computing device to perform the method provided above.
[0297] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0298] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0299] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0300] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0301] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0302] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0303] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0304] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A text retrieval method, characterized in that, include: Identify the text retrieval requirements to be processed; The text retrieval request is input into the first small language model to obtain the retrieval results. The first small language model is obtained by training the second small language model based on training data and loss values. The training data includes input text, output text, and task vector. The loss value is determined by the large language model based on the input vector, the task vector, and the output vector. The input vector is a vector generated by the second small language model to represent the input text, and the output vector is a vector generated by the large language model based on the standard answer to represent the output text.
2. The method according to claim 1, characterized in that, The loss value is determined by the large language model based on the input vector, the task vector, and the output vector, and includes: The loss value is determined by the large language model based on a first compression result and / or a second compression result, wherein the first compression result is obtained by compressing the sequence length of the input vector, and the second compression result is obtained by compressing the sequence length of the task vector.
3. The method according to claim 2, characterized in that, The first compression result includes the product of a first length and a compression dimension, where the first length is the result of compressing the sequence length of the input vector; The second compression result includes the product of a second length and the compression dimension, wherein the second length is the result of compressing the sequence length of the task vector based on a reference vector generated by the large language model; The compression dimension is determined based on the decoder embedding dimension of the large language model.
4. The method according to claim 2 or 3, characterized in that, The text retrieval request is input into the first small language model to obtain the retrieval results, including: The first small language model generates a first input vector based on the text retrieval requirements; The first compression model outputs a second input vector based on the first input vector, where the second input vector is the compression result of the first input vector; The first small language model processes the text retrieval request based on the second input vector. The first compression model is obtained by training a second compression model based on the training data and the loss value, and the second compression model is used to determine the first compression result and / or the second compression result.
5. The method according to any one of claims 1 to 4, characterized in that, The loss value includes the loss corresponding to the output vector.
6. The method according to any one of claims 1 to 5, characterized in that, The training data includes question-and-answer data and / or summary data.
7. A method for training a model, characterized in that, Applied to small language models, including: Receive input text, wherein the input text, output text, and task vector belong to a set of training data; Generate an input vector to represent the input text; The small language model is trained based on a loss value, which is determined by the large language model based on the input vector, the task vector, and the output vector. The output vector is a vector generated by the large language model based on the standard answer to represent the output text.
8. The method according to claim 7, characterized in that, The loss value is also used to train the compression model; The compression model is used to determine the product of a first length and a compression dimension, where the first length is the result of compressing the sequence length of the input vector; And / or, The compression model is used to determine the product of the second length and the compression dimension, wherein the second length is the result of compressing the sequence length of the task vector based on a reference vector generated by the large language model; The compression dimension is determined based on the decoder embedding dimension of the large language model.
9. A method for compressing vectors, characterized in that, include: Compress or increase the encoder dimension of the input vector to a compressed dimension, wherein the compressed dimension is determined based on the decoder embedding dimension of the large language model, and the input vector is generated by the small language model; The sequence length of the input vector is compressed to a first length, and the product of the compressed dimension and the first length is used to train the small language model.
10. The method according to claim 9, characterized in that, The method further includes: The sequence length of the task vector is compressed to a second length based on a reference vector generated by the large language model, and the product of the compression dimension and the second length is used to train the small language model.
11. The method according to claim 10, characterized in that, The input vector is generated based on the input text; The input text, output text, and task vector belong to a set of training data. The output vector is a vector generated by the large language model based on the standard answer to represent the output text. The product of the compressed dimension and the first length, the product of the compressed dimension and the second length, and the output vector are used to determine the loss value for the large language model, and the loss value is used to train the small language model.
12. A text retrieval device, characterized in that, include: The processing unit is used to determine the text retrieval requirements to be processed. The processing unit is also used to input the text retrieval request into the first small language model to obtain retrieval results. The first small language model is obtained by training the second small language model based on training data and loss values. The training data includes input text, output text, and task vector. The loss value is determined by the large language model based on the input vector, the task vector, and the output vector. The input vector is a vector generated by the second small language model to represent the input text, and the output vector is a vector generated by the large language model based on the standard answer to represent the output text.
13. A device for model training, characterized in that, include: A receiving unit is used to receive input text, wherein the input text, output text, and task vector belong to a set of training data. A processing unit is configured to generate an input vector that characterizes the input text; The processing unit is also used to train a small language model based on a loss value, which is determined by the large language model based on the input vector, the task vector, and the output vector. The output vector is a vector generated by the large language model based on the standard answer to represent the output text.
14. An apparatus for compressing vectors, characterized in that, include: A compression unit is used to compress or increase the encoder dimension of the input vector to a compressed dimension, wherein the compressed dimension is determined based on the decoder embedding dimension of the large language model, and the input vector is generated by the small language model. The compression unit is also used to compress the sequence length of the input vector to a first length, and the product of the compression dimension and the first length is used to train the small language model.
15. A chip or chip system, characterized in that, Includes: a circuit for performing the method as described in any one of claims 1 to 6, or the method as described in claim 7 or 8, or the method as described in any one of claims 9 to 11.
16. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1 to 6, or the method as described in claim 7 or 8, or the method as described in any one of claims 9 to 11.
17. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster performs the method as described in any one of claims 1 to 6, or the method as described in claim 7 or 8, or the method as described in any one of claims 9 to 11.
18. A computer-readable storage medium, characterized in that, It includes computer program instructions that, when executed by a cluster of computing devices, perform the method as described in any one of claims 1 to 6, or the method as described in claim 7 or 8, or the method as described in any one of claims 9 to 11.